Skip to main content
Glama
aksdrx

mcp-human-search

by aksdrx
README.md
# mcp-human-search

Human-like web search as a [Model Context Protocol](https://modelcontextprotocol.io) server — a cross-harness port of [dsh-human-search](https://github.com/aksdrx/dsh-human-search). A **real browser drives real search engines**, with an ordered fallback chain across **Google, DuckDuckGo, Bing, Baidu, and Sogou**, CAPTCHA/account-login handoff to you, and cookies that persist per engine — owned by this server only.

Any MCP-capable harness can use it: **DeepSeek Harness**, **pi.dev**, **Claude Code**, **Cursor**, **VS Code**, and anything else that speaks MCP over stdio.

No harness installation is modified. Everything the server writes lives under one state root — `$HUMAN_SEARCH_HOME` or `~/.human-search/` (profiles, cookies, optional server-managed Chromium). Remove that directory and the machine is exactly as it was.

## How it works

- The harness launches the server over **stdio** and gets three tools: `human_search`, `human_search_sign_in`, `human_search_status`.
- Each search launches (or reuses) a **persistent, server-private Chromium profile per engine**, opens the engine's home page, types the query like a person (per-character jitter, small pauses), and reads the organic results.
- The typed submit is **verified, not assumed**: the search box's value is read back (a page whose JavaScript has not hydrated yet silently swallows keystrokes), the autosuggest panel is dismissed before Enter (with it open, Enter submits a trending suggestion instead of your query), and the settled page is checked — a SERP for a different query, a "Loading…" bot-check stub, or results sharing no word with the query (a decoy SERP under IP-reputation pressure) is never trusted. Anything unusable retries once through the engine's results URL, then fails over.
- Engines are tried strictly in your configured order. Any failure — CAPTCHA or bot wall, timeout, parse failure, zero results, an engine busy with your sign-in window — fails over to the next engine. The returned result notes which engine served it and why others were skipped.
- When an engine blocks with a CAPTCHA:
  1. the search **fails over immediately** (the other engines keep answering), and
  2. the server opens a **headed browser window on the machine running the server** with that engine's profile, so you can solve the CAPTCHA and optionally sign in to your account (Google account for Google, Microsoft account for Bing, Baidu passport, …). An MCP logging notification is emitted for harnesses that surface them (pi.dev appends them to `~/.pi/agent/mcp.log`).
  3. once the block clears, cookies are persisted (in the profile and as a portable `storageState` snapshot) and the engine rejoins the chain automatically.
- A blocked engine is skipped for a 10-minute cooldown after a failed sign-in attempt so the chain's budget isn't burned on it; call `human_search_sign_in` any time to re-open the window.

## Install

Prerequisites: Node ≥ 20.

```sh
# run directly with npx (no install step)
npx -y mcp-human-search

# or from a local checkout (development)
pnpm install && pnpm build
node lib/cli.js
```

### Browser setup

The server uses, in order:

1. an explicit **browser executable** path (`executablePath` in the config file, or `HUMAN_SEARCH_BROWSER`),
2. your system **Google Chrome / Microsoft Edge / Chromium**,
3. a **server-managed Chromium** under `<state root>/browsers/`.

For 3, install once per machine:

```sh
npx -y mcp-human-search install-browser      # or: node lib/cli.js install-browser
```

Headless searches prefer the server-managed **full Chromium** (new-headless mode — a far less bot-flagged fingerprint than the dedicated headless shell; Google serves the shell a CAPTCHA wall even signed in), with the lighter headless shell as fallback. Whichever binary wins, a `HeadlessChrome` user-agent marker — an instant bot signal — is probed once and rewritten to the headful-equivalent string automatically.

Check what the server sees without connecting a harness:

```sh
npx -y mcp-human-search status
```

### First-run warm-up (recommended)

Fresh machines and datacenter/VPN IPs start with no engine reputation, which means CAPTCHAs or decoy results. Warm the profiles once, as a human, from a desktop session (WSLg/X11 on Windows counts):

```sh
npx -y mcp-human-search warm google duckduckgo bing   # any subset / order
```

One headed window per engine opens on the server's shared profile: run one search, solve any challenge that appears, optionally sign in (Google/Microsoft/Baidu accounts all raise the trust floor), then close the window to advance. Cookies persist in the profile the headless chain reuses for every future search. Set `HUMAN_SEARCH_BROWSER=/path/to/chrome` to warm with a specific binary.

## Harness configuration

### DeepSeek Harness (DSH)

Add one row via [`dsh-mcp-client`](https://github.com/deepseek-ai/deepseek-harness) to your composition:

```yaml
- id: mcp-human-search
  name: '@deepseek-ai/dsh-mcp-client'
  config:
    serverName: human-search
    transport: stdio
    command: npx
    args: ['-y', 'mcp-human-search']
```

The tools appear as `mcp__human-search__human_search` etc. Note this **adds** a tool alongside the stock `web_search` — steer the model with prompting (or disable the stock tool) if you want human search preferred. The default 50 s chain budget fits DSH's 60 s `toolCallTimeoutMs`. If you want the stock `web_search` tool itself backed by this engine chain, use the native [dsh-human-search](https://github.com/aksdrx/dsh-human-search) plugin instead.

### pi.dev

```sh
pi mcp add human-search -- npx -y mcp-human-search
```

or edit `~/.pi/agent/mcp.json` (project-level: `.pi/mcp.json`):

```json
{
  "mcpServers": {
    "human-search": {
      "command": "npx",
      "args": ["-y", "mcp-human-search"],
      "exposure": "direct",
      "timeout": 90,
      "description": "Human-like web search via a real browser with engine fallback (Google, DuckDuckGo, Bing, Baidu, Sogou)"
    }
  }
}
```

`exposure: "direct"` declares the three tools like built-ins — the default `codemode` exposure would hide them behind script discovery, which rarely pays off for a three-tool server. `timeout` (seconds, default 60) gives headroom over the 50 s chain budget. Run `/mcp` in a session (or `pi mcp list` in a shell) to inspect the connection.

### Claude Code

```sh
claude mcp add human-search -- npx -y mcp-human-search
```

### Cursor / VS Code

Standard `mcpServers` entry:

```json
{
  "mcpServers": {
    "human-search": { "command": "npx", "args": ["-y", "mcp-human-search"] }
  }
}
```

## Tools

### `human_search`

Input: `{ query: string, maxResults?: number (1–50, default 15) }`.

Runs the ordered fallback chain. Success returns a provenance note ("Human web search served by Google. Skipped: DuckDuckGo: blocked (…).") followed by a numbered result list, plus `structuredContent` (`servedBy`, `sources`, `truncated`, `attempts`) for clients that consume it. When every engine fails, the tool returns an error with the per-engine trail and a fix hint.

### `human_search_sign_in`

Input: `{ engine: "google" | "duckduckgo" | "bing" | "baidu" | "sogou" }`.

Opens a headed browser window **on the machine running this server**, on that engine's persistent profile, so you can solve a CAPTCHA and/or sign in. You type credentials into the engine's own pages; the server never sees them. Returns immediately: window opened / already open / cannot open (with headless-host guidance). While the window is open, searches on that engine fail over instantly instead of queueing; the episode times out after 10 minutes.

### `human_search_status`

No input. Reports — without any network calls — whether a usable browser was found, per-engine health (ok / blocked / signing-in, with reasons), sign-in windows currently open, the state root, and the resolved configuration.

## Configuration

Configuration is resolved **once at server start** (restart to reload): defaults → `<state root>/config.json` → environment variables.

`<state root>/config.json` (all keys optional):

```json
{
  "engines": [{ "id": "google", "enabled": true }, { "id": "bing", "enabled": true }],
  "headless": true,
  "locale": "",
  "executablePath": "",
  "perEngineTimeoutMs": 15000,
  "chainBudgetMs": 50000,
  "idleCloseMs": 300000
}
```

- **engines** — the array order is the fallback order; engines not listed are appended (enabled) in the default order, so the chain never silently loses one. Set `"enabled": false` to exclude one.
- **headless** — `false` shows every search in a visible window (debugging).
- **locale** — overrides every engine's default locale.
- **executablePath** — explicit Chrome/Edge/Chromium binary.
- **perEngineTimeoutMs** — per-engine budget (1 000–60 000).
- **chainBudgetMs** — whole-chain budget (10 000–120 000); keep it below your harness's tool-call timeout (DSH and pi.dev default to 60 s).
- **idleCloseMs** — close idle engine browsers after this long; `0` keeps them alive.

Environment overrides (win over the file):

| Variable | Meaning |
|---|---|
| `HUMAN_SEARCH_HOME` | State root (default `~/.human-search`) |
| `HUMAN_SEARCH_ENGINES` | Comma list: the enabled set **and** its order (`bing,google` = only those two, Bing first) |
| `HUMAN_SEARCH_BROWSER` | Explicit browser executable path |
| `HUMAN_SEARCH_HEADLESS` | `0`/`false`/`no` runs searches headed |
| `HUMAN_SEARCH_LOCALE` | Locale override |
| `HUMAN_SEARCH_PER_ENGINE_TIMEOUT_MS` | Per-engine budget |
| `HUMAN_SEARCH_CHAIN_BUDGET_MS` | Whole-chain budget |
| `HUMAN_SEARCH_IDLE_CLOSE_MS` | Idle browser close delay (`0` = keep alive) |

A malformed `config.json` is a stderr warning and defaults — never a crash.

Engine notes, learned from live validation:

- Result links wrapped in engine redirects (Google `/url?q=`, Bing `/ck/a` base64 payloads) are unwrapped to their targets; Baidu and Sogou redirect links are kept as-is (they resolve for the reader).
- Extracted results must share at least one word with the query (for Latin-script queries). An engine under IP-reputation pressure sometimes serves a perfectly formed SERP whose results are unrelated decoys; those are treated as "no results" and the chain fails over instead of citing junk.
- Baidu and Sogou are Chinese engines and can be slow outside China; raise `perEngineTimeoutMs` if they time out on your network.
- If an engine keeps failing on your IP, run the [warm-up](#first-run-warm-up-recommended) once — a minute of real usage builds more trust than any amount of headless retrying.

## CAPTCHA and account login

- The interactive window appears **on the machine running the server**. When that's your desktop or laptop, it's your screen. On a headless server there is no display — the server logs why, keeps failing over, and you can warm the profiles on a desktop with `mcp-human-search warm` against the same `HUMAN_SEARCH_HOME`, then copy the state root over (or point `HUMAN_SEARCH_HOME` at a shared location). SSH with X forwarding also works.
- You never have to give the server credentials: you type them into the engine's own pages, in a real browser profile owned by this server. Cookies persist only there.
- While a sign-in window is open, that engine's searches fail over instantly instead of queueing.

## Migrating from dsh-human-search

The profiles are interchangeable. To reuse your warmed cookies, copy the DSH plugin's state over the new root before first use:

```sh
cp -a ~/.dsh/web-human-search/. ~/.human-search/
```

## Uninstall

Remove the MCP server entry from your harness configuration, then delete the state root:

```sh
rm -rf ~/.human-search
```

## Privacy & footprint

- Cookies, local storage, and exported `storageState` snapshots live in the state root (mode 0700) and are never sent anywhere except the engines they belong to. The model sees search results only — never cookies or credentials.
- The server performs searches exactly as rendered to a human in a browser; it does not use scraper endpoints or APIs.
- `PLAYWRIGHT_BROWSERS_PATH` is pointed at the server's own browsers directory inside the server process; a server-managed Chromium never touches other tooling's caches.

## Development

```sh
pnpm install
pnpm build        # lib/server.js + lib/cli.js (committed; git installs don't run build scripts)
pnpm test         # unit tests (chain logic, config, engine extraction, MCP tool surface)
pnpm typecheck
```

Layout: `src/engines/` per-engine adapters (pure markup descriptions), `src/chain.ts` the fallback chain and human-like driver, `src/browser.ts` the per-engine persistent-browser pool, `src/login.ts` the sign-in episode coordinator, `src/search.ts` the search tool logic, `src/server.ts` the MCP server, `src/config.ts`/`src/state.ts` startup configuration and on-disk state, `src/cli.ts` the `mcp-human-search` binary. `tests/` runs without a browser via scripted fake pages, and includes an in-memory-transport MCP client test of the full tool surface.

## License

MIT

TDQS

A4.3/5.0

Scored across 3 tools

Disambiguation5/5

Each tool targets a distinct concern: human_search_status inspects server state without networking, human_search performs the actual search, and human_search_sign_in handles CAPTCHA/credential solving. The descriptions explicitly cross-reference each other (e.g. 'Use human_search_status to inspect engine health'), leaving no realistic misselection.

Naming Consistency4/5

All three tools share the human_search_ prefix and snake_case, giving a clear family. The main action tool drops the suffix entirely (human_search instead of human_search_run/search), a minor deviation from the otherwise predictable prefix+qualifier pattern.

Tool Count5/5

Three tools cleanly cover the server's narrow purpose of browser-driven human-like search: search, inspect status, and solve sign-in. Nothing is redundant and nothing feels missing by way of extra surface area.

Completeness4/5

The search/status/sign-in lifecycle is well covered, including failover and provenance reporting. Minor gaps exist—no tool to close a pending sign-in window, clear/logout cookies, or list warm profiles (some of this is only reachable via the documented CLI).

Maintenance

ActivityMaintained
ResponsivenessNo issues