Skip to main content
Glama
README.md
# websearch-mcp

Semantic web search as an MCP server, using **no general search engine** — no Google, Bing, DuckDuckGo, or Exa. A cheap LLM navigates the live web, a local embedding model indexes everything it walks past, and the index answers the next similar question by itself.

## The idea

Building a general web index means crawling billions of pages. This does the opposite: **it looks things up only when asked, and remembers what it found.**

```
query → local vector index
  ├─ good hit  → return immediately (milliseconds, no API cost)
  └─ no hit    → navigate the live web:
        LLM picks entry sites → fetch page → LLM scores the page's links
        against the goal → expand only the most promising ones → repeat
        → index every page walked past
```

So it **gets faster the more you use it**: a brand-new topic takes a minute or two, and anything in that neighbourhood is a cache hit afterwards.

The load-bearing trick is that **the model prunes at every node**. A page has 50–150 outgoing links, so expanding all of them is exponential and hopeless. Scoring them against the goal turns the tree into a guided beam search — affordable precisely because the model is cheap.

## Entry points

With no search engine, navigation starts from **each site's own search** — what a person does when they know which site holds the answer but not which page. Every entry below was probed live:

| Entry point | Notes |
|---|---|
| Wikipedia (en/zh) | Article URLs directly — its `/w/` search is robots-disallowed. An article on X links to the authoritative sources on X, which makes it the best general-purpose hub |
| GitHub repos / issues | Site search works |
| Hacker News | Official Algolia JSON API |
| Stack Exchange | Official JSON API (the HTML site rejects non-browser user agents) |
| MDN / PyPI / PaperMC javadoc | Site search or a fixed entry page |

Sites whose search is robots-disallowed (Reddit, ArchWiki, Baidu Baike) or that are JavaScript-only shells (devdocs, the Hacker News web UI) are **deliberately absent** — where an official API exists, that is used instead.

Add your own with `kb_add_hub`, or edit `websearch/hubs.py`.

## Tools

| Tool | Purpose |
|---|---|
| `web_search` | Semantic search. Checks the local index first, navigates the live web on a miss |
| `web_read` | Fetch one URL's main text (and index it) |
| `kb_status` | Index size, domains covered, cache hit rate |
| `kb_add_hub` | Register a site as a future entry point |

## What gets stored

Per page: **URL, title, summary, tags, and a 384-dimension vector** — about 2KB. Page text is never retained. A hundred thousand pages fit in roughly 200MB.

Retrieval is brute-force cosine over the whole matrix (~50ms at 100k pages), so there is no ANN index to build or tune until this grows by another order of magnitude.

- `index.sqlite3` — metadata
- `vectors.npy` + `vector_ids.json` — the vector matrix

## Precision

Vector similarity alone is not enough. Multilingual e5 packs unrelated text into a narrow band, so an off-topic page can still score 0.80 — early versions happily answered a question about kimchi with Minecraft documentation. Retrieval is therefore followed by a cheap LLM relevance pass, and when nothing survives it says so instead of returning whatever the index happens to hold.

## Being a good citizen

- **robots.txt** fully respected, plus a 1 req/s per-domain rate limit
- **SSRF protection** — hosts are resolved and private/loopback/link-local addresses refused
- **Prompt injection** — crawled content is always wrapped in data tags and explicitly marked as data, never instructions
- URL normalisation and content hashing to avoid re-indexing the same page

## Running it

```bash
python3 -m venv .venv
.venv/bin/pip install --extra-index-url https://download.pytorch.org/whl/cpu torch
.venv/bin/pip install -r requirements.txt

cp .env.example .env      # fill in OPENROUTER_API_KEY
OPENROUTER_API_KEY=... WEBSEARCH_DATA=/var/lib/websearch .venv/bin/python server.py
```

Wire it into any MCP client:

```json
{
  "websearch": {
    "command": "/opt/websearch-mcp/.venv/bin/python",
    "args": ["/opt/websearch-mcp/server.py"],
    "env": {
      "HOME": "/var/lib/websearch",
      "HF_HOME": "/var/lib/websearch/hf",
      "WEBSEARCH_DATA": "/var/lib/websearch"
    }
  }
}
```

Everything in `websearch/config.py` can be overridden by environment variable — navigation budget (`NAV_MAX_PAGES`, `NAV_MAX_DEPTH`, `NAV_MAX_SECONDS`), embedding model (`EMBED_MODEL`), politeness (`PER_DOMAIN_DELAY`), and so on. See `.env.example`.

Queries and summaries work in any language the embedding model covers; the LLM is asked to answer in the language of the query.

## Cost

Embeddings run locally, so indexing is free. The LLM is used only for choosing entry sites, scoring links, and judging relevance — roughly **$0.005 per cold query** and nothing at all for a cache hit, using a cheap model such as DeepSeek V4 Flash through OpenRouter.

## Honest limitations

- **This does not replace a general search engine for arbitrary queries.** It is strong in the areas you have actually used it in, and empty everywhere else.
- A brand-new topic costs one to two minutes on the first query. That is the price of navigating rather than querying a pre-built index.
- When a site redesigns or changes its robots.txt, the corresponding entry point breaks and `hubs.py` needs updating.

## License

MIT — see [LICENSE](LICENSE).

Maintenance

ActivitySlowing
ResponsivenessNo issues