Skip to main content
Glama
rthomas24

WebVector MCP Server

by rthomas24

WebVector

npm: webvector npm: webvector-mcp npm: webvector-cli CI License: MIT

Web research for AI agents in one call: search → read the full pages → rank → cited passages. No API keys, no model download, ~12 MB. Ships as an MCP server, a library and a CLI.

npx -y webvector-cli search "what changed in the MCP spec in 2026?"
**[1]** Streamable HTTP — Model Context Protocol — <https://modelcontextprotocol.io/specification/2026-07-28/…>
> Protocol versions 2025-03-26 through 2025-11-25 also used the Streamable HTTP transport, but in a
> different shape: servers could assign a session via the Mcp-Session-Id header … None of these
> mechanisms are part of this revision.
## Sources
- Streamable HTTP — Model Context Protocol — <https://…> [1]

WebVector CLI demo

Quick start: MCP server

Cursor, Claude Code, Claude Desktop, Windsurf, VS Code, Zed — paste into your MCP config (mcp.json / claude_desktop_config.json / .vscode/mcp.json):

{
  "mcpServers": {
    "webvector": {
      "command": "npx",
      "args": ["-y", "webvector-mcp"]
    }
  }
}

Or via Claude Code CLI:

claude mcp add webvector -- npx -y webvector-mcp

What you get: webvector_research (one-call web research with citations), webvector_fetch (read any URL), webvector_search (result list only), webvector_verify (citation checker), webvector_status (diagnostics). Zero config, no API keys required. Free OSS; you only pay upstream providers you opt into (Brave, Serper, OpenAI embeddings, etc. — DuckDuckGo is the default and requires no key).

Optional semantic tier: Add local ONNX embeddings (offline) or set OPENAI_API_KEY / VOYAGE_API_KEY / GEMINI_API_KEY — ranking upgrades from BM25 to hybrid automatically.

Run it

MCP server (Claude Code, Claude Desktop, Cursor, Windsurf, VS Code, Zed …):

claude mcp add webvector -- npx -y webvector-mcp
{ "mcpServers": { "webvector": { "command": "npx", "args": ["-y", "webvector-mcp"] } } }

Library:

import { WebVector } from 'webvector';
const wv = new WebVector();
const res = await wv.research('reciprocal rank fusion k constant', { relatedQueries: ['RRF formula'] });
console.log(res.markdown);          // cited passages, ready for a prompt
console.log(res.evidence?.level);   // 'strong' | 'weak' | 'none' + suggestedQueries

CLI: npm i -g webvector-cliwebvector search "…" -k 8, webvector fetch <url> --query "…", webvector doctor.

Semantic tier (optional): npm i @huggingface/transformers (local ONNX embeddings, offline) or set OPENAI_API_KEY / VOYAGE_API_KEY / GEMINI_API_KEY … — ranking upgrades from BM25 to hybrid automatically. webvector doctor shows the active tier.

What it does

Capability

Example

One-call research — search, fetch every result (HTML, PDF, served Markdown), chunk, rank, cite

wv.research(q) · MCP webvector_research · webvector search q

Hybrid ranking that works keyless — BM25F (title/heading/body fields, proximity, identifiers like AbortSignal.any), vectors when available, relative-score fusion, per-source/domain diversity, adjacent-chunk merge

retrieval.bm25.*, retrieval.fusion, retrieval.maxPerDomain

Sub-questions covered — pass related_queries; xQuAD-style selection guarantees every aspect gets passages

{ relatedQueries: ['UDP connectionless', 'TCP handshake'] }res.coverage

Evidence gate + follow-ups — LLM-free verdict (strong / weak / none), suggested queries, optional one in-call retry

res.evidence, { autoRetry: 1 }, MCP auto_retry

Highlights, token budgets, deep links — best sentence per passage, packing into max_tokens with an explicit "N omitted" footer, url#:~:text= citations, #page=N for PDFs

output.passageMode: 'highlight', max_tokens, output.deepLinks

Verify citations — classify each sentence of an answer as verbatim / paraphrase / unsupported against the cited passages; flags numbers not in the source

wv.verifyCitations(answer, { sessionId }) · MCP webvector_verify · webvector verify

Read one page well — pagination (start_index), CSS selector, link lists, query-focused passages

MCP webvector_fetch · wv.fetch(url, { selector })

Fetch more pages, cleaner — markdown-first content negotiation (10–100× smaller on docs sites), fast paths (arXiv HTML, GitHub README/issues, Hacker News & Stack Exchange APIs, Google Docs), extractor ensemble with a recall guard, JS-shell detection (PARSE_NEEDS_JS) + optional render hook, __NEXT_DATA__ recovery, boilerplate suppression

ingestion.acceptMarkdown, ingestion.fastPaths, ingestion.html.strategy, ingestion.render

Fast on repeat — SQLite page cache with ETag revalidation (second run: 0 requests), persistent embedding cache, single-flight, per-call max_age_ms / cache_mode

~/.cache/webvector/pages.sqlite, webvector cache stats

Sessions & stores — pages read once are reused across calls; memory / sqlite / Chroma / Qdrant / pgvector

store.mode: session, store.provider: sqlite

Providers — 11 search (DuckDuckGo default, Brave, Serper, Tavily, Exa, SearXNG …), 9 embedding, 5 rerankers, custom in one function

docs/PROVIDERS.md

Agent-ready MCP — namespaced tools, ≤2 KB instructions, concise/detailed output, depth presets, objective, sessions, --max-uses / --allowed-domains guardrails, research & verify_claim prompts; adapters for Anthropic (search_result blocks), OpenAI, Vercel AI SDK, LangChain

packages/mcp

Polite & safe — robots.txt + Content-Signal, per-host pacing, honest UA, SSRF guard, bot-wall detection (never retried), size/time caps, no telemetry, secrets redacted

SECURITY.md

Measured — offline eval over 32 recorded cases + 40-fixture extraction corpus run in CI; ranking changes are gated on it

npm run eval · eval/

Markets (opt-in) — ticker/market news from free feeds (deduped, event-tagged), SEC EDGAR filings + full-text search, macro/Fed/earnings calendar, StockTwits + FINRA short volume, VIX/yields pulse; every source classified open/feed/gray, gray off by default

wv.markets.news(...) · MCP --tools markets · docs/MARKETS.md

Why WebVector

WebVector

Firecrawl

Jina Reader

Tavily / Exa

Playwright / Browserbase

Cursor/Claude WebSearch/WebFetch

Research pipeline

Search + fetch full pages + rank + cited passages in one call

Manual orchestration of crawl → LLM

Single-page read or search

Search only (no full pages) or API fetch

Manual browser scripting

Search returns snippets; fetch returns full page dump

Wedge

The finished research call, not a step

Deep site crawling (we don't)

Clean single-page markdown

Hosted search API (we can use as provider)

JS-heavy SPAs (we detect PARSE_NEEDS_JS, need render hook)

Built-in convenience; no ranking or citations

API keys

None (DuckDuckGo default); opt into providers

Required

Free tier, then key

Required

Required (+ browser infra)

Built into client (key implicit)

Runs where

Local Node process

Hosted service

Hosted service

Hosted service

Local or hosted browser

Client MCP or built-in

Output

Ranked cited passages, evidence gate, token budgets

Raw crawled content or LLM-processed

Clean markdown of one page

Search results with snippets

Full page content + JS state

Search snippets or raw HTML/markdown

Best for

Agents researching the live web with citations

Crawling entire sites, sitemaps, dynamic content

Reading one clean page

Hosted search when you need a key-based API

SPAs, forms, auth flows, screenshots

Quick built-in search or page fetch

Fair comparison: Firecrawl crawls sites (we don't); Jina Reader excels at single-page markdown (we use markdown-first content negotiation but focus on multi-page research); Tavily/Exa are hosted search we can use as providers; Playwright handles JS-heavy SPAs (we don't, unless you plug in a renderer); Cursor/Claude built-in tools are snippets-or-dump vs our ranked cited passages with evidence gating.

Configure

Zero config works. Otherwise webvector.config.yaml (with editor autocomplete via $schema) or WEBVECTOR_* env vars — every key in docs/CONFIGURATION.md. webvector init writes a starter file.

Docs

Full guide · Configuration · Providers · Markets · Architecture · MCP server · CLI · Security · Contributing · Eval

Develop

git clone https://github.com/rthomas24/web-vector && cd web-vector
npm install && npm run build && npm test && npm run eval

Requires Node ≥ 22.12. MIT © Ryan Thomas.

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/rthomas24/web-vector'

If you have feedback or need assistance with the MCP directory API, please join our Discord server