Skip to main content
Glama
AceAtDev

Odysseus Web MCP

by AceAtDev

Odysseus Web MCP — Secure Web Search and Fetch Server for AI Assistants

Odysseus Web MCP is a standalone Model Context Protocol (MCP) server for safe public-web search and URL fetching. It runs locally over stdio and gives MCP-compatible AI assistants two retrieval tools: web_search to discover sources and web_fetch to retrieve and extract public URLs.

Built for clients such as Claude Code, Cursor, and Codex, it combines search-provider fallback, readable HTML/PDF/text extraction, optional JavaScript rendering, and SSRF protections including DNS validation and redirect rechecks.

Live web search terminal demo

Live web fetch terminal demo

Features

  • Search the public web with provider fallback and ranked, attributed sources.

  • Fetch and extract HTML, PDF, and text content from public URLs.

  • Protect against SSRF with public-network checks, DNS validation, and redirect revalidation.

  • Return bounded, evidence-oriented output with cursors, quality signals, and discovered links.

  • Optionally render JavaScript-heavy pages in isolated Playwright.

Related MCP server: pyaireader

Install in minutes

Requirements: Python 3.11+ and uv.

# after downloading/extracting this folder (or cloning your copy)
cd odysseus-web-mcp
uv venv .venv
uv pip install -e '.[dev]'
./run-web-mcp.sh

The server communicates over stdio, so it does not open a web port and does not need to be installed into your host application's Python environment. Register the absolute launcher path in your MCP client:

{
  "name": "odysseus-web-mcp",
  "command": "/absolute/path/to/odysseus-web-mcp/run-web-mcp.sh",
  "args": [],
  "cwd": "/absolute/path/to/odysseus-web-mcp"
}

The launcher automatically uses the package's .venv. State defaults to ~/.local/share/odysseus-web-mcp; set WEB_MCP_DATA_DIR to place it elsewhere. No API key is required for the default fallback path, though Brave, Tavily, and Serper keys can be added when you want those providers.

The two tools

Use it to discover sources for a focused question. It accepts one to three queries plus optional mode, vertical, and freshness controls.

{
  "queries": "Model Context Protocol Python SDK",
  "mode": "discovery",
  "vertical": "general"
}

The response contains ranked URLs, titles, snippets, provider attempts, cache state, a plain-text display projection, and an evidence_id. A host can take any returned URL directly into web_fetch.

web_fetch

Use it to read a known public URL or a bounded batch of URLs.

{
  "url": "https://example.com",
  "focus": "the page's purpose",
  "render": "auto"
}

It returns extracted text, title and document kind, content quality, link discovery, redirect history, HTTP status, truncation/continuation metadata, and an evidence_id. Private and special-use destinations are rejected before transport by default.

Example: how an agent uses the MCP

An agent normally uses the tools as a two-step retrieval loop: search first, then fetch the source it wants to inspect. The payloads below show the shape of a real MCP interaction; IDs and result text are abbreviated for readability.

1. Agent searches for sources

{
  "name": "web_search",
  "arguments": {
    "queries": "official Model Context Protocol architecture",
    "mode": "grounding",
    "vertical": "general"
  }
}

The MCP returns a text content block containing structured JSON:

{
  "status": "ok",
  "query": "official Model Context Protocol architecture",
  "sources": [
    {
      "title": "Architecture - Model Context Protocol",
      "url": "https://modelcontextprotocol.io/docs/concepts/architecture",
      "snippet": "Understand the architecture and communication model...",
      "provider": "duckduckgo",
      "relevance_score": 1.0
    }
  ],
  "provider_attempts": {
    "searxng": "empty",
    "duckduckgo": "ok"
  },
  "evidence_id": "a1b2c3d4...",
  "exit_code": 0
}

2. Agent fetches the selected source

The agent takes the returned URL and calls the second tool:

{
  "name": "web_fetch",
  "arguments": {
    "url": "https://modelcontextprotocol.io/docs/concepts/architecture",
    "focus": "How do clients and servers communicate?",
    "render": "auto"
  }
}

The MCP returns bounded, extracted evidence:

{
  "success": true,
  "url": "https://modelcontextprotocol.io/docs/concepts/architecture",
  "final_url": "https://modelcontextprotocol.io/docs/concepts/architecture",
  "http_status": 200,
  "document_kind": "html",
  "content_quality": "good",
  "content": "The Model Context Protocol defines how clients and servers...",
  "links": [
    {
      "url": "https://modelcontextprotocol.io/docs/concepts/transports",
      "text": "Transports"
    }
  ],
  "evidence_id": "e5f6g7h8...",
  "exit_code": 0
}

The agent can now answer the user from the extracted content, preserve the evidence_id for traceability, and continue with another web_fetch using a returned cursor if the page was longer than the output budget.

How it works locally

MCP host ──stdio──▶ mcp_server.py
                       ├─ web_search → provider chain → ranked evidence
                       └─ web_fetch  → security → HTTP/extract/render → evidence

All persistent state is rooted under WEB_MCP_DATA_DIR. The package has no runtime imports from Odysseus and no access to its credentials, database, memory, browser profiles, scheduler, or agent loop.

Read the full local system design in docs/TECHNICAL_DESIGN.md, and see how the GIFs were recorded in docs/INTERACTIVE_DEMO.md.

Search providers and configuration

The default provider chain is:

SearXNG → Brave → Tavily → Serper → DuckDuckGo → Wikipedia → Bing

Configure it with WEB_MCP_SEARCH_PROVIDER_CHAIN. Optional credentials are DATA_BRAVE_API_KEY, TAVILY_API_KEY, and SERPER_API_KEY. Copy .env.example as a reference, but keep secrets in the host environment rather than committing them.

The optional browser path is disabled by default:

uv pip install -e '.[render]'
./.venv/bin/python -m playwright install chromium
export WEB_MCP_RENDER_ENABLED=true

Distribution and discovery

The server is published in the official MCP Registry under io.github.AceAtDev/odysseus-web-mcp.

For Claude Desktop and other MCPB-compatible clients, download the validated MCPB release bundle from the v0.1.0 GitHub Release. The bundle uses the uv runtime to resolve the declared Python dependencies without shipping a machine-specific virtual environment.

Verify it yourself

The project has a focused test suite and a live qualification runner:

./.venv/bin/python -m pytest -q
./.venv/bin/python tests/live_20_cases.py --output reports/live-20-cases.json

The live qualification runs 10 searches and 10 fetches through the real MCP launcher with disposable state. The latest verification record is in VERIFICATION.md.

To re-record the terminal previews from fresh live calls (requires ImageMagick's convert command):

./.venv/bin/python demos/record_terminal_demos.py

Each GIF is intentionally under ten seconds and shows a real MCP handshake and result shape, not a static product mockup.

Project boundaries

This package is a retrieval primitive, not an agent loop, general-purpose crawler, scheduler, memory store, browser-profile manager, or credential vault. It is designed to be downloaded and connected as an independent MCP server.

License and status

This is the standalone extraction workspace for the Odysseus web search/fetch capability. See MIGRATION_MAP.md for the source-to-module mapping and VERIFICATION.md for the current evidence-based status.

Available Tools

2 tools
web_fetchA
Read-only

Read one known public URL or a bounded batch of URLs. Enforces public-URL SSRF checks, redirect and byte limits, returns extracted text, quality signals, discovered links, evidence_id, and cursor continuation when content is long.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoOne known http or https URL or bare domain.
fullNoRaise the download cap to the configured hard maximum.
urlsNoBounded batch of known URLs.
focusNoQuestion or section to prioritize in the projection.
cursorNoParagraph cursor returned by an earlier fetch.
renderNoWhether to use the optional isolated browser renderer.auto

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark readOnlyHint and openWorldHint, and the description adds rich behavioral detail: SSRF checks, redirect/byte limits, and the exact return payload (extracted text, quality signals, discovered links, evidence_id, cursor). This goes well beyond the annotations and fully discloses operational constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, dense sentence front-loads the core purpose and then lists behaviors and outputs in a logical order. Every clause adds value; there is no fluff or redundancy, making it highly efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only tool with no output schema, the description enumerates all return components (extracted text, quality signals, discovered links, evidence_id, cursor) and key limits (SSRF, redirects, bytes). The agent has enough to invoke it correctly without missing critical information.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all parameters are documented. The description mentions cursor continuation and bounded batch, but these are already captured in the schema (e.g., cursor, maxItems). No additional semantic nuance beyond the schema is provided, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: reading one known public URL or a bounded batch, with explicit mention of SSRF checks and output elements. This distinguishes it from the sibling web_search, which is for discovery, by emphasizing 'known' URLs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage via 'known public URL' but does not explicitly name web_search as the alternative for finding URLs. It provides context (known URLs) but lacks explicit when-to-use versus when-not-to-use guidance, so it falls at the implied level.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 2 tool updatesv0.1.0
    • First observedweb_fetch
    • First observedweb_search

TDQS

A4.2/5.0

Scored across 2 tools

Disambiguation5/5

web_search and web_fetch are clearly distinct: one discovers pages via queries, the other retrieves a known URL. No functional overlap.

Naming Consistency5/5

Both tools follow a consistent web_verb pattern, making the surface predictable and easy to navigate.

Tool Count4/5

Only two tools are exposed, which is slightly below the typical 3-15 range, but each tool is essential and the set is appropriately narrow for a web search/fetch server.

Completeness4/5

Search and fetch cover the core web access workflow, including pagination/cursor handling, but the surface is minimal and may lack advanced retrieval features such as site-specific extraction or structured data parsing.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    C
    quality
    B
    maintenance
    MCP server for safely reading public URLs for AI agents, providing tools to fetch, extract, cache, and inspect web content as evidence.
    15
    MIT
  • A
    license
    B
    quality
    B
    maintenance
    A local MCP server providing web search, page extraction, and safe browser automation tools for Hermes, Claude Code, and other MCP clients.
    10
    MIT
  • F
    license
    A
    quality
    C
    maintenance
    An MCP server that brings Parallel web search and URL extraction to Codex and other Model Context Protocol clients.
    2
    21
    -

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/AceAtDev/odysseus-web-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server