Skip to main content
Glama
README.md
# EnriWeb

EnriWeb is a **Model Context Protocol (MCP)** server over `stdio` that exposes **web search** and **URL fetching** tools by delegating execution to **EnriProxy**.

If your MCP client can call MCP tools, it can do web search / fetch in a consistent way without implementing provider-specific scraping logic.

## What this project is

- An MCP server process your MCP host launches (OpenCode, Claude Code, Codex, etc.)
- A thin client for EnriProxy (input validation + structured output)

## Requirements

- Node.js `>= 22` (recommended: Node 24 LTS)
- A reachable EnriProxy server with:
  - `POST /v1/tools/web_search`
  - `POST /v1/tools/web_fetch`
- An EnriProxy API key (configured on the EnriProxy side)

## Install

```powershell
# Global install
npm install -g @bedolla/enriweb

# Or run without installing
npx -y @bedolla/enriweb@latest --help
```

## Build

```powershell
npm install
npm run typecheck
npm run build
```

## Usage

### 1) Configure your MCP host

EnriWeb runs as an MCP server over `stdio`. Your MCP host is responsible for launching the process.

Example: global install

```jsonc
{
  "EnriWeb": {
    "type": "stdio",
    "command": "enriweb",
    "args": [],
    "env": {
      "ENRIPROXY_URL": "http://127.0.0.1:8787",
      "ENRIPROXY_API_KEY": "YOUR_ENRIPROXY_API_KEY"
    }
  }
}
```

Example: no install (always uses whatever npm currently tags as `latest`)

```jsonc
{
  "EnriWeb": {
    "type": "stdio",
    "command": "npx",
    "args": ["-y", "@bedolla/enriweb@latest"],
    "env": {
      "ENRIPROXY_URL": "http://127.0.0.1:8787",
      "ENRIPROXY_API_KEY": "YOUR_ENRIPROXY_API_KEY"
    }
  }
}
```

<details>
<summary>Use a local dev checkout</summary>

```jsonc
{
  "EnriWeb": {
    "type": "stdio",
    "command": "node",
    "args": ["C:\\\\Users\\\\Administrator\\\\Projects\\\\EnriWeb\\\\dist\\\\index.js"],
    "env": {
      "ENRIPROXY_URL": "http://127.0.0.1:8787",
      "ENRIPROXY_API_KEY": "YOUR_ENRIPROXY_API_KEY"
    }
  }
}
```

</details>

## Configuration

EnriWeb is configured via environment variables:

- `ENRIPROXY_URL` (`string`, optional, default: `http://127.0.0.1:8787`)
- `ENRIPROXY_API_KEY` (`string`, required)
- `ENRIWEB_TIMEOUT_MS` (`string`, optional, default: `60000`)
  - Parsed as an integer (milliseconds).
- `ENRIWEB_WEB_FETCH_DEFAULT_MAX_CHARS` (`string`, optional, default: `200000`)
  - Parsed as an integer.
- `ENRIWEB_GITHUB_TOKEN` (`string`, optional)
  - Used for GitHub API enrichment to improve rate limits.

## MCP tools

EnriWeb exposes these MCP tools:

- `web_search`
- `web_fetch`

<details>
<summary>Tool inputs (option-by-option)</summary>

General notes:

- All tools accept a single JSON object as their input (the MCP `arguments` for that tool).
- EnriWeb returns both:
  - a short human-readable preview (`content`)
  - the full result payload (`structuredContent`)

---

### `web_search`

Search the web via EnriProxy.

Inputs:

- `query` (`string`, required): search query string.
- `max_results` (`number`, optional)
  - Must be `>= 1`.
  - If omitted, EnriProxy uses its configured default.
  - The upper limit is enforced server-side (EnriWeb does not hardcode a max).
- `recency` (`string`, optional, default: `noLimit`)
  - One of: `oneDay` | `oneWeek` | `oneMonth` | `oneYear` | `noLimit`
- `allowed_domains` (`string[]`, optional): allowlist of domains to include.
- `blocked_domains` (`string[]`, optional): blocklist of domains to exclude.
- `search_prompt` (`string`, optional): extra context to refine the search intent.

Example `arguments` object:

```jsonc
{
  "query": "qdrant docker compose autostart systemd",
  "max_results": 10,
  "recency": "oneMonth"
}
```

---

### `web_fetch`

Fetch and read content from a URL via EnriProxy.

Inputs:

- `url` (`string`, required unless `cursor` is provided): full URL (`http://` or `https://`).
- `cursor` (`string`, optional): opaque cursor returned by a previous `web_fetch` call.
- `offset_chars` (`number`, optional, default: `0`): cursor read offset in characters (`offset` is a legacy alias).
- `limit_chars` (`number`, optional): cursor read limit in characters (default: `max_chars`; `limit` is a legacy alias).
- `prompt` (`string`, optional): extraction hint (what to focus on).
- `max_chars` (`number`, optional): maximum content length (default: `ENRIWEB_WEB_FETCH_DEFAULT_MAX_CHARS`).
- `format` (`string`, optional): content flavor for HTML pages — `"text"` (default, lightweight structured text), `"markdown"` (full markdown with links, emphasis, code fences, images, and tables), or `"html"` (sanitized markup for DOM inspection — scripts/styles stripped, tags intact). Use markdown only when the exact page structure matters; text is cheaper for factual lookups.
- `content` (`string`, optional): HTML scope — `"full"` (default, whole page) or `"main"` (article/main container only; drops nav, sidebars, cookie banners, and footers, typically saving 60-80% of tokens).
- `include_links` (`boolean`, optional): append the `ENLACES DE LA PÁGINA` inventory with every unique link (label + URL, up to 200) — useful for informed crawling or handing image URLs to URL-capable media analysis tools.
- `include_metadata` (`boolean`, optional): append the `METADATOS DE LA PÁGINA` block with language, author, published date, and `og:image`.
- `anchor` (`string`, optional): section selector — element id (with or without `#`) or exact heading text; returns only that section up to the next same-or-higher heading. When the section is missing, the response says so and returns the full document.

Notes:

- If the response includes a `cursor`, you can page through the captured content by calling `web_fetch` again with `cursor` + `offset_chars` + `limit_chars`.

Example `arguments` object:

```jsonc
{
  "url": "https://example.com/docs",
  "max_chars": 200000
}
```

</details>

TDQS

A4.6/5.0

Scored across 2 tools

Disambiguation5/5

web_search and web_fetch have clearly distinct purposes: one discovers URLs and information from the web, the other retrieves content from a specific URL. There is no overlap or ambiguity in choosing between them.

Naming Consistency5/5

Both tools consistently follow the web_verb pattern (web_search, web_fetch), making the domain and action immediately clear and predictable.

Tool Count4/5

Two tools is minimal, but it fully covers the core need of a web access server: searching and fetching. It is slightly under the typical 3-15 range yet remains well-scoped and not bloated.

Completeness5/5

For the stated purpose of web search and retrieval, the pair forms a complete workflow: search returns candidate URLs and web_fetch reads the chosen page, including pagination and content-shaping options. There are no obvious dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues