Skip to main content
Glama
RocwoDev

MCP Web Utilities Server

by RocwoDev
README.md
### MCP Web Utilities Server

Lightweight server exposing web search and page fetching through:
- MCP (`stdio`) for MCP clients
- OpenAPI HTTP (`FastAPI`)

#### Features
- `search_on_web` and `search_on_website` using `ddgs` (DDGS | Dux Distributed Global Search, a multi-source search engine).
- `fetch_webpage` that returns simplified Markdown using `crawl4ai` with stealth settings.

#### Requirements
- Python 3.13+
- `uv` installed



#### Setup

Install `uv` on Windows (PowerShell):
```
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"
```

```
uv sync
```

Then activate the virtual environment and run the crawler setup:
```
.venv\Scripts\activate
crawl4ai-setup
```

#### Run the server

MCP (`stdio`):
```
uv run src\main.py
```

Or when developing:
```
start_mcp_server.cmd
```

OpenAPI HTTP (`127.0.0.1:8000`):
```
uv run uvicorn src.mainhttp:app --host 127.0.0.1 --port 8000
```

Or when developing:
```
start_mcp_server_http.cmd
```

OpenAPI schema:
```
http://127.0.0.1:8000/openapi.json
```

#### LM Studio Configuration
To use this server in LM Studio, add the following to your MCP settings (`mcp.json`):

```json
{
  "mcpServers": {
    "web-utilities": {
      "command": "uv",
      "args": [
        "--directory",
        "D:\\Dev\\McpServer",
        "run",
        "src\\main.py"
      ]
    }
  }
}
```
*Note: Replace `D:\\Dev\\McpServer` with the actual path to your project.*

#### Tools
`search_on_web(query: str, results: int = 10) -> str`
- Returns results formatted as:
```
[title](url)
description
```

`search_on_website(query: str, sites: list[str], results: int = 10) -> str`
- Same format, restricted to the provided `sites`.

`fetch_webpage(target_url: str) -> str`
- Returns simplified Markdown for the target page.

#### OpenAPI Endpoints
- `GET /search/web?query=...&results=10`
- `GET /search/website?query=...&sites=example.com&sites=docs.python.org&results=10`
- `GET /webpage?target_url=https://example.com`
- `GET /date`

#### Tests
```
python -m unittest src.tests
```

#### Notes
- Avoid writing to `STDOUT` (e.g., `print`) when the server is running; it will break JSON RPC communication.
- Network-dependent tests may fail if external services are blocked in the current environment.

TDQS

B3.1/5.0

Scored across 4 tools

Disambiguation2/5

There is significant overlap between 'search_on_web' and 'search_on_website' - both perform web searches with nearly identical descriptions and return formats, differing only in site restriction. 'fetch_webpage' is distinct for content retrieval, but 'get_current_date' feels disconnected from the web utilities theme, creating conceptual ambiguity about the server's purpose.

Naming Consistency4/5

Three of four tools follow a consistent verb_noun pattern with snake_case ('fetch_webpage', 'search_on_web', 'search_on_website'), which is good. However, 'get_current_date' uses a different verb style ('get' vs 'fetch/search'), breaking the pattern slightly but maintaining readability.

Tool Count3/5

With only 4 tools, the count feels thin for a 'Web Utilities Server' - one would expect more comprehensive web-related functionality. However, the tools do cover basic web operations (fetching, searching), so it's borderline rather than severely inadequate.

Completeness2/5

For a web utilities server, there are significant gaps: no tools for analyzing webpage content (beyond fetching), no HTTP request utilities, no cookie/session management, and no web scraping capabilities beyond basic fetching. The inclusion of 'get_current_date' feels out of scope, further highlighting the incomplete coverage of web-related operations.

Maintenance

ActivityInactive
ResponsivenessNo issues