MCP URL Fetcher
# MCP URL Format Converter
A Model Context Protocol (MCP) server that fetches content from any URL and converts it to your desired output format.
## Overview
MCP URL Format Converter provides tools for retrieving content from any web URL and transforming it into various formats (HTML, JSON, Markdown, or plain text), regardless of the original content type. It's designed to work with any MCP-compatible client, including Claude for Desktop, enabling LLMs to access, transform, and analyze web content in a consistent format.
## Features
- 🔄 **Format Conversion**: Transform any web content to HTML, JSON, Markdown, or plain text
- 🌐 **Universal Input Support**: Handle websites, APIs, raw files, and more
- 🔍 **Automatic Content Detection**: Intelligently identifies source format
- 🧰 **Robust Library Support**: Uses industry-standard libraries:
- Cheerio for HTML parsing
- Marked for Markdown processing
- Fast-XML-Parser for XML handling
- CSVtoJSON for CSV conversion
- SanitizeHTML for security
- Turndown for HTML-to-Markdown conversion
- 🔧 **Advanced Format Processing**:
- HTML parsing with metadata extraction
- JSON pretty-printing and structure preservation
- Markdown rendering with styling
- CSV-to-table conversion
- XML-to-JSON transformation
- 📜 **History Tracking**: Maintains logs of recently fetched URLs
- 🛡️ **Security Focus**: Content sanitization to prevent XSS attacks
## Installation
### Prerequisites
- Node.js 16.x or higher
- npm or yarn
### Quick Start
1. Clone the repository:
```bash
git clone https://github.com/yourusername/mcp-url-converter.git
cd mcp-url-converter
```
2. Install dependencies:
```bash
npm install
```
3. Build the project:
```bash
npm run build
```
4. Run the server:
```bash
npm start
```
## Integration with Claude for Desktop
1. Open your Claude for Desktop configuration file:
- macOS: `~/Library/Application Support/Claude/claude_desktop_config.json`
- Windows: `%APPDATA%\Claude\claude_desktop_config.json`
2. Add the URL converter server to your configuration:
```json
{
"mcpServers": {
"url-converter": {
"command": "node",
"args": ["/absolute/path/to/mcp-url-converter/build/index.js"]
}
}
}
```
3. Restart Claude for Desktop
## Available Tools
### `fetch`
Fetches content from any URL and automatically detects the best output format.
**Parameters:**
- `url` (string, required): The URL to fetch content from
- `format` (string, optional): Format to convert to (`auto`, `html`, `json`, `markdown`, `text`). Default: `auto`
**Example:**
```
Can you fetch https://example.com and choose the best format to display it?
```
### `fetch-json`
Fetches content from any URL and converts it to JSON format.
**Parameters:**
- `url` (string, required): The URL to fetch content from
- `prettyPrint` (boolean, optional): Whether to pretty-print the JSON. Default: `true`
**Example:**
```
Can you fetch https://example.com and convert it to JSON format?
```
### `fetch-html`
Fetches content from any URL and converts it to HTML format.
**Parameters:**
- `url` (string, required): The URL to fetch content from
- `extractText` (boolean, optional): Whether to extract text content only. Default: `false`
**Example:**
```
Can you fetch https://api.example.com/users and convert it to HTML?
```
### `fetch-markdown`
Fetches content from any URL and converts it to Markdown format.
**Parameters:**
- `url` (string, required): The URL to fetch content from
**Example:**
```
Can you fetch https://example.com and convert it to Markdown?
```
### `fetch-text`
Fetches content from any URL and converts it to plain text format.
**Parameters:**
- `url` (string, required): The URL to fetch content from
**Example:**
```
Can you fetch https://example.com and convert it to plain text?
```
### `web-search` and `deep-research`
These tools provide interfaces to Perplexity search capabilities (when supported by the MCP host).
## Available Resources
### `recent-urls://list`
Returns a list of recently fetched URLs with timestamps and output formats.
**Example:**
```
What URLs have I fetched recently?
```
## Security
This server implements several security measures:
- HTML sanitization using `sanitize-html` to prevent XSS attacks
- Content validation before processing
- Error handling and safe defaults
- Input parameter validation with Zod
- Safe output encoding
## Testing
You can test the server using the MCP Inspector:
```bash
npm run test
```
## Troubleshooting
### Common Issues
1. **Connection errors**: Verify that the URL is accessible and correctly formatted
2. **Conversion errors**: Some complex content may not convert cleanly between formats
3. **Cross-origin issues**: Some websites may block requests from unknown sources
### Debug Mode
For additional debugging information, set the `DEBUG` environment variable:
```bash
DEBUG=mcp:* npm start
```
## License
This project is licensed under the MIT License - see the LICENSE file for details.
## Acknowledgments
- Built with the [Model Context Protocol](https://modelcontextprotocol.io/)
- Uses modern, actively maintained libraries with security focus
- Sanitization approach based on OWASP recommendations
---
Last updated: 29 March 2025
TDQS
Scored across 5 tools
The tools have significant overlap in purpose, as all fetch content from URLs, differing only in output format. An agent could easily misselect between fetch-html, fetch-json, fetch-markdown, and fetch-text, since the descriptions don't clarify when to use one format over another. The generic 'fetch' tool with automatic detection further confuses boundaries, as it might duplicate or conflict with the format-specific tools.
Tool names follow a highly consistent verb-noun pattern throughout, with all tools using 'fetch' as the verb followed by a hyphen and format descriptor (e.g., fetch-html, fetch-json). There are no deviations in style or convention, making the naming predictable and easy to parse for an agent.
With 5 tools, the count is reasonable for a URL fetching server, but it feels borderline due to redundancy. The tools could potentially be consolidated into fewer, more flexible tools (e.g., a single fetch tool with a format parameter), making the current set feel slightly over-specified for the simple domain of fetching URLs.
For the domain of fetching URL content, the tool set is nearly complete, covering automatic detection and common output formats (HTML, JSON, Markdown, plain text). A minor gap exists in handling errors or advanced configurations (e.g., headers, timeouts), but agents can likely work around this with the provided tools for basic fetching tasks.