Skip to main content
Glama
tolik-unicornrider

Website Scraper MCP Server

README.md
# Website Scraper

A command-line tool and MCP server for scraping websites and converting HTML to Markdown.

## Features

- Extracts meaningful content from web pages using Mozilla's [Readability](https://github.com/mozilla/readability) library (the same engine used in Firefox's Reader View)
- Converts clean HTML to high-quality Markdown with TurndownService
- Securely handles HTML by removing potentially harmful script tags
- Works as both a command-line tool and an MCP server
- Supports direct conversion of local HTML files to Markdown

## Installation

```bash
# Install dependencies
npm install

# Build the project
npm run build

# Optionally, install globally
npm install -g .
```

## Usage

### CLI Mode

```bash
# Print output to console
scrape https://example.com

# Save output to a file
scrape https://example.com output.md

# Convert a local HTML file to Markdown
scrape --html-file input.html

# Convert a local HTML file and save output to a file
scrape --html-file input.html output.md

# Show help
scrape --help

# Or run via npm script
npm run start:cli -- https://example.com
```

### MCP Server Mode

This tool can be used as a Model Context Protocol (MCP) server:

```bash
# Start in MCP server mode
npm start
```

## Code Structure

- `src/index.ts` - Core functionality and MCP server implementation
- `src/cli.ts` - Command-line interface implementation
- `src/data_processing.ts` - HTML to Markdown conversion functionality

## API

The tool exports the following functions:

```typescript
// Scrape a website and convert to Markdown
import { scrapeToMarkdown } from './build/index.js';

// Convert HTML string to Markdown directly
import { htmlToMarkdown } from './build/data_processing.js';

async function example() {
  // Web scraping
  const markdown = await scrapeToMarkdown('https://example.com');
  console.log(markdown);
  
  // Direct HTML conversion
  const html = '<h1>Hello World</h1><p>This is <strong>bold</strong> text.</p>';
  const md = htmlToMarkdown(html);
  console.log(md);
}
```

## License

ISC 

TDQS

D1.7/5.0

Scored across 1 tool

Disambiguation5/5

With only one tool, there is no possibility of ambiguity or overlap between tools. The single tool 'scrape-to-markdown' stands alone with no other tools to confuse it with.

Naming Consistency5/5

Since there is only one tool, naming consistency is inherently perfect. The tool name 'scrape-to-markdown' follows a clear verb-noun pattern, and there are no other tools to deviate from this pattern.

Tool Count2/5

A single tool for a 'Website Scraper MCP Server' is too few for the apparent scope, as scraping typically involves multiple operations like configuring, parsing, or handling errors. This minimal set feels incomplete and thin for the domain.

Completeness1/5

The tool set is severely incomplete for a scraping server. It lacks essential operations such as fetching URLs, handling different content types, error management, or data extraction beyond markdown. This single tool creates significant gaps that will cause agent failures in real-world scraping tasks.

Maintenance

ActivityInactive
ResponsivenessNo issues