Integrations
Utilizes Mozilla's Readability library (the same engine used in Firefox's Reader View) to extract meaningful content from web pages for conversion to Markdown
Converts clean HTML to high-quality Markdown with TurndownService, supporting both web scraping and direct conversion of local HTML files
Leverages Mozilla's Readability library to extract the main content from web pages while removing clutter and navigation elements
Website Scraper
A command-line tool and MCP server for scraping websites and converting HTML to Markdown.
Features
- Extracts meaningful content from web pages using Mozilla's Readability library (the same engine used in Firefox's Reader View)
- Converts clean HTML to high-quality Markdown with TurndownService
- Securely handles HTML by removing potentially harmful script tags
- Works as both a command-line tool and an MCP server
- Supports direct conversion of local HTML files to Markdown
Installation
Usage
CLI Mode
MCP Server Mode
This tool can be used as a Model Context Protocol (MCP) server:
Code Structure
src/index.ts
- Core functionality and MCP server implementationsrc/cli.ts
- Command-line interface implementationsrc/data_processing.ts
- HTML to Markdown conversion functionality
API
The tool exports the following functions:
License
ISC
This server cannot be installed
hybrid server
The server is able to function both locally and remotely, depending on the configuration or use case.
An MCP server that extracts meaningful content from websites and converts HTML to high-quality Markdown, using Mozilla's Readability engine.
Related MCP Servers
- AsecurityAlicenseAqualityA powerful MCP server for fetching and transforming web content into various formats (HTML, JSON, Markdown, Plain Text) with ease.Last updated -414612TypeScriptMIT License
- AsecurityAlicenseAqualityAn MCP server for fetching and transforming web content into various formats.Last updated -44PythonMIT License
- -securityAlicense-qualityA Python-based MCP server that crawls websites to extract and save content as markdown files, with features for mapping website structure and links.Last updated -1PythonMIT License
- -securityAlicense-qualityA Python implementation of an MCP server that extracts webpage content, removes ads and non-essential elements, and transforms it into clean, LLM-optimized Markdown.Last updated -1PythonMIT License