Skip to main content
Glama
emzimmer

Mozilla Readability Parser MCP Server

by emzimmer

Mozilla 가독성 파서 MCP 서버

웹페이지 콘텐츠를 추출하여 깔끔하고 LLM 최적화된 마크다운으로 변환하는 모델 컨텍스트 프로토콜(MCP) 서버입니다. 기사 제목, 본문, 발췌문, 작성자 이름, 사이트 이름을 반환합니다. Mozilla의 가독성 알고리즘을 사용하여 핵심 콘텐츠 구조는 유지하면서 광고, 탐색 창, 푸터 및 불필요한 요소를 제거합니다. MCP에 대해 자세히 알아보세요 .

특징

  • 광고, 탐색, 바닥글 및 기타 필수적이지 않은 콘텐츠를 제거합니다.

  • 깔끔한 HTML을 잘 포맷된 Markdown으로 변환합니다(Turndown도 사용함)

  • 기사 메타데이터(제목, 발췌문, 작성자, 사이트 이름)를 반환합니다.

  • 오류를 우아하게 처리합니다

Related MCP server: cleanfetch

그냥 가져오면 되지 않을까?

간단한 가져오기 요청과 달리 이 서버는 다음을 수행합니다.

  • Mozilla의 가독성 알고리즘을 사용하여 관련 콘텐츠만 추출합니다.

  • 광고, 팝업, 탐색 메뉴 등의 노이즈를 제거합니다.

  • 불필요한 HTML/CSS를 제거하여 토큰 사용량을 줄입니다.

  • 더 나은 LLM 처리를 위해 일관된 Markdown 형식을 제공합니다.

  • 콘텐츠에 대한 유용한 메타데이터가 포함되어 있습니다.

설치

Smithery를 통해 설치

Smithery를 통해 Claude Desktop용 Mozilla Readability Parser를 자동으로 설치하려면:

지엑스피1

수동 설치

npm install server-moz-readability

도구 참조

parse

웹페이지 콘텐츠를 가져와서 깔끔한 마크다운으로 변환합니다.

인수:

{
  "url": {
    "type": "string",
    "description": "The website URL to parse",
    "required": true
  }
}

보고:

{
  "title": "Article title",
  "content": "Markdown content...",
  "metadata": {
    "excerpt": "Brief summary",
    "byline": "Author information",
    "siteName": "Source website name"
  }
}

Claude Desktop과 함께 사용

claude_desktop_config.json 에 다음을 추가하세요:

{
  "mcpServers": {
    "readability": {
      "command": "npx",
      "args": ["-y", "server-moz-readability"]
    }
  }
}

종속성

  • @mozilla/readability - 콘텐츠 추출

  • 턴다운 - HTML에서 마크다운으로 변환

  • jsdom - DOM 파싱

  • axios - HTTP 요청

특허

MIT

Available Tools

1 tool
parseA

Extracts and transforms webpage content into clean, LLM-optimized Markdown. Returns article title, main content, excerpt, byline and site name. Uses Mozilla's Readability algorithm to remove ads, navigation, footers and non-essential elements while preserving the core content structure.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe website URL to parse

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes key behaviors: the transformation process ('extracts and transforms'), the algorithm used ('Mozilla's Readability algorithm'), what gets removed ('ads, navigation, footers and non-essential elements'), and what is preserved ('core content structure'). However, it doesn't mention potential limitations like rate limits, authentication needs, or error conditions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded, with two sentences that efficiently convey the tool's purpose, output, and key behavioral traits. Every sentence adds value without redundancy, making it easy to understand at a glance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (single parameter, no output schema, no annotations), the description is largely complete. It explains what the tool does, how it processes content, and what it returns. However, without an output schema, it could benefit from more detail on the return structure (e.g., format of the Markdown), and it lacks information on error handling or edge cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, with the parameter 'url' clearly documented as 'The website URL to parse'. The description doesn't add any additional meaning or context about the parameter beyond what the schema provides, such as URL format requirements or examples. With high schema coverage, the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with specific verbs ('extracts and transforms') and resources ('webpage content'), specifying the output format ('clean, LLM-optimized Markdown') and what it returns ('article title, main content, excerpt, byline and site name'). It distinguishes itself by mentioning the algorithm used ('Mozilla's Readability algorithm') and what it removes ('ads, navigation, footers and non-essential elements').

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for extracting structured content from webpages, but does not explicitly state when to use this tool versus alternatives, nor provide exclusions or prerequisites. With no sibling tools, the lack of explicit guidelines is less critical, but it still doesn't offer clear when/when-not instructions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev1.0.0
    • First observedparse

TDQS

A3.9/5.0

Scored across 1 tool

Disambiguation5/5

With only one tool, there is no possibility of ambiguity or overlap between tools. The single tool has a clear and distinct purpose focused on parsing webpage content into clean Markdown.

Naming Consistency5/5

A single tool inherently has perfect naming consistency, as there are no other tools to compare against. The tool name 'parse' is straightforward and appropriate for its function.

Tool Count2/5

A single tool is too few for a server's purpose, even if that purpose is narrow. This limits functionality and makes the server feel thin, as it lacks complementary operations like configuration, validation, or batch processing that might be expected in a parsing domain.

Completeness2/5

The tool surface is severely incomplete for a parsing server. While the 'parse' tool covers the core extraction function, there are obvious gaps such as no tools for handling errors, validating inputs, managing configurations, or providing metadata about the parsing process, which could lead to agent failures in real-world scenarios.

Maintenance

ActivityInactive
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Fetches web pages and converts them to clean, readable markdown format by extracting main content while removing navigation, ads, and other non-essential elements to minimize token usage.
    4
    -
  • A
    license
    A
    quality
    C
    maintenance
    Enables AI agents to read web pages reliably, returning clean markdown content, hyperlinks, and metadata without navigation or ad noise.
    3
    9 npm
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Converts any webpage into clean, LLM-ready Markdown, removing noise and supporting JavaScript rendering.
    MIT