Skip to main content
Glama
Calandrel

Trafilatura MCP Server

by Calandrel

Trafilatura MCP Server

This repository contains a Model Context Protocol (MCP) server that provides a tool-based interface to the Trafilatura library, a powerful tool for web scraping. It is designed for use with MCP-compatible clients, allowing developers and models to extract main content and metadata from web pages programmatically.

Features

  • Web Scraping: Utilizes Trafilatura to extract the main text content from a given URL.

  • Metadata Extraction: Retrieves metadata such as title, author, date, and more.

  • Configurable Extraction: Options to include or exclude comments and tables from the output.

  • Simple Tool: Exposes a single, easy-to-use fetch_and_extract tool.

  • Asynchronous: Built with an asynchronous architecture for efficient I/O operations.

  • MCP Standard: Communicates over standard I/O, making it compatible with various MCP clients.

Related MCP server: article-scraper-mcp

Prerequisites

Before running the server, you need to have Python 3.12+ and uv installed. You will also need Node.js and npx to run the MCP Inspector tool for testing.

Installation

  1. Clone the repository.

    git clone <repository-url>
    cd trafilatura_mcp
  2. Create a virtual environment and install the required dependencies:

    # Create a virtual environment
    uv venv
    
    # Activate the virtual environment
    source .venv/bin/activate
    
    # Install the dependencies
    uv sync

Running and Testing the Server

The MCP server is a command-line application that communicates over standard I/O. To use it, a client (like an IDE, a coding agent, or an inspector tool) must launch the server process.

Running for Diagnostics

You can run the script directly from your terminal to see if it starts without errors. This is a quick way to validate your Python environment and the script's basic syntax.

python3 trafilatura_mcp.py

The server will start and wait for input, but you won't be able to interact with it directly from your terminal.

Testing with MCP Inspector

The recommended way to test the server interactively is with MCP Inspector. It provides an interactive shell for sending requests to your server.

  1. Launch the Inspector: You can run the inspector without a permanent installation using npx. The inspector will launch your MCP server script for you. From your project directory, run:

    npx @modelcontextprotocol/inspector uv run -- python3 trafilatura_mcp.py
  2. Interact with the Server: Once the inspector starts, you can connect to the server and use commands like list_tools and call_tool.

    Example session:

    # List the available tool
    > list_tools
    
    # Call the 'fetch_and_extract' tool with a URL
    > call_tool fetch_and_extract '''{"url": "https://www.theguardian.com/us-news/2025/sep/28/mass-shootings-north-carolina-texas-new-orleans"}'''

Configuration

This server does not require any external API keys or configuration files.

Usage with an MCP Client (VS Code Example)

You can connect to this server from any standard MCP client. Here’s how to do it in a VS Code environment that supports MCP:

  1. Configure Your MCP Client: In your IDE's MCP client settings (e.g., in mcp.json for VS Code), configure a new MCP server that points to the script.

    Example mcp.json entry:

    {
      "servers": {
        "trafilatura_scraper": {
          "command": "uv",
          "args": [
            "run",
            "python3",
            "trafilatura_mcp.py"
          ],
          "cwd": "/path/to/your/project/trafilatura_mcp"
        }
      }
    }

    Note: Replace /path/to/your/project/trafilatura_mcp with the absolute path to the project directory.

  2. Use the Tool: Once connected, you can use the exposed tool in your chat or agent interactions. For example, to extract content from a news article, you could send the following structured tool call:

    {
      "tool": "fetch_and_extract",
      "arguments": {
        "url": "https://apnews.com/article/elon-musk-x-twitter-hate-speech-antisemitism-0d35c5a69fd5c6183b729f7f3c87064a",
        "include_comments": false,
        "include_tables": true
      }
    }

    The server will fetch the URL, extract the main content and metadata, and return it as a JSON object.

Available Tools

1 tool
fetch_and_extractB

Fetches a URL and extracts the main content, metadata, and comments. Returns a JSON object with the extracted data.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL of the web page to process.
include_commentsNoWhether to include comment sections at the bottom of articles.
include_tablesNoExtract text from HTML <table> elements.

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions that the tool returns a JSON object with extracted data, but does not cover important aspects like error handling, rate limits, authentication needs, or what happens if the URL is inaccessible. For a tool that fetches external content, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded, consisting of two sentences that directly state the tool's function and output. There is no wasted language, and every sentence earns its place by providing essential information efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (fetching and extracting web content) and the lack of annotations and output schema, the description is somewhat incomplete. It covers the basic purpose and output format but misses behavioral details like error handling or performance considerations. However, it is adequate as a minimum viable description for a tool with no siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, so the schema already documents all parameters (url, include_comments, include_tables) with clear descriptions. The description does not add any additional meaning or context beyond what the schema provides, such as examples or edge cases. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: fetching a URL and extracting main content, metadata, and comments. It specifies the verb ('fetches' and 'extracts') and resource ('URL'), but since there are no sibling tools, it cannot distinguish from alternatives. The description is not tautological and provides a clear action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention any prerequisites, exclusions, or specific contexts for usage. Without sibling tools, there is no explicit comparison, but it lacks any usage context or limitations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 1 tool updatev1.0.0
    • Changedfetch_and_extract2 fields changed
      • addedInput schema / description
        Added value: +"Input model for the fetch_and_extract tool."
      • addedInput schema / title
        Added value: +"TrafilaturaInput"
  2. 1 tool update
    • First observedfetch_and_extract

TDQS

B3.2/5.0
Disambiguation5/5

With only one tool, there is no possibility of ambiguity or overlap between tools. The tool's purpose is clearly defined as fetching and extracting content from a URL, making it distinct by default.

Naming Consistency5/5

Since there is only one tool, naming consistency is inherently perfect. The tool name 'fetch_and_extract' follows a clear verb_noun pattern, and there are no other tools to compare it against for inconsistency.

Tool Count2/5

A single tool is generally too few for a server's purpose, as it limits functionality and may indicate an incomplete surface. For a content extraction server, one tool feels thin and could benefit from additional operations like batch processing or configuration options.

Completeness2/5

The tool surface is severely incomplete for a content extraction domain. While the single tool handles basic fetching and extraction, there are obvious gaps such as no ability to extract from local files, handle different extraction modes, or manage extraction settings, which could lead to agent failures in more complex scenarios.

Maintenance

ActivityInactive
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/Calandrel/trafilatura_mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server