LLMSTXT-MCP Server
Allows fetching and processing Modal's documentation via the llms.txt standard, converting content to clean Markdown.
Provides access to Next.js documentation by fetching and converting its llms.txt file into Markdown.
Provides access to Node.js documentation by fetching and converting its llms.txt file into Markdown.
Allows fetching and processing NVIDIA's documentation via the llms.txt standard, converting content to clean Markdown.
Provides access to React documentation by fetching and converting its llms.txt file into Markdown.
Allows fetching and processing Supabase's documentation via the llms.txt standard, converting content to clean Markdown.
LLMSTXT-MCP Server
A Model Context Protocol (MCP) server that provides access to LLMS.TXT documentation files. This server allows AI agents to fetch and process documentation from various sources.
Features
Multiple Documentation Sources: Access documentation from React, Next.js, Node.js and more
HTTP Fetching: Fetch documentation from any HTTPS URL with domain restrictions
HTML to Markdown: Automatically converts HTML content to clean Markdown
Environment Configuration: Customize behavior via environment variables
MCP Protocol: Full Model Context Protocol compliance
Related MCP server: MCP LLMS-TXT Documentation Server
Installation
npm install -g @pinkpixel/llmstxt-mcpUsage
The server starts automatically and provides two tools:
Available Tools
list_doc_sources- Lists all configured documentation sourcesfetch_docs- Fetches documentation from a URL and converts to Markdown
Default Configuration
By default, the server is configured with:
React documentation:
https://react.dev/llms.txtNext.js documentation:
https://nextjs.org/llms.txtNode.js documentation:
https://nodejs.org/llms.txtLLMSTXT Directory (Cloud):
https://directory.llmstxt.cloud/- Curated directory of companies using llms.txtLLMSTXT Site Directory:
https://llmstxt.site/- Comprehensive list with token counts and stats
Discovery Directories
The two directory sources provide access to thousands of websites that have adopted the llms.txt standard:
directory.llmstxt.cloud - Curated directory with companies like Anthropic, Supabase, Modal, NVIDIA, and many others across AI, developer tools, finance, and products categories
llmstxt.site - Comprehensive searchable directory with detailed statistics and token counts for each llms.txt file
These directories are invaluable for discovering documentation from the growing ecosystem of companies adopting the llms.txt standard.
Environment Variables
Customize behavior with these environment variables:
LLMSTXT_CONFIG: Path to custom configuration file (not implemented yet)LLMSTXT_FOLLOW_REDIRECTS: Whether to follow HTTP redirects (true/false)LLMSTXT_TIMEOUT: Request timeout in seconds (default:10)LLMSTXT_ALLOWED_DOMAINS: Comma-separated list of allowed domains (default:*for all)
Example:
export LLMSTXT_FOLLOW_REDIRECTS=true
export LLMSTXT_TIMEOUT=30
export LLMSTXT_ALLOWED_DOMAINS="react.dev,nextjs.org,nodejs.org"MCP Client Configuration
You can install globally with npm i -g @pinkpixel/llmstxt-mcp and then it can be ran with "llmstxt-mcp"
Add to your mcp_config.json:
{
"mcpServers": {
"llmstxt": {
"command": "llmstxt-mcp",
"env": {
"LLMSTXT_FOLLOW_REDIRECTS": "true",
"LLMSTXT_TIMEOUT": "30"
}
}
}
}OR use with npx
{
"mcpServers": {
"llmstxt": {
"command": "npx",
"args": ["-y", "@pinkpixel/llmstxt-mcp", "llmstxt-mcp"],
"env": {
"LLMSTXT_FOLLOW_REDIRECTS": "true",
"LLMSTXT_TIMEOUT": "30"
}
}
}
}Testing
Test with MCP Inspector
# Build first
npm run build
# Test with inspector
npm run inspector
# Or directly
npx @modelcontextprotocol/inspector ./build/index.jsManual Testing
# List available tools
npx @modelcontextprotocol/inspector --cli ./build/index.js --method tools/list
# List doc sources
npx @modelcontextprotocol/inspector --cli ./build/index.js --method tools/call --tool-name list_doc_sources
# Fetch documentation
npx @modelcontextprotocol/inspector --cli ./build/index.js --method tools/call --tool-name fetch_docs --tool-arg url="https://example.com"Security
Domain Restrictions: Configure allowed domains via environment variables
HTTPS Recommended: Always use HTTPS URLs when possible
No Local Files: Current version only supports HTTP/HTTPS URLs
Development
# Clone the repository
git clone https://github.com/pinkpixel-dev/llmstxt-mcp.git
cd llmstxt-mcp
# Install dependencies
npm install
# Build the server
npm run build
# Test with inspector
npm run inspectorError Handling
The server provides detailed error messages for:
Invalid or unreachable URLs
Domain restriction violations
Timeout errors
HTTP errors (404, 500, etc.)
License
MIT License - see LICENSE file for details.
Made with ❤️ by Pink Pixel
Available Tools
2 toolsfetch_docsA
Fetch and parse documentation from a given URL or local file.
Use this tool after list_doc_sources to:
First fetch the llms.txt file from a documentation source
Analyze the URLs listed in the llms.txt file
Then fetch specific documentation pages relevant to the user's question
Args: url: The URL to fetch documentation from.
Returns: The fetched documentation content converted to markdown, or an error message if the request fails or the URL is not from an allowed domain.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL or file path to fetch documentation from |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the conversion to markdown, and error conditions (request failure, disallowed domains), which are useful. It does not explicitly state that the operation is read-only, but the fetch nature implies it. It also lacks details on rate limits or size constraints, but these are minor.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a summary, workflow list, Args, and Returns sections. It is slightly long but each part adds value, and the main purpose is front-loaded. It avoids unnecessary fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with no output schema, the description covers all essential aspects: purpose, usage sequence, input description, return behavior, and error handling. An agent has sufficient information to call it correctly without additional guidance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the parameter is already fully documented. The description repeats the parameter meaning and adds usage context but does not introduce new constraints or format details, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: to fetch and parse documentation from a URL or local file. It distinguishes itself from its sibling list_doc_sources by explicitly positioning itself as a follow-up step, making its role unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage instructions: it should be used after list_doc_sources, and outlines a three-step workflow (fetch llms.txt, analyze URLs, fetch specific pages). This leaves no ambiguity about when and how to invoke the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_doc_sourcesB
List all configured documentation sources
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full burden of behavioral disclosure. It implies a read-only operation via the verb 'list' but does not state any side effects, permission requirements, or output format. For a simple listing tool this is somewhat acceptable, but it lacks explicit transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that conveys the core action without any wasted words. It is appropriately sized for a zero-parameter tool and is front-loaded with the verb.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate for a trivial list operation but lacks context about the output (no output schema) and does not clarify the relationship with the sibling tool. An agent might benefit from knowing whether this returns metadata, full documents, or just names, but the minimal nature of the tool partially mitigates this gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline is 4. The schema is empty and fully described (100% coverage), so there is nothing for the description to add. The description appropriately avoids redundant parameter details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('List') and resource ('all configured documentation sources'), making the tool's purpose unambiguous. However, it does not explicitly distinguish it from the sibling 'fetch_docs', though the name implies a read-only listing versus fetching content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus the sibling 'fetch_docs'. The description gives no context about selection criteria or alternatives, leaving the agent to infer usage from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v1.0.4- First observed
fetch_docs - First observed
list_doc_sources
TDQS
Scored across 2 tools
The two tools serve clearly distinct purposes: listing configured sources versus fetching a specific documentation URL. There is no overlap or ambiguity about when to use each.
Both tool names follow a consistent verb_noun pattern (list_doc_sources, fetch_docs) using snake_case. The naming is predictable and readable.
With only two tools, the server feels minimal, though it may be appropriate for its narrow scope of listing sources and fetching docs. According to calibration, 1-2 tools is borderline thin, so this is not a strong score.
The tool surface covers the core workflow: list sources, fetch llms.txt or specific pages, and convert to markdown. A minor gap is the lack of a dedicated tool for managing sources (add/remove), but that appears to be configuration outside the server's purpose.
Maintenance
Related MCP Connectors
Turn any public website into an MCP server for agents to search, read and navigate.
MCP server (stdio): fetch web pages as clean readable markdown via the AgentForge API
Document-to-Markdown MCP server — convert PDF, Office and HTML into LLM-ready Markdown.
MCP server for Firecrawl — web search, scraping, and biomedical/arXiv paper search.
Related MCP Servers
- AlicenseAqualityCmaintenanceAn MCP server that fetches web pages and extracts clean, AI-friendly Markdown content using Mozilla Readability. It provides secure web access for LLMs with built-in SSRF protection and automated content cleaning for improved context retrieval and summarization.178 npmMIT
- AlicenseNot gradedqualityDmaintenanceAn MCP server that enables users to fetch and audit documentation from user-defined llms.txt index files. It provides tools to list documentation sources and retrieve content from specific URLs with built-in domain access controls for secure context retrieval.MIT
- AlicenseAqualityCmaintenanceMCP server that converts URLs to clean Markdown/Text for LLM agents.2553 npm5MIT
- AlicenseNot gradedqualityCmaintenanceA documentation MCP server that crawls websites and Git repositories, stores them as Markdown, and provides tools to search and retrieve documentation for local LLMs and AI agents.Apache 2.0