mcp-pdf2md
MCP-PDF2MD
MCP-PDF2MD服务
基于 MCP 的高性能 PDF 转 Markdown 转换服务,由 MinerU API 提供支持,支持本地文件和 URL 链接的批量处理,并进行结构化输出。
主要特点
格式转换:将PDF文件转换为结构化的Markdown格式。
多源支持:同时处理本地 PDF 文件和 URL 链接。
智能处理:自动选择最佳处理方法。
批量处理:支持多文件批量转换,高效处理大量PDF文件。
MCP 集成:与 Claude Desktop 等 LLM 客户端无缝集成。
结构保存:维护原始文档结构,包括标题、段落、列表等。
智能布局:以人类可读的顺序输出文本,适用于单列、多列和复杂布局。
公式转换:自动识别文档中的公式并转换为 LaTeX 格式。
表格提取:自动识别文档中的表格并将其转换为结构化格式。
清理优化:删除页眉、页脚、脚注、页码等,确保语义一致性。
高质量提取:从 PDF 文档中高质量提取文本、图像和布局信息。
Related MCP server: pdf2md-mcp
系统要求
软件:Python 3.10+
快速入门
克隆仓库并进入目录:
git clone https://github.com/FutureUnreal/mcp-pdf2md.git cd mcp-pdf2md创建虚拟环境并安装依赖项:
Linux/macOS :
uv venv source .venv/bin/activate uv pip install -e .窗户:
uv venv .venv\Scripts\activate uv pip install -e .配置环境变量:
在项目根目录下创建
.env文件,并设置以下环境变量:MINERU_API_BASE=https://mineru.net/api/v4/extract/task MINERU_BATCH_API=https://mineru.net/api/v4/extract/task/batch MINERU_BATCH_RESULTS_API=https://mineru.net/api/v4/extract-results/batch MINERU_API_KEY=your_api_key_here启动服务:
uv run pdf2md
命令行参数
服务器支持以下命令行参数:
Claude桌面配置
在Claude Desktop中添加以下配置:
窗户:
{
"mcpServers": {
"pdf2md": {
"command": "uv",
"args": [
"--directory",
"C:\\path\\to\\mcp-pdf2md",
"run",
"pdf2md",
"--output-dir",
"C:\\path\\to\\output"
],
"env": {
"MINERU_API_KEY": "your_api_key_here"
}
}
}
}Linux/macOS :
{
"mcpServers": {
"pdf2md": {
"command": "uv",
"args": [
"--directory",
"/path/to/mcp-pdf2md",
"run",
"pdf2md",
"--output-dir",
"/path/to/output"
],
"env": {
"MINERU_API_KEY": "your_api_key_here"
}
}
}
}**关于 API 密钥配置的注意事项:**您可以通过两种方式设置 API 密钥:
在项目目录内的
.env文件中(推荐用于开发)在Claude Desktop配置如上图(建议常规使用)
如果您在两个地方都设置了 API 密钥,则 Claude Desktop 配置中的密钥将优先。
MCP 工具
该服务器提供以下 MCP 工具:
convert_pdf_url :将 PDF URL 转换为 Markdown
convert_pdf_file :将本地 PDF 文件转换为 Markdown 文件
获取 MinerU API 密钥
本项目依赖 MinerU API 进行 PDF 内容提取。获取 API 密钥:
访问MinerU官网并注册账号
登录后,通过此链接申请API测试资质
您的申请获得批准后,您可以访问API 管理页面
按照提供的说明生成您的 API 密钥
复制生成的 API 密钥
使用此字符串作为
MINERU_API_KEY的值
请注意,MinerU API 的访问目前处于测试阶段,需要获得 MinerU 团队的批准。审批流程可能需要一些时间,请根据实际情况做好规划。
演示
输入 PDF

输出 Markdown

执照
MIT 许可证 - 有关详细信息,请参阅 LICENSE 文件。
致谢
该项目基于MinerU的 API。
Available Tools
2 toolsconvert_pdf_fileC
Convert local PDF file to Markdown, supports single file or file list
Args:
file_path: PDF file local path or path list, can be separated by spaces, commas, or newlines
enable_ocr: Whether to enable OCR (default: True)
Returns:
dict: Conversion result information
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | ||
| enable_ocr | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It mentions the OCR capability and return format (dict with conversion result information), but lacks critical details: whether this is a read-only operation, what happens with invalid files, if there are size/time limitations, what specific information the result dict contains, or error handling behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is reasonably concise with clear sections (Args, Returns) and front-loaded purpose statement. However, the 'Args' and 'Returns' labels add some redundancy since this information is partially available in the schema, and some sentences could be more efficiently worded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a file conversion tool with 2 parameters, no annotations, and no output schema, the description is insufficient. It lacks information about file format requirements, conversion quality, error conditions, output structure details, performance characteristics, or how the tool differs from its sibling. The return value description ('dict: Conversion result information') is particularly vague given no output schema exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description provides basic parameter information in the Args section, explaining that file_path accepts local paths or lists with various separators, and enable_ocr defaults to True. However, with 0% schema description coverage, it doesn't fully compensate by explaining path format requirements, file accessibility constraints, or what OCR actually does in this context beyond the boolean toggle.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: converting PDF files to Markdown format, with support for single files or lists. It specifies the resource (PDF files) and action (convert to Markdown), though it doesn't explicitly differentiate from the sibling tool 'convert_pdf_url' which likely handles URL-based PDFs rather than local files.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. While it mentions support for single files or lists, it doesn't explain when to choose this over 'convert_pdf_url' or other potential conversion tools. There's no mention of prerequisites, limitations, or typical use cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
convert_pdf_urlB
Convert PDF URL to Markdown, supports single URL or URL list
Args:
url: PDF file URL or URL list, can be separated by spaces, commas, or newlines
enable_ocr: Whether to enable OCR (default: True)
Returns:
dict: Conversion result information
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| enable_ocr | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions OCR support with a default setting, which adds some context, but fails to describe critical behaviors such as rate limits, authentication requirements, error handling, or what the conversion result information includes. For a tool that processes external URLs and performs conversion, this is a significant gap in transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized and front-loaded, starting with the core purpose followed by parameter details in a structured format. Every sentence adds value, with no redundant information. However, the use of 'dict' in the returns section is slightly vague, though this is mitigated by the lack of an output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (processing PDF URLs with OCR options) and the absence of annotations and output schema, the description is minimally adequate. It covers the basic purpose and parameters but lacks details on behavioral traits, error cases, and output structure. This leaves gaps that could hinder an agent's ability to use the tool effectively in varied contexts.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds meaningful semantics beyond the input schema, which has 0% description coverage. It explains that 'url' can be a single URL or a list separated by spaces, commas, or newlines, and clarifies the default value and purpose of 'enable_ocr'. This compensates well for the schema's lack of descriptions, making the parameters understandable without relying on the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: converting PDF URLs to Markdown format. It specifies the resource (PDF URLs) and the action (convert to Markdown), which is specific and actionable. However, it doesn't explicitly differentiate from its sibling tool 'convert_pdf_file' beyond mentioning URL vs. file handling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by mentioning support for single URLs or URL lists, but it doesn't provide explicit guidance on when to use this tool versus alternatives like 'convert_pdf_file'. No when-not-to-use scenarios or prerequisites are mentioned, leaving the agent to infer context from the tool name and description alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
- First observed
convert_pdf_file - First observed
convert_pdf_url
TDQS
Scored across 2 tools
The two tools have clearly distinct purposes: one handles local file paths, the other handles URLs. The naming and descriptions make it impossible to confuse which tool to use for a given input source.
Both tools follow an identical verb_noun pattern (convert_pdf_file and convert_pdf_url) with consistent snake_case formatting. The naming is perfectly predictable across the toolset.
With only two tools, the server feels minimal but functional. While it covers the core conversion task for both local files and URLs, the count is borderline thin for a PDF-to-Markdown domain that could potentially include more operations like batch processing, format options, or metadata extraction.
The server covers the essential conversion operation for both local files and remote URLs, which are the two main input sources for PDFs. The minor gap is the lack of additional PDF manipulation or output customization tools, but agents can perform basic conversions without dead ends.
Maintenance
Related MCP Connectors
PDF, Word, PowerPoint, Excel, HTML, EPUB to Markdown: OCR, page ranges, tables, RAG chunking
Convert documents and web pages to clean Markdown: PDF, DOCX, XLSX, EPUB, scanned files, any URL.
High-fidelity PDF to structured Markdown conversion and document field extraction.
Convert PDF, DOCX, HTML, and URLs to clean, LLM-ready markdown with tables preserved
Related MCP Servers
- AlicenseAqualityDmaintenanceConverts various file types and web content to Markdown format. It provides a set of tools to transform PDFs, images, audio files, web pages, and more into easily readable and shareable Markdown text.10345 npm2,997MIT
- AlicenseNot gradedqualityDmaintenanceConverts PDF files to Markdown format using AI sampling capabilities.MIT
- AlicenseNot gradedqualityFmaintenanceConverts markdown files into professional PDF documents with automatic table of contents and interactive navigation.8MIT
- AlicenseNot gradedqualityCmaintenanceConvert PDF documents to Markdown and query them using AI with source attribution and confidence scoring, supporting multiple LLM providers.MIT