Skip to main content
Glama
sanhua1
by sanhua1

pdf-agent-mcp

中文

pdf-agent-mcp 是一个本地 MCP 服务,给 AI agent 提供 PDF 文本层读取能力。

工具列表

  • inspect_pdf:检查 PDF 基本信息、页数、是否可能有文本层

  • extract_pdf_text:按 raw / lines / blocks 抽取文本

  • extract_pdf_outline:提取 PDF 目录(书签)

  • extract_pdf_page:提取单页文本项和坐标

使用说明

环境要求:Node.js 22+

npm install
npm run dev
npm run lint
npm test
npm run build

推荐直接用 npx 启动:

npx -y github:sanhua1/pdf-agent-mcp

Agent 自然语言交互示例

在 Claude/Codex 里可直接说:

  1. 先帮我 inspect 这个 PDF:/path/to/doc.pdf

  2. 把 1-5 页按 lines 模式提取出来

  3. 第 10 页排版乱,改用 blocks 模式再提取一次

  4. 先读取目录,再按章节整理成 Markdown 摘要

Claude Code 配置方法

{
  "mcpServers": {
    "pdf-agent-mcp": {
      "command": "npx",
      "args": ["-y", "github:sanhua1/pdf-agent-mcp"]
    }
  }
}

Codex 配置方法

[mcp_servers.pdf-agent-mcp]
command = "npx"
args = ["-y", "github:sanhua1/pdf-agent-mcp"]

English

pdf-agent-mcp is a local MCP server for extracting text-layer content from PDF files.

Tools

  • inspect_pdf: inspect metadata, page count, and text-layer hints

  • extract_pdf_text: extract text in raw / lines / blocks modes

  • extract_pdf_outline: read PDF bookmarks/outlines

  • extract_pdf_page: extract text items with coordinates from a single page

Quick Start

Requirement: Node.js 22+

npm install
npm run dev

Run with npx:

npx -y github:sanhua1/pdf-agent-mcp

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/sanhua1/pdf-agent-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server