pdf4vllm
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@pdf4vllmread the quarterly report PDF in my documents folder"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
pdf4vllm
PDF reading MCP server optimized for vision LLMs.
문제
방식 | 문제점 |
텍스트 추출 | 인코딩 깨짐 → 쓰레기 출력, 이미지-텍스트 순서 뒤섞임 |
이미지 변환 | 토큰 폭발 (특히 페이지 많을 때) |
Related MCP server: PDF Reader MCP Server
해결
pdf4vllm은 PDF가 지저분하다고 가정합니다.
텍스트 손상 자동 감지 → 이미지로 자동 전환
읽기 순서 보존 (텍스트 → 표 → 이미지 블록 순서대로)
페이지 제한으로 컨텍스트 오버플로우 방지
불필요한 이미지 자동 필터링 (로고, 선, 헤더/푸터)
설치
pip install pdf4vllm-mcp
# 또는
uvx pdf4vllm-mcpClaude Desktop 설정
git clone https://github.com/PyJudge/pdf4vllm-mcp.git
cd pdf4vllm-mcp
python scripts/install_mcp.py또는 직접 설정 (~/Library/Application Support/Claude/claude_desktop_config.json):
{
"mcpServers": {
"pdf4vllm": {
"command": "/python/경로",
"args": ["/pdf4vllm-mcp/경로/src/server.py"]
}
}
}도구
도구 | 설명 |
| PDF 파일 찾기 (glob 패턴 |
| PDF 내용 블록으로 추출 |
| PDF 내 텍스트 검색 ( |
추출 모드
모드 | 설명 |
| 텍스트 추출 시도 → 손상 감지 시 이미지로 전환 |
| 텍스트/표만 추출, 이미지 없음 |
| 페이지를 이미지로만 렌더링 |
Problem
Approach | Issue |
Text extraction | Encoding corruption → garbage output, mixed text-image ordering |
Image conversion | Token explosion (especially with many pages) |
Solution
pdf4vllm assumes PDFs are messy.
Auto-detects text corruption → switches to image automatically
Preserves reading order (text → table → image blocks in sequence)
Page limits prevent context overflow
Filters unnecessary images (logos, lines, headers/footers)
PDF Input
↓
Corruption Detection (pdfminer.six + pattern analysis)
↓
┌─────────────┬─────────────┐
│ Corrupted │ Clean │
│ → Image │ → Text + │
│ only │ Tables + │
│ │ Images │
└─────────────┴─────────────┘
↓
Ordered Blocks (JSON)Install
pip install pdf4vllm-mcp
# or run without installing
uvx pdf4vllm-mcpClaude Desktop Setup
git clone https://github.com/PyJudge/pdf4vllm-mcp.git
cd pdf4vllm-mcp
python scripts/install_mcp.pyOr manually edit ~/Library/Application Support/Claude/claude_desktop_config.json:
{
"mcpServers": {
"pdf4vllm": {
"command": "/path/to/python",
"args": ["/path/to/pdf4vllm-mcp/src/server.py"]
}
}
}Claude Code Setup
Create .mcp.json in your project:
{
"mcpServers": {
"pdf4vllm": {
"command": "uvx",
"args": ["pdf4vllm-mcp"]
}
}
}Tools
Tool | Description |
| Find PDF files with glob filtering ( |
| Extract PDF content as ordered blocks |
| Search text in PDFs using pdfgrep (requires |
Extraction Modes
Mode | Description |
| Try text extraction → switch to image if corrupted |
| Text/tables only, no images |
| Render pages as images only |
Output Format
{
"pages": [
{
"page_number": 1,
"content_blocks": [
{"type": "text", "content": "..."},
{"type": "table", "content": "| A | B |"},
{"type": "image", "content": "[IMAGE_0]"}
]
}
]
}When text is corrupted:
{
"page_number": 2,
"content_blocks": [],
"text_corrupted": true,
"page_image": "[IMAGE_1]"
}Configuration
config.json or environment variables:
{
"max_pages_per_request": 10,
"max_image_dimension": 842,
"page_image_dpi": 100
}export PDF_MAX_PAGES=20
export PDF_PAGE_IMAGE_DPI=150Test Server
pip install pdf4vllm-mcp[test]
python test_server.py
# → http://localhost:8000License
MIT
Available Tools
3 toolsgrep_pdfA
Search text in PDFs. Standard grep/rg does NOT work on PDFs (binary format). Use this tool instead. Returns matching lines with page numbers. NOTE: No page limit (unlike read_pdf's 10-page limit).
| Name | Required | Description | Default |
|---|---|---|---|
| pattern | Yes | Search pattern (regex by default) | |
| file_path | No | Specific PDF file. If not provided, searches ALL PDFs in working_directory | |
| working_directory | No | Base directory for search (only used when file_path is not provided) | . |
| ignore_case | No | Case-insensitive search | |
| fixed_strings | No | Treat pattern as literal string, not regex | |
| context | No | Lines of context before/after match (0-5, default: 2) | |
| max_count | No | Maximum matches to return (1-100) | |
| recursive | No | Include subdirectories when searching directory | |
| start_page | No | Start page (1-indexed, inclusive) | |
| end_page | No | End page (1-indexed, inclusive). None = last page |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It successfully describes key behavioral traits: that it returns 'matching lines with page numbers' and has 'No page limit (unlike read_pdf's 10-page limit).' However, it doesn't mention performance characteristics, error conditions, or what happens with invalid PDFs. For a tool with no annotations, this is good but not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise (3 sentences) and front-loaded with the core purpose. Every sentence earns its place: first states what it does, second explains why to use it over alternatives, third distinguishes from sibling tool. Zero wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (10 parameters, no output schema, no annotations), the description does well by explaining the core behavior and key differentiators. However, it doesn't describe the return format in detail (beyond 'matching lines with page numbers') or error handling. For a search tool with rich parameters but no output schema, this is good but could be more complete about results structure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 10 parameters thoroughly. The description doesn't add any parameter-specific information beyond what's in the schema. According to scoring rules, with high schema coverage (>80%), the baseline is 3 even with no param info in description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Search text in PDFs' (verb+resource). It explicitly distinguishes from standard grep/rg tools that don't work on PDFs, and distinguishes from sibling 'read_pdf' by noting the lack of page limit. This is specific and clearly differentiates from alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use this tool: 'Standard grep/rg does NOT work on PDFs (binary format). Use this tool instead.' It also distinguishes from sibling 'read_pdf' by noting 'No page limit (unlike read_pdf's 10-page limit).' This gives clear when-to-use and when-not-to-use guidance with named alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_pdfsA
Find PDF files in a directory. Use name_pattern for glob filtering (e.g., 'report'). Returns name, path, pages for each PDF. Use the returned 'path' directly with read_pdf.
| Name | Required | Description | Default |
|---|---|---|---|
| working_directory | No | Working directory to search (relative or absolute path, default: current directory) | . |
| recursive | No | Whether to include subdirectories | |
| max_depth | No | Maximum recursion depth (default: 2) | |
| name_pattern | No | Glob pattern for filename filtering (e.g., '*report*', 'doc_202?.pdf') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes the tool's read-only nature (implied by 'Find' and returns data) and output format ('Returns name, path, pages'), but lacks details on error handling, performance limits, or permissions required for directory access. It adds value beyond the schema by explaining the return structure and integration with 'read_pdf'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is highly concise and front-loaded, with three sentences that each serve a clear purpose: stating the tool's function, providing a parameter example, and explaining the output and integration. There is no wasted text, making it efficient and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (searching files with filtering), no annotations, and no output schema, the description is largely complete. It covers purpose, usage, and output structure, but lacks details on error cases or performance constraints (e.g., large directories). The integration hint with 'read_pdf' enhances context, though some behavioral gaps remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents all 4 parameters. The description adds minimal semantic value beyond the schema, only mentioning 'name_pattern' with an example ('e.g., '*report*''), which is already covered in the schema description. No additional parameter insights are provided, meeting the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Find PDF files'), target resource ('in a directory'), and scope ('Returns name, path, pages for each PDF'). It distinguishes from sibling tools by explicitly mentioning 'read_pdf' for a different purpose, avoiding overlap.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit guidance on when to use this tool ('Find PDF files in a directory') and when to use an alternative ('Use the returned 'path' directly with read_pdf'), clearly differentiating from the sibling 'read_pdf' tool. It also includes a practical example ('e.g., '*report*'') to illustrate usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_pdfA
Read PDF content. Always prefer this over cat or file read for PDF files. Limits: 10 pages per request. Works with both text and scanned documents. Use 'image_only' to see actual page layout, or 'text_only' for pure text.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | PDF file path (relative or absolute path) | |
| start_page | No | Start page (1-indexed, inclusive) | |
| end_page | No | End page (1-indexed, inclusive). None = last page | |
| extraction_mode | No | Content extraction mode: - 'auto' (default): Smart detection - extract text/tables, add page image only if corrupted - 'text_only': Extract text/tables only, no images - 'image_only': Skip text extraction, provide only full page images | auto |
| filter_header_footer | No | Whether to filter out header/footer images (top/bottom 6% of page) | |
| crop_images | No | Whether to crop images to max_image_dimension | |
| max_image_dimension | No | Maximum image dimension in pixels (default: 842, A4 height) | |
| page_image_dpi | No | DPI for page image rendering (default: 100) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes key behavioral traits: the 10-page limit per request, compatibility with both text and scanned documents, and the distinction between extraction modes. However, it doesn't mention error handling, performance characteristics, or authentication requirements, which would be valuable for a tool with 8 parameters.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized with four concise sentences that each serve a distinct purpose: stating the core function, providing usage preference, specifying limits/compatibility, and explaining mode options. It's front-loaded with the most important information. The only minor improvement would be integrating the page limit more naturally with the parameter explanations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (8 parameters, no output schema, no annotations), the description does a good job covering the essential context. It explains the tool's purpose, when to use it, key limitations, and high-level mode options. However, for a tool with this many parameters and no output schema, it could benefit from mentioning what the return format looks like or providing more guidance on parameter combinations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, so the schema already documents all 8 parameters thoroughly. The description adds minimal value beyond the schema - it mentions the 'image_only' and 'text_only' modes which are already in the schema's enum, but doesn't provide additional context about parameter interactions or usage patterns. This meets the baseline expectation when schema coverage is complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs ('Read PDF content') and distinguishes it from alternatives ('Always prefer this over cat or file read for PDF files'). It identifies the resource (PDF files) and differentiates from sibling tools like 'grep_pdf' and 'list_pdfs' by focusing on content extraction rather than searching or listing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use this tool vs. alternatives ('Always prefer this over cat or file read for PDF files') and offers specific usage contexts ('Use 'image_only' to see actual page layout, or 'text_only' for pure text'). It clearly directs the agent away from generic file operations for PDFs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
Each tool has a clearly distinct purpose: grep_pdf searches text within PDFs, list_pdfs finds PDF files in directories, and read_pdf extracts content from PDFs. There is no overlap in functionality, making it easy for an agent to select the correct tool without confusion.
All tool names follow a consistent verb_noun pattern with snake_case (grep_pdf, list_pdfs, read_pdf). The naming is predictable and readable, with no deviations or mixed conventions, ensuring clarity and ease of use.
With 3 tools, the server is well-scoped for its purpose of PDF processing. Each tool earns its place by covering essential operations: listing, reading, and searching PDFs, without being overly sparse or bloated.
The tool set covers core PDF operations (list, read, search) effectively, with no dead ends. A minor gap exists in lacking explicit CRUD operations like create or delete PDFs, but this is reasonable given the server's focus on reading and searching rather than full lifecycle management.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
High-fidelity PDF to structured Markdown conversion and document field extraction.
Extract tables, text and formulas from PDFs, including scanned pages and broken text layers.
Turn any PDF into structured JSON via AI + OCR: invoices, bank statements, contracts.
OCR and document understanding: extract text from images, then summarize or translate it.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceProvides intelligent OCR and PDF processing capabilities that automatically detect whether PDFs contain digital text or scanned images and apply appropriate extraction methods. Supports text extraction, OCR processing, structure analysis, and batch operations.MIT
- AlicenseAqualityDmaintenanceEnables reading, searching, and metadata extraction from PDF files without loading the entire content into the context window. It provides efficient tools for text cleaning, page-specific extraction, and context-aware search results.3581MIT
- AlicenseAqualityDmaintenanceEnables LLMs to read and extract content from PDF files with high-fidelity LaTeX recognition and layout awareness using a Python-based extraction engine. It includes a robust Node.js fallback and supports page range filtering for efficient processing of large documents.189MIT
- FlicenseAqualityDmaintenanceAgentic PDF search via MCP, enabling intelligent document retrieval through LLM reasoning instead of vector similarity.24
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/PyJudge/pdf4vllm-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server