Skip to main content
Glama

pdf4vllm

License: MIT Python 3.10+ PyPI Open in Gitpod

PDF reading MCP server optimized for vision LLMs.

문제

방식

문제점

텍스트 추출

인코딩 깨짐 → 쓰레기 출력, 이미지-텍스트 순서 뒤섞임

이미지 변환

토큰 폭발 (특히 페이지 많을 때)

Related MCP server: PDF Reader MCP Server

해결

pdf4vllm은 PDF가 지저분하다고 가정합니다.

  • 텍스트 손상 자동 감지 → 이미지로 자동 전환

  • 읽기 순서 보존 (텍스트 → 표 → 이미지 블록 순서대로)

  • 페이지 제한으로 컨텍스트 오버플로우 방지

  • 불필요한 이미지 자동 필터링 (로고, 선, 헤더/푸터)

설치

pip install pdf4vllm-mcp
# 또는
uvx pdf4vllm-mcp

Claude Desktop 설정

git clone https://github.com/PyJudge/pdf4vllm-mcp.git
cd pdf4vllm-mcp
python scripts/install_mcp.py

또는 직접 설정 (~/Library/Application Support/Claude/claude_desktop_config.json):

{
  "mcpServers": {
    "pdf4vllm": {
      "command": "/python/경로",
      "args": ["/pdf4vllm-mcp/경로/src/server.py"]
    }
  }
}

도구

도구

설명

list_pdfs

PDF 파일 찾기 (glob 패턴 name_pattern 지원)

read_pdf

PDF 내용 블록으로 추출

grep_pdf

PDF 내 텍스트 검색 (pdfgrep 설치 필요)

추출 모드

모드

설명

auto (기본)

텍스트 추출 시도 → 손상 감지 시 이미지로 전환

text_only

텍스트/표만 추출, 이미지 없음

image_only

페이지를 이미지로만 렌더링


Problem

Approach

Issue

Text extraction

Encoding corruption → garbage output, mixed text-image ordering

Image conversion

Token explosion (especially with many pages)

Solution

pdf4vllm assumes PDFs are messy.

  • Auto-detects text corruption → switches to image automatically

  • Preserves reading order (text → table → image blocks in sequence)

  • Page limits prevent context overflow

  • Filters unnecessary images (logos, lines, headers/footers)

PDF Input
    ↓
Corruption Detection (pdfminer.six + pattern analysis)
    ↓
┌─────────────┬─────────────┐
│  Corrupted  │    Clean    │
│  → Image    │  → Text +   │
│    only     │    Tables + │
│             │    Images   │
└─────────────┴─────────────┘
    ↓
Ordered Blocks (JSON)

Install

pip install pdf4vllm-mcp
# or run without installing
uvx pdf4vllm-mcp

Claude Desktop Setup

git clone https://github.com/PyJudge/pdf4vllm-mcp.git
cd pdf4vllm-mcp
python scripts/install_mcp.py

Or manually edit ~/Library/Application Support/Claude/claude_desktop_config.json:

{
  "mcpServers": {
    "pdf4vllm": {
      "command": "/path/to/python",
      "args": ["/path/to/pdf4vllm-mcp/src/server.py"]
    }
  }
}

Claude Code Setup

Create .mcp.json in your project:

{
  "mcpServers": {
    "pdf4vllm": {
      "command": "uvx",
      "args": ["pdf4vllm-mcp"]
    }
  }
}

Tools

Tool

Description

list_pdfs

Find PDF files with glob filtering (name_pattern)

read_pdf

Extract PDF content as ordered blocks

grep_pdf

Search text in PDFs using pdfgrep (requires pdfgrep installed)

Extraction Modes

Mode

Description

auto (default)

Try text extraction → switch to image if corrupted

text_only

Text/tables only, no images

image_only

Render pages as images only

Output Format

{
  "pages": [
    {
      "page_number": 1,
      "content_blocks": [
        {"type": "text", "content": "..."},
        {"type": "table", "content": "| A | B |"},
        {"type": "image", "content": "[IMAGE_0]"}
      ]
    }
  ]
}

When text is corrupted:

{
  "page_number": 2,
  "content_blocks": [],
  "text_corrupted": true,
  "page_image": "[IMAGE_1]"
}

Configuration

config.json or environment variables:

{
  "max_pages_per_request": 10,
  "max_image_dimension": 842,
  "page_image_dpi": 100
}
export PDF_MAX_PAGES=20
export PDF_PAGE_IMAGE_DPI=150

Test Server

pip install pdf4vllm-mcp[test]
python test_server.py
# → http://localhost:8000

License

MIT


GitHub · PyPI

Available Tools

3 tools
grep_pdfA

Search text in PDFs. Standard grep/rg does NOT work on PDFs (binary format). Use this tool instead. Returns matching lines with page numbers. NOTE: No page limit (unlike read_pdf's 10-page limit).

ParametersJSON Schema
NameRequiredDescriptionDefault
patternYesSearch pattern (regex by default)
file_pathNoSpecific PDF file. If not provided, searches ALL PDFs in working_directory
working_directoryNoBase directory for search (only used when file_path is not provided).
ignore_caseNoCase-insensitive search
fixed_stringsNoTreat pattern as literal string, not regex
contextNoLines of context before/after match (0-5, default: 2)
max_countNoMaximum matches to return (1-100)
recursiveNoInclude subdirectories when searching directory
start_pageNoStart page (1-indexed, inclusive)
end_pageNoEnd page (1-indexed, inclusive). None = last page

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It successfully describes key behavioral traits: that it returns 'matching lines with page numbers' and has 'No page limit (unlike read_pdf's 10-page limit).' However, it doesn't mention performance characteristics, error conditions, or what happens with invalid PDFs. For a tool with no annotations, this is good but not comprehensive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise (3 sentences) and front-loaded with the core purpose. Every sentence earns its place: first states what it does, second explains why to use it over alternatives, third distinguishes from sibling tool. Zero wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (10 parameters, no output schema, no annotations), the description does well by explaining the core behavior and key differentiators. However, it doesn't describe the return format in detail (beyond 'matching lines with page numbers') or error handling. For a search tool with rich parameters but no output schema, this is good but could be more complete about results structure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 10 parameters thoroughly. The description doesn't add any parameter-specific information beyond what's in the schema. According to scoring rules, with high schema coverage (>80%), the baseline is 3 even with no param info in description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Search text in PDFs' (verb+resource). It explicitly distinguishes from standard grep/rg tools that don't work on PDFs, and distinguishes from sibling 'read_pdf' by noting the lack of page limit. This is specific and clearly differentiates from alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use this tool: 'Standard grep/rg does NOT work on PDFs (binary format). Use this tool instead.' It also distinguishes from sibling 'read_pdf' by noting 'No page limit (unlike read_pdf's 10-page limit).' This gives clear when-to-use and when-not-to-use guidance with named alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_pdfsA

Find PDF files in a directory. Use name_pattern for glob filtering (e.g., 'report'). Returns name, path, pages for each PDF. Use the returned 'path' directly with read_pdf.

ParametersJSON Schema
NameRequiredDescriptionDefault
working_directoryNoWorking directory to search (relative or absolute path, default: current directory).
recursiveNoWhether to include subdirectories
max_depthNoMaximum recursion depth (default: 2)
name_patternNoGlob pattern for filename filtering (e.g., '*report*', 'doc_202?.pdf')

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes the tool's read-only nature (implied by 'Find' and returns data) and output format ('Returns name, path, pages'), but lacks details on error handling, performance limits, or permissions required for directory access. It adds value beyond the schema by explaining the return structure and integration with 'read_pdf'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is highly concise and front-loaded, with three sentences that each serve a clear purpose: stating the tool's function, providing a parameter example, and explaining the output and integration. There is no wasted text, making it efficient and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (searching files with filtering), no annotations, and no output schema, the description is largely complete. It covers purpose, usage, and output structure, but lacks details on error cases or performance constraints (e.g., large directories). The integration hint with 'read_pdf' enhances context, though some behavioral gaps remain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents all 4 parameters. The description adds minimal semantic value beyond the schema, only mentioning 'name_pattern' with an example ('e.g., '*report*''), which is already covered in the schema description. No additional parameter insights are provided, meeting the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Find PDF files'), target resource ('in a directory'), and scope ('Returns name, path, pages for each PDF'). It distinguishes from sibling tools by explicitly mentioning 'read_pdf' for a different purpose, avoiding overlap.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit guidance on when to use this tool ('Find PDF files in a directory') and when to use an alternative ('Use the returned 'path' directly with read_pdf'), clearly differentiating from the sibling 'read_pdf' tool. It also includes a practical example ('e.g., '*report*'') to illustrate usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_pdfA

Read PDF content. Always prefer this over cat or file read for PDF files. Limits: 10 pages per request. Works with both text and scanned documents. Use 'image_only' to see actual page layout, or 'text_only' for pure text.

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathYesPDF file path (relative or absolute path)
start_pageNoStart page (1-indexed, inclusive)
end_pageNoEnd page (1-indexed, inclusive). None = last page
extraction_modeNoContent extraction mode: - 'auto' (default): Smart detection - extract text/tables, add page image only if corrupted - 'text_only': Extract text/tables only, no images - 'image_only': Skip text extraction, provide only full page imagesauto
filter_header_footerNoWhether to filter out header/footer images (top/bottom 6% of page)
crop_imagesNoWhether to crop images to max_image_dimension
max_image_dimensionNoMaximum image dimension in pixels (default: 842, A4 height)
page_image_dpiNoDPI for page image rendering (default: 100)

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes key behavioral traits: the 10-page limit per request, compatibility with both text and scanned documents, and the distinction between extraction modes. However, it doesn't mention error handling, performance characteristics, or authentication requirements, which would be valuable for a tool with 8 parameters.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized with four concise sentences that each serve a distinct purpose: stating the core function, providing usage preference, specifying limits/compatibility, and explaining mode options. It's front-loaded with the most important information. The only minor improvement would be integrating the page limit more naturally with the parameter explanations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (8 parameters, no output schema, no annotations), the description does a good job covering the essential context. It explains the tool's purpose, when to use it, key limitations, and high-level mode options. However, for a tool with this many parameters and no output schema, it could benefit from mentioning what the return format looks like or providing more guidance on parameter combinations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, so the schema already documents all 8 parameters thoroughly. The description adds minimal value beyond the schema - it mentions the 'image_only' and 'text_only' modes which are already in the schema's enum, but doesn't provide additional context about parameter interactions or usage patterns. This meets the baseline expectation when schema coverage is complete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with specific verbs ('Read PDF content') and distinguishes it from alternatives ('Always prefer this over cat or file read for PDF files'). It identifies the resource (PDF files) and differentiates from sibling tools like 'grep_pdf' and 'list_pdfs' by focusing on content extraction rather than searching or listing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use this tool vs. alternatives ('Always prefer this over cat or file read for PDF files') and offers specific usage contexts ('Use 'image_only' to see actual page layout, or 'text_only' for pure text'). It clearly directs the agent away from generic file operations for PDFs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

A4.4/5.0
Disambiguation5/5

Each tool has a clearly distinct purpose: grep_pdf searches text within PDFs, list_pdfs finds PDF files in directories, and read_pdf extracts content from PDFs. There is no overlap in functionality, making it easy for an agent to select the correct tool without confusion.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern with snake_case (grep_pdf, list_pdfs, read_pdf). The naming is predictable and readable, with no deviations or mixed conventions, ensuring clarity and ease of use.

Tool Count5/5

With 3 tools, the server is well-scoped for its purpose of PDF processing. Each tool earns its place by covering essential operations: listing, reading, and searching PDFs, without being overly sparse or bloated.

Completeness4/5

The tool set covers core PDF operations (list, read, search) effectively, with no dead ends. A minor gap exists in lacking explicit CRUD operations like create or delete PDFs, but this is reasonable given the server's focus on reading and searching rather than full lifecycle management.

Maintenance

ActivityInactive
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides intelligent OCR and PDF processing capabilities that automatically detect whether PDFs contain digital text or scanned images and apply appropriate extraction methods. Supports text extraction, OCR processing, structure analysis, and batch operations.
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    Enables reading, searching, and metadata extraction from PDF files without loading the entire content into the context window. It provides efficient tools for text cleaning, page-specific extraction, and context-aware search results.
    3
    58
    1
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    Enables LLMs to read and extract content from PDF files with high-fidelity LaTeX recognition and layout awareness using a Python-based extraction engine. It includes a robust Node.js fallback and supports page range filtering for efficient processing of large documents.
    1
    89
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/PyJudge/pdf4vllm-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server