Skip to main content
Glama

Lizeur - PDF Content Extraction MCP Server

Lizeur is a Model Context Protocol (MCP) server that enables AI assistants to extract and read content from PDF documents using Mistral AI's OCR capabilities. It provides a simple interface for converting PDF files to markdown text that can be easily consumed by AI models.

Features

  • PDF OCR Processing: Uses Mistral AI's latest OCR model to extract text from PDF documents

  • Intelligent Caching: Automatically caches processed documents to avoid re-processing

  • Markdown Output: Returns clean markdown text for easy integration with AI workflows

  • FastMCP Integration: Built with FastMCP for optimal performance and ease of use

Related MCP server: PDF Reader MCP Server

Prerequisites

  • Python 3.10

  • UV package manager

  • Mistral AI API key

Installation

From pypi

pip install lizeur

And add the following configuration to your mcp.json file:

Note: Lizeur will be installed in the python3.10 folder. If this folder is not in your system PATH, your IDE may not be able to detect the lizeur binary.

Solution: You can add the full path to the lizeur binary in the command field to ensure your IDE can locate it.

{
  "mcpServers": {
    "lizeur": {
      "command": "lizeur",
      "env": {
        "MISTRAL_API_KEY": "your-mistral-api-key-here",
        "CACHE_PATH": "your cache path",
      }
    }
  }
}

Manual

1. Clone the Repository

git clone https://github.com/SilverBzH/lizeur
cd lizeur

2. Create and Activate Virtual Environment

# Create a virtual environment
uv venv --python 3.10

# Activate the virtual environment
# On macOS/Linux:
source .venv/bin/activate

# On Windows:
# .venv\Scripts\activate

3. Install Dependencies and Build

# Install dependencies
uv sync

# Build the package
uv build

4. Install System-Wide

# Install the package system-wide
uv pip install --system .

This will install the lizeur command globally on your system.

Usage

Once configured, the MCP server provides two tools that can be used by AI assistants:

Available Functions

read_pdf

  • Function: read_pdf

  • Parameter: absolute_path (string) - The absolute path to the PDF file

  • Returns: Complete OCR response including all pages with markdown content, bounding boxes, and other OCR metadata

read_pdf_text

  • Function: read_pdf_text

  • Parameter: absolute_path (string) - The absolute path to the PDF file

  • Returns: Markdown text content from all pages without the full OCR metadata (simpler for agents to process)

Example Usage in AI Assistant

The AI assistant can now use the tools like this:

What the OP command looks like for this specific controller, here is the doc /path/to/document.pdf

The MCP server will:

  1. Check if the document is already cached

  2. If not cached, upload the PDF to Mistral AI for OCR processing This will use your MISTRAL API key and cost money

  3. Extract the text and convert it to markdown

  4. Cache the result for future use

  5. Return the markdown content

Note: Use read_pdf_text when you only need the text content, or read_pdf when you need the complete OCR response with metadata. read_pdf can be confusion for some agent if the pdf file is big.

Development

Local Development Setup

# Install in development mode
uv pip install -e .

# Run the server directly
python main.py

Project Structure

  • main.py - Main server implementation with FastMCP integration

  • pyproject.toml - Project configuration and dependencies

  • uv.lock - Locked dependency versions

Dependencies

  • mcp[cli]>=1.12.4 - Model Context Protocol implementation

  • mistralai>=0.0.10 - Mistral AI Python client

License

This project is licensed under the MIT License.

Support

For issues and questions, please refer to the project repository or contact the maintainers.

Available Tools

2 tools
read_pdfB

Read a PDF document and return the complete OCRResponse as a dictionary.

Returns the full OCR response including all pages, not just the first page. The response includes pages with markdown content, bounding boxes, and other OCR metadata.

ParametersJSON Schema
NameRequiredDescriptionDefault
absolute_pathYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It adds some context: it describes the return format ('complete OCRResponse as a dictionary'), specifies that it includes all pages and metadata like markdown content and bounding boxes. However, it lacks details on permissions, error handling, performance, or other operational traits that would be helpful for an agent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized with three sentences that are front-loaded: the first sentence states the core purpose, and the following sentences add important details about the response. There's minimal waste, though it could be slightly more structured by explicitly separating purpose from output details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (reading PDFs with OCR), no annotations, no output schema, and low schema coverage, the description is partially complete. It covers the output format and scope well but misses parameter explanations and behavioral context like error cases or limitations. It's adequate as a minimum viable description but has clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 1 parameter with 0% description coverage, so the description must compensate. It doesn't mention the 'absolute_path' parameter at all, failing to explain what it represents or provide any usage context. Since schema coverage is low, the description adds no value beyond what the schema provides, resulting in a baseline score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Read a PDF document and return the complete OCRResponse as a dictionary.' It specifies the verb ('read'), resource ('PDF document'), and output format ('OCRResponse as a dictionary'). However, it doesn't explicitly differentiate from its sibling 'read_pdf_text' beyond mentioning the full OCR response, leaving some ambiguity about the distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It mentions returning 'the complete OCR response' and not just the first page, but doesn't explain when to choose this over 'read_pdf_text' or other potential tools. There are no explicit when/when-not instructions or named alternatives beyond the sibling tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_pdf_textA

Read a PDF document and return only the markdown text content from all pages.

This is a simpler alternative to read_pdf that returns just the text content
without the full OCR metadata, which can be easier for agents to process.
ParametersJSON Schema
NameRequiredDescriptionDefault
absolute_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the tool's behavior by stating it reads PDFs and returns markdown text, but lacks details on error handling, performance, or limitations (e.g., file size, supported PDF formats). The description adds some value but doesn't fully cover behavioral traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose in the first sentence, followed by a concise explanation of its advantage over the sibling tool. Every sentence adds value without redundancy, making it efficient and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has an output schema (which handles return values), no annotations, and low complexity, the description is mostly complete. It clearly states the purpose, usage guidelines, and output type. However, it could benefit from more behavioral details (e.g., error cases) to be fully comprehensive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 0%, so the description must compensate. It doesn't explicitly mention the 'absolute_path' parameter, but implies it by referring to reading 'a PDF document.' Since there's only one parameter, the baseline is 4, as the description provides enough context to infer the parameter's purpose without detailed semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb ('Read') and resource ('a PDF document'), specifies the output ('markdown text content from all pages'), and explicitly distinguishes it from its sibling tool ('read_pdf') by noting it's a simpler alternative that returns just text without full OCR metadata. This provides specific differentiation and purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use this tool vs. its alternative: 'This is a simpler alternative to read_pdf that returns just the text content without the full OCR metadata, which can be easier for agents to process.' It clearly defines the context for choosing this tool over its sibling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

B3.4/5.0
Disambiguation2/5

The two tools have overlapping purposes: both read PDF documents and extract text content. While read_pdf returns full OCR metadata and read_pdf_text returns only markdown text, an agent might struggle to choose between them when only text is needed, as both could technically serve that purpose. The descriptions help clarify the difference, but the core functionality is very similar.

Naming Consistency5/5

Both tools follow a consistent verb_noun naming pattern with snake_case: read_pdf and read_pdf_text. The naming is clear and predictable, making it easy for agents to understand the action (read) and target (pdf or pdf_text). There are no deviations or mixed conventions in this small set.

Tool Count2/5

With only 2 tools, this server feels under-scoped for a PDF processing domain. While the tools cover reading and extracting text, there are obvious gaps like creating, editing, or converting PDFs. A typical PDF server would benefit from more operations, making this count too low for comprehensive functionality.

Completeness2/5

The tool surface is severely incomplete for PDF processing. It only includes reading operations (two variants of the same basic function) and lacks essential capabilities such as creating PDFs, merging/splitting files, converting formats, or editing content. This will likely cause agent failures when more complex tasks are required.

Maintenance

ActivityInactive
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/SilverBzH/lizeur'

If you have feedback or need assistance with the MCP directory API, please join our Discord server