Skip to main content
Glama

Mistral OCR MCP Server

A Model Context Protocol (MCP) server that provides tools for extracting text and images from PDF and image files using the Mistral OCR API.

Features

  • Local & URL Extraction: Extract markdown from local files or remote URLs

  • Image Handling on Demand: Optionally save embedded images to disk with proper relative links

  • Advanced OCR: Page selection, table format control, model selection

  • Health Check: Built-in API status endpoint

  • Security Sandbox: Restricts file writes to a configured allowed directory

  • Zero-Install Deployment: Run with uvx without prior installation

  • Supported Formats: PDF (.pdf), PNG (.png), JPEG (.jpg, .jpeg), WebP (.webp), GIF (.gif)


Related MCP server: MCP-PDF2MD

Client Configuration

Claude Desktop

Add this to your claude_desktop_config.json:

  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json

  • Windows: %APPDATA%\Claude\claude_desktop_config.json

{
  "mcpServers": {
    "mistral-ocr": {
      "command": "uvx",
      "args": ["mistral-ocr-mcp"],
      "env": {
        "MISTRAL_API_KEY": "your-api-key-here",
        "MISTRAL_OCR_ALLOWED_DIR": "/absolute/path/to/allowed/directory"
      }
    }
  }
}

OpenCode

Add this to the mcp section of your configuration file:

{
  "mcp": {
    "mistral-ocr": {
      "type": "local",
      "command": ["uvx", "mistral-ocr-mcp"],
      "enabled": true,
      "environment": {
        "MISTRAL_API_KEY": "your-api-key-here",
        "MISTRAL_OCR_ALLOWED_DIR": "/absolute/path/to/allowed/directory"
      }
    }
  }
}

Codex

If you use the Codex CLI, you can add the server with:

codex mcp add mistral-ocr -- uvx mistral-ocr-mcp

Make sure the environment variables MISTRAL_API_KEY and MISTRAL_OCR_ALLOWED_DIR are set in your shell environment.


Configuration

Required Environment Variables

Variable

Description

Example

MISTRAL_API_KEY

Your Mistral API key (never logged)

sk-abc123...

MISTRAL_OCR_ALLOWED_DIR

Absolute path to allowed write directory

/Users/username/workdir

Security Sandbox

The server enforces a write directory sandbox to prevent unauthorized file writes. When include_images is True, the output_dir parameter must be within MISTRAL_OCR_ALLOWED_DIR. Text-only extraction (include_images=False) is read-only and has no sandbox restrictions.

Validation Examples:

MISTRAL_OCR_ALLOWED_DIR

output_dir

Result

/Users/username/workdir

/Users/username/workdir/project/output

✅ Allowed

/Users/username/workdir

/Users/username/workdir

✅ Allowed (exact match)

/Users/username/workdir

/Users/username/documents

❌ Rejected

/Users/username/workdir

/Users/username/workdir/../documents

❌ Rejected (resolves outside)

Security Notes:

  • All paths are canonicalized (symlinks resolved, .. eliminated) before validation

  • Image filenames are sanitized to prevent path traversal attacks


Tool Reference

Tool 1: extract_markdown

Extract markdown from a local file, optionally saving embedded images to disk.

Arguments:

Parameter

Type

Required

Default

Description

file_path

string

Yes

Absolute path to input file (PDF or image)

output_dir

string

No

null

Absolute path to output parent directory. Required when include_images is True.

include_images

boolean

No

false

When True and output_dir is set, saves images to disk

Returns (text-only):

{"result": "# Document Title\n\nExtracted markdown content..."}

Returns (with images):

{
  "output_directory": "/absolute/path/to/output/report",
  "markdown_file": "/absolute/path/to/output/report/content.md",
  "images": ["img_abc123.png", "img_def456.jpeg"]
}

Behavior (with images):

  1. Creates a subdirectory named after the input file stem (e.g., report for report.pdf)

  2. If the subdirectory already exists, appends a timestamp: report_20260102_143022

  3. Saves all extracted images as <sanitized_id>.<ext> (e.g., img_abc123.png)

  4. Saves markdown to content.md with relative image links (e.g., ![](./img_abc123.png))

Output Structure:

/Users/username/workdir/extracted/
  quarterly-report/
    content.md          # Markdown with relative image links
    img_abc123.png      # First extracted image
    img_def456.jpeg     # Second extracted image

Tool 2: extract_markdown_from_url

Extract markdown from a publicly accessible URL, optionally saving embedded images to disk.

Arguments:

Parameter

Type

Required

Default

Description

file_url

string

Yes

Public URL to a PDF or image

output_dir

string

No

null

Absolute path to output parent directory. Required when include_images is True.

include_images

boolean

No

false

When True and output_dir is set, saves images to disk

Returns (text-only):

{"result": "# Document Title\n\nExtracted markdown content..."}

Returns (with images):

{
  "output_directory": "/absolute/path/to/output/doc",
  "markdown_file": "/absolute/path/to/output/doc/content.md",
  "images": ["img_abc123.png"]
}

Tool 3: extract_markdown_advanced

Extract markdown with advanced OCR options.

Arguments:

Parameter

Type

Required

Default

Description

file_path

string

Yes

Absolute path to input file (PDF or image)

pages

array[integer]

No

null

Page numbers to process (1-indexed, e.g. [1, 3, 5])

table_format_

string

No

null

Table output format ("markdown" or "html")

model

string

No

"mistral-ocr-latest"

OCR model to use

Returns:

{"result": "# Document\n\n| Col 1 | Col 2 |\n|-------|-------|\n..."}

Tool 4: ocr_status

Check API connectivity and key validity.

Arguments: none

Returns:

{
  "status": "ok",
  "message": "API key is working"
}

Example Client Usage

Here's a minimal Python example using the MCP SDK to call the tools:

import asyncio
from mcp import ClientSession, StdioServerParameters
from mcp.client.stdio import stdio_client

async def extract_document():
    server_params = StdioServerParameters(
        command="mistral-ocr-mcp",
        env={
            "MISTRAL_API_KEY": "your-api-key",
            "MISTRAL_OCR_ALLOWED_DIR": "/Users/username/workdir"
        }
    )
    
    async with stdio_client(server_params) as (read, write):
        async with ClientSession(read, write) as session:
            await session.initialize()
            
            # Text-only extraction
            result = await session.call_tool(
                "extract_markdown",
                arguments={"file_path": "/path/to/document.pdf"}
            )
            print(result.content[0].text)
            
            # Extraction with images
            result = await session.call_tool(
                "extract_markdown",
                arguments={
                    "file_path": "/path/to/document.pdf",
                    "include_images": True,
                    "output_dir": "/Users/username/workdir/output"
                }
            )
            print(result.content[0].text)
            
            # Extract from URL
            result = await session.call_tool(
                "extract_markdown_from_url",
                arguments={"file_url": "https://example.com/doc.pdf"}
            )
            print(result.content[0].text)
            
            # Check API status
            result = await session.call_tool("ocr_status", arguments={})
            print(result.content[0].text)

asyncio.run(extract_document())

Troubleshooting

Error

Cause

Solution

Missing required environment variable: MISTRAL_API_KEY

MISTRAL_API_KEY not set

Set the environment variable before running the server

Missing required environment variable: MISTRAL_OCR_ALLOWED_DIR

MISTRAL_OCR_ALLOWED_DIR not set

Set the environment variable to an absolute path

MISTRAL_OCR_ALLOWED_DIR must be an absolute path

Relative path provided (e.g., ~/documents)

Use an absolute path (e.g., /Users/username/documents)

MISTRAL_OCR_ALLOWED_DIR does not exist

Directory does not exist on filesystem

Create the directory first: mkdir -p /path/to/dir

MISTRAL_OCR_ALLOWED_DIR is not a directory

Path points to a file, not a directory

Ensure the path is a directory

validate file_path: must be an absolute path: {path}

Relative path provided for input file

Use an absolute path (e.g., /Users/username/file.pdf)

validate file_path: resolve failed, path does not exist: {path}

Input file does not exist

Check the file path and ensure the file exists

validate file_path: unsupported file type '{suffix}'. Supported types: ...

File extension not supported

Use .pdf, .png, .jpg, .jpeg, .webp, or .gif

validate output_dir: resolve failed, path does not exist: {path}

Output directory does not exist

Create the directory first: mkdir -p {path}

validate output_dir: path is not a directory: {path}

Path points to a file, not a directory

Ensure the path is a directory

validate output_dir: writability check failed, directory not writable: {path}

Output directory exists but is not writable

Check directory permissions: chmod u+w {path}

output_dir must be within the allowed directory

output_dir is outside MISTRAL_OCR_ALLOWED_DIR

Use a path within the allowed directory

Mistral OCR request failed (status=401): {message}

Invalid API key

Check your MISTRAL_API_KEY

Mistral OCR request failed (status=429): {message}

Rate limit exceeded

Wait and retry, or check your API quota


Development

Setup

Clone the repository and install with development dependencies:

git clone https://github.com/ORDIS-Co-Ltd/mistral-ocr-mcp
cd mistral-ocr-mcp
pip install -e '.[dev]'

Run the server locally:

MISTRAL_API_KEY="your-key" \
MISTRAL_OCR_ALLOWED_DIR="/path/to/allowed/dir" \
python -m mistral_ocr_mcp

Run Tests

pytest

Project Structure

mistral-ocr-mcp/
├── src/
│   └── mistral_ocr_mcp/
│       ├── __init__.py
│       ├── __main__.py          # Entry point
│       ├── server.py            # MCP server and tool definitions
│       ├── config.py            # Configuration loading and validation
│       ├── extraction.py        # OCR orchestration logic
│       ├── mistral_client.py    # Mistral API client
│       ├── images.py            # Image parsing and saving
│       ├── markdown_rewrite.py  # Markdown link rewriting
│       └── path_sandbox.py      # Path validation and sandbox enforcement
├── tests/                       # Unit tests
├── pyproject.toml              # Package configuration
└── README.md                   # This file

License

MIT


Contributing

Contributions are welcome! Please open an issue or submit a pull request.


Available Tools

4 tools
extract_markdownA

Extract markdown text from a PDF or image file.

When output_dir is provided, saves the extracted markdown to content.md inside a named subdirectory. When include_images is also True, saves embedded images alongside the markdown file. Otherwise returns the markdown text inline.

Args: file_path: Absolute path to the input file (PDF or image) output_dir: Absolute path to an existing output directory (must be within allowed dir). When set, saves markdown to disk at <output_dir>/<file_stem>/content.md. include_images: When True (requires output_dir), save images to disk and rewrite markdown with relative image links.

Returns: When output_dir is not set: result: Extracted markdown content When output_dir is set (with or without images): output_directory: Absolute path to the output subdirectory markdown_file: Absolute path to the content.md file images: List of saved image filenames (empty if include_images is False)

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathYes
output_dirNo
include_imagesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
imagesYes
resultYes
markdown_fileYes
output_directoryYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations, the description thoroughly discloses side effects: saving markdown to disk, saving embedded images, rewriting relative links, and requiring output_dir for include_images. It also documents the conditional return structure, giving the agent full visibility into tool behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The docstring is well-organized into a summary, Args, and Returns sections, with no unnecessary filler. Every sentence adds essential information about the tool's behavior and return values.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity and lack of output schema in view, the description provides a complete account of all possible return values and parameter interactions. It covers both modes (inline vs. file output) and the image handling behavior, leaving minimal ambiguity for the agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description fully compensates by explaining file_path as an absolute path, output_dir as an existing directory within allowed bounds, and include_images as a dependent flag. This adds significant meaning beyond the bare schema type and defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description begins with a clear action ('Extract markdown text') and specific source ('PDF or image file'), establishing the tool's purpose. It implicitly distinguishes from sibling tools like extract_markdown_from_url by emphasizing local file paths.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context for when to use the tool (local PDF or image files) and detailed guidance on parameter combinations (output_dir, include_images). However, it does not explicitly mention when to prefer alternative sibling tools like extract_markdown_from_url or extract_markdown_advanced.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

extract_markdown_advancedA

Extract markdown with advanced OCR options.

Args: file_path: Absolute path to the input file (PDF or image) pages: Specific page numbers to process (1-indexed, e.g. [1, 3, 5]) table_format_: Output format for tables ("markdown" or "html") model: OCR model to use (default: "mistral-ocr-latest")

Returns: Extracted markdown content as a string

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNomistral-ocr-latest
pagesNo
file_pathYes
table_format_No

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the full burden. It discloses the return type and key parameters, but does not mention access permissions, rate limits, or side effects. Since extraction is inherently read-only, the description is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a concise one-sentence summary followed by a structured argument list and return statement. No unnecessary text, and the structure is front-loaded and easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers all parameters and the return value, which is sufficient for a tool of this complexity. It lacks guidance on choosing this over simpler alternatives, but is otherwise complete for an extraction task.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description fully compensates by explaining each parameter: file_path (absolute path, PDF/image), pages (1-indexed), table_format_ (markdown/html), model (with default). This adds clear meaning beyond the raw schema types.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Extract markdown with advanced OCR options', providing a specific verb and resource. The 'advanced' qualifier distinguishes it from the basic 'extract_markdown' sibling tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for scenarios requiring advanced OCR (model selection, pagination, table formats) but does not explicitly state when to use this tool over alternatives like 'extract_markdown' or 'extract_markdown_from_url'. No exclusions or alternative comparison is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

extract_markdown_from_urlA

Extract markdown text from a publicly accessible URL.

Processes a PDF or image directly from a URL without uploading a local file first. When output_dir is provided, saves the extracted markdown to content.md inside a named subdirectory. When include_images is also True, saves embedded images alongside the markdown file.

Args: file_url: Publicly accessible URL to a PDF or image output_dir: Absolute path to an existing output directory (must be within allowed dir). When set, saves markdown to disk at <output_dir>/<url_stem>/content.md. include_images: When True (requires output_dir), save images to disk and rewrite markdown with relative image links.

Returns: When output_dir is not set: result: Extracted markdown content When output_dir is set (with or without images): output_directory: Absolute path to the output subdirectory markdown_file: Absolute path to the content.md file images: List of saved image filenames (empty if include_images is False)

ParametersJSON Schema
NameRequiredDescriptionDefault
file_urlYes
output_dirNo
include_imagesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
imagesYes
resultYes
markdown_fileYes
output_directoryYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden and it succeeds: it discloses file-saving side effects (output_dir creates a subdirectory with content.md), conditional image handling, the requirement that output_dir exist and be within an allowed directory, and the return value shapes for each mode. This goes well beyond the minimal schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with a clear one-sentence summary, followed by structured Args and Returns sections. Every sentence conveys necessary information, and the formatting makes conditional behavior easy to parse. It is appropriately sized for a tool with three parameters and two operating modes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has moderate complexity with conditional behavior depending on output_dir and include_images. The description fully covers all modes, return types, and disk interactions, making it complete even without an explicit output schema. The return value explanation is sufficiently detailed and aligns with the listed output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description is the sole source of parameter meaning. It provides rich semantics: file_url must be publicly accessible; output_dir must be an existing absolute path within an allowed dir and determines disk-saving path; include_images requires output_dir and triggers image saving plus markdown link rewriting. This fully compensates for the empty schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Extract markdown text from a publicly accessible URL.' It clearly states it processes a PDF or image directly from a URL and contrasts with a local-file flow by saying 'without uploading a local file first,' distinguishing it from sibling extract_markdown.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies when to use this tool (for URL-based PDFs/images) and explains conditional behaviors for output_dir and include_images. It doesn't explicitly name alternatives or provide 'when not to use' exclusions, but the URL-vs-local context is enough to infer appropriate usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ocr_statusA

Check Mistral API connectivity and key validity.

Makes a lightweight API call to verify the configured API key is working correctly.

Returns: Dictionary with status ("ok" or "error") and message

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the behavioral disclosure burden. It transparently notes that the tool makes a lightweight API call, checks key validity, and returns a status dictionary. It could add more details on error handling or timeouts, but the essentials are covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured: the first sentence states the core purpose, followed by a brief behavioral explanation and a clear 'Returns' section. Every sentence provides useful information without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple status check tool with no input parameters and an output schema present, the description is fully complete. It covers the tool's purpose, what it does, and what it returns, making it sufficient for an agent to select and invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters, so there is nothing to describe beyond the schema. The baseline for zero parameters is 4, and the description appropriately adds no irrelevant parameter details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with a specific verb and resource: 'Check Mistral API connectivity and key validity.' It also clarifies the action ('Makes a lightweight API call to verify the configured API key is working correctly'), which fully distinguishes it from sibling extraction tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool (when needing to verify API connectivity/key validity), but it does not explicitly state usage context, exclusions, or alternatives. The sibling tools are contextually different, but no direct comparison is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.2.4
    • First observedextract_markdown
    • First observedextract_markdown_advanced
    • First observedextract_markdown_from_url
    • First observedocr_status

TDQS

A4.3/5.0

Scored across 4 tools

Disambiguation4/5

extract_markdown and extract_markdown_advanced both process local files with overlapping functionality, but the 'advanced' variant clearly adds page selection and table formatting options; extract_markdown_from_url is distinct by input source, and ocr_status is clearly separate. Mostly distinct with one potential confusion.

Naming Consistency4/5

Three tools follow the verb_noun pattern (extract_markdown, extract_markdown_from_url, extract_markdown_advanced), but ocr_status deviates by using a noun phrase instead of an action verb, creating minor inconsistency.

Tool Count5/5

Four tools cover the core OCR extraction workflows (local file, URL, advanced options) plus a connectivity check, making a well-scoped and appropriate set for this server.

Completeness4/5

The server covers the primary extraction methods and includes advanced options for page selection and table format. Minor gaps exist such as no batch processing or format listing, but core workflows are complete.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers