Mistral OCR MCP Server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Mistral OCR MCP Serverextract markdown from /Users/me/documents/report.pdf"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Mistral OCR MCP Server
A Model Context Protocol (MCP) server that provides tools for extracting text and images from PDF and image files using the Mistral OCR API.
Features
Local & URL Extraction: Extract markdown from local files or remote URLs
Image Handling on Demand: Optionally save embedded images to disk with proper relative links
Advanced OCR: Page selection, table format control, model selection
Health Check: Built-in API status endpoint
Security Sandbox: Restricts file writes to a configured allowed directory
Zero-Install Deployment: Run with
uvxwithout prior installationSupported Formats: PDF (
.pdf), PNG (.png), JPEG (.jpg,.jpeg), WebP (.webp), GIF (.gif)
Related MCP server: MCP-PDF2MD
Client Configuration
Claude Desktop
Add this to your claude_desktop_config.json:
macOS:
~/Library/Application Support/Claude/claude_desktop_config.jsonWindows:
%APPDATA%\Claude\claude_desktop_config.json
{
"mcpServers": {
"mistral-ocr": {
"command": "uvx",
"args": ["mistral-ocr-mcp"],
"env": {
"MISTRAL_API_KEY": "your-api-key-here",
"MISTRAL_OCR_ALLOWED_DIR": "/absolute/path/to/allowed/directory"
}
}
}
}OpenCode
Add this to the mcp section of your configuration file:
{
"mcp": {
"mistral-ocr": {
"type": "local",
"command": ["uvx", "mistral-ocr-mcp"],
"enabled": true,
"environment": {
"MISTRAL_API_KEY": "your-api-key-here",
"MISTRAL_OCR_ALLOWED_DIR": "/absolute/path/to/allowed/directory"
}
}
}
}Codex
If you use the Codex CLI, you can add the server with:
codex mcp add mistral-ocr -- uvx mistral-ocr-mcpMake sure the environment variables MISTRAL_API_KEY and MISTRAL_OCR_ALLOWED_DIR are set in your shell environment.
Configuration
Required Environment Variables
Variable | Description | Example |
| Your Mistral API key (never logged) |
|
| Absolute path to allowed write directory |
|
Security Sandbox
The server enforces a write directory sandbox to prevent unauthorized file writes.
When include_images is True, the output_dir parameter must be within
MISTRAL_OCR_ALLOWED_DIR. Text-only extraction (include_images=False)
is read-only and has no sandbox restrictions.
Validation Examples:
|
| Result |
|
| ✅ Allowed |
|
| ✅ Allowed (exact match) |
|
| ❌ Rejected |
|
| ❌ Rejected (resolves outside) |
Security Notes:
All paths are canonicalized (symlinks resolved,
..eliminated) before validationImage filenames are sanitized to prevent path traversal attacks
Tool Reference
Tool 1: extract_markdown
Extract markdown from a local file, optionally saving embedded images to disk.
Arguments:
Parameter | Type | Required | Default | Description |
|
| Yes | — | Absolute path to input file (PDF or image) |
|
| No |
| Absolute path to output parent directory. Required when |
|
| No |
| When |
Returns (text-only):
{"result": "# Document Title\n\nExtracted markdown content..."}Returns (with images):
{
"output_directory": "/absolute/path/to/output/report",
"markdown_file": "/absolute/path/to/output/report/content.md",
"images": ["img_abc123.png", "img_def456.jpeg"]
}Behavior (with images):
Creates a subdirectory named after the input file stem (e.g.,
reportforreport.pdf)If the subdirectory already exists, appends a timestamp:
report_20260102_143022Saves all extracted images as
<sanitized_id>.<ext>(e.g.,img_abc123.png)Saves markdown to
content.mdwith relative image links (e.g.,)
Output Structure:
/Users/username/workdir/extracted/
quarterly-report/
content.md # Markdown with relative image links
img_abc123.png # First extracted image
img_def456.jpeg # Second extracted imageTool 2: extract_markdown_from_url
Extract markdown from a publicly accessible URL, optionally saving embedded images to disk.
Arguments:
Parameter | Type | Required | Default | Description |
|
| Yes | — | Public URL to a PDF or image |
|
| No |
| Absolute path to output parent directory. Required when |
|
| No |
| When |
Returns (text-only):
{"result": "# Document Title\n\nExtracted markdown content..."}Returns (with images):
{
"output_directory": "/absolute/path/to/output/doc",
"markdown_file": "/absolute/path/to/output/doc/content.md",
"images": ["img_abc123.png"]
}Tool 3: extract_markdown_advanced
Extract markdown with advanced OCR options.
Arguments:
Parameter | Type | Required | Default | Description |
|
| Yes | — | Absolute path to input file (PDF or image) |
|
| No |
| Page numbers to process (1-indexed, e.g. |
|
| No |
| Table output format ( |
|
| No |
| OCR model to use |
Returns:
{"result": "# Document\n\n| Col 1 | Col 2 |\n|-------|-------|\n..."}Tool 4: ocr_status
Check API connectivity and key validity.
Arguments: none
Returns:
{
"status": "ok",
"message": "API key is working"
}Example Client Usage
Here's a minimal Python example using the MCP SDK to call the tools:
import asyncio
from mcp import ClientSession, StdioServerParameters
from mcp.client.stdio import stdio_client
async def extract_document():
server_params = StdioServerParameters(
command="mistral-ocr-mcp",
env={
"MISTRAL_API_KEY": "your-api-key",
"MISTRAL_OCR_ALLOWED_DIR": "/Users/username/workdir"
}
)
async with stdio_client(server_params) as (read, write):
async with ClientSession(read, write) as session:
await session.initialize()
# Text-only extraction
result = await session.call_tool(
"extract_markdown",
arguments={"file_path": "/path/to/document.pdf"}
)
print(result.content[0].text)
# Extraction with images
result = await session.call_tool(
"extract_markdown",
arguments={
"file_path": "/path/to/document.pdf",
"include_images": True,
"output_dir": "/Users/username/workdir/output"
}
)
print(result.content[0].text)
# Extract from URL
result = await session.call_tool(
"extract_markdown_from_url",
arguments={"file_url": "https://example.com/doc.pdf"}
)
print(result.content[0].text)
# Check API status
result = await session.call_tool("ocr_status", arguments={})
print(result.content[0].text)
asyncio.run(extract_document())Troubleshooting
Error | Cause | Solution |
|
| Set the environment variable before running the server |
|
| Set the environment variable to an absolute path |
| Relative path provided (e.g., | Use an absolute path (e.g., |
| Directory does not exist on filesystem | Create the directory first: |
| Path points to a file, not a directory | Ensure the path is a directory |
| Relative path provided for input file | Use an absolute path (e.g., |
| Input file does not exist | Check the file path and ensure the file exists |
| File extension not supported | Use |
| Output directory does not exist | Create the directory first: |
| Path points to a file, not a directory | Ensure the path is a directory |
| Output directory exists but is not writable | Check directory permissions: |
|
| Use a path within the allowed directory |
| Invalid API key | Check your |
| Rate limit exceeded | Wait and retry, or check your API quota |
Development
Setup
Clone the repository and install with development dependencies:
git clone https://github.com/ORDIS-Co-Ltd/mistral-ocr-mcp
cd mistral-ocr-mcp
pip install -e '.[dev]'Run the server locally:
MISTRAL_API_KEY="your-key" \
MISTRAL_OCR_ALLOWED_DIR="/path/to/allowed/dir" \
python -m mistral_ocr_mcpRun Tests
pytestProject Structure
mistral-ocr-mcp/
├── src/
│ └── mistral_ocr_mcp/
│ ├── __init__.py
│ ├── __main__.py # Entry point
│ ├── server.py # MCP server and tool definitions
│ ├── config.py # Configuration loading and validation
│ ├── extraction.py # OCR orchestration logic
│ ├── mistral_client.py # Mistral API client
│ ├── images.py # Image parsing and saving
│ ├── markdown_rewrite.py # Markdown link rewriting
│ └── path_sandbox.py # Path validation and sandbox enforcement
├── tests/ # Unit tests
├── pyproject.toml # Package configuration
└── README.md # This fileLicense
MIT
Contributing
Contributions are welcome! Please open an issue or submit a pull request.
Links
GitHub Repository: https://github.com/ORDIS-Co-Ltd/mistral-ocr-mcp
MCP Specification: https://modelcontextprotocol.io
Mistral AI: https://mistral.ai
Available Tools
4 toolsextract_markdownA
Extract markdown text from a PDF or image file.
When output_dir is provided, saves the extracted markdown to
content.md inside a named subdirectory. When include_images
is also True, saves embedded images alongside the markdown file.
Otherwise returns the markdown text inline.
Args:
file_path: Absolute path to the input file (PDF or image)
output_dir: Absolute path to an existing output directory (must be
within allowed dir). When set, saves markdown to disk at
<output_dir>/<file_stem>/content.md.
include_images: When True (requires output_dir), save images to
disk and rewrite markdown with relative image links.
Returns: When output_dir is not set: result: Extracted markdown content When output_dir is set (with or without images): output_directory: Absolute path to the output subdirectory markdown_file: Absolute path to the content.md file images: List of saved image filenames (empty if include_images is False)
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | ||
| output_dir | No | ||
| include_images | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| images | Yes | |
| result | Yes | |
| markdown_file | Yes | |
| output_directory | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without annotations, the description thoroughly discloses side effects: saving markdown to disk, saving embedded images, rewriting relative links, and requiring output_dir for include_images. It also documents the conditional return structure, giving the agent full visibility into tool behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The docstring is well-organized into a summary, Args, and Returns sections, with no unnecessary filler. Every sentence adds essential information about the tool's behavior and return values.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity and lack of output schema in view, the description provides a complete account of all possible return values and parameter interactions. It covers both modes (inline vs. file output) and the image handling behavior, leaving minimal ambiguity for the agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description fully compensates by explaining file_path as an absolute path, output_dir as an existing directory within allowed bounds, and include_images as a dependent flag. This adds significant meaning beyond the bare schema type and defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description begins with a clear action ('Extract markdown text') and specific source ('PDF or image file'), establishing the tool's purpose. It implicitly distinguishes from sibling tools like extract_markdown_from_url by emphasizing local file paths.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context for when to use the tool (local PDF or image files) and detailed guidance on parameter combinations (output_dir, include_images). However, it does not explicitly mention when to prefer alternative sibling tools like extract_markdown_from_url or extract_markdown_advanced.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract_markdown_advancedA
Extract markdown with advanced OCR options.
Args: file_path: Absolute path to the input file (PDF or image) pages: Specific page numbers to process (1-indexed, e.g. [1, 3, 5]) table_format_: Output format for tables ("markdown" or "html") model: OCR model to use (default: "mistral-ocr-latest")
Returns: Extracted markdown content as a string
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | mistral-ocr-latest | |
| pages | No | ||
| file_path | Yes | ||
| table_format_ | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the full burden. It discloses the return type and key parameters, but does not mention access permissions, rate limits, or side effects. Since extraction is inherently read-only, the description is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a concise one-sentence summary followed by a structured argument list and return statement. No unnecessary text, and the structure is front-loaded and easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers all parameters and the return value, which is sufficient for a tool of this complexity. It lacks guidance on choosing this over simpler alternatives, but is otherwise complete for an extraction task.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description fully compensates by explaining each parameter: file_path (absolute path, PDF/image), pages (1-indexed), table_format_ (markdown/html), model (with default). This adds clear meaning beyond the raw schema types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Extract markdown with advanced OCR options', providing a specific verb and resource. The 'advanced' qualifier distinguishes it from the basic 'extract_markdown' sibling tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for scenarios requiring advanced OCR (model selection, pagination, table formats) but does not explicitly state when to use this tool over alternatives like 'extract_markdown' or 'extract_markdown_from_url'. No exclusions or alternative comparison is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract_markdown_from_urlA
Extract markdown text from a publicly accessible URL.
Processes a PDF or image directly from a URL without uploading
a local file first. When output_dir is provided, saves the
extracted markdown to content.md inside a named subdirectory.
When include_images is also True, saves embedded images
alongside the markdown file.
Args:
file_url: Publicly accessible URL to a PDF or image
output_dir: Absolute path to an existing output directory (must be
within allowed dir). When set, saves markdown to disk at
<output_dir>/<url_stem>/content.md.
include_images: When True (requires output_dir), save images to
disk and rewrite markdown with relative image links.
Returns: When output_dir is not set: result: Extracted markdown content When output_dir is set (with or without images): output_directory: Absolute path to the output subdirectory markdown_file: Absolute path to the content.md file images: List of saved image filenames (empty if include_images is False)
| Name | Required | Description | Default |
|---|---|---|---|
| file_url | Yes | ||
| output_dir | No | ||
| include_images | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| images | Yes | |
| result | Yes | |
| markdown_file | Yes | |
| output_directory | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden and it succeeds: it discloses file-saving side effects (output_dir creates a subdirectory with content.md), conditional image handling, the requirement that output_dir exist and be within an allowed directory, and the return value shapes for each mode. This goes well beyond the minimal schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with a clear one-sentence summary, followed by structured Args and Returns sections. Every sentence conveys necessary information, and the formatting makes conditional behavior easy to parse. It is appropriately sized for a tool with three parameters and two operating modes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has moderate complexity with conditional behavior depending on output_dir and include_images. The description fully covers all modes, return types, and disk interactions, making it complete even without an explicit output schema. The return value explanation is sufficiently detailed and aligns with the listed output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description is the sole source of parameter meaning. It provides rich semantics: file_url must be publicly accessible; output_dir must be an existing absolute path within an allowed dir and determines disk-saving path; include_images requires output_dir and triggers image saving plus markdown link rewriting. This fully compensates for the empty schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Extract markdown text from a publicly accessible URL.' It clearly states it processes a PDF or image directly from a URL and contrasts with a local-file flow by saying 'without uploading a local file first,' distinguishing it from sibling extract_markdown.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use this tool (for URL-based PDFs/images) and explains conditional behaviors for output_dir and include_images. It doesn't explicitly name alternatives or provide 'when not to use' exclusions, but the URL-vs-local context is enough to infer appropriate usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ocr_statusA
Check Mistral API connectivity and key validity.
Makes a lightweight API call to verify the configured API key is working correctly.
Returns: Dictionary with status ("ok" or "error") and message
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the behavioral disclosure burden. It transparently notes that the tool makes a lightweight API call, checks key validity, and returns a status dictionary. It could add more details on error handling or timeouts, but the essentials are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured: the first sentence states the core purpose, followed by a brief behavioral explanation and a clear 'Returns' section. Every sentence provides useful information without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple status check tool with no input parameters and an output schema present, the description is fully complete. It covers the tool's purpose, what it does, and what it returns, making it sufficient for an agent to select and invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, so there is nothing to describe beyond the schema. The baseline for zero parameters is 4, and the description appropriately adds no irrelevant parameter details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb and resource: 'Check Mistral API connectivity and key validity.' It also clarifies the action ('Makes a lightweight API call to verify the configured API key is working correctly'), which fully distinguishes it from sibling extraction tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool (when needing to verify API connectivity/key validity), but it does not explicitly state usage context, exclusions, or alternatives. The sibling tools are contextually different, but no direct comparison is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.2.4- First observed
extract_markdown - First observed
extract_markdown_advanced - First observed
extract_markdown_from_url - First observed
ocr_status
TDQS
Scored across 4 tools
extract_markdown and extract_markdown_advanced both process local files with overlapping functionality, but the 'advanced' variant clearly adds page selection and table formatting options; extract_markdown_from_url is distinct by input source, and ocr_status is clearly separate. Mostly distinct with one potential confusion.
Three tools follow the verb_noun pattern (extract_markdown, extract_markdown_from_url, extract_markdown_advanced), but ocr_status deviates by using a noun phrase instead of an action verb, creating minor inconsistency.
Four tools cover the core OCR extraction workflows (local file, URL, advanced options) plus a connectivity check, making a well-scoped and appropriate set for this server.
The server covers the primary extraction methods and includes advanced options for page selection and table format. Minor gaps exist such as no batch processing or format listing, but core workflows are complete.
Maintenance
Related MCP Connectors
OCR, transcription, file extraction, and image generation for AI agents via MCP.
Parse PDF/Word/PPT/HTML to Markdown; tables as JSON, image extraction, RAG chunking, page ranges.
OCR and document understanding: extract text from images, then summarize or translate it.
Read PDFs and images as markdown or text, with exact costs and hard spend caps. $0.75/1k pages.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceOCR images or pdfs, locally or by URLs by using Mistral OCR API (paid)39MIT
- AlicenseAqualityDmaintenanceConverts PDF files from local storage or URLs to structured Markdown format using Mistral AI's OCR API, preserving document structure and extracting images.21MIT
- FlicenseAqualityDmaintenanceExtracts text content from PDFs and images using Mistral's OCR API, enabling OCR capabilities in MCP-compatible clients like Cursor and Claude Desktop.18-
- AlicenseNot gradedqualityAmaintenanceConverts documents and images to Markdown using Mistral AI's OCR, enabling AI-powered document processing via MCP-compatible clients like Claude Desktop.45 npm2MIT