Mistral MCP — Document Extraction
This server lets you use Mistral AI to chat, code, analyze images, OCR documents, transcribe audio, extract structured document data, and run Mistral workflows.
Chat: Generate completions with Mistral models, including structured JSON/schema output and token usage.
Code: Use Codestral fill-in-the-middle completions for editor autocomplete or refactoring.
Vision: Chat with images or document URLs using vision-capable Mistral models.
OCR: Convert PDFs/images to markdown with tables, annotations, confidence scores, headers/footers, and bounding boxes.
Audio: Transcribe speech with Voxtral, with optional diarization, timestamps, context bias, and language hints.
Document extraction: With the core profile,
process_documentturns text/Markdown or OCR into schema-validated JSON for invoices, contracts, ID documents, or generic text, with classification and cache controls.Workflows: Execute, monitor, and signal/query/update running Mistral workflow executions.
Profiles: Expose different tool sets (
core,metier-docs,workflows,admin,self-hosted), including admin tools for Files, Batch, Conversations, and Libraries.
Enables durable workflow execution with Temporal, supporting starting, querying, and signaling long-running Mistral Workflows, including human-in-the-loop checkpoints.
Mistral MCP server for document extraction
Turn text or Markdown invoices into typed JSON through MCP. mistral-mcp
uses Mistral chat to extract vendors, totals, line items and due dates, then
validates the response schema. Optional Mistral OCR handles PDF and image inputs.
process_document also supports contracts, identity documents and automatic
classification. Six tools are available by default, including chat, vision,
transcription and code completion.
Français · Migration guide · Examples · Deployment
npm package · 1.0.0 release notes · GitHub releases · Mistral API docs
Version 1.0.0 — breaking changes from 0.11.0: core now exposes six tools;
existing orchestration clients must choose an explicit profile. Document results
require extraction_source, and ocr_confidence / page_count can be null.
Migration and rollback instructions.
Install in an MCP client
Requires Node.js 20+, npm and a Mistral API key with access and quota for the
requested model. For clients using mcpServers JSON, configure the stdio server:
{
"mcpServers": {
"mistral": {
"command": "npx",
"args": ["-y", "mistral-mcp@1.0.0"],
"env": {
"MISTRAL_API_KEY": "your_key_here",
"MISTRAL_DEFAULT_MODEL": "ministral-3b-latest",
"MISTRAL_MCP_PROFILE": "core"
}
}
}
}This runs npx -y mistral-mcp@1.0.0. Restart the client and refresh its tool
catalog. The server reads the environment supplied by the client; it does not
load .env automatically. Use your client's secret configuration for the key.
ministral-3b-latest was verified on the test account; model access and free
quota depend on your account. Check your limits
before making calls. Local MCP hosting still sends extraction requests to Mistral.
Related MCP server: MCP Server TypeScript
Quick start: an existing text or Markdown invoice
The invoice script and fixtures are source examples, not included in the npm package. Check out the release tag and build from the repository root:
git clone https://github.com/Swih/mistral-mcp.git
cd mistral-mcp
git checkout v1.0.0
npm ci
npm run buildSet the key and chat model in your environment or in a local .env file:
MISTRAL_API_KEY=your_key_here
MISTRAL_DEFAULT_MODEL=ministral-3b-latestThe example loads .env with dotenv. Keep the key out of version control.
Choose a chat model with quota on your account; the default may have zero quota.
npm run example:invoice -- test/fixtures/invoice-text.md --output invoice-result.jsonThis uses the synthetic Markdown invoice.
For your own existing UTF-8 .txt or .md invoice, the command syntax is:
node examples/invoice.mjs <local-file.txt|local-file.md> [--output result.json]The script reads the text locally and calls process_document with
source: { type: "text", text: "..." }, kind: "invoice" and
options.cache: "bypass" through the local server's core profile. The text must
contain non-whitespace content and fit within 60,000 UTF-16 code units (JavaScript
string length). Markdown and whitespace are preserved unchanged. There is no
Files upload or OCR call; invoice extraction sends the text to Mistral chat.
The example uses Mistral Cloud only and still requires a key and chat quota;
processing is not entirely local or guaranteed free. Check your account's
limits.
Without --output, it prints validated JSON; with it, it writes to a new file and
prints that path. Existing files are not overwritten. The output path is reserved
before API calls and may remain empty after failure; remove it or choose a new
path before retrying.
On 2026-09-28, the live text test and the CLI
example succeeded with ministral-3b-latest, using chat only. The verified fields
were vendor ACME SAS, total 12960 EUR, due date 2026-09-11 and the quantities,
unit prices and amounts of all three invoice lines. This verifies one synthetic
invoice, not a general accuracy or reliability score.
For a PDF or image, the existing OCR route remains available:
npm run example:invoice -- test/fixtures/corpus/invoice-fr-table.pdf --output invoice-ocr-result.jsonPDF, PNG, JPEG and WebP files up to 20 MiB require Files, OCR and chat access and
quota. The script uploads the file, calls process_document and attempts to
delete the upload in finally, including after extraction failure. There is no
separate OCR readiness probe. Upload and cleanup use the Files API without
exposing admin tools in core. Known limitation: the test account's HTTP 429
/ zero OCR quota blocked live OCR validation. The successful text run does not
validate OCR extraction.
Compare any extracted result against its source. The synthetic PDF and fixture ground truth describe expected document content, not captured live output. Schema validation checks the shape and types of the response; it does not verify factual accuracy, invoice arithmetic, tax treatment or accounting correctness. Review extracted fields against the source before using them.
Profiles
MISTRAL_MCP_PROFILE selects one of five profiles. The default is core for
Mistral Cloud; a custom MISTRAL_BASE_URL infers self-hosted unless you set a
profile explicitly.
Profile | Tools | Scope in |
| 6 | Documents, chat, vision, transcription and code completion |
| 17 | Preserved legacy profile: the six core tools plus all 11 orchestration tools; a superset of the old 16-tool core |
| 11 | Workflows, connectors and search-index discovery |
| 46 | All tools implemented by this server, including Files, Batch, Conversations and Libraries |
| 5 | Chat, streaming chat, embeddings, function calling and vision on a compatible endpoint |
full remains a deprecated alias of admin, not a sixth profile. Set
MISTRAL_MCP_PROFILE=metier-docs to preserve the old core tool set after upgrading;
choose workflows for orchestration alone or admin for the complete tool set.
Restart the server and refresh tool discovery after changing profiles.
npx -y mistral-mcp@1.0.0 --doctor reports the local profile and tool list without API
calls. The mistral://capabilities resource reports the active endpoint, tool
families and reasons for omitted tools. Neither proves account access or quota.
Core tools and document behavior
Tool | Purpose |
| Supplied text/Markdown or OCR, optional classification and schema-validated extraction for invoices, contracts, identity documents or generic text |
| Raw OCR text, tables, annotations and optional blocks from PDFs or images |
| Chat with images supplied by URL or base64 |
| Chat completion, including structured response formats |
| Audio transcription with optional speaker diarization |
| Fill-in-the-middle code completion |
process_document accepts source: { type: "text", text: string }, a document
URL, a base64 image or an uploaded file ID. Text may contain Markdown and is
preserved unchanged; blank or whitespace-only strings and strings above 60,000
UTF-16 code units are rejected for every kind, including generic.
kind is auto (default), invoice, contract, id_document or generic.
Successful calls return readable content and JSON structuredContent; failures
return isError: true.
For example, these are tool arguments using synthetic input, not a live result:
{
"source": {
"type": "text",
"text": "# Invoice DEMO-001\nVendor: Example Studio\nService: 2 hours at EUR 50\nTotal due: EUR 100\nDue date: 2026-10-15\n"
},
"kind": "invoice",
"options": { "cache": "bypass" }
}On a cache miss or with cache: "bypass", auto calls chat to classify even a
text source; invoice, contract and id_document use chat for typed extraction.
Explicit kind: "generic" with a text source makes no API calls and returns
the supplied text as both ocr_text and structured_text.
Every successful result includes these fields:
Field | Provided text / Markdown | Successful OCR source |
|
|
|
| Original input text, unchanged | OCR text |
|
| Number from 0 to 1 |
|
| Number of processed pages |
options.maxPagesandoptions.minOcrConfidenceapply only to OCR sources. They do not paginate or score supplied text. For OCR, missing, incomplete or invalid confidence scores cause an error, as do scores below the requested minimum. The default0.3is unmeasured; OCR confidence does not establish extraction accuracy.For OCR sources,
options.maxPagesdefaults to 50 (maximum 200). Typed extraction rejects OCR text above 60,000 UTF-16 code units: split the document or usegenericfor OCR text. The text-source input limit still applies togeneric.options.languageHintsguides typed extraction, not the OCR model.options.cache: "bypass"skips cache reads and writes. Other modes areread_onlyandread_write. Identity documents bypass the cache by default, including afterautoclassification; explicitread_writeopts them in.Cache files contain extracted content.
MISTRAL_MCP_CACHE_DIRsets the location;MISTRAL_MCP_CACHE_TTL_HOURSdefaults to 168 hours (0disables reuse and new writes). Cleanup is opportunistic during cache operations. Bypass does not erase older entries, and expiration does not guarantee deletion at a set time. Pipeline versionv1.0.0-text.1invalidates reuse of older cache entries; it does not guarantee their immediate deletion.
The synthetic corpus separates required OCR text
from expected extracted invoice fields. npm run eval:docs evaluates these
separately through real API calls. Fixture truth is not a live accuracy result;
text-based synthetic PDFs do not establish accuracy on degraded scans.
Development and evaluation guidance.
API and deployment references
mistral://capabilities describes the active tool set. mistral://models reads the
upstream catalog and reports fallback if the API call fails. mistral://voices
is available in admin; mistral://workflows is available in metier-docs,
workflows and admin. Catalog presence does not establish access or quota.
You can host the MCP process and configure its upstream endpoint, credentials, tool exposure and cache policy. By default, requests go to Mistral Cloud: local MCP hosting does not make document inference local. These controls alone do not establish data residency or regulatory compliance.
A custom MISTRAL_BASE_URL infers self-hosted: chat, streaming chat, embeddings,
function calling and vision, subject to endpoint/model support. It does not
include OCR or process_document. An explicit profile overrides inference but
does not add missing APIs to a backend.
Reference | Contents |
Removed core tools, explicit profiles, pinned | |
Local invoices, transcription and library-backed conversations | |
Tool families and MCP tool input schemas | Complete tool membership and argument reference |
Meeting minutes, email replies, commits, legal summaries, invoice reminders and code review | |
Docker, Compose, Kubernetes, custom endpoints, cache and HTTP settings | |
HTTPS deployment; public connector calls are not established as end-to-end validated here | |
Optional plugin with 11 skills, pinned to | |
Build, tests, evaluation and release checks | |
Changes and security reporting |
stdio is the default transport. --http or MCP_TRANSPORT=http enables
Streamable HTTP at 127.0.0.1:3333/mcp by default, with configurable bearer
authentication and allowed origins. Integrated OAuth is not provided.
Tool audit records go to stderr and omit arguments and result payloads;
MISTRAL_MCP_AUDIT=off disables them.
The protocol-era tests cover MCP 2026-07-28
and the 2025 handshake using the same registrations. npm run check:release
checks the build and local tests, including the installed package against an API
stub. Live API validation is separate; skipped tests do not count as success.
Package pinning does not guarantee future upstream availability or compatibility.
MIT license — Copyright Dayan Decamp.
Available Tools
6 toolscodestral_fimCodestral fill-in-the-middle completionARead-only
Fill-in-the-middle code completion with Codestral.
Given prompt (code preceding the cursor) and suffix (code after the cursor),
Codestral writes the middle. Use for editor autocomplete scenarios, code-patching
agents, or structured refactors where you know the target boundaries.
Default stop tokens: [] — let the model decide. Override with stop if needed.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Random seed for deterministic sampling. Maps to Mistral's `random_seed`. | |
| stop | No | Stop generation when any of these text sequences is encountered. | |
| model | No | FIM (fill-in-the-middle) model identifier. Any identifier your endpoint serves is accepted — read mistral://models for the live catalog. Known Mistral aliases: codestral-latest. | |
| top_p | No | Nucleus sampling probability mass: 0.1 considers tokens in the top 10% of probability mass. Prefer adjusting this or temperature, not both. | |
| prompt | Yes | Code preceding the cursor. | |
| suffix | Yes | Code after the cursor. Can be empty string. | |
| max_tokens | No | Maximum number of tokens to generate. Input tokens plus this limit must fit within the model's context length. | |
| temperature | No | Sampling temperature: higher values make output more random; lower values make it more focused. Prefer adjusting this or top_p, not both. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | Yes | |
| model | Yes | |
| usage | No | |
| finish_reason | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=false, and destructiveHint=false. The description adds useful generation behavior by noting the default stop tokens are empty and that the model decides when to stop unless `stop` is overridden.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with a one-line purpose, followed by cursor semantics, use cases, and stop-token behavior. Every sentence earns its place and no filler is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given rich annotations, 100% schema coverage, and an output schema, the description supplies the missing conceptual model: prompt before cursor, suffix after cursor, and Codestral writes the middle. It is complete enough for correct invocation without redundant return-value explanation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so prompt/suffix meanings are already documented. The description adds extra semantics for `stop`: default is [] and the model decides, with explicit override guidance, going beyond the schema's brief stop description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: fill-in-the-middle code completion with Codestral, plus the prompt/suffix cursor model. The FIM and editor-autocomplete framing distinguishes it from sibling chat, vision, OCR, transcription, and document tools without needing schema inspection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear usage contexts: editor autocomplete, code-patching agents, and structured refactors where target boundaries are known. It does not explicitly name when to avoid this tool or route to a sibling like mistral_chat for general generation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mistral_chatMistral chat completionARead-only
Generate a chat completion using a Mistral model.
When to use:
Drafting French (or any European-language) content where Mistral shines.
Codestral for code-specific generation/review.
Ministral for cheap / low-latency classification.
Returns structured content with the assistant text and token usage. Does NOT stream — use mistral_chat_stream for long outputs with progress updates.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Random seed for deterministic sampling. Maps to Mistral's `random_seed`. | |
| model | No | Chat model. Default: mistral-medium-latest (override with MISTRAL_DEFAULT_MODEL). | |
| top_p | No | Nucleus sampling probability mass: 0.1 considers tokens in the top 10% of probability mass. Prefer adjusting this or temperature, not both. | |
| messages | Yes | Chat messages in role/content form. | |
| max_tokens | No | Maximum number of tokens to generate. Input tokens plus this limit must fit within the model's context length. | |
| temperature | No | Sampling temperature: higher values make output more random; lower values make it more focused. Prefer adjusting this or top_p, not both. | |
| response_format | No | Force a structured output: `{type:"json_object"}` for JSON mode, `{type:"json_schema", json_schema:{...}}` for strict schema mode. | |
| reasoning_effort | No | Reasoning effort; supported values depend on the selected model. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | Yes | |
| model | Yes | |
| usage | No | |
| finish_reason | No | |
| reasoning_content | No | Reasoning trace returned by Magistral models. Absent for non-reasoning models. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover safety (readOnlyHint=true, destructiveHint=false, openWorldHint=true), so the bar is lower. The description still adds useful non-safety behavior: it does not stream, and it returns structured content with assistant text and token usage. It does not mention rate limits, latency, or error behavior, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action, organizes guidance as scannable bullets, and closes with the return shape and the streaming exclusion. No filler sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description needn't detail return values, and annotations cover safety. It still supplies the streaming caveat and model-family guidance. Minor gap: it never states default model behavior or auth expectations, but that is largely handled by the schema's model description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents seed, model, top_p, temperature, response_format, and reasoning_effort thoroughly. The description adds only indirect model-selection guidance ('Codestral', 'Ministral') rather than explaining parameters, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening line states a specific verb and resource: 'Generate a chat completion using a Mistral model.' The bullets further differentiate sub-cases (French/European drafting, Codestral for code, Ministral for cheap classification), letting an agent distinguish this from codestral_fim and other siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit 'When to use' bullets plus a named exclusion: 'Does NOT stream — use mistral_chat_stream for long outputs with progress updates.' This gives both positive triggers and a concrete alternative with the condition that selects it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mistral_ocrMistral OCR (document to markdown)ARead-onlyIdempotent
Run Mistral OCR on a PDF or image, returning structured markdown per page.
Input document is one of:
{ type: "document_url", documentUrl: "https://...pdf" }
{ type: "image_url", imageUrl: "https://..." | "data:image/..." }
{ type: "file", fileId: "" }
Options:
pages: array of 0-indexed page numbers or string like "0-5,7".tableFormat: 'markdown' (default) or 'html'.extractHeader/extractFooter: include page header/footer when present.includeImageBase64: embed extracted image bytes as base64 in the response.document_annotation_format: JSON schema for whole-document structured extraction.bbox_annotation_format: JSON schema for extracted image / bbox annotations.confidence_scores_granularity: 'page', 'word', or 'block'. 'block' adds per-block content/type confidence underpages[].blocks[].confidence_scoresand requires OCR 4.1 or newer.includeBlocks: return paragraph-level blocks (bounding box + type) in reading order — titles, lists, tables, images, equations, captions, code, references, aside text, header, footer, signature. Requires OCR 4 (mistral-ocr-4-0) or newer; older models accept the flag but return an emptyblocksarray.
Returns pages[].markdown plus optional pages[].hyperlinks, header, footer,
images bounding boxes, blocks, annotations, confidence scores, and dimensions.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | OCR model. Default: mistral-ocr-latest. | |
| pages | No | Zero-based page numbers to process, as an array or comma-separated numbers and ranges such as "0-5,7". | |
| document | Yes | Document to process, supplied as a document URL, an image URL or data URI, or an uploaded file ID. | |
| imageLimit | No | Maximum number of images to extract from the document. | |
| tableFormat | No | Format for extracted tables: Markdown or HTML. | |
| imageMinSize | No | Minimum height and width of an image to extract. | |
| extractFooter | No | Extract each page footer into its footer field and remove it from the page markdown. | |
| extractHeader | No | Extract each page header into its header field and remove it from the page markdown. | |
| includeBlocks | No | Return paragraph-level blocks (bounding box + type) per page. Requires OCR 4 (mistral-ocr-4-0) or newer. | |
| includeImageBase64 | No | Include base64-encoded data for extracted images in the response. | |
| bbox_annotation_format | No | JSON Schema for structured annotations of each extracted bounding box or image. | |
| document_annotation_format | No | JSON Schema for a structured annotation extracted from the entire document. | |
| document_annotation_prompt | No | Instructions for whole-document structured extraction. Requires document_annotation_format. | |
| confidence_scores_granularity | No | Confidence granularity. 'block' also fills pages[].blocks[].confidence_scores and requires OCR 4.1 (mistral-ocr-4-1) or newer. |
Output Schema
| Name | Required | Description |
|---|---|---|
| model | Yes | |
| pages | Yes | |
| usage | No | |
| annotations | No | |
| pages_count | Yes | |
| document_annotation | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (which only declare read-only, idempotent, non-destructive, open-world), the description discloses real behavioral preconditions: includeBlocks requires OCR 4+, confidence_scores_granularity='block' requires OCR 4.1+, and older models silently accept includeBlocks but return an empty blocks array. That silent-failure warning is exactly the kind of context an agent cannot get from the schema or annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with a one-line purpose, then cleanly sectioned into input shapes, options, and returns for a 14-parameter tool. The length is proportionate to the surface area and each bullet maps to a decision the caller must make.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only OCR tool with a full output schema and 100% schema coverage, the description supplies everything else needed: the three mutually exclusive document input forms, option semantics, model-version gates, and the shape of the response. Nothing an agent needs to invoke it correctly is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds genuine value on several parameters the schema documents only tersely: the annotated string form of `pages` ('0-5,7'), the default of tableFormat, the distinction between document_annotation_format and bbox_annotation_format, and the version requirements tied to includeBlocks and confidence granularity. Some parameters (imageLimit, imageMinSize, model) are untouched, which keeps it from a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('Run Mistral OCR on a PDF or image') and states the core output ('returning structured markdown per page'). It is clear, but it never distinguishes itself from siblings like mistral_vision or process_document, so an agent cannot tell from the description alone which of those overlapping tools to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the enumerated input shapes and option list, but there is no explicit when-to-use or when-not-to-use guidance, and no routing to alternative siblings such as mistral_vision for image-only OCR. The reader must infer the fit from the parameter list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mistral_visionMistral multimodal chat (vision)ARead-only
Chat completion with multimodal input: text + image_url parts.
Requires a vision-capable model. Accepted:
pixtral-large-latest
pixtral-12b-latest
mistral-large-latest
mistral-medium-latest
mistral-small-latest
Each message's content is either a plain string (pure text) or an array of
parts { type: 'text', text } / { type: 'image_url', imageUrl }. The image URL
can be an https URL or a data: URI base64 payload.
Returns the assistant text + token usage. For non-visual requests, prefer mistral_chat.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Random seed for deterministic sampling. Maps to Mistral's `random_seed`. | |
| model | No | Vision-capable Mistral model. Default: pixtral-large-latest. | |
| top_p | No | Nucleus sampling probability mass: 0.1 considers tokens in the top 10% of probability mass. Prefer adjusting this or temperature, not both. | |
| messages | Yes | Chat messages. Pure-text requests are accepted, but this tool is intended primarily for multimodal prompts containing image parts. | |
| max_tokens | No | Maximum number of tokens to generate. Input tokens plus this limit must fit within the model's context length. | |
| temperature | No | Sampling temperature: higher values make output more random; lower values make it more focused. Prefer adjusting this or top_p, not both. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | Yes | |
| model | Yes | |
| usage | No | |
| finish_reason | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered. The description adds useful non-structural context: the vision-model requirement, accepted model IDs, and that it returns assistant text plus token usage. It does not cover cost/latency or image size limits, so not a full 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then model list, then payload format, then routing guidance. The model enumeration is slightly verbose but each line is actionable and the structure is scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists so return values needn't be explained, and the description still notes the assistant-text + token-usage return. Model constraints, payload shapes, and sibling routing are all present; nothing needed to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description goes beyond by enumerating the accepted vision model IDs (the schema only says 'Vision-capable Mistral model') and clarifying the content part shapes and image URL formats.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Chat completion with multimodal input: text + image_url parts') and explicitly differentiates from the sibling mistral_chat for non-visual requests. An agent can distinguish this from mistral_ocr and mistral_chat without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes usage: 'Requires a vision-capable model' with an enumerated accepted-model list, and states 'For non-visual requests, prefer mistral_chat.' This names both the condition to use it and the alternative to use otherwise.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
process_documentProcess a business document end-to-endARead-onlyIdempotent
Single-call pipeline: provided text/Markdown or Mistral OCR → classify (if kind=auto) → typed extraction → validation. source.type=text skips OCR and Files uploads. Text with kind=generic makes no API calls; classification and typed extraction use Mistral chat. Results expose extraction_source. Provided text has null ocr_confidence and page_count; ocr_text contains the supplied text unchanged. Replaces the manual chain of mistral_ocr + mistral_chat + JSON parsing.
Kinds: contract | invoice | id_document | generic. Use kind=auto to let the server classify.
Returns a discriminated union — switch on kind to access typed fields.
Validation checks schema and, for OCR sources, OCR confidence; not factual or accounting accuracy.
Typed extraction rejects text longer than 60000 characters rather than truncating it.
Cache keys include source, kind, page limit, endpoint, models and pipeline version. Override location with MISTRAL_MCP_CACHE_DIR. Override mode with options.cache. Default cache mode is 'read_write' EXCEPT for kind=id_document (auto-bypass to avoid persisting PII). Set options.cache='read_write' explicitly to opt in for id documents.
options.maxPages and options.minOcrConfidence apply only to OCR sources. The confidence floor defaults to 0.3. Below the floor the
tool returns isError. Missing or partial confidence scores also return isError;
use mistral_ocr directly if you need raw OCR without a confidence guarantee.
0.3 is a conservative starting point, not a measured one: calibrate it for your
corpus with npm run eval:docs.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | Extraction task. auto classifies the document; generic returns text without typed extraction. Text source with generic makes no API calls. | auto |
| source | Yes | Already extracted text/Markdown, or an OCR source: remote URL, uploaded file ID, or inline image. | |
| options | No | Page selection, OCR confidence floor and local cache policy. |
Output Schema
| Name | Required | Description |
|---|---|---|
| dob | No | |
| kind | Yes | |
| name | No | |
| total | No | |
| expiry | No | |
| vendor | No | |
| clauses | No | |
| country | No | |
| parties | No | |
| summary | No | |
| currency | No | |
| due_date | No | |
| ocr_text | Yes | Text used for extraction: provided text unchanged, or Markdown returned by Mistral OCR. |
| anomalies | No | |
| cache_hit | Yes | |
| key_dates | No | |
| source_id | Yes | |
| line_items | No | |
| page_count | Yes | Pages processed by Mistral OCR. Null for provided text, whose pagination is unknown. |
| risk_score | No | |
| document_type | No | |
| ocr_confidence | Yes | Mean Mistral OCR page confidence. Null for provided text; never an extraction accuracy score. |
| structured_text | No | |
| pipeline_version | Yes | |
| extraction_source | Yes | How the input text was obtained. provided_text is supplied by the caller, not verified by OCR. |
| total_duration_ms | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, but the description goes well beyond them: cache key composition and the id_document PII auto-bypass, the isError conditions around OCR confidence, that validation is schema-only (not factual), and the 60000-character rejection behavior. This is unusually rich behavioral disclosure for a read-only processing tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the pipeline summary, then grouped by concerns (sources, kinds, returns, validation, cache, options). Length is defensible for a 3-param nested pipeline tool, but a few statements restate schema facts (the 60000 limit, the 0.3 floor, maxPages/minOcrConfidence scoping), which is mild redundancy given a fully described schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description still usefully explains the discriminated-union return (switch on kind, extraction_source, ocr_confidence semantics). Combined with the caveats about confidence floors and cache bypass, an agent has everything needed to call and interpret the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds framing the schema does not: that maxPages/minOcrConfidence apply only to OCR sources (and how they interact with provided text), that text sources yield null ocr_confidence/page_count, and that the 0.3 floor is a placeholder to calibrate. It adds value beyond the field-level docs, though several details (60000 chars, 0.3 default) are duplicated from the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete pipeline with specific stages (classify → typed extraction → validation) and the input modalities it accepts. It also explicitly positions itself against siblings: 'Replaces the manual chain of mistral_ocr + mistral_chat + JSON parsing.' An agent can distinguish it from mistral_ocr and mistral_chat without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit routing guidance: kind=auto lets the server classify, source.type=text skips OCR and Files uploads, text+generic makes no API calls, and 'use mistral_ocr directly if you need raw OCR without a confidence guarantee.' It names the alternative and the condition that selects it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voxtral_transcribeVoxtral speech-to-textARead-onlyIdempotent
Transcribe an audio file to text using Mistral Voxtral.
Accepted models:
voxtral-mini-latest
voxtral-small-latest
Audio source is one of:
{ type: "file_url", fileUrl: "https://..." } (public URL)
{ type: "file", fileId: "" }
Options:
language: ISO-639-1 hint (e.g. 'fr', 'en'). Boosts accuracy when known.temperature: sampling temperature.diarize: return per-speaker segments (default false).timestampGranularities: ['segment'] to return per-segment timestamps.contextBias: list of phrases/terms that should bias the decoder.
Returns plain text, detected language, optional segments[], and token usage.
| Name | Required | Description | Default |
|---|---|---|---|
| audio | Yes | Audio to transcribe, supplied as a public URL or an uploaded file ID. | |
| model | No | STT model. Default: voxtral-mini-latest. | |
| diarize | No | Identify speakers in the returned transcription segments. Defaults to false. | |
| language | No | ISO-639-1 language hint (e.g. 'fr', 'en'). | |
| contextBias | No | Words or phrases to favor when decoding the audio. | |
| temperature | No | Sampling temperature for transcription. | |
| timestampGranularities | No | Only 'segment' is currently supported. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | Yes | |
| model | Yes | |
| usage | No | |
| language | Yes | |
| segments | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds useful behavioral context beyond that: accepted model IDs, default values for diarize and model, the only-supported granularity, and the shape of the response (text, language, segments, token usage). It does not mention auth or rate limits, but that is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action, then organizes models, source options, options, and return values into scannable bullet groups. It is efficient, though the option bullets partially duplicate what the schema already documents with 100% coverage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema, full annotation coverage, and 100% schema description coverage, the description supplies everything an agent needs: models, input modes, option semantics, and defaults. No material gap remains for invoking this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description nonetheless adds meaning beyond the schema by explaining intent ('language ... boosts accuracy when known', 'contextBias: phrases that should bias the decoder', 'diarize: return per-speaker segments'), which helps the agent use the parameters correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence gives a specific verb and resource ('Transcribe an audio file to text'), which is unambiguous and clearly distinct from the sibling tools (chat, FIM, vision, OCR, document processing). An agent can immediately tell this is the audio speech-to-text tool without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly lays out the two mutually exclusive audio source modes (public URL vs. uploaded file ID) and when each applies, which is exactly the routing decision an agent must make. It does not name alternatives or exclusions for when *not* to transcribe, but the context is otherwise clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.8.3- Changed
codestral_fim14 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / max_tokens / descriptionAdded value: +"Maximum number of tokens to generate. Input tokens plus this limit must fit within the model's context length." - added
Input schema / properties / max_tokens / maximumAdded value: +9007199254740991 - added
Input schema / properties / model / descriptionAdded value: +"FIM (fill-in-the-middle) model identifier. Any identifier your endpoint serves is accepted — read mistral://models for the live catalog. Known Mistral aliases: codestral-latest." - removed
Input schema / properties / model / enumRemoved value: -[ - "codestral-latest" -] - added
Input schema / properties / model / maxLengthAdded value: +200 - added
Input schema / properties / model / minLengthAdded value: +1 - added
Input schema / properties / seed / maximumAdded value: +9007199254740991 - added
Input schema / properties / seed / minimumAdded value: +-9007199254740991 - added
Input schema / properties / stop / descriptionAdded value: +"Stop generation when any of these text sequences is encountered." - added
Input schema / properties / temperature / descriptionAdded value: +"Sampling temperature: higher values make output more random; lower values make it more focused. Prefer adjusting this or top_p, not both." - added
Input schema / properties / top_p / descriptionAdded value: +"Nucleus sampling probability mass: 0.1 considers tokens in the top 10% of probability mass. Prefer adjusting this or temperature, not both." - changed
Output schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mistral_chat19 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / max_tokens / descriptionAdded value: +"Maximum number of tokens to generate. Input tokens plus this limit must fit within the model's context length." - added
Input schema / properties / max_tokens / maximumAdded value: +9007199254740991 - removed
Input schema / properties / messages / items / additionalPropertiesRemoved value: -false - added
Input schema / properties / messages / items / properties / content / descriptionAdded value: +"Text of the message." - added
Input schema / properties / messages / items / properties / role / descriptionAdded value: +"Message author: system for instructions, user for requests, or assistant for prior replies." - changed
Input schema / properties / model / descriptionPrevious value: -"Mistral chat model alias. Allowed: mistral-large-latest, mistral-medium-latest, mistral-small-latest, ministral-3b-latest, ministral-8b-latest, ministral-14b-latest, magistral-medium-latest, magistral-small-latest, devstral-latest, devstral-small-latest, codestral-latest, voxtral-small-latest. Default: mistral-medium-latest."New value: +"Chat model. Default: mistral-medium-latest (override with MISTRAL_DEFAULT_MODEL)." - removed
Input schema / properties / model / enumRemoved value: -[ - "mistral-large-latest", - "mistral-medium-latest", - "mistral-small-latest", - "ministral-3b-latest", - "ministral-8b-latest", - "ministral-14b-latest", - "magistral-medium-latest", - "magistral-small-latest", - "devstral-latest", - "devstral-small-latest", - "codestral-latest", - "voxtral-small-latest" -] - added
Input schema / properties / model / maxLengthAdded value: +200 - added
Input schema / properties / model / minLengthAdded value: +1 - changed
Input schema / properties / reasoning_effort / descriptionPrevious value: -"Controls reasoning depth for Magistral models. 'high' enables full chain-of-thought; 'none' disables it. Ignored on non-reasoning models."New value: +"Reasoning effort; supported values depend on the selected model." - changed
Input schema / properties / reasoning_effort / enumPrevious value: -[ - "none", - "high" -]New value: +[ + "none", + "minimal", + "low", + "medium", + "high", + "xhigh" +] - changed
Input schema / properties / response_format / anyOfPrevious value: -[ - { - "additionalProperties": false, - "properties": { - "type": { - "const": "text", - "type": "string" - } - }, - "required": [ - "type" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "type": { - "const": "json_object", - "type": "string" - } - }, - "required": [ - "type" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "json_schema": { - "additionalProperties": false, - "properties": { - "description": { - "type": "string" - }, - "name": { - "description": "Identifier for the schema; surfaced in API errors.", - "maxLength": 64, - "minLength": 1, - "type": "string" - }, - "schema": { - "additionalProperties": {}, - "description": "JSON Schema object the response must conform to.", - "type": "object" - }, - "strict": { - "description": "If true, the API rejects responses that do not strictly match the schema.", - "type": "boolean" - } - }, - "required": [ - "name", - "schema" - ], - "type": "object" - }, - "type": { - "const": "json_schema", - "type": "string" - } - }, - "required": [ - "type", - "json_schema" - ], - "type": "object" - } -]New value: +[ + { + "properties": { + "type": { + "const": "text", + "description": "Generate plain text without a JSON format constraint.", + "type": "string" + } + }, + "required": [ + "type" + ], + "type": "object" + }, + { + "properties": { + "type": { + "const": "json_object", + "description": "Generate JSON. Also instruct the model to produce JSON in a system or user message.", + "type": "string" + } + }, + "required": [ + "type" + ], + "type": "object" + }, + { + "properties": { + "json_schema": { + "description": "Named JSON Schema and optional strictness for the generated response.", + "properties": { + "description": { + "description": "Description of the response the schema defines.", + "type": "string" + }, + "name": { + "description": "Identifier for the schema; surfaced in API errors.", + "maxLength": 64, + "minLength": 1, + "type": "string" + }, + "schema": { + "additionalProperties": {}, + "description": "JSON Schema object the response must conform to.", + "propertyNames": { + "type": "string" + }, + "type": "object" + }, + "strict": { + "description": "If true, the API rejects responses that do not strictly match the schema.", + "type": "boolean" + } + }, + "required": [ + "name", + "schema" + ], + "type": "object" + }, + "type": { + "const": "json_schema", + "description": "Generate JSON conforming to the supplied json_schema.", + "type": "string" + } + }, + "required": [ + "type", + "json_schema" + ], + "type": "object" + } +] - added
Input schema / properties / seed / maximumAdded value: +9007199254740991 - added
Input schema / properties / seed / minimumAdded value: +-9007199254740991 - added
Input schema / properties / temperature / descriptionAdded value: +"Sampling temperature: higher values make output more random; lower values make it more focused. Prefer adjusting this or top_p, not both." - added
Input schema / properties / top_p / descriptionAdded value: +"Nucleus sampling probability mass: 0.1 considers tokens in the top 10% of probability mass. Prefer adjusting this or temperature, not both." - changed
Output schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mistral_ocr49 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false - removed
Input schema / properties / bbox_annotation_format / additionalPropertiesRemoved value: -false - added
Input schema / properties / bbox_annotation_format / descriptionAdded value: +"JSON Schema for structured annotations of each extracted bounding box or image." - removed
Input schema / properties / bbox_annotation_format / properties / json_schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / bbox_annotation_format / properties / json_schema / descriptionAdded value: +"Named JSON Schema and optional strictness for the extracted annotation." - added
Input schema / properties / bbox_annotation_format / properties / json_schema / properties / description / descriptionAdded value: +"Description of the annotation to extract." - added
Input schema / properties / bbox_annotation_format / properties / json_schema / properties / name / descriptionAdded value: +"Name identifying the annotation schema." - added
Input schema / properties / bbox_annotation_format / properties / json_schema / properties / schema / descriptionAdded value: +"JSON Schema object defining the fields to extract into the annotation." - added
Input schema / properties / bbox_annotation_format / properties / json_schema / properties / schema / propertyNamesAdded value: +{ + "type": "string" +} - added
Input schema / properties / bbox_annotation_format / properties / json_schema / properties / strict / descriptionAdded value: +"Whether the annotation must strictly follow the supplied JSON Schema." - added
Input schema / properties / confidence_scores_granularity / descriptionAdded value: +"Confidence granularity. 'block' also fills pages[].blocks[].confidence_scores and requires OCR 4.1 (mistral-ocr-4-1) or newer." - changed
Input schema / properties / confidence_scores_granularity / enumPrevious value: -[ - "page", - "word" -]New value: +[ + "page", + "word", + "block" +] - changed
Input schema / properties / document / anyOfPrevious value: -[ - { - "additionalProperties": false, - "properties": { - "documentName": { - "type": "string" - }, - "documentUrl": { - "description": "HTTPS URL to a PDF or image.", - "type": "string" - }, - "type": { - "const": "document_url", - "type": "string" - } - }, - "required": [ - "type", - "documentUrl" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "imageUrl": { - "description": "HTTPS URL or data:image/...;base64,... payload.", - "type": "string" - }, - "type": { - "const": "image_url", - "type": "string" - } - }, - "required": [ - "type", - "imageUrl" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "fileId": { - "description": "ID of a file previously uploaded via the Files API.", - "type": "string" - }, - "type": { - "const": "file", - "type": "string" - } - }, - "required": [ - "type", - "fileId" - ], - "type": "object" - } -]New value: +[ + { + "properties": { + "documentName": { + "description": "Filename of the referenced document.", + "type": "string" + }, + "documentUrl": { + "description": "HTTPS URL to a PDF or image.", + "type": "string" + }, + "type": { + "const": "document_url", + "description": "Read the document from documentUrl.", + "type": "string" + } + }, + "required": [ + "type", + "documentUrl" + ], + "type": "object" + }, + { + "properties": { + "imageUrl": { + "description": "HTTPS URL or data:image/...;base64,... payload.", + "type": "string" + }, + "type": { + "const": "image_url", + "description": "Read the image from the URL or data URI in imageUrl.", + "type": "string" + } + }, + "required": [ + "type", + "imageUrl" + ], + "type": "object" + }, + { + "properties": { + "fileId": { + "description": "ID of a file previously uploaded via the Files API.", + "type": "string" + }, + "type": { + "const": "file", + "description": "Read a previously uploaded file identified by fileId.", + "type": "string" + } + }, + "required": [ + "type", + "fileId" + ], + "type": "object" + } +] - added
Input schema / properties / document / descriptionAdded value: +"Document to process, supplied as a document URL, an image URL or data URI, or an uploaded file ID." - removed
Input schema / properties / document_annotation_format / $refRemoved value: -"#/properties/bbox_annotation_format" - added
Input schema / properties / document_annotation_format / descriptionAdded value: +"JSON Schema for a structured annotation extracted from the entire document." - added
Input schema / properties / document_annotation_format / propertiesAdded value: +{ + "json_schema": { + "description": "Named JSON Schema and optional strictness for the extracted annotation.", + "properties": { + "description": { + "description": "Description of the annotation to extract.", + "type": "string" + }, + "name": { + "description": "Name identifying the annotation schema.", + "minLength": 1, + "type": "string" + }, + "schema": { + "additionalProperties": {}, + "description": "JSON Schema object defining the fields to extract into the annotation.", + "propertyNames": { + "type": "string" + }, + "type": "object" + }, + "strict": { + "description": "Whether the annotation must strictly follow the supplied JSON Schema.", + "type": "boolean" + } + }, + "required": [ + "name", + "schema" + ], + "type": "object" + }, + "type": { + "const": "json_schema", + "description": "Only json_schema is accepted by OCR annotation formats.", + "type": "string" + } +} - added
Input schema / properties / document_annotation_format / requiredAdded value: +[ + "type", + "json_schema" +] - added
Input schema / properties / document_annotation_format / typeAdded value: +"object" - added
Input schema / properties / document_annotation_prompt / descriptionAdded value: +"Instructions for whole-document structured extraction. Requires document_annotation_format." - added
Input schema / properties / extractFooter / descriptionAdded value: +"Extract each page footer into its footer field and remove it from the page markdown." - added
Input schema / properties / extractHeader / descriptionAdded value: +"Extract each page header into its header field and remove it from the page markdown." - added
Input schema / properties / imageLimit / descriptionAdded value: +"Maximum number of images to extract from the document." - added
Input schema / properties / imageLimit / maximumAdded value: +9007199254740991 - added
Input schema / properties / imageMinSize / descriptionAdded value: +"Minimum height and width of an image to extract." - added
Input schema / properties / imageMinSize / maximumAdded value: +9007199254740991 - added
Input schema / properties / includeBlocksAdded value: +{ + "description": "Return paragraph-level blocks (bounding box + type) per page. Requires OCR 4 (mistral-ocr-4-0) or newer.", + "type": "boolean" +} - added
Input schema / properties / includeImageBase64 / descriptionAdded value: +"Include base64-encoded data for extracted images in the response." - removed
Input schema / properties / model / enumRemoved value: -[ - "mistral-ocr-latest" -] - added
Input schema / properties / model / maxLengthAdded value: +200 - added
Input schema / properties / model / minLengthAdded value: +1 - changed
Input schema / properties / pages / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "items": { - "minimum": 0, - "type": "integer" - }, - "type": "array" - } -]New value: +[ + { + "type": "string" + }, + { + "items": { + "maximum": 9007199254740991, + "minimum": 0, + "type": "integer" + }, + "type": "array" + } +] - added
Input schema / properties / pages / descriptionAdded value: +"Zero-based page numbers to process, as an array or comma-separated numbers and ranges such as \"0-5,7\"." - added
Input schema / properties / tableFormat / descriptionAdded value: +"Format for extracted tables: Markdown or HTML." - changed
Output schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - added
Output schema / properties / annotations / properties / image_annotations / items / properties / page_index / maximumAdded value: +9007199254740991 - added
Output schema / properties / annotations / properties / image_annotations / items / properties / page_index / minimumAdded value: +-9007199254740991 - added
Output schema / properties / pages / items / properties / blocksAdded value: +{ + "description": "Paragraph-level blocks in reading order. Populated when `includeBlocks: true` and the model is OCR 4 or newer.", + "items": { + "additionalProperties": false, + "properties": { + "bottom_right_x": { + "type": "number" + }, + "bottom_right_y": { + "type": "number" + }, + "confidence_scores": { + "additionalProperties": false, + "description": "Populated only when `confidence_scores_granularity: 'block'`. Fields are null when the signal is absent (an image-only block has no content to score).", + "properties": { + "average_content_confidence_score": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ] + }, + "block_type_confidence_score": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ] + }, + "minimum_content_confidence_score": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ] + } + }, + "type": "object" + }, + "content": { + "type": "string" + }, + "image_id": { + "description": "Set on type:image — references the matching entry in `images[]`.", + "type": "string" + }, + "table_id": { + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "description": "Set on type:table — references the matching entry in `tables[]`." + }, + "top_left_x": { + "type": "number" + }, + "top_left_y": { + "type": "number" + }, + "type": { + "enum": [ + "text", + "title", + "list", + "table", + "image", + "equation", + "caption", + "code", + "references", + "aside_text", + "header", + "footer", + "signature" + ], + "type": "string" + } + }, + "required": [ + "type", + "top_left_x", + "top_left_y", + "bottom_right_x", + "bottom_right_y", + "content" + ], + "type": "object" + }, + "type": "array" +} - added
Output schema / properties / pages / items / properties / confidence_scores / properties / word_confidence_scores / items / properties / start_index / maximumAdded value: +9007199254740991 - added
Output schema / properties / pages / items / properties / confidence_scores / properties / word_confidence_scores / items / properties / start_index / minimumAdded value: +-9007199254740991 - added
Output schema / properties / pages / items / properties / footer / anyOfAdded value: +[ + { + "type": "string" + }, + { + "type": "null" + } +] - removed
Output schema / properties / pages / items / properties / footer / typeRemoved value: -[ - "string", - "null" -] - added
Output schema / properties / pages / items / properties / header / anyOfAdded value: +[ + { + "type": "string" + }, + { + "type": "null" + } +] - removed
Output schema / properties / pages / items / properties / header / typeRemoved value: -[ - "string", - "null" -] - added
Output schema / properties / pages / items / properties / index / maximumAdded value: +9007199254740991 - added
Output schema / properties / pages / items / properties / index / minimumAdded value: +-9007199254740991 - added
Output schema / properties / pages_count / maximumAdded value: +9007199254740991 - added
Output schema / properties / pages_count / minimumAdded value: +-9007199254740991
- Changed
mistral_vision16 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / max_tokens / descriptionAdded value: +"Maximum number of tokens to generate. Input tokens plus this limit must fit within the model's context length." - added
Input schema / properties / max_tokens / maximumAdded value: +9007199254740991 - removed
Input schema / properties / messages / items / additionalPropertiesRemoved value: -false - changed
Input schema / properties / messages / items / properties / content / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "items": { - "anyOf": [ - { - "additionalProperties": false, - "properties": { - "text": { - "type": "string" - }, - "type": { - "const": "text", - "type": "string" - } - }, - "required": [ - "type", - "text" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "imageUrl": { - "anyOf": [ - { - "description": "https URL or data:image/...;base64,... payload", - "type": "string" - }, - { - "additionalProperties": false, - "properties": { - "detail": { - "enum": [ - "auto", - "low", - "high" - ], - "type": "string" - }, - "url": { - "type": "string" - } - }, - "required": [ - "url" - ], - "type": "object" - } - ] - }, - "type": { - "const": "image_url", - "type": "string" - } - }, - "required": [ - "type", - "imageUrl" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "documentName": { - "type": "string" - }, - "documentUrl": { - "type": "string" - }, - "type": { - "const": "document_url", - "type": "string" - } - }, - "required": [ - "type", - "documentUrl" - ], - "type": "object" - } - ] - }, - "minItems": 1, - "type": "array" - } -]New value: +[ + { + "type": "string" + }, + { + "items": { + "anyOf": [ + { + "properties": { + "text": { + "description": "Text to include in the message.", + "type": "string" + }, + "type": { + "const": "text", + "description": "Identifies a text content part.", + "type": "string" + } + }, + "required": [ + "type", + "text" + ], + "type": "object" + }, + { + "properties": { + "imageUrl": { + "anyOf": [ + { + "description": "https URL or data:image/...;base64,... payload", + "type": "string" + }, + { + "properties": { + "detail": { + "description": "Image detail hint: automatic, low, or high.", + "enum": [ + "auto", + "low", + "high" + ], + "type": "string" + }, + "url": { + "description": "HTTPS URL or data:image/...;base64,... payload.", + "type": "string" + } + }, + "required": [ + "url" + ], + "type": "object" + } + ], + "description": "Image source as a URL or base64 data URI, optionally with a detail hint." + }, + "type": { + "const": "image_url", + "description": "Identifies an image content part.", + "type": "string" + } + }, + "required": [ + "type", + "imageUrl" + ], + "type": "object" + }, + { + "properties": { + "documentName": { + "description": "Filename of the referenced document.", + "type": "string" + }, + "documentUrl": { + "description": "URL of the PDF or document to include in the message.", + "type": "string" + }, + "type": { + "const": "document_url", + "description": "Identifies a document content part.", + "type": "string" + } + }, + "required": [ + "type", + "documentUrl" + ], + "type": "object" + } + ] + }, + "minItems": 1, + "type": "array" + } +] - added
Input schema / properties / messages / items / properties / content / descriptionAdded value: +"Message text or an ordered list of text, image, and document content parts." - added
Input schema / properties / messages / items / properties / role / descriptionAdded value: +"Message author: system for instructions, user for requests, or assistant for prior replies." - removed
Input schema / properties / model / enumRemoved value: -[ - "pixtral-large-latest", - "pixtral-12b-latest", - "mistral-large-latest", - "mistral-medium-latest", - "mistral-small-latest" -] - added
Input schema / properties / model / maxLengthAdded value: +200 - added
Input schema / properties / model / minLengthAdded value: +1 - added
Input schema / properties / seed / maximumAdded value: +9007199254740991 - added
Input schema / properties / seed / minimumAdded value: +-9007199254740991 - added
Input schema / properties / temperature / descriptionAdded value: +"Sampling temperature: higher values make output more random; lower values make it more focused. Prefer adjusting this or top_p, not both." - added
Input schema / properties / top_p / descriptionAdded value: +"Nucleus sampling probability mass: 0.1 considers tokens in the top 10% of probability mass. Prefer adjusting this or temperature, not both." - changed
Output schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Added
process_document - Changed
voxtral_transcribe13 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false - changed
Input schema / properties / audio / anyOfPrevious value: -[ - { - "additionalProperties": false, - "properties": { - "fileUrl": { - "description": "HTTPS URL to an audio file (mp3/wav/flac/ogg/webm/m4a).", - "type": "string" - }, - "type": { - "const": "file_url", - "type": "string" - } - }, - "required": [ - "type", - "fileUrl" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "fileId": { - "description": "ID of an audio file previously uploaded via the Files API (purpose=audio).", - "type": "string" - }, - "type": { - "const": "file", - "type": "string" - } - }, - "required": [ - "type", - "fileId" - ], - "type": "object" - } -]New value: +[ + { + "properties": { + "fileUrl": { + "description": "HTTPS URL to an audio file (mp3/wav/flac/ogg/webm/m4a).", + "type": "string" + }, + "type": { + "const": "file_url", + "description": "Transcribe audio from the URL in fileUrl.", + "type": "string" + } + }, + "required": [ + "type", + "fileUrl" + ], + "type": "object" + }, + { + "properties": { + "fileId": { + "description": "ID of an audio file previously uploaded via the Files API (purpose=audio).", + "type": "string" + }, + "type": { + "const": "file", + "description": "Transcribe an uploaded audio file identified by fileId.", + "type": "string" + } + }, + "required": [ + "type", + "fileId" + ], + "type": "object" + } +] - added
Input schema / properties / audio / descriptionAdded value: +"Audio to transcribe, supplied as a public URL or an uploaded file ID." - added
Input schema / properties / contextBias / descriptionAdded value: +"Words or phrases to favor when decoding the audio." - added
Input schema / properties / diarize / descriptionAdded value: +"Identify speakers in the returned transcription segments. Defaults to false." - removed
Input schema / properties / model / enumRemoved value: -[ - "voxtral-mini-latest", - "voxtral-small-latest" -] - added
Input schema / properties / model / maxLengthAdded value: +200 - added
Input schema / properties / model / minLengthAdded value: +1 - added
Input schema / properties / temperature / descriptionAdded value: +"Sampling temperature for transcription." - changed
Output schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - added
Output schema / properties / language / anyOfAdded value: +[ + { + "type": "string" + }, + { + "type": "null" + } +] - removed
Output schema / properties / language / typeRemoved value: -[ - "string", - "null" -]
- Removed
workflow_execute - Removed
workflow_interact - Removed
workflow_status
24 tool updates
v0.7.0- Removed
batch_cancel - Removed
batch_create - Removed
batch_get - Removed
batch_list - Changed
codestral_fim1 field changed- added
Input schema / properties / seedAdded value: +{ + "description": "Random seed for deterministic sampling. Maps to Mistral's `random_seed`.", + "type": "integer" +}
- Removed
files_delete - Removed
files_get - Removed
files_list - Removed
files_signed_url - Removed
files_upload - Removed
mcp_sample - Removed
mistral_agent - Changed
mistral_chat4 fields changed- added
Input schema / properties / reasoning_effortAdded value: +{ + "description": "Controls reasoning depth for Magistral models. 'high' enables full chain-of-thought; 'none' disables it. Ignored on non-reasoning models.", + "enum": [ + "none", + "high" + ], + "type": "string" +} - added
Input schema / properties / response_formatAdded value: +{ + "anyOf": [ + { + "additionalProperties": false, + "properties": { + "type": { + "const": "text", + "type": "string" + } + }, + "required": [ + "type" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "type": { + "const": "json_object", + "type": "string" + } + }, + "required": [ + "type" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "json_schema": { + "additionalProperties": false, + "properties": { + "description": { + "type": "string" + }, + "name": { + "description": "Identifier for the schema; surfaced in API errors.", + "maxLength": 64, + "minLength": 1, + "type": "string" + }, + "schema": { + "additionalProperties": {}, + "description": "JSON Schema object the response must conform to.", + "type": "object" + }, + "strict": { + "description": "If true, the API rejects responses that do not strictly match the schema.", + "type": "boolean" + } + }, + "required": [ + "name", + "schema" + ], + "type": "object" + }, + "type": { + "const": "json_schema", + "type": "string" + } + }, + "required": [ + "type", + "json_schema" + ], + "type": "object" + } + ], + "description": "Force a structured output: `{type:\"json_object\"}` for JSON mode, `{type:\"json_schema\", json_schema:{...}}` for strict schema mode." +} - added
Input schema / properties / seedAdded value: +{ + "description": "Random seed for deterministic sampling. Maps to Mistral's `random_seed`.", + "type": "integer" +} - added
Output schema / properties / reasoning_contentAdded value: +{ + "description": "Reasoning trace returned by Magistral models. Absent for non-reasoning models.", + "type": "string" +}
- Removed
mistral_chat_stream - Removed
mistral_classify - Removed
mistral_embed - Removed
mistral_moderate - Changed
mistral_ocr7 fields changed- added
Input schema / properties / bbox_annotation_formatAdded value: +{ + "additionalProperties": false, + "properties": { + "json_schema": { + "additionalProperties": false, + "properties": { + "description": { + "type": "string" + }, + "name": { + "minLength": 1, + "type": "string" + }, + "schema": { + "additionalProperties": {}, + "type": "object" + }, + "strict": { + "type": "boolean" + } + }, + "required": [ + "name", + "schema" + ], + "type": "object" + }, + "type": { + "const": "json_schema", + "description": "Only json_schema is accepted by OCR annotation formats.", + "type": "string" + } + }, + "required": [ + "type", + "json_schema" + ], + "type": "object" +} - added
Input schema / properties / confidence_scores_granularityAdded value: +{ + "enum": [ + "page", + "word" + ], + "type": "string" +} - added
Input schema / properties / document_annotation_formatAdded value: +{ + "$ref": "#/properties/bbox_annotation_format" +} - added
Input schema / properties / document_annotation_promptAdded value: +{ + "type": "string" +} - added
Output schema / properties / annotationsAdded value: +{ + "additionalProperties": false, + "properties": { + "document_annotation": { + "type": "string" + }, + "image_annotations": { + "items": { + "additionalProperties": false, + "properties": { + "annotation": { + "type": "string" + }, + "bbox": { + "additionalProperties": false, + "properties": { + "bottom_right_x": { + "type": "number" + }, + "bottom_right_y": { + "type": "number" + }, + "top_left_x": { + "type": "number" + }, + "top_left_y": { + "type": "number" + } + }, + "type": "object" + }, + "image_id": { + "type": "string" + }, + "page_index": { + "type": "integer" + } + }, + "required": [ + "page_index", + "annotation" + ], + "type": "object" + }, + "type": "array" + } + }, + "type": "object" +} - added
Output schema / properties / pages / items / properties / confidence_scoresAdded value: +{ + "additionalProperties": false, + "properties": { + "average_page_confidence_score": { + "type": "number" + }, + "minimum_page_confidence_score": { + "type": "number" + }, + "word_confidence_scores": { + "items": { + "additionalProperties": false, + "properties": { + "confidence": { + "type": "number" + }, + "start_index": { + "type": "integer" + }, + "text": { + "type": "string" + } + }, + "required": [ + "text", + "confidence", + "start_index" + ], + "type": "object" + }, + "type": "array" + } + }, + "required": [ + "average_page_confidence_score", + "minimum_page_confidence_score" + ], + "type": "object" +} - added
Output schema / properties / pages / items / properties / images / items / properties / image_annotationAdded value: +{ + "type": "string" +}
- Removed
mistral_tool_call - Changed
mistral_vision1 field changed- added
Input schema / properties / seedAdded value: +{ + "description": "Random seed for deterministic sampling. Maps to Mistral's `random_seed`.", + "type": "integer" +}
- Removed
voxtral_speak - Added
workflow_execute - Added
workflow_interact - Added
workflow_status
22 tool updates
v0.4.0- First observed
batch_cancel - First observed
batch_create - First observed
batch_get - First observed
batch_list - First observed
codestral_fim - First observed
files_delete - First observed
files_get - First observed
files_list - First observed
files_signed_url - First observed
files_upload - First observed
mcp_sample - First observed
mistral_agent - First observed
mistral_chat - First observed
mistral_chat_stream - First observed
mistral_classify - First observed
mistral_embed - First observed
mistral_moderate - First observed
mistral_ocr - First observed
mistral_tool_call - First observed
mistral_vision - First observed
voxtral_speak - First observed
voxtral_transcribe
TDQS
Scored across 6 tools
Each tool targets a distinct modality/action: code FIM, text chat, OCR, vision chat, document pipeline, and audio transcription. There is minor overlap between mistral_chat and mistral_vision (both chat) and between process_document and mistral_ocr, but the descriptions explicitly steer selection, so confusion is limited.
Names mix provider-prefixed conventions (codestral_fim, mistral_chat, mistral_ocr, mistral_vision, voxtral_transcribe) with a plain action name (process_document). The pattern is somewhat readable but not a predictable verb_noun scheme and the provider prefixes fragment consistency.
Six tools is well-scoped for a multimodal Mistral wrapper, with each tool covering a distinct capability rather than redundancy. Slightly broad for a server named 'Document Extraction' since it also includes code completion and chat, but not excessive.
Covers OCR, vision, chat, transcription, and a combined pipeline, but several tools reference a Files API (fileId) that is not exposed as a tool, creating a dead-end for uploading files. No model-listing or file-management operations leave workflows partially blocked.
Maintenance
Related MCP Connectors
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…
MCP server for progressive tool usage at any scale (see https://klavis.ai)
- QuallaaOAuthcom.quallaa
Talk to your public-facing AI from any MCP client — Claude, ChatGPT, Cursor, Cline, Windsurf.
Nifty's MCP server — exposes tasks, projects, messages, and files as tools for AI agents.
Related MCP Servers
- -licenseNot gradedqualityDmaintenanceA TypeScript implementation of a Model Context Protocol server and client that enables interaction with language models (specifically Mistral running on Ollama).-
- -licenseNot gradedqualityNot gradedmaintenanceA production-ready TypeScript MCP server providing basic tools (add, echo, timestamp), resources (server info, greetings, data access), and prompt templates (analyze, code-review, summarize). Serves as a foundation for building custom MCP servers with extensible architecture.398 npm-
- AlicenseBqualityCmaintenanceA TypeScript-based MCP server that provides tools to interact with local Codex and Gemini CLIs via stdio transport. It enables users to execute prompts through the ask_codex and ask_gemini tools, supporting custom models and timeout configurations.110 npm6MIT
- FlicenseAqualityDmaintenanceA TypeScript MCP server template with Zod validation, dual transport (stdio/HTTP), and modular architecture for building MCP-compatible tools, resources, and prompts.11-