Mistral MCP — Document Extraction
mistral-mcp
Mistral AI の機能をあらゆる MCP クライアント(Claude Code、Cursor、Zed、Windsurf、Claude Desktop など)に提供する MCP サーバー
フランス語版: README.fr.md
目的
Mistral はフランス語、コード、OCR、モデレーション、音声、エージェント型ワークフローに優れたモデルを持っていますが、ほとんどの MCP 対応 IDE はデフォルトで Anthropic や OpenAI を使用します。mistral-mcp は、それらの Mistral の機能をクリーンな MCP インターフェースとして提供するため、エージェントループを再構築することなく、適切なサブタスクを適切なモデルにルーティングできます。
このリポジトリの目的は「単なる薄いラッパー」を作ることではありません。明確なスキーマ、予測可能な出力、トランスポートの柔軟性、そして優れたテストカバレッジを備えた、堅牢で保守性の高い MCP サーバーを目指しています。
Related MCP server: MCP Server TypeScript
現在の機能 (v0.4.0)
ツール (22)
コア生成:
mistral_chatmistral_chat_streammistral_embedmistral_tool_callcodestral_fim
ビジョンと音声:
mistral_visionmistral_ocrvoxtral_transcribevoxtral_speak
エージェントと分類器:
mistral_agentmistral_moderatemistral_classify
ファイルとバッチ:
files_uploadfiles_listfiles_getfiles_deletefiles_signed_urlbatch_createbatch_listbatch_getbatch_cancel
MCP ネイティブユーティリティ:
mcp_sample- MCP サンプリングを介してクライアントモデルに生成を委任します
リソース (2)
mistral://models- 利用可能なエイリアスとライブモデルカタログmistral://voices- Voxtral TTS 用のライブ音声カタログ
プロンプト (6)
フランス語の厳選プロンプト:
french_invoice_reminderfrench_meeting_minutesfrench_email_replyfrench_commit_messagefrench_legal_summary
英語の厳選プロンプト:
codestral_review
プロンプトの列挙型引数は completable() でラップされているため、MCP クライアントは completion/complete を介してプロンプト引数の補完を呼び出すことができます。
特徴
すべてのツールに
inputSchema、outputSchema、アノテーションを備えた高レベルなMcpServerAPIデュアルトランスポートサポート: デフォルトの stdio、およびリモートデプロイ用のストリーム配信可能な HTTP
あらゆる場所での構造化出力:
structuredContentとテキストフォールバックmcp_sampleによる MCP サンプリングのサポート列挙型のようなプロンプト引数に対するプロンプト補完のサポート
ツールと並行して登録されるリソースとプロンプト(後付けではありません)
Mistral SDK クライアントでのリトライ/バックオフおよびリクエストタイムアウト
トランスポート
Stdio
デフォルトモード。Claude Code やほとんどのローカル MCP クライアントで使用されます。
node dist/index.jsストリーム配信可能な HTTP
--http または MCP_TRANSPORT=http で有効にします。
MCP_TRANSPORT=http node dist/index.js関連する環境変数:
MCP_HTTP_HOST- デフォルト127.0.0.1MCP_HTTP_PORT- デフォルト3333MCP_HTTP_PATH- デフォルト/mcpMCP_HTTP_TOKEN- オプションのベアラートークンMCP_HTTP_ALLOWED_ORIGINS- オプションのカンマ区切り許可リストMCP_HTTP_STATELESS=1- ステートレスセッションモード
/healthz は意図的に公開されており、MCP サーバーにはアクセスしません。
インストール
git clone https://github.com/Swih/mistral-mcp.git
cd mistral-mcp
npm install
npm run buildAPI キーを設定します:
export MISTRAL_API_KEY=your_key_hereまたは、リポジトリルートの .env を使用してください。決してコミットしないでください。
Claude Code での使用
claude mcp add mistral -- node /absolute/path/to/mistral-mcp/dist/index.jsプロンプトの例:
この PDF に対して
mistral_ocrを使用し、抽出されたテキストに対してfrench_meeting_minutesを実行してください。
開発
npm run dev
npm run build
npm run lint
npm test
npm run inspectorテスト戦略
スイートには現在、4 つのレイヤーにわたる 148 個のテストが含まれています:
ツール、リソース、プロンプト、トランスポート、音声、エージェント、ファイル、バッチ、サンプリングのユニットテスト
ツールメタデータと MCP 準拠の保証に関するコントラクトテスト
MISTRAL_API_KEYが設定されている場合の実際の Mistral API に対するライブ API テストビルドされたサーバーに対する Stdio エンドツーエンドテスト
MISTRAL_API_KEY がない場合、ローカルのデフォルトは 139 passing と 9 gated(ライブ/stdio テスト)になります。
プロジェクト構成
mistral-mcp/
|-- src/
| |-- index.ts
| |-- transport.ts
| |-- tools.ts
| |-- tools-fn.ts
| |-- tools-vision.ts
| |-- tools-audio.ts
| |-- tools-agents.ts
| |-- tools-files.ts
| |-- tools-batch.ts
| |-- tools-sampling.ts
| |-- resources.ts
| `-- prompts.ts
|-- test/
|-- examples/
|-- .github/workflows/ci.yml
|-- package.json
`-- tsconfig.test.jsonステータス
v0.4.0 — リリース済み。v0.3.0 からの完全な差分については CHANGELOG.md を参照してください:
共有ヘルパー、ライブモデル + 音声カタログ、コントラクトテスト
ビジョン + OCR
音声文字起こし + 音声合成
エージェント + モデレーション + 分類
ファイル + バッチ API
ストリーム配信可能な HTTP トランスポート + MCP サンプリング
5 つのフランス語厳選プロンプト + 1 つの英語プロンプト + プロンプト引数補完
例
実行可能なスクリプトは examples/ にあります。examples/README.md を参照してください。
ライセンス
MIT Copyright Dayan Decamp
Available Tools
6 toolscodestral_fimCodestral fill-in-the-middle completionARead-only
Fill-in-the-middle code completion with Codestral.
Given prompt (code preceding the cursor) and suffix (code after the cursor),
Codestral writes the middle. Use for editor autocomplete scenarios, code-patching
agents, or structured refactors where you know the target boundaries.
Default stop tokens: [] — let the model decide. Override with stop if needed.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Random seed for deterministic sampling. Maps to Mistral's `random_seed`. | |
| stop | No | Stop generation when any of these text sequences is encountered. | |
| model | No | FIM (fill-in-the-middle) model identifier. Any identifier your endpoint serves is accepted — read mistral://models for the live catalog. Known Mistral aliases: codestral-latest. | |
| top_p | No | Nucleus sampling probability mass: 0.1 considers tokens in the top 10% of probability mass. Prefer adjusting this or temperature, not both. | |
| prompt | Yes | Code preceding the cursor. | |
| suffix | Yes | Code after the cursor. Can be empty string. | |
| max_tokens | No | Maximum number of tokens to generate. Input tokens plus this limit must fit within the model's context length. | |
| temperature | No | Sampling temperature: higher values make output more random; lower values make it more focused. Prefer adjusting this or top_p, not both. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | Yes | |
| model | Yes | |
| usage | No | |
| finish_reason | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=false, and destructiveHint=false. The description adds useful generation behavior by noting the default stop tokens are empty and that the model decides when to stop unless `stop` is overridden.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with a one-line purpose, followed by cursor semantics, use cases, and stop-token behavior. Every sentence earns its place and no filler is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given rich annotations, 100% schema coverage, and an output schema, the description supplies the missing conceptual model: prompt before cursor, suffix after cursor, and Codestral writes the middle. It is complete enough for correct invocation without redundant return-value explanation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so prompt/suffix meanings are already documented. The description adds extra semantics for `stop`: default is [] and the model decides, with explicit override guidance, going beyond the schema's brief stop description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: fill-in-the-middle code completion with Codestral, plus the prompt/suffix cursor model. The FIM and editor-autocomplete framing distinguishes it from sibling chat, vision, OCR, transcription, and document tools without needing schema inspection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear usage contexts: editor autocomplete, code-patching agents, and structured refactors where target boundaries are known. It does not explicitly name when to avoid this tool or route to a sibling like mistral_chat for general generation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mistral_chatMistral chat completionARead-only
Generate a chat completion using a Mistral model.
When to use:
Drafting French (or any European-language) content where Mistral shines.
Codestral for code-specific generation/review.
Ministral for cheap / low-latency classification.
Returns structured content with the assistant text and token usage. Does NOT stream — use mistral_chat_stream for long outputs with progress updates.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Random seed for deterministic sampling. Maps to Mistral's `random_seed`. | |
| model | No | Chat model. Default: mistral-medium-latest (override with MISTRAL_DEFAULT_MODEL). | |
| top_p | No | Nucleus sampling probability mass: 0.1 considers tokens in the top 10% of probability mass. Prefer adjusting this or temperature, not both. | |
| messages | Yes | Chat messages in role/content form. | |
| max_tokens | No | Maximum number of tokens to generate. Input tokens plus this limit must fit within the model's context length. | |
| temperature | No | Sampling temperature: higher values make output more random; lower values make it more focused. Prefer adjusting this or top_p, not both. | |
| response_format | No | Force a structured output: `{type:"json_object"}` for JSON mode, `{type:"json_schema", json_schema:{...}}` for strict schema mode. | |
| reasoning_effort | No | Reasoning effort; supported values depend on the selected model. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | Yes | |
| model | Yes | |
| usage | No | |
| finish_reason | No | |
| reasoning_content | No | Reasoning trace returned by Magistral models. Absent for non-reasoning models. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover safety (readOnlyHint=true, destructiveHint=false, openWorldHint=true), so the bar is lower. The description still adds useful non-safety behavior: it does not stream, and it returns structured content with assistant text and token usage. It does not mention rate limits, latency, or error behavior, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action, organizes guidance as scannable bullets, and closes with the return shape and the streaming exclusion. No filler sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description needn't detail return values, and annotations cover safety. It still supplies the streaming caveat and model-family guidance. Minor gap: it never states default model behavior or auth expectations, but that is largely handled by the schema's model description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents seed, model, top_p, temperature, response_format, and reasoning_effort thoroughly. The description adds only indirect model-selection guidance ('Codestral', 'Ministral') rather than explaining parameters, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening line states a specific verb and resource: 'Generate a chat completion using a Mistral model.' The bullets further differentiate sub-cases (French/European drafting, Codestral for code, Ministral for cheap classification), letting an agent distinguish this from codestral_fim and other siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit 'When to use' bullets plus a named exclusion: 'Does NOT stream — use mistral_chat_stream for long outputs with progress updates.' This gives both positive triggers and a concrete alternative with the condition that selects it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mistral_ocrMistral OCR (document to markdown)ARead-onlyIdempotent
Run Mistral OCR on a PDF or image, returning structured markdown per page.
Input document is one of:
{ type: "document_url", documentUrl: "https://...pdf" }
{ type: "image_url", imageUrl: "https://..." | "data:image/..." }
{ type: "file", fileId: "" }
Options:
pages: array of 0-indexed page numbers or string like "0-5,7".tableFormat: 'markdown' (default) or 'html'.extractHeader/extractFooter: include page header/footer when present.includeImageBase64: embed extracted image bytes as base64 in the response.document_annotation_format: JSON schema for whole-document structured extraction.bbox_annotation_format: JSON schema for extracted image / bbox annotations.confidence_scores_granularity: 'page', 'word', or 'block'. 'block' adds per-block content/type confidence underpages[].blocks[].confidence_scoresand requires OCR 4.1 or newer.includeBlocks: return paragraph-level blocks (bounding box + type) in reading order — titles, lists, tables, images, equations, captions, code, references, aside text, header, footer, signature. Requires OCR 4 (mistral-ocr-4-0) or newer; older models accept the flag but return an emptyblocksarray.
Returns pages[].markdown plus optional pages[].hyperlinks, header, footer,
images bounding boxes, blocks, annotations, confidence scores, and dimensions.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | OCR model. Default: mistral-ocr-latest. | |
| pages | No | Zero-based page numbers to process, as an array or comma-separated numbers and ranges such as "0-5,7". | |
| document | Yes | Document to process, supplied as a document URL, an image URL or data URI, or an uploaded file ID. | |
| imageLimit | No | Maximum number of images to extract from the document. | |
| tableFormat | No | Format for extracted tables: Markdown or HTML. | |
| imageMinSize | No | Minimum height and width of an image to extract. | |
| extractFooter | No | Extract each page footer into its footer field and remove it from the page markdown. | |
| extractHeader | No | Extract each page header into its header field and remove it from the page markdown. | |
| includeBlocks | No | Return paragraph-level blocks (bounding box + type) per page. Requires OCR 4 (mistral-ocr-4-0) or newer. | |
| includeImageBase64 | No | Include base64-encoded data for extracted images in the response. | |
| bbox_annotation_format | No | JSON Schema for structured annotations of each extracted bounding box or image. | |
| document_annotation_format | No | JSON Schema for a structured annotation extracted from the entire document. | |
| document_annotation_prompt | No | Instructions for whole-document structured extraction. Requires document_annotation_format. | |
| confidence_scores_granularity | No | Confidence granularity. 'block' also fills pages[].blocks[].confidence_scores and requires OCR 4.1 (mistral-ocr-4-1) or newer. |
Output Schema
| Name | Required | Description |
|---|---|---|
| model | Yes | |
| pages | Yes | |
| usage | No | |
| annotations | No | |
| pages_count | Yes | |
| document_annotation | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (which only declare read-only, idempotent, non-destructive, open-world), the description discloses real behavioral preconditions: includeBlocks requires OCR 4+, confidence_scores_granularity='block' requires OCR 4.1+, and older models silently accept includeBlocks but return an empty blocks array. That silent-failure warning is exactly the kind of context an agent cannot get from the schema or annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with a one-line purpose, then cleanly sectioned into input shapes, options, and returns for a 14-parameter tool. The length is proportionate to the surface area and each bullet maps to a decision the caller must make.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only OCR tool with a full output schema and 100% schema coverage, the description supplies everything else needed: the three mutually exclusive document input forms, option semantics, model-version gates, and the shape of the response. Nothing an agent needs to invoke it correctly is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds genuine value on several parameters the schema documents only tersely: the annotated string form of `pages` ('0-5,7'), the default of tableFormat, the distinction between document_annotation_format and bbox_annotation_format, and the version requirements tied to includeBlocks and confidence granularity. Some parameters (imageLimit, imageMinSize, model) are untouched, which keeps it from a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('Run Mistral OCR on a PDF or image') and states the core output ('returning structured markdown per page'). It is clear, but it never distinguishes itself from siblings like mistral_vision or process_document, so an agent cannot tell from the description alone which of those overlapping tools to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the enumerated input shapes and option list, but there is no explicit when-to-use or when-not-to-use guidance, and no routing to alternative siblings such as mistral_vision for image-only OCR. The reader must infer the fit from the parameter list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mistral_visionMistral multimodal chat (vision)ARead-only
Chat completion with multimodal input: text + image_url parts.
Requires a vision-capable model. Accepted:
pixtral-large-latest
pixtral-12b-latest
mistral-large-latest
mistral-medium-latest
mistral-small-latest
Each message's content is either a plain string (pure text) or an array of
parts { type: 'text', text } / { type: 'image_url', imageUrl }. The image URL
can be an https URL or a data: URI base64 payload.
Returns the assistant text + token usage. For non-visual requests, prefer mistral_chat.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Random seed for deterministic sampling. Maps to Mistral's `random_seed`. | |
| model | No | Vision-capable Mistral model. Default: pixtral-large-latest. | |
| top_p | No | Nucleus sampling probability mass: 0.1 considers tokens in the top 10% of probability mass. Prefer adjusting this or temperature, not both. | |
| messages | Yes | Chat messages. Pure-text requests are accepted, but this tool is intended primarily for multimodal prompts containing image parts. | |
| max_tokens | No | Maximum number of tokens to generate. Input tokens plus this limit must fit within the model's context length. | |
| temperature | No | Sampling temperature: higher values make output more random; lower values make it more focused. Prefer adjusting this or top_p, not both. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | Yes | |
| model | Yes | |
| usage | No | |
| finish_reason | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered. The description adds useful non-structural context: the vision-model requirement, accepted model IDs, and that it returns assistant text plus token usage. It does not cover cost/latency or image size limits, so not a full 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then model list, then payload format, then routing guidance. The model enumeration is slightly verbose but each line is actionable and the structure is scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists so return values needn't be explained, and the description still notes the assistant-text + token-usage return. Model constraints, payload shapes, and sibling routing are all present; nothing needed to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description goes beyond by enumerating the accepted vision model IDs (the schema only says 'Vision-capable Mistral model') and clarifying the content part shapes and image URL formats.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Chat completion with multimodal input: text + image_url parts') and explicitly differentiates from the sibling mistral_chat for non-visual requests. An agent can distinguish this from mistral_ocr and mistral_chat without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes usage: 'Requires a vision-capable model' with an enumerated accepted-model list, and states 'For non-visual requests, prefer mistral_chat.' This names both the condition to use it and the alternative to use otherwise.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
process_documentProcess a business document end-to-endARead-onlyIdempotent
Single-call pipeline: provided text/Markdown or Mistral OCR → classify (if kind=auto) → typed extraction → validation. source.type=text skips OCR and Files uploads. Text with kind=generic makes no API calls; classification and typed extraction use Mistral chat. Results expose extraction_source. Provided text has null ocr_confidence and page_count; ocr_text contains the supplied text unchanged. Replaces the manual chain of mistral_ocr + mistral_chat + JSON parsing.
Kinds: contract | invoice | id_document | generic. Use kind=auto to let the server classify.
Returns a discriminated union — switch on kind to access typed fields.
Validation checks schema and, for OCR sources, OCR confidence; not factual or accounting accuracy.
Typed extraction rejects text longer than 60000 characters rather than truncating it.
Cache keys include source, kind, page limit, endpoint, models and pipeline version. Override location with MISTRAL_MCP_CACHE_DIR. Override mode with options.cache. Default cache mode is 'read_write' EXCEPT for kind=id_document (auto-bypass to avoid persisting PII). Set options.cache='read_write' explicitly to opt in for id documents.
options.maxPages and options.minOcrConfidence apply only to OCR sources. The confidence floor defaults to 0.3. Below the floor the
tool returns isError. Missing or partial confidence scores also return isError;
use mistral_ocr directly if you need raw OCR without a confidence guarantee.
0.3 is a conservative starting point, not a measured one: calibrate it for your
corpus with npm run eval:docs.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | Extraction task. auto classifies the document; generic returns text without typed extraction. Text source with generic makes no API calls. | auto |
| source | Yes | Already extracted text/Markdown, or an OCR source: remote URL, uploaded file ID, or inline image. | |
| options | No | Page selection, OCR confidence floor and local cache policy. |
Output Schema
| Name | Required | Description |
|---|---|---|
| dob | No | |
| kind | Yes | |
| name | No | |
| total | No | |
| expiry | No | |
| vendor | No | |
| clauses | No | |
| country | No | |
| parties | No | |
| summary | No | |
| currency | No | |
| due_date | No | |
| ocr_text | Yes | Text used for extraction: provided text unchanged, or Markdown returned by Mistral OCR. |
| anomalies | No | |
| cache_hit | Yes | |
| key_dates | No | |
| source_id | Yes | |
| line_items | No | |
| page_count | Yes | Pages processed by Mistral OCR. Null for provided text, whose pagination is unknown. |
| risk_score | No | |
| document_type | No | |
| ocr_confidence | Yes | Mean Mistral OCR page confidence. Null for provided text; never an extraction accuracy score. |
| structured_text | No | |
| pipeline_version | Yes | |
| extraction_source | Yes | How the input text was obtained. provided_text is supplied by the caller, not verified by OCR. |
| total_duration_ms | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, but the description goes well beyond them: cache key composition and the id_document PII auto-bypass, the isError conditions around OCR confidence, that validation is schema-only (not factual), and the 60000-character rejection behavior. This is unusually rich behavioral disclosure for a read-only processing tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the pipeline summary, then grouped by concerns (sources, kinds, returns, validation, cache, options). Length is defensible for a 3-param nested pipeline tool, but a few statements restate schema facts (the 60000 limit, the 0.3 floor, maxPages/minOcrConfidence scoping), which is mild redundancy given a fully described schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description still usefully explains the discriminated-union return (switch on kind, extraction_source, ocr_confidence semantics). Combined with the caveats about confidence floors and cache bypass, an agent has everything needed to call and interpret the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds framing the schema does not: that maxPages/minOcrConfidence apply only to OCR sources (and how they interact with provided text), that text sources yield null ocr_confidence/page_count, and that the 0.3 floor is a placeholder to calibrate. It adds value beyond the field-level docs, though several details (60000 chars, 0.3 default) are duplicated from the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete pipeline with specific stages (classify → typed extraction → validation) and the input modalities it accepts. It also explicitly positions itself against siblings: 'Replaces the manual chain of mistral_ocr + mistral_chat + JSON parsing.' An agent can distinguish it from mistral_ocr and mistral_chat without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit routing guidance: kind=auto lets the server classify, source.type=text skips OCR and Files uploads, text+generic makes no API calls, and 'use mistral_ocr directly if you need raw OCR without a confidence guarantee.' It names the alternative and the condition that selects it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voxtral_transcribeVoxtral speech-to-textARead-onlyIdempotent
Transcribe an audio file to text using Mistral Voxtral.
Accepted models:
voxtral-mini-latest
voxtral-small-latest
Audio source is one of:
{ type: "file_url", fileUrl: "https://..." } (public URL)
{ type: "file", fileId: "" }
Options:
language: ISO-639-1 hint (e.g. 'fr', 'en'). Boosts accuracy when known.temperature: sampling temperature.diarize: return per-speaker segments (default false).timestampGranularities: ['segment'] to return per-segment timestamps.contextBias: list of phrases/terms that should bias the decoder.
Returns plain text, detected language, optional segments[], and token usage.
| Name | Required | Description | Default |
|---|---|---|---|
| audio | Yes | Audio to transcribe, supplied as a public URL or an uploaded file ID. | |
| model | No | STT model. Default: voxtral-mini-latest. | |
| diarize | No | Identify speakers in the returned transcription segments. Defaults to false. | |
| language | No | ISO-639-1 language hint (e.g. 'fr', 'en'). | |
| contextBias | No | Words or phrases to favor when decoding the audio. | |
| temperature | No | Sampling temperature for transcription. | |
| timestampGranularities | No | Only 'segment' is currently supported. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | Yes | |
| model | Yes | |
| usage | No | |
| language | Yes | |
| segments | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds useful behavioral context beyond that: accepted model IDs, default values for diarize and model, the only-supported granularity, and the shape of the response (text, language, segments, token usage). It does not mention auth or rate limits, but that is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action, then organizes models, source options, options, and return values into scannable bullet groups. It is efficient, though the option bullets partially duplicate what the schema already documents with 100% coverage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema, full annotation coverage, and 100% schema description coverage, the description supplies everything an agent needs: models, input modes, option semantics, and defaults. No material gap remains for invoking this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description nonetheless adds meaning beyond the schema by explaining intent ('language ... boosts accuracy when known', 'contextBias: phrases that should bias the decoder', 'diarize: return per-speaker segments'), which helps the agent use the parameters correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence gives a specific verb and resource ('Transcribe an audio file to text'), which is unambiguous and clearly distinct from the sibling tools (chat, FIM, vision, OCR, document processing). An agent can immediately tell this is the audio speech-to-text tool without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly lays out the two mutually exclusive audio source modes (public URL vs. uploaded file ID) and when each applies, which is exactly the routing decision an agent must make. It does not name alternatives or exclusions for when *not* to transcribe, but the context is otherwise clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.8.3- Changed
codestral_fim14 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / max_tokens / descriptionAdded value: +"Maximum number of tokens to generate. Input tokens plus this limit must fit within the model's context length." - added
Input schema / properties / max_tokens / maximumAdded value: +9007199254740991 - added
Input schema / properties / model / descriptionAdded value: +"FIM (fill-in-the-middle) model identifier. Any identifier your endpoint serves is accepted — read mistral://models for the live catalog. Known Mistral aliases: codestral-latest." - removed
Input schema / properties / model / enumRemoved value: -[ - "codestral-latest" -] - added
Input schema / properties / model / maxLengthAdded value: +200 - added
Input schema / properties / model / minLengthAdded value: +1 - added
Input schema / properties / seed / maximumAdded value: +9007199254740991 - added
Input schema / properties / seed / minimumAdded value: +-9007199254740991 - added
Input schema / properties / stop / descriptionAdded value: +"Stop generation when any of these text sequences is encountered." - added
Input schema / properties / temperature / descriptionAdded value: +"Sampling temperature: higher values make output more random; lower values make it more focused. Prefer adjusting this or top_p, not both." - added
Input schema / properties / top_p / descriptionAdded value: +"Nucleus sampling probability mass: 0.1 considers tokens in the top 10% of probability mass. Prefer adjusting this or temperature, not both." - changed
Output schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mistral_chat19 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / max_tokens / descriptionAdded value: +"Maximum number of tokens to generate. Input tokens plus this limit must fit within the model's context length." - added
Input schema / properties / max_tokens / maximumAdded value: +9007199254740991 - removed
Input schema / properties / messages / items / additionalPropertiesRemoved value: -false - added
Input schema / properties / messages / items / properties / content / descriptionAdded value: +"Text of the message." - added
Input schema / properties / messages / items / properties / role / descriptionAdded value: +"Message author: system for instructions, user for requests, or assistant for prior replies." - changed
Input schema / properties / model / descriptionPrevious value: -"Mistral chat model alias. Allowed: mistral-large-latest, mistral-medium-latest, mistral-small-latest, ministral-3b-latest, ministral-8b-latest, ministral-14b-latest, magistral-medium-latest, magistral-small-latest, devstral-latest, devstral-small-latest, codestral-latest, voxtral-small-latest. Default: mistral-medium-latest."New value: +"Chat model. Default: mistral-medium-latest (override with MISTRAL_DEFAULT_MODEL)." - removed
Input schema / properties / model / enumRemoved value: -[ - "mistral-large-latest", - "mistral-medium-latest", - "mistral-small-latest", - "ministral-3b-latest", - "ministral-8b-latest", - "ministral-14b-latest", - "magistral-medium-latest", - "magistral-small-latest", - "devstral-latest", - "devstral-small-latest", - "codestral-latest", - "voxtral-small-latest" -] - added
Input schema / properties / model / maxLengthAdded value: +200 - added
Input schema / properties / model / minLengthAdded value: +1 - changed
Input schema / properties / reasoning_effort / descriptionPrevious value: -"Controls reasoning depth for Magistral models. 'high' enables full chain-of-thought; 'none' disables it. Ignored on non-reasoning models."New value: +"Reasoning effort; supported values depend on the selected model." - changed
Input schema / properties / reasoning_effort / enumPrevious value: -[ - "none", - "high" -]New value: +[ + "none", + "minimal", + "low", + "medium", + "high", + "xhigh" +] - changed
Input schema / properties / response_format / anyOfPrevious value: -[ - { - "additionalProperties": false, - "properties": { - "type": { - "const": "text", - "type": "string" - } - }, - "required": [ - "type" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "type": { - "const": "json_object", - "type": "string" - } - }, - "required": [ - "type" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "json_schema": { - "additionalProperties": false, - "properties": { - "description": { - "type": "string" - }, - "name": { - "description": "Identifier for the schema; surfaced in API errors.", - "maxLength": 64, - "minLength": 1, - "type": "string" - }, - "schema": { - "additionalProperties": {}, - "description": "JSON Schema object the response must conform to.", - "type": "object" - }, - "strict": { - "description": "If true, the API rejects responses that do not strictly match the schema.", - "type": "boolean" - } - }, - "required": [ - "name", - "schema" - ], - "type": "object" - }, - "type": { - "const": "json_schema", - "type": "string" - } - }, - "required": [ - "type", - "json_schema" - ], - "type": "object" - } -]New value: +[ + { + "properties": { + "type": { + "const": "text", + "description": "Generate plain text without a JSON format constraint.", + "type": "string" + } + }, + "required": [ + "type" + ], + "type": "object" + }, + { + "properties": { + "type": { + "const": "json_object", + "description": "Generate JSON. Also instruct the model to produce JSON in a system or user message.", + "type": "string" + } + }, + "required": [ + "type" + ], + "type": "object" + }, + { + "properties": { + "json_schema": { + "description": "Named JSON Schema and optional strictness for the generated response.", + "properties": { + "description": { + "description": "Description of the response the schema defines.", + "type": "string" + }, + "name": { + "description": "Identifier for the schema; surfaced in API errors.", + "maxLength": 64, + "minLength": 1, + "type": "string" + }, + "schema": { + "additionalProperties": {}, + "description": "JSON Schema object the response must conform to.", + "propertyNames": { + "type": "string" + }, + "type": "object" + }, + "strict": { + "description": "If true, the API rejects responses that do not strictly match the schema.", + "type": "boolean" + } + }, + "required": [ + "name", + "schema" + ], + "type": "object" + }, + "type": { + "const": "json_schema", + "description": "Generate JSON conforming to the supplied json_schema.", + "type": "string" + } + }, + "required": [ + "type", + "json_schema" + ], + "type": "object" + } +] - added
Input schema / properties / seed / maximumAdded value: +9007199254740991 - added
Input schema / properties / seed / minimumAdded value: +-9007199254740991 - added
Input schema / properties / temperature / descriptionAdded value: +"Sampling temperature: higher values make output more random; lower values make it more focused. Prefer adjusting this or top_p, not both." - added
Input schema / properties / top_p / descriptionAdded value: +"Nucleus sampling probability mass: 0.1 considers tokens in the top 10% of probability mass. Prefer adjusting this or temperature, not both." - changed
Output schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mistral_ocr49 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false - removed
Input schema / properties / bbox_annotation_format / additionalPropertiesRemoved value: -false - added
Input schema / properties / bbox_annotation_format / descriptionAdded value: +"JSON Schema for structured annotations of each extracted bounding box or image." - removed
Input schema / properties / bbox_annotation_format / properties / json_schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / bbox_annotation_format / properties / json_schema / descriptionAdded value: +"Named JSON Schema and optional strictness for the extracted annotation." - added
Input schema / properties / bbox_annotation_format / properties / json_schema / properties / description / descriptionAdded value: +"Description of the annotation to extract." - added
Input schema / properties / bbox_annotation_format / properties / json_schema / properties / name / descriptionAdded value: +"Name identifying the annotation schema." - added
Input schema / properties / bbox_annotation_format / properties / json_schema / properties / schema / descriptionAdded value: +"JSON Schema object defining the fields to extract into the annotation." - added
Input schema / properties / bbox_annotation_format / properties / json_schema / properties / schema / propertyNamesAdded value: +{ + "type": "string" +} - added
Input schema / properties / bbox_annotation_format / properties / json_schema / properties / strict / descriptionAdded value: +"Whether the annotation must strictly follow the supplied JSON Schema." - added
Input schema / properties / confidence_scores_granularity / descriptionAdded value: +"Confidence granularity. 'block' also fills pages[].blocks[].confidence_scores and requires OCR 4.1 (mistral-ocr-4-1) or newer." - changed
Input schema / properties / confidence_scores_granularity / enumPrevious value: -[ - "page", - "word" -]New value: +[ + "page", + "word", + "block" +] - changed
Input schema / properties / document / anyOfPrevious value: -[ - { - "additionalProperties": false, - "properties": { - "documentName": { - "type": "string" - }, - "documentUrl": { - "description": "HTTPS URL to a PDF or image.", - "type": "string" - }, - "type": { - "const": "document_url", - "type": "string" - } - }, - "required": [ - "type", - "documentUrl" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "imageUrl": { - "description": "HTTPS URL or data:image/...;base64,... payload.", - "type": "string" - }, - "type": { - "const": "image_url", - "type": "string" - } - }, - "required": [ - "type", - "imageUrl" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "fileId": { - "description": "ID of a file previously uploaded via the Files API.", - "type": "string" - }, - "type": { - "const": "file", - "type": "string" - } - }, - "required": [ - "type", - "fileId" - ], - "type": "object" - } -]New value: +[ + { + "properties": { + "documentName": { + "description": "Filename of the referenced document.", + "type": "string" + }, + "documentUrl": { + "description": "HTTPS URL to a PDF or image.", + "type": "string" + }, + "type": { + "const": "document_url", + "description": "Read the document from documentUrl.", + "type": "string" + } + }, + "required": [ + "type", + "documentUrl" + ], + "type": "object" + }, + { + "properties": { + "imageUrl": { + "description": "HTTPS URL or data:image/...;base64,... payload.", + "type": "string" + }, + "type": { + "const": "image_url", + "description": "Read the image from the URL or data URI in imageUrl.", + "type": "string" + } + }, + "required": [ + "type", + "imageUrl" + ], + "type": "object" + }, + { + "properties": { + "fileId": { + "description": "ID of a file previously uploaded via the Files API.", + "type": "string" + }, + "type": { + "const": "file", + "description": "Read a previously uploaded file identified by fileId.", + "type": "string" + } + }, + "required": [ + "type", + "fileId" + ], + "type": "object" + } +] - added
Input schema / properties / document / descriptionAdded value: +"Document to process, supplied as a document URL, an image URL or data URI, or an uploaded file ID." - removed
Input schema / properties / document_annotation_format / $refRemoved value: -"#/properties/bbox_annotation_format" - added
Input schema / properties / document_annotation_format / descriptionAdded value: +"JSON Schema for a structured annotation extracted from the entire document." - added
Input schema / properties / document_annotation_format / propertiesAdded value: +{ + "json_schema": { + "description": "Named JSON Schema and optional strictness for the extracted annotation.", + "properties": { + "description": { + "description": "Description of the annotation to extract.", + "type": "string" + }, + "name": { + "description": "Name identifying the annotation schema.", + "minLength": 1, + "type": "string" + }, + "schema": { + "additionalProperties": {}, + "description": "JSON Schema object defining the fields to extract into the annotation.", + "propertyNames": { + "type": "string" + }, + "type": "object" + }, + "strict": { + "description": "Whether the annotation must strictly follow the supplied JSON Schema.", + "type": "boolean" + } + }, + "required": [ + "name", + "schema" + ], + "type": "object" + }, + "type": { + "const": "json_schema", + "description": "Only json_schema is accepted by OCR annotation formats.", + "type": "string" + } +} - added
Input schema / properties / document_annotation_format / requiredAdded value: +[ + "type", + "json_schema" +] - added
Input schema / properties / document_annotation_format / typeAdded value: +"object" - added
Input schema / properties / document_annotation_prompt / descriptionAdded value: +"Instructions for whole-document structured extraction. Requires document_annotation_format." - added
Input schema / properties / extractFooter / descriptionAdded value: +"Extract each page footer into its footer field and remove it from the page markdown." - added
Input schema / properties / extractHeader / descriptionAdded value: +"Extract each page header into its header field and remove it from the page markdown." - added
Input schema / properties / imageLimit / descriptionAdded value: +"Maximum number of images to extract from the document." - added
Input schema / properties / imageLimit / maximumAdded value: +9007199254740991 - added
Input schema / properties / imageMinSize / descriptionAdded value: +"Minimum height and width of an image to extract." - added
Input schema / properties / imageMinSize / maximumAdded value: +9007199254740991 - added
Input schema / properties / includeBlocksAdded value: +{ + "description": "Return paragraph-level blocks (bounding box + type) per page. Requires OCR 4 (mistral-ocr-4-0) or newer.", + "type": "boolean" +} - added
Input schema / properties / includeImageBase64 / descriptionAdded value: +"Include base64-encoded data for extracted images in the response." - removed
Input schema / properties / model / enumRemoved value: -[ - "mistral-ocr-latest" -] - added
Input schema / properties / model / maxLengthAdded value: +200 - added
Input schema / properties / model / minLengthAdded value: +1 - changed
Input schema / properties / pages / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "items": { - "minimum": 0, - "type": "integer" - }, - "type": "array" - } -]New value: +[ + { + "type": "string" + }, + { + "items": { + "maximum": 9007199254740991, + "minimum": 0, + "type": "integer" + }, + "type": "array" + } +] - added
Input schema / properties / pages / descriptionAdded value: +"Zero-based page numbers to process, as an array or comma-separated numbers and ranges such as \"0-5,7\"." - added
Input schema / properties / tableFormat / descriptionAdded value: +"Format for extracted tables: Markdown or HTML." - changed
Output schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - added
Output schema / properties / annotations / properties / image_annotations / items / properties / page_index / maximumAdded value: +9007199254740991 - added
Output schema / properties / annotations / properties / image_annotations / items / properties / page_index / minimumAdded value: +-9007199254740991 - added
Output schema / properties / pages / items / properties / blocksAdded value: +{ + "description": "Paragraph-level blocks in reading order. Populated when `includeBlocks: true` and the model is OCR 4 or newer.", + "items": { + "additionalProperties": false, + "properties": { + "bottom_right_x": { + "type": "number" + }, + "bottom_right_y": { + "type": "number" + }, + "confidence_scores": { + "additionalProperties": false, + "description": "Populated only when `confidence_scores_granularity: 'block'`. Fields are null when the signal is absent (an image-only block has no content to score).", + "properties": { + "average_content_confidence_score": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ] + }, + "block_type_confidence_score": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ] + }, + "minimum_content_confidence_score": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ] + } + }, + "type": "object" + }, + "content": { + "type": "string" + }, + "image_id": { + "description": "Set on type:image — references the matching entry in `images[]`.", + "type": "string" + }, + "table_id": { + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "description": "Set on type:table — references the matching entry in `tables[]`." + }, + "top_left_x": { + "type": "number" + }, + "top_left_y": { + "type": "number" + }, + "type": { + "enum": [ + "text", + "title", + "list", + "table", + "image", + "equation", + "caption", + "code", + "references", + "aside_text", + "header", + "footer", + "signature" + ], + "type": "string" + } + }, + "required": [ + "type", + "top_left_x", + "top_left_y", + "bottom_right_x", + "bottom_right_y", + "content" + ], + "type": "object" + }, + "type": "array" +} - added
Output schema / properties / pages / items / properties / confidence_scores / properties / word_confidence_scores / items / properties / start_index / maximumAdded value: +9007199254740991 - added
Output schema / properties / pages / items / properties / confidence_scores / properties / word_confidence_scores / items / properties / start_index / minimumAdded value: +-9007199254740991 - added
Output schema / properties / pages / items / properties / footer / anyOfAdded value: +[ + { + "type": "string" + }, + { + "type": "null" + } +] - removed
Output schema / properties / pages / items / properties / footer / typeRemoved value: -[ - "string", - "null" -] - added
Output schema / properties / pages / items / properties / header / anyOfAdded value: +[ + { + "type": "string" + }, + { + "type": "null" + } +] - removed
Output schema / properties / pages / items / properties / header / typeRemoved value: -[ - "string", - "null" -] - added
Output schema / properties / pages / items / properties / index / maximumAdded value: +9007199254740991 - added
Output schema / properties / pages / items / properties / index / minimumAdded value: +-9007199254740991 - added
Output schema / properties / pages_count / maximumAdded value: +9007199254740991 - added
Output schema / properties / pages_count / minimumAdded value: +-9007199254740991
- Changed
mistral_vision16 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / max_tokens / descriptionAdded value: +"Maximum number of tokens to generate. Input tokens plus this limit must fit within the model's context length." - added
Input schema / properties / max_tokens / maximumAdded value: +9007199254740991 - removed
Input schema / properties / messages / items / additionalPropertiesRemoved value: -false - changed
Input schema / properties / messages / items / properties / content / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "items": { - "anyOf": [ - { - "additionalProperties": false, - "properties": { - "text": { - "type": "string" - }, - "type": { - "const": "text", - "type": "string" - } - }, - "required": [ - "type", - "text" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "imageUrl": { - "anyOf": [ - { - "description": "https URL or data:image/...;base64,... payload", - "type": "string" - }, - { - "additionalProperties": false, - "properties": { - "detail": { - "enum": [ - "auto", - "low", - "high" - ], - "type": "string" - }, - "url": { - "type": "string" - } - }, - "required": [ - "url" - ], - "type": "object" - } - ] - }, - "type": { - "const": "image_url", - "type": "string" - } - }, - "required": [ - "type", - "imageUrl" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "documentName": { - "type": "string" - }, - "documentUrl": { - "type": "string" - }, - "type": { - "const": "document_url", - "type": "string" - } - }, - "required": [ - "type", - "documentUrl" - ], - "type": "object" - } - ] - }, - "minItems": 1, - "type": "array" - } -]New value: +[ + { + "type": "string" + }, + { + "items": { + "anyOf": [ + { + "properties": { + "text": { + "description": "Text to include in the message.", + "type": "string" + }, + "type": { + "const": "text", + "description": "Identifies a text content part.", + "type": "string" + } + }, + "required": [ + "type", + "text" + ], + "type": "object" + }, + { + "properties": { + "imageUrl": { + "anyOf": [ + { + "description": "https URL or data:image/...;base64,... payload", + "type": "string" + }, + { + "properties": { + "detail": { + "description": "Image detail hint: automatic, low, or high.", + "enum": [ + "auto", + "low", + "high" + ], + "type": "string" + }, + "url": { + "description": "HTTPS URL or data:image/...;base64,... payload.", + "type": "string" + } + }, + "required": [ + "url" + ], + "type": "object" + } + ], + "description": "Image source as a URL or base64 data URI, optionally with a detail hint." + }, + "type": { + "const": "image_url", + "description": "Identifies an image content part.", + "type": "string" + } + }, + "required": [ + "type", + "imageUrl" + ], + "type": "object" + }, + { + "properties": { + "documentName": { + "description": "Filename of the referenced document.", + "type": "string" + }, + "documentUrl": { + "description": "URL of the PDF or document to include in the message.", + "type": "string" + }, + "type": { + "const": "document_url", + "description": "Identifies a document content part.", + "type": "string" + } + }, + "required": [ + "type", + "documentUrl" + ], + "type": "object" + } + ] + }, + "minItems": 1, + "type": "array" + } +] - added
Input schema / properties / messages / items / properties / content / descriptionAdded value: +"Message text or an ordered list of text, image, and document content parts." - added
Input schema / properties / messages / items / properties / role / descriptionAdded value: +"Message author: system for instructions, user for requests, or assistant for prior replies." - removed
Input schema / properties / model / enumRemoved value: -[ - "pixtral-large-latest", - "pixtral-12b-latest", - "mistral-large-latest", - "mistral-medium-latest", - "mistral-small-latest" -] - added
Input schema / properties / model / maxLengthAdded value: +200 - added
Input schema / properties / model / minLengthAdded value: +1 - added
Input schema / properties / seed / maximumAdded value: +9007199254740991 - added
Input schema / properties / seed / minimumAdded value: +-9007199254740991 - added
Input schema / properties / temperature / descriptionAdded value: +"Sampling temperature: higher values make output more random; lower values make it more focused. Prefer adjusting this or top_p, not both." - added
Input schema / properties / top_p / descriptionAdded value: +"Nucleus sampling probability mass: 0.1 considers tokens in the top 10% of probability mass. Prefer adjusting this or temperature, not both." - changed
Output schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Added
process_document - Changed
voxtral_transcribe13 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false - changed
Input schema / properties / audio / anyOfPrevious value: -[ - { - "additionalProperties": false, - "properties": { - "fileUrl": { - "description": "HTTPS URL to an audio file (mp3/wav/flac/ogg/webm/m4a).", - "type": "string" - }, - "type": { - "const": "file_url", - "type": "string" - } - }, - "required": [ - "type", - "fileUrl" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "fileId": { - "description": "ID of an audio file previously uploaded via the Files API (purpose=audio).", - "type": "string" - }, - "type": { - "const": "file", - "type": "string" - } - }, - "required": [ - "type", - "fileId" - ], - "type": "object" - } -]New value: +[ + { + "properties": { + "fileUrl": { + "description": "HTTPS URL to an audio file (mp3/wav/flac/ogg/webm/m4a).", + "type": "string" + }, + "type": { + "const": "file_url", + "description": "Transcribe audio from the URL in fileUrl.", + "type": "string" + } + }, + "required": [ + "type", + "fileUrl" + ], + "type": "object" + }, + { + "properties": { + "fileId": { + "description": "ID of an audio file previously uploaded via the Files API (purpose=audio).", + "type": "string" + }, + "type": { + "const": "file", + "description": "Transcribe an uploaded audio file identified by fileId.", + "type": "string" + } + }, + "required": [ + "type", + "fileId" + ], + "type": "object" + } +] - added
Input schema / properties / audio / descriptionAdded value: +"Audio to transcribe, supplied as a public URL or an uploaded file ID." - added
Input schema / properties / contextBias / descriptionAdded value: +"Words or phrases to favor when decoding the audio." - added
Input schema / properties / diarize / descriptionAdded value: +"Identify speakers in the returned transcription segments. Defaults to false." - removed
Input schema / properties / model / enumRemoved value: -[ - "voxtral-mini-latest", - "voxtral-small-latest" -] - added
Input schema / properties / model / maxLengthAdded value: +200 - added
Input schema / properties / model / minLengthAdded value: +1 - added
Input schema / properties / temperature / descriptionAdded value: +"Sampling temperature for transcription." - changed
Output schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - added
Output schema / properties / language / anyOfAdded value: +[ + { + "type": "string" + }, + { + "type": "null" + } +] - removed
Output schema / properties / language / typeRemoved value: -[ - "string", - "null" -]
- Removed
workflow_execute - Removed
workflow_interact - Removed
workflow_status
24 tool updates
v0.7.0- Removed
batch_cancel - Removed
batch_create - Removed
batch_get - Removed
batch_list - Changed
codestral_fim1 field changed- added
Input schema / properties / seedAdded value: +{ + "description": "Random seed for deterministic sampling. Maps to Mistral's `random_seed`.", + "type": "integer" +}
- Removed
files_delete - Removed
files_get - Removed
files_list - Removed
files_signed_url - Removed
files_upload - Removed
mcp_sample - Removed
mistral_agent - Changed
mistral_chat4 fields changed- added
Input schema / properties / reasoning_effortAdded value: +{ + "description": "Controls reasoning depth for Magistral models. 'high' enables full chain-of-thought; 'none' disables it. Ignored on non-reasoning models.", + "enum": [ + "none", + "high" + ], + "type": "string" +} - added
Input schema / properties / response_formatAdded value: +{ + "anyOf": [ + { + "additionalProperties": false, + "properties": { + "type": { + "const": "text", + "type": "string" + } + }, + "required": [ + "type" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "type": { + "const": "json_object", + "type": "string" + } + }, + "required": [ + "type" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "json_schema": { + "additionalProperties": false, + "properties": { + "description": { + "type": "string" + }, + "name": { + "description": "Identifier for the schema; surfaced in API errors.", + "maxLength": 64, + "minLength": 1, + "type": "string" + }, + "schema": { + "additionalProperties": {}, + "description": "JSON Schema object the response must conform to.", + "type": "object" + }, + "strict": { + "description": "If true, the API rejects responses that do not strictly match the schema.", + "type": "boolean" + } + }, + "required": [ + "name", + "schema" + ], + "type": "object" + }, + "type": { + "const": "json_schema", + "type": "string" + } + }, + "required": [ + "type", + "json_schema" + ], + "type": "object" + } + ], + "description": "Force a structured output: `{type:\"json_object\"}` for JSON mode, `{type:\"json_schema\", json_schema:{...}}` for strict schema mode." +} - added
Input schema / properties / seedAdded value: +{ + "description": "Random seed for deterministic sampling. Maps to Mistral's `random_seed`.", + "type": "integer" +} - added
Output schema / properties / reasoning_contentAdded value: +{ + "description": "Reasoning trace returned by Magistral models. Absent for non-reasoning models.", + "type": "string" +}
- Removed
mistral_chat_stream - Removed
mistral_classify - Removed
mistral_embed - Removed
mistral_moderate - Changed
mistral_ocr7 fields changed- added
Input schema / properties / bbox_annotation_formatAdded value: +{ + "additionalProperties": false, + "properties": { + "json_schema": { + "additionalProperties": false, + "properties": { + "description": { + "type": "string" + }, + "name": { + "minLength": 1, + "type": "string" + }, + "schema": { + "additionalProperties": {}, + "type": "object" + }, + "strict": { + "type": "boolean" + } + }, + "required": [ + "name", + "schema" + ], + "type": "object" + }, + "type": { + "const": "json_schema", + "description": "Only json_schema is accepted by OCR annotation formats.", + "type": "string" + } + }, + "required": [ + "type", + "json_schema" + ], + "type": "object" +} - added
Input schema / properties / confidence_scores_granularityAdded value: +{ + "enum": [ + "page", + "word" + ], + "type": "string" +} - added
Input schema / properties / document_annotation_formatAdded value: +{ + "$ref": "#/properties/bbox_annotation_format" +} - added
Input schema / properties / document_annotation_promptAdded value: +{ + "type": "string" +} - added
Output schema / properties / annotationsAdded value: +{ + "additionalProperties": false, + "properties": { + "document_annotation": { + "type": "string" + }, + "image_annotations": { + "items": { + "additionalProperties": false, + "properties": { + "annotation": { + "type": "string" + }, + "bbox": { + "additionalProperties": false, + "properties": { + "bottom_right_x": { + "type": "number" + }, + "bottom_right_y": { + "type": "number" + }, + "top_left_x": { + "type": "number" + }, + "top_left_y": { + "type": "number" + } + }, + "type": "object" + }, + "image_id": { + "type": "string" + }, + "page_index": { + "type": "integer" + } + }, + "required": [ + "page_index", + "annotation" + ], + "type": "object" + }, + "type": "array" + } + }, + "type": "object" +} - added
Output schema / properties / pages / items / properties / confidence_scoresAdded value: +{ + "additionalProperties": false, + "properties": { + "average_page_confidence_score": { + "type": "number" + }, + "minimum_page_confidence_score": { + "type": "number" + }, + "word_confidence_scores": { + "items": { + "additionalProperties": false, + "properties": { + "confidence": { + "type": "number" + }, + "start_index": { + "type": "integer" + }, + "text": { + "type": "string" + } + }, + "required": [ + "text", + "confidence", + "start_index" + ], + "type": "object" + }, + "type": "array" + } + }, + "required": [ + "average_page_confidence_score", + "minimum_page_confidence_score" + ], + "type": "object" +} - added
Output schema / properties / pages / items / properties / images / items / properties / image_annotationAdded value: +{ + "type": "string" +}
- Removed
mistral_tool_call - Changed
mistral_vision1 field changed- added
Input schema / properties / seedAdded value: +{ + "description": "Random seed for deterministic sampling. Maps to Mistral's `random_seed`.", + "type": "integer" +}
- Removed
voxtral_speak - Added
workflow_execute - Added
workflow_interact - Added
workflow_status
22 tool updates
v0.4.0- First observed
batch_cancel - First observed
batch_create - First observed
batch_get - First observed
batch_list - First observed
codestral_fim - First observed
files_delete - First observed
files_get - First observed
files_list - First observed
files_signed_url - First observed
files_upload - First observed
mcp_sample - First observed
mistral_agent - First observed
mistral_chat - First observed
mistral_chat_stream - First observed
mistral_classify - First observed
mistral_embed - First observed
mistral_moderate - First observed
mistral_ocr - First observed
mistral_tool_call - First observed
mistral_vision - First observed
voxtral_speak - First observed
voxtral_transcribe
TDQS
Scored across 6 tools
Each tool targets a distinct modality/action: code FIM, text chat, OCR, vision chat, document pipeline, and audio transcription. There is minor overlap between mistral_chat and mistral_vision (both chat) and between process_document and mistral_ocr, but the descriptions explicitly steer selection, so confusion is limited.
Names mix provider-prefixed conventions (codestral_fim, mistral_chat, mistral_ocr, mistral_vision, voxtral_transcribe) with a plain action name (process_document). The pattern is somewhat readable but not a predictable verb_noun scheme and the provider prefixes fragment consistency.
Six tools is well-scoped for a multimodal Mistral wrapper, with each tool covering a distinct capability rather than redundancy. Slightly broad for a server named 'Document Extraction' since it also includes code completion and chat, but not excessive.
Covers OCR, vision, chat, transcription, and a combined pipeline, but several tools reference a Files API (fileId) that is not exposed as a tool, creating a dead-end for uploading files. No model-listing or file-management operations leave workflows partially blocked.
Maintenance
Related MCP Connectors
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…
MCP server for progressive tool usage at any scale (see https://klavis.ai)
- QuallaaOAuthcom.quallaa
Talk to your public-facing AI from any MCP client — Claude, ChatGPT, Cursor, Cline, Windsurf.
Nifty's MCP server — exposes tasks, projects, messages, and files as tools for AI agents.
Related MCP Servers
- -licenseNot gradedqualityDmaintenanceA TypeScript implementation of a Model Context Protocol server and client that enables interaction with language models (specifically Mistral running on Ollama).-
- -licenseNot gradedqualityNot gradedmaintenanceA production-ready TypeScript MCP server providing basic tools (add, echo, timestamp), resources (server info, greetings, data access), and prompt templates (analyze, code-review, summarize). Serves as a foundation for building custom MCP servers with extensible architecture.398 npm-
- AlicenseBqualityCmaintenanceA TypeScript-based MCP server that provides tools to interact with local Codex and Gemini CLIs via stdio transport. It enables users to execute prompts through the ask_codex and ask_gemini tools, supporting custom models and timeout configurations.110 npm6MIT
- FlicenseAqualityDmaintenanceA TypeScript MCP server template with Zod validation, dual transport (stdio/HTTP), and modular architecture for building MCP-compatible tools, resources, and prompts.11-