Mistral MCP — Document Extraction
mistral-mcp
Mistral AI 기능을 모든 MCP 클라이언트(Claude Code, Cursor, Zed, Windsurf, Claude Desktop)에 노출하는 MCP 서버
프랑스어 버전: README.fr.md
목적
Mistral은 프랑스어, 코드, OCR, 조정, 오디오 및 에이전트 스타일 워크플로우를 위한 강력한 모델을 보유하고 있지만, 대부분의 MCP 지원 IDE는 Anthropic이나 OpenAI를 기본으로 사용합니다. mistral-mcp는 이러한 Mistral 기능을 깔끔한 MCP 인터페이스로 제공하여, 에이전트 루프를 다시 구축할 필요 없이 적절한 하위 작업을 적절한 모델로 라우팅할 수 있게 합니다.
이 저장소의 목표는 "또 하나의 얇은 래퍼"가 되는 것이 아닙니다. 명확한 스키마, 예측 가능한 출력, 전송 유연성 및 우수한 테스트 커버리지를 갖춘 강력하고 유지 관리 가능한 MCP 서버를 지향합니다.
Related MCP server: MCP Server TypeScript
현재 인터페이스 (v0.4.0)
도구 (22)
핵심 생성:
mistral_chatmistral_chat_streammistral_embedmistral_tool_callcodestral_fim
비전 및 오디오:
mistral_visionmistral_ocrvoxtral_transcribevoxtral_speak
에이전트 및 분류기:
mistral_agentmistral_moderatemistral_classify
파일 및 배치:
files_uploadfiles_listfiles_getfiles_deletefiles_signed_urlbatch_createbatch_listbatch_getbatch_cancel
MCP 네이티브 유틸리티:
mcp_sample- MCP 샘플링을 통해 클라이언트 모델에 생성 위임
리소스 (2)
mistral://models- 허용된 별칭 및 라이브 모델 카탈로그mistral://voices- Voxtral TTS를 위한 라이브 음성 카탈로그
프롬프트 (6)
프랑스어 큐레이팅 프롬프트:
french_invoice_reminderfrench_meeting_minutesfrench_email_replyfrench_commit_messagefrench_legal_summary
영어 큐레이팅 프롬프트:
codestral_review
프롬프트 열거형 인수는 completable()로 래핑되어 있어, MCP 클라이언트가 completion/complete를 통해 프롬프트 인수 완성을 호출할 수 있습니다.
주요 특징
모든 도구에
inputSchema,outputSchema및 주석이 포함된 고수준McpServerAPI이중 전송 지원: 기본적으로 stdio, 원격 배포를 위한 Streamable HTTP
모든 곳에서 구조화된 출력:
structuredContent및 텍스트 대체mcp_sample을 통한 MCP 샘플링 지원열거형 프롬프트 인수에 대한 프롬프트 완성 지원
나중에 추가된 것이 아니라 도구와 함께 등록된 리소스 및 프롬프트
Mistral SDK 클라이언트의 재시도/백오프 및 요청 시간 초과
전송
Stdio
기본 모드입니다. Claude Code 및 대부분의 로컬 MCP 클라이언트가 사용하는 방식입니다.
node dist/index.jsStreamable HTTP
--http 또는 MCP_TRANSPORT=http로 활성화합니다.
MCP_TRANSPORT=http node dist/index.js관련 환경 변수:
MCP_HTTP_HOST- 기본값127.0.0.1MCP_HTTP_PORT- 기본값3333MCP_HTTP_PATH- 기본값/mcpMCP_HTTP_TOKEN- 선택적 베어러 토큰MCP_HTTP_ALLOWED_ORIGINS- 선택적 쉼표로 구분된 허용 목록MCP_HTTP_STATELESS=1- 상태 비저장 세션 모드
/healthz는 의도적으로 공개되어 있으며 MCP 서버에 접근하지 않습니다.
설치
git clone https://github.com/Swih/mistral-mcp.git
cd mistral-mcp
npm install
npm run buildAPI 키 설정:
export MISTRAL_API_KEY=your_key_here또는 저장소 루트의 .env를 사용하세요. 절대 커밋하지 마십시오.
Claude Code에서 사용
claude mcp add mistral -- node /absolute/path/to/mistral-mcp/dist/index.js예시 프롬프트:
이 PDF에
mistral_ocr을 사용한 다음, 추출된 텍스트에french_meeting_minutes를 실행하세요.
개발
npm run dev
npm run build
npm run lint
npm test
npm run inspector테스트 전략
현재 테스트 모음은 4개 계층에 걸쳐 148개의 테스트를 포함합니다:
도구, 리소스, 프롬프트, 전송, 오디오, 에이전트, 파일, 배치 및 샘플링에 대한 단위 테스트
도구 메타데이터 및 MCP 보장에 대한 계약 테스트
MISTRAL_API_KEY가 설정되었을 때 실제 Mistral API에 대한 라이브 API 테스트빌드된 서버에 대한 Stdio 엔드투엔드 테스트
MISTRAL_API_KEY가 없으면 로컬 기본값은 139개 통과 및 9개 게이트 라이브/stdio 테스트입니다.
프로젝트 레이아웃
mistral-mcp/
|-- src/
| |-- index.ts
| |-- transport.ts
| |-- tools.ts
| |-- tools-fn.ts
| |-- tools-vision.ts
| |-- tools-audio.ts
| |-- tools-agents.ts
| |-- tools-files.ts
| |-- tools-batch.ts
| |-- tools-sampling.ts
| |-- resources.ts
| `-- prompts.ts
|-- test/
|-- examples/
|-- .github/workflows/ci.yml
|-- package.json
`-- tsconfig.test.json상태
v0.4.0 — 출시됨. v0.3.0 대비 전체 변경 사항은 CHANGELOG.md를 참조하세요:
공유 헬퍼, 라이브 모델 + 음성 카탈로그, 계약 테스트
비전 + OCR
오디오 전사 + 음성
에이전트 + 조정 + 분류
파일 + 배치 API
Streamable HTTP 전송 + MCP 샘플링
프랑스어 큐레이팅 프롬프트 5개 + 영어 프롬프트 1개 + 프롬프트 인수 완성
예시
실행 가능한 스크립트는 examples/에 있습니다. examples/README.md를 참조하세요.
라이선스
MIT Copyright Dayan Decamp
Available Tools
6 toolscodestral_fimCodestral fill-in-the-middle completionARead-only
Fill-in-the-middle code completion with Codestral.
Given prompt (code preceding the cursor) and suffix (code after the cursor),
Codestral writes the middle. Use for editor autocomplete scenarios, code-patching
agents, or structured refactors where you know the target boundaries.
Default stop tokens: [] — let the model decide. Override with stop if needed.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Random seed for deterministic sampling. Maps to Mistral's `random_seed`. | |
| stop | No | Stop generation when any of these text sequences is encountered. | |
| model | No | FIM (fill-in-the-middle) model identifier. Any identifier your endpoint serves is accepted — read mistral://models for the live catalog. Known Mistral aliases: codestral-latest. | |
| top_p | No | Nucleus sampling probability mass: 0.1 considers tokens in the top 10% of probability mass. Prefer adjusting this or temperature, not both. | |
| prompt | Yes | Code preceding the cursor. | |
| suffix | Yes | Code after the cursor. Can be empty string. | |
| max_tokens | No | Maximum number of tokens to generate. Input tokens plus this limit must fit within the model's context length. | |
| temperature | No | Sampling temperature: higher values make output more random; lower values make it more focused. Prefer adjusting this or top_p, not both. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | Yes | |
| model | Yes | |
| usage | No | |
| finish_reason | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=false, and destructiveHint=false. The description adds useful generation behavior by noting the default stop tokens are empty and that the model decides when to stop unless `stop` is overridden.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with a one-line purpose, followed by cursor semantics, use cases, and stop-token behavior. Every sentence earns its place and no filler is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given rich annotations, 100% schema coverage, and an output schema, the description supplies the missing conceptual model: prompt before cursor, suffix after cursor, and Codestral writes the middle. It is complete enough for correct invocation without redundant return-value explanation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so prompt/suffix meanings are already documented. The description adds extra semantics for `stop`: default is [] and the model decides, with explicit override guidance, going beyond the schema's brief stop description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: fill-in-the-middle code completion with Codestral, plus the prompt/suffix cursor model. The FIM and editor-autocomplete framing distinguishes it from sibling chat, vision, OCR, transcription, and document tools without needing schema inspection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear usage contexts: editor autocomplete, code-patching agents, and structured refactors where target boundaries are known. It does not explicitly name when to avoid this tool or route to a sibling like mistral_chat for general generation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mistral_chatMistral chat completionARead-only
Generate a chat completion using a Mistral model.
When to use:
Drafting French (or any European-language) content where Mistral shines.
Codestral for code-specific generation/review.
Ministral for cheap / low-latency classification.
Returns structured content with the assistant text and token usage. Does NOT stream — use mistral_chat_stream for long outputs with progress updates.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Random seed for deterministic sampling. Maps to Mistral's `random_seed`. | |
| model | No | Chat model. Default: mistral-medium-latest (override with MISTRAL_DEFAULT_MODEL). | |
| top_p | No | Nucleus sampling probability mass: 0.1 considers tokens in the top 10% of probability mass. Prefer adjusting this or temperature, not both. | |
| messages | Yes | Chat messages in role/content form. | |
| max_tokens | No | Maximum number of tokens to generate. Input tokens plus this limit must fit within the model's context length. | |
| temperature | No | Sampling temperature: higher values make output more random; lower values make it more focused. Prefer adjusting this or top_p, not both. | |
| response_format | No | Force a structured output: `{type:"json_object"}` for JSON mode, `{type:"json_schema", json_schema:{...}}` for strict schema mode. | |
| reasoning_effort | No | Reasoning effort; supported values depend on the selected model. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | Yes | |
| model | Yes | |
| usage | No | |
| finish_reason | No | |
| reasoning_content | No | Reasoning trace returned by Magistral models. Absent for non-reasoning models. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover safety (readOnlyHint=true, destructiveHint=false, openWorldHint=true), so the bar is lower. The description still adds useful non-safety behavior: it does not stream, and it returns structured content with assistant text and token usage. It does not mention rate limits, latency, or error behavior, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action, organizes guidance as scannable bullets, and closes with the return shape and the streaming exclusion. No filler sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description needn't detail return values, and annotations cover safety. It still supplies the streaming caveat and model-family guidance. Minor gap: it never states default model behavior or auth expectations, but that is largely handled by the schema's model description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents seed, model, top_p, temperature, response_format, and reasoning_effort thoroughly. The description adds only indirect model-selection guidance ('Codestral', 'Ministral') rather than explaining parameters, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening line states a specific verb and resource: 'Generate a chat completion using a Mistral model.' The bullets further differentiate sub-cases (French/European drafting, Codestral for code, Ministral for cheap classification), letting an agent distinguish this from codestral_fim and other siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit 'When to use' bullets plus a named exclusion: 'Does NOT stream — use mistral_chat_stream for long outputs with progress updates.' This gives both positive triggers and a concrete alternative with the condition that selects it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mistral_ocrMistral OCR (document to markdown)ARead-onlyIdempotent
Run Mistral OCR on a PDF or image, returning structured markdown per page.
Input document is one of:
{ type: "document_url", documentUrl: "https://...pdf" }
{ type: "image_url", imageUrl: "https://..." | "data:image/..." }
{ type: "file", fileId: "" }
Options:
pages: array of 0-indexed page numbers or string like "0-5,7".tableFormat: 'markdown' (default) or 'html'.extractHeader/extractFooter: include page header/footer when present.includeImageBase64: embed extracted image bytes as base64 in the response.document_annotation_format: JSON schema for whole-document structured extraction.bbox_annotation_format: JSON schema for extracted image / bbox annotations.confidence_scores_granularity: 'page', 'word', or 'block'. 'block' adds per-block content/type confidence underpages[].blocks[].confidence_scoresand requires OCR 4.1 or newer.includeBlocks: return paragraph-level blocks (bounding box + type) in reading order — titles, lists, tables, images, equations, captions, code, references, aside text, header, footer, signature. Requires OCR 4 (mistral-ocr-4-0) or newer; older models accept the flag but return an emptyblocksarray.
Returns pages[].markdown plus optional pages[].hyperlinks, header, footer,
images bounding boxes, blocks, annotations, confidence scores, and dimensions.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | OCR model. Default: mistral-ocr-latest. | |
| pages | No | Zero-based page numbers to process, as an array or comma-separated numbers and ranges such as "0-5,7". | |
| document | Yes | Document to process, supplied as a document URL, an image URL or data URI, or an uploaded file ID. | |
| imageLimit | No | Maximum number of images to extract from the document. | |
| tableFormat | No | Format for extracted tables: Markdown or HTML. | |
| imageMinSize | No | Minimum height and width of an image to extract. | |
| extractFooter | No | Extract each page footer into its footer field and remove it from the page markdown. | |
| extractHeader | No | Extract each page header into its header field and remove it from the page markdown. | |
| includeBlocks | No | Return paragraph-level blocks (bounding box + type) per page. Requires OCR 4 (mistral-ocr-4-0) or newer. | |
| includeImageBase64 | No | Include base64-encoded data for extracted images in the response. | |
| bbox_annotation_format | No | JSON Schema for structured annotations of each extracted bounding box or image. | |
| document_annotation_format | No | JSON Schema for a structured annotation extracted from the entire document. | |
| document_annotation_prompt | No | Instructions for whole-document structured extraction. Requires document_annotation_format. | |
| confidence_scores_granularity | No | Confidence granularity. 'block' also fills pages[].blocks[].confidence_scores and requires OCR 4.1 (mistral-ocr-4-1) or newer. |
Output Schema
| Name | Required | Description |
|---|---|---|
| model | Yes | |
| pages | Yes | |
| usage | No | |
| annotations | No | |
| pages_count | Yes | |
| document_annotation | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (which only declare read-only, idempotent, non-destructive, open-world), the description discloses real behavioral preconditions: includeBlocks requires OCR 4+, confidence_scores_granularity='block' requires OCR 4.1+, and older models silently accept includeBlocks but return an empty blocks array. That silent-failure warning is exactly the kind of context an agent cannot get from the schema or annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with a one-line purpose, then cleanly sectioned into input shapes, options, and returns for a 14-parameter tool. The length is proportionate to the surface area and each bullet maps to a decision the caller must make.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only OCR tool with a full output schema and 100% schema coverage, the description supplies everything else needed: the three mutually exclusive document input forms, option semantics, model-version gates, and the shape of the response. Nothing an agent needs to invoke it correctly is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds genuine value on several parameters the schema documents only tersely: the annotated string form of `pages` ('0-5,7'), the default of tableFormat, the distinction between document_annotation_format and bbox_annotation_format, and the version requirements tied to includeBlocks and confidence granularity. Some parameters (imageLimit, imageMinSize, model) are untouched, which keeps it from a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('Run Mistral OCR on a PDF or image') and states the core output ('returning structured markdown per page'). It is clear, but it never distinguishes itself from siblings like mistral_vision or process_document, so an agent cannot tell from the description alone which of those overlapping tools to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the enumerated input shapes and option list, but there is no explicit when-to-use or when-not-to-use guidance, and no routing to alternative siblings such as mistral_vision for image-only OCR. The reader must infer the fit from the parameter list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mistral_visionMistral multimodal chat (vision)ARead-only
Chat completion with multimodal input: text + image_url parts.
Requires a vision-capable model. Accepted:
pixtral-large-latest
pixtral-12b-latest
mistral-large-latest
mistral-medium-latest
mistral-small-latest
Each message's content is either a plain string (pure text) or an array of
parts { type: 'text', text } / { type: 'image_url', imageUrl }. The image URL
can be an https URL or a data: URI base64 payload.
Returns the assistant text + token usage. For non-visual requests, prefer mistral_chat.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Random seed for deterministic sampling. Maps to Mistral's `random_seed`. | |
| model | No | Vision-capable Mistral model. Default: pixtral-large-latest. | |
| top_p | No | Nucleus sampling probability mass: 0.1 considers tokens in the top 10% of probability mass. Prefer adjusting this or temperature, not both. | |
| messages | Yes | Chat messages. Pure-text requests are accepted, but this tool is intended primarily for multimodal prompts containing image parts. | |
| max_tokens | No | Maximum number of tokens to generate. Input tokens plus this limit must fit within the model's context length. | |
| temperature | No | Sampling temperature: higher values make output more random; lower values make it more focused. Prefer adjusting this or top_p, not both. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | Yes | |
| model | Yes | |
| usage | No | |
| finish_reason | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered. The description adds useful non-structural context: the vision-model requirement, accepted model IDs, and that it returns assistant text plus token usage. It does not cover cost/latency or image size limits, so not a full 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then model list, then payload format, then routing guidance. The model enumeration is slightly verbose but each line is actionable and the structure is scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists so return values needn't be explained, and the description still notes the assistant-text + token-usage return. Model constraints, payload shapes, and sibling routing are all present; nothing needed to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description goes beyond by enumerating the accepted vision model IDs (the schema only says 'Vision-capable Mistral model') and clarifying the content part shapes and image URL formats.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Chat completion with multimodal input: text + image_url parts') and explicitly differentiates from the sibling mistral_chat for non-visual requests. An agent can distinguish this from mistral_ocr and mistral_chat without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes usage: 'Requires a vision-capable model' with an enumerated accepted-model list, and states 'For non-visual requests, prefer mistral_chat.' This names both the condition to use it and the alternative to use otherwise.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
process_documentProcess a business document end-to-endARead-onlyIdempotent
Single-call pipeline: provided text/Markdown or Mistral OCR → classify (if kind=auto) → typed extraction → validation. source.type=text skips OCR and Files uploads. Text with kind=generic makes no API calls; classification and typed extraction use Mistral chat. Results expose extraction_source. Provided text has null ocr_confidence and page_count; ocr_text contains the supplied text unchanged. Replaces the manual chain of mistral_ocr + mistral_chat + JSON parsing.
Kinds: contract | invoice | id_document | generic. Use kind=auto to let the server classify.
Returns a discriminated union — switch on kind to access typed fields.
Validation checks schema and, for OCR sources, OCR confidence; not factual or accounting accuracy.
Typed extraction rejects text longer than 60000 characters rather than truncating it.
Cache keys include source, kind, page limit, endpoint, models and pipeline version. Override location with MISTRAL_MCP_CACHE_DIR. Override mode with options.cache. Default cache mode is 'read_write' EXCEPT for kind=id_document (auto-bypass to avoid persisting PII). Set options.cache='read_write' explicitly to opt in for id documents.
options.maxPages and options.minOcrConfidence apply only to OCR sources. The confidence floor defaults to 0.3. Below the floor the
tool returns isError. Missing or partial confidence scores also return isError;
use mistral_ocr directly if you need raw OCR without a confidence guarantee.
0.3 is a conservative starting point, not a measured one: calibrate it for your
corpus with npm run eval:docs.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | Extraction task. auto classifies the document; generic returns text without typed extraction. Text source with generic makes no API calls. | auto |
| source | Yes | Already extracted text/Markdown, or an OCR source: remote URL, uploaded file ID, or inline image. | |
| options | No | Page selection, OCR confidence floor and local cache policy. |
Output Schema
| Name | Required | Description |
|---|---|---|
| dob | No | |
| kind | Yes | |
| name | No | |
| total | No | |
| expiry | No | |
| vendor | No | |
| clauses | No | |
| country | No | |
| parties | No | |
| summary | No | |
| currency | No | |
| due_date | No | |
| ocr_text | Yes | Text used for extraction: provided text unchanged, or Markdown returned by Mistral OCR. |
| anomalies | No | |
| cache_hit | Yes | |
| key_dates | No | |
| source_id | Yes | |
| line_items | No | |
| page_count | Yes | Pages processed by Mistral OCR. Null for provided text, whose pagination is unknown. |
| risk_score | No | |
| document_type | No | |
| ocr_confidence | Yes | Mean Mistral OCR page confidence. Null for provided text; never an extraction accuracy score. |
| structured_text | No | |
| pipeline_version | Yes | |
| extraction_source | Yes | How the input text was obtained. provided_text is supplied by the caller, not verified by OCR. |
| total_duration_ms | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, but the description goes well beyond them: cache key composition and the id_document PII auto-bypass, the isError conditions around OCR confidence, that validation is schema-only (not factual), and the 60000-character rejection behavior. This is unusually rich behavioral disclosure for a read-only processing tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the pipeline summary, then grouped by concerns (sources, kinds, returns, validation, cache, options). Length is defensible for a 3-param nested pipeline tool, but a few statements restate schema facts (the 60000 limit, the 0.3 floor, maxPages/minOcrConfidence scoping), which is mild redundancy given a fully described schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description still usefully explains the discriminated-union return (switch on kind, extraction_source, ocr_confidence semantics). Combined with the caveats about confidence floors and cache bypass, an agent has everything needed to call and interpret the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds framing the schema does not: that maxPages/minOcrConfidence apply only to OCR sources (and how they interact with provided text), that text sources yield null ocr_confidence/page_count, and that the 0.3 floor is a placeholder to calibrate. It adds value beyond the field-level docs, though several details (60000 chars, 0.3 default) are duplicated from the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete pipeline with specific stages (classify → typed extraction → validation) and the input modalities it accepts. It also explicitly positions itself against siblings: 'Replaces the manual chain of mistral_ocr + mistral_chat + JSON parsing.' An agent can distinguish it from mistral_ocr and mistral_chat without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit routing guidance: kind=auto lets the server classify, source.type=text skips OCR and Files uploads, text+generic makes no API calls, and 'use mistral_ocr directly if you need raw OCR without a confidence guarantee.' It names the alternative and the condition that selects it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voxtral_transcribeVoxtral speech-to-textARead-onlyIdempotent
Transcribe an audio file to text using Mistral Voxtral.
Accepted models:
voxtral-mini-latest
voxtral-small-latest
Audio source is one of:
{ type: "file_url", fileUrl: "https://..." } (public URL)
{ type: "file", fileId: "" }
Options:
language: ISO-639-1 hint (e.g. 'fr', 'en'). Boosts accuracy when known.temperature: sampling temperature.diarize: return per-speaker segments (default false).timestampGranularities: ['segment'] to return per-segment timestamps.contextBias: list of phrases/terms that should bias the decoder.
Returns plain text, detected language, optional segments[], and token usage.
| Name | Required | Description | Default |
|---|---|---|---|
| audio | Yes | Audio to transcribe, supplied as a public URL or an uploaded file ID. | |
| model | No | STT model. Default: voxtral-mini-latest. | |
| diarize | No | Identify speakers in the returned transcription segments. Defaults to false. | |
| language | No | ISO-639-1 language hint (e.g. 'fr', 'en'). | |
| contextBias | No | Words or phrases to favor when decoding the audio. | |
| temperature | No | Sampling temperature for transcription. | |
| timestampGranularities | No | Only 'segment' is currently supported. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | Yes | |
| model | Yes | |
| usage | No | |
| language | Yes | |
| segments | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds useful behavioral context beyond that: accepted model IDs, default values for diarize and model, the only-supported granularity, and the shape of the response (text, language, segments, token usage). It does not mention auth or rate limits, but that is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action, then organizes models, source options, options, and return values into scannable bullet groups. It is efficient, though the option bullets partially duplicate what the schema already documents with 100% coverage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema, full annotation coverage, and 100% schema description coverage, the description supplies everything an agent needs: models, input modes, option semantics, and defaults. No material gap remains for invoking this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description nonetheless adds meaning beyond the schema by explaining intent ('language ... boosts accuracy when known', 'contextBias: phrases that should bias the decoder', 'diarize: return per-speaker segments'), which helps the agent use the parameters correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence gives a specific verb and resource ('Transcribe an audio file to text'), which is unambiguous and clearly distinct from the sibling tools (chat, FIM, vision, OCR, document processing). An agent can immediately tell this is the audio speech-to-text tool without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly lays out the two mutually exclusive audio source modes (public URL vs. uploaded file ID) and when each applies, which is exactly the routing decision an agent must make. It does not name alternatives or exclusions for when *not* to transcribe, but the context is otherwise clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.8.3- Changed
codestral_fim14 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / max_tokens / descriptionAdded value: +"Maximum number of tokens to generate. Input tokens plus this limit must fit within the model's context length." - added
Input schema / properties / max_tokens / maximumAdded value: +9007199254740991 - added
Input schema / properties / model / descriptionAdded value: +"FIM (fill-in-the-middle) model identifier. Any identifier your endpoint serves is accepted — read mistral://models for the live catalog. Known Mistral aliases: codestral-latest." - removed
Input schema / properties / model / enumRemoved value: -[ - "codestral-latest" -] - added
Input schema / properties / model / maxLengthAdded value: +200 - added
Input schema / properties / model / minLengthAdded value: +1 - added
Input schema / properties / seed / maximumAdded value: +9007199254740991 - added
Input schema / properties / seed / minimumAdded value: +-9007199254740991 - added
Input schema / properties / stop / descriptionAdded value: +"Stop generation when any of these text sequences is encountered." - added
Input schema / properties / temperature / descriptionAdded value: +"Sampling temperature: higher values make output more random; lower values make it more focused. Prefer adjusting this or top_p, not both." - added
Input schema / properties / top_p / descriptionAdded value: +"Nucleus sampling probability mass: 0.1 considers tokens in the top 10% of probability mass. Prefer adjusting this or temperature, not both." - changed
Output schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mistral_chat19 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / max_tokens / descriptionAdded value: +"Maximum number of tokens to generate. Input tokens plus this limit must fit within the model's context length." - added
Input schema / properties / max_tokens / maximumAdded value: +9007199254740991 - removed
Input schema / properties / messages / items / additionalPropertiesRemoved value: -false - added
Input schema / properties / messages / items / properties / content / descriptionAdded value: +"Text of the message." - added
Input schema / properties / messages / items / properties / role / descriptionAdded value: +"Message author: system for instructions, user for requests, or assistant for prior replies." - changed
Input schema / properties / model / descriptionPrevious value: -"Mistral chat model alias. Allowed: mistral-large-latest, mistral-medium-latest, mistral-small-latest, ministral-3b-latest, ministral-8b-latest, ministral-14b-latest, magistral-medium-latest, magistral-small-latest, devstral-latest, devstral-small-latest, codestral-latest, voxtral-small-latest. Default: mistral-medium-latest."New value: +"Chat model. Default: mistral-medium-latest (override with MISTRAL_DEFAULT_MODEL)." - removed
Input schema / properties / model / enumRemoved value: -[ - "mistral-large-latest", - "mistral-medium-latest", - "mistral-small-latest", - "ministral-3b-latest", - "ministral-8b-latest", - "ministral-14b-latest", - "magistral-medium-latest", - "magistral-small-latest", - "devstral-latest", - "devstral-small-latest", - "codestral-latest", - "voxtral-small-latest" -] - added
Input schema / properties / model / maxLengthAdded value: +200 - added
Input schema / properties / model / minLengthAdded value: +1 - changed
Input schema / properties / reasoning_effort / descriptionPrevious value: -"Controls reasoning depth for Magistral models. 'high' enables full chain-of-thought; 'none' disables it. Ignored on non-reasoning models."New value: +"Reasoning effort; supported values depend on the selected model." - changed
Input schema / properties / reasoning_effort / enumPrevious value: -[ - "none", - "high" -]New value: +[ + "none", + "minimal", + "low", + "medium", + "high", + "xhigh" +] - changed
Input schema / properties / response_format / anyOfPrevious value: -[ - { - "additionalProperties": false, - "properties": { - "type": { - "const": "text", - "type": "string" - } - }, - "required": [ - "type" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "type": { - "const": "json_object", - "type": "string" - } - }, - "required": [ - "type" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "json_schema": { - "additionalProperties": false, - "properties": { - "description": { - "type": "string" - }, - "name": { - "description": "Identifier for the schema; surfaced in API errors.", - "maxLength": 64, - "minLength": 1, - "type": "string" - }, - "schema": { - "additionalProperties": {}, - "description": "JSON Schema object the response must conform to.", - "type": "object" - }, - "strict": { - "description": "If true, the API rejects responses that do not strictly match the schema.", - "type": "boolean" - } - }, - "required": [ - "name", - "schema" - ], - "type": "object" - }, - "type": { - "const": "json_schema", - "type": "string" - } - }, - "required": [ - "type", - "json_schema" - ], - "type": "object" - } -]New value: +[ + { + "properties": { + "type": { + "const": "text", + "description": "Generate plain text without a JSON format constraint.", + "type": "string" + } + }, + "required": [ + "type" + ], + "type": "object" + }, + { + "properties": { + "type": { + "const": "json_object", + "description": "Generate JSON. Also instruct the model to produce JSON in a system or user message.", + "type": "string" + } + }, + "required": [ + "type" + ], + "type": "object" + }, + { + "properties": { + "json_schema": { + "description": "Named JSON Schema and optional strictness for the generated response.", + "properties": { + "description": { + "description": "Description of the response the schema defines.", + "type": "string" + }, + "name": { + "description": "Identifier for the schema; surfaced in API errors.", + "maxLength": 64, + "minLength": 1, + "type": "string" + }, + "schema": { + "additionalProperties": {}, + "description": "JSON Schema object the response must conform to.", + "propertyNames": { + "type": "string" + }, + "type": "object" + }, + "strict": { + "description": "If true, the API rejects responses that do not strictly match the schema.", + "type": "boolean" + } + }, + "required": [ + "name", + "schema" + ], + "type": "object" + }, + "type": { + "const": "json_schema", + "description": "Generate JSON conforming to the supplied json_schema.", + "type": "string" + } + }, + "required": [ + "type", + "json_schema" + ], + "type": "object" + } +] - added
Input schema / properties / seed / maximumAdded value: +9007199254740991 - added
Input schema / properties / seed / minimumAdded value: +-9007199254740991 - added
Input schema / properties / temperature / descriptionAdded value: +"Sampling temperature: higher values make output more random; lower values make it more focused. Prefer adjusting this or top_p, not both." - added
Input schema / properties / top_p / descriptionAdded value: +"Nucleus sampling probability mass: 0.1 considers tokens in the top 10% of probability mass. Prefer adjusting this or temperature, not both." - changed
Output schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mistral_ocr49 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false - removed
Input schema / properties / bbox_annotation_format / additionalPropertiesRemoved value: -false - added
Input schema / properties / bbox_annotation_format / descriptionAdded value: +"JSON Schema for structured annotations of each extracted bounding box or image." - removed
Input schema / properties / bbox_annotation_format / properties / json_schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / bbox_annotation_format / properties / json_schema / descriptionAdded value: +"Named JSON Schema and optional strictness for the extracted annotation." - added
Input schema / properties / bbox_annotation_format / properties / json_schema / properties / description / descriptionAdded value: +"Description of the annotation to extract." - added
Input schema / properties / bbox_annotation_format / properties / json_schema / properties / name / descriptionAdded value: +"Name identifying the annotation schema." - added
Input schema / properties / bbox_annotation_format / properties / json_schema / properties / schema / descriptionAdded value: +"JSON Schema object defining the fields to extract into the annotation." - added
Input schema / properties / bbox_annotation_format / properties / json_schema / properties / schema / propertyNamesAdded value: +{ + "type": "string" +} - added
Input schema / properties / bbox_annotation_format / properties / json_schema / properties / strict / descriptionAdded value: +"Whether the annotation must strictly follow the supplied JSON Schema." - added
Input schema / properties / confidence_scores_granularity / descriptionAdded value: +"Confidence granularity. 'block' also fills pages[].blocks[].confidence_scores and requires OCR 4.1 (mistral-ocr-4-1) or newer." - changed
Input schema / properties / confidence_scores_granularity / enumPrevious value: -[ - "page", - "word" -]New value: +[ + "page", + "word", + "block" +] - changed
Input schema / properties / document / anyOfPrevious value: -[ - { - "additionalProperties": false, - "properties": { - "documentName": { - "type": "string" - }, - "documentUrl": { - "description": "HTTPS URL to a PDF or image.", - "type": "string" - }, - "type": { - "const": "document_url", - "type": "string" - } - }, - "required": [ - "type", - "documentUrl" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "imageUrl": { - "description": "HTTPS URL or data:image/...;base64,... payload.", - "type": "string" - }, - "type": { - "const": "image_url", - "type": "string" - } - }, - "required": [ - "type", - "imageUrl" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "fileId": { - "description": "ID of a file previously uploaded via the Files API.", - "type": "string" - }, - "type": { - "const": "file", - "type": "string" - } - }, - "required": [ - "type", - "fileId" - ], - "type": "object" - } -]New value: +[ + { + "properties": { + "documentName": { + "description": "Filename of the referenced document.", + "type": "string" + }, + "documentUrl": { + "description": "HTTPS URL to a PDF or image.", + "type": "string" + }, + "type": { + "const": "document_url", + "description": "Read the document from documentUrl.", + "type": "string" + } + }, + "required": [ + "type", + "documentUrl" + ], + "type": "object" + }, + { + "properties": { + "imageUrl": { + "description": "HTTPS URL or data:image/...;base64,... payload.", + "type": "string" + }, + "type": { + "const": "image_url", + "description": "Read the image from the URL or data URI in imageUrl.", + "type": "string" + } + }, + "required": [ + "type", + "imageUrl" + ], + "type": "object" + }, + { + "properties": { + "fileId": { + "description": "ID of a file previously uploaded via the Files API.", + "type": "string" + }, + "type": { + "const": "file", + "description": "Read a previously uploaded file identified by fileId.", + "type": "string" + } + }, + "required": [ + "type", + "fileId" + ], + "type": "object" + } +] - added
Input schema / properties / document / descriptionAdded value: +"Document to process, supplied as a document URL, an image URL or data URI, or an uploaded file ID." - removed
Input schema / properties / document_annotation_format / $refRemoved value: -"#/properties/bbox_annotation_format" - added
Input schema / properties / document_annotation_format / descriptionAdded value: +"JSON Schema for a structured annotation extracted from the entire document." - added
Input schema / properties / document_annotation_format / propertiesAdded value: +{ + "json_schema": { + "description": "Named JSON Schema and optional strictness for the extracted annotation.", + "properties": { + "description": { + "description": "Description of the annotation to extract.", + "type": "string" + }, + "name": { + "description": "Name identifying the annotation schema.", + "minLength": 1, + "type": "string" + }, + "schema": { + "additionalProperties": {}, + "description": "JSON Schema object defining the fields to extract into the annotation.", + "propertyNames": { + "type": "string" + }, + "type": "object" + }, + "strict": { + "description": "Whether the annotation must strictly follow the supplied JSON Schema.", + "type": "boolean" + } + }, + "required": [ + "name", + "schema" + ], + "type": "object" + }, + "type": { + "const": "json_schema", + "description": "Only json_schema is accepted by OCR annotation formats.", + "type": "string" + } +} - added
Input schema / properties / document_annotation_format / requiredAdded value: +[ + "type", + "json_schema" +] - added
Input schema / properties / document_annotation_format / typeAdded value: +"object" - added
Input schema / properties / document_annotation_prompt / descriptionAdded value: +"Instructions for whole-document structured extraction. Requires document_annotation_format." - added
Input schema / properties / extractFooter / descriptionAdded value: +"Extract each page footer into its footer field and remove it from the page markdown." - added
Input schema / properties / extractHeader / descriptionAdded value: +"Extract each page header into its header field and remove it from the page markdown." - added
Input schema / properties / imageLimit / descriptionAdded value: +"Maximum number of images to extract from the document." - added
Input schema / properties / imageLimit / maximumAdded value: +9007199254740991 - added
Input schema / properties / imageMinSize / descriptionAdded value: +"Minimum height and width of an image to extract." - added
Input schema / properties / imageMinSize / maximumAdded value: +9007199254740991 - added
Input schema / properties / includeBlocksAdded value: +{ + "description": "Return paragraph-level blocks (bounding box + type) per page. Requires OCR 4 (mistral-ocr-4-0) or newer.", + "type": "boolean" +} - added
Input schema / properties / includeImageBase64 / descriptionAdded value: +"Include base64-encoded data for extracted images in the response." - removed
Input schema / properties / model / enumRemoved value: -[ - "mistral-ocr-latest" -] - added
Input schema / properties / model / maxLengthAdded value: +200 - added
Input schema / properties / model / minLengthAdded value: +1 - changed
Input schema / properties / pages / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "items": { - "minimum": 0, - "type": "integer" - }, - "type": "array" - } -]New value: +[ + { + "type": "string" + }, + { + "items": { + "maximum": 9007199254740991, + "minimum": 0, + "type": "integer" + }, + "type": "array" + } +] - added
Input schema / properties / pages / descriptionAdded value: +"Zero-based page numbers to process, as an array or comma-separated numbers and ranges such as \"0-5,7\"." - added
Input schema / properties / tableFormat / descriptionAdded value: +"Format for extracted tables: Markdown or HTML." - changed
Output schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - added
Output schema / properties / annotations / properties / image_annotations / items / properties / page_index / maximumAdded value: +9007199254740991 - added
Output schema / properties / annotations / properties / image_annotations / items / properties / page_index / minimumAdded value: +-9007199254740991 - added
Output schema / properties / pages / items / properties / blocksAdded value: +{ + "description": "Paragraph-level blocks in reading order. Populated when `includeBlocks: true` and the model is OCR 4 or newer.", + "items": { + "additionalProperties": false, + "properties": { + "bottom_right_x": { + "type": "number" + }, + "bottom_right_y": { + "type": "number" + }, + "confidence_scores": { + "additionalProperties": false, + "description": "Populated only when `confidence_scores_granularity: 'block'`. Fields are null when the signal is absent (an image-only block has no content to score).", + "properties": { + "average_content_confidence_score": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ] + }, + "block_type_confidence_score": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ] + }, + "minimum_content_confidence_score": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ] + } + }, + "type": "object" + }, + "content": { + "type": "string" + }, + "image_id": { + "description": "Set on type:image — references the matching entry in `images[]`.", + "type": "string" + }, + "table_id": { + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "description": "Set on type:table — references the matching entry in `tables[]`." + }, + "top_left_x": { + "type": "number" + }, + "top_left_y": { + "type": "number" + }, + "type": { + "enum": [ + "text", + "title", + "list", + "table", + "image", + "equation", + "caption", + "code", + "references", + "aside_text", + "header", + "footer", + "signature" + ], + "type": "string" + } + }, + "required": [ + "type", + "top_left_x", + "top_left_y", + "bottom_right_x", + "bottom_right_y", + "content" + ], + "type": "object" + }, + "type": "array" +} - added
Output schema / properties / pages / items / properties / confidence_scores / properties / word_confidence_scores / items / properties / start_index / maximumAdded value: +9007199254740991 - added
Output schema / properties / pages / items / properties / confidence_scores / properties / word_confidence_scores / items / properties / start_index / minimumAdded value: +-9007199254740991 - added
Output schema / properties / pages / items / properties / footer / anyOfAdded value: +[ + { + "type": "string" + }, + { + "type": "null" + } +] - removed
Output schema / properties / pages / items / properties / footer / typeRemoved value: -[ - "string", - "null" -] - added
Output schema / properties / pages / items / properties / header / anyOfAdded value: +[ + { + "type": "string" + }, + { + "type": "null" + } +] - removed
Output schema / properties / pages / items / properties / header / typeRemoved value: -[ - "string", - "null" -] - added
Output schema / properties / pages / items / properties / index / maximumAdded value: +9007199254740991 - added
Output schema / properties / pages / items / properties / index / minimumAdded value: +-9007199254740991 - added
Output schema / properties / pages_count / maximumAdded value: +9007199254740991 - added
Output schema / properties / pages_count / minimumAdded value: +-9007199254740991
- Changed
mistral_vision16 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / max_tokens / descriptionAdded value: +"Maximum number of tokens to generate. Input tokens plus this limit must fit within the model's context length." - added
Input schema / properties / max_tokens / maximumAdded value: +9007199254740991 - removed
Input schema / properties / messages / items / additionalPropertiesRemoved value: -false - changed
Input schema / properties / messages / items / properties / content / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "items": { - "anyOf": [ - { - "additionalProperties": false, - "properties": { - "text": { - "type": "string" - }, - "type": { - "const": "text", - "type": "string" - } - }, - "required": [ - "type", - "text" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "imageUrl": { - "anyOf": [ - { - "description": "https URL or data:image/...;base64,... payload", - "type": "string" - }, - { - "additionalProperties": false, - "properties": { - "detail": { - "enum": [ - "auto", - "low", - "high" - ], - "type": "string" - }, - "url": { - "type": "string" - } - }, - "required": [ - "url" - ], - "type": "object" - } - ] - }, - "type": { - "const": "image_url", - "type": "string" - } - }, - "required": [ - "type", - "imageUrl" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "documentName": { - "type": "string" - }, - "documentUrl": { - "type": "string" - }, - "type": { - "const": "document_url", - "type": "string" - } - }, - "required": [ - "type", - "documentUrl" - ], - "type": "object" - } - ] - }, - "minItems": 1, - "type": "array" - } -]New value: +[ + { + "type": "string" + }, + { + "items": { + "anyOf": [ + { + "properties": { + "text": { + "description": "Text to include in the message.", + "type": "string" + }, + "type": { + "const": "text", + "description": "Identifies a text content part.", + "type": "string" + } + }, + "required": [ + "type", + "text" + ], + "type": "object" + }, + { + "properties": { + "imageUrl": { + "anyOf": [ + { + "description": "https URL or data:image/...;base64,... payload", + "type": "string" + }, + { + "properties": { + "detail": { + "description": "Image detail hint: automatic, low, or high.", + "enum": [ + "auto", + "low", + "high" + ], + "type": "string" + }, + "url": { + "description": "HTTPS URL or data:image/...;base64,... payload.", + "type": "string" + } + }, + "required": [ + "url" + ], + "type": "object" + } + ], + "description": "Image source as a URL or base64 data URI, optionally with a detail hint." + }, + "type": { + "const": "image_url", + "description": "Identifies an image content part.", + "type": "string" + } + }, + "required": [ + "type", + "imageUrl" + ], + "type": "object" + }, + { + "properties": { + "documentName": { + "description": "Filename of the referenced document.", + "type": "string" + }, + "documentUrl": { + "description": "URL of the PDF or document to include in the message.", + "type": "string" + }, + "type": { + "const": "document_url", + "description": "Identifies a document content part.", + "type": "string" + } + }, + "required": [ + "type", + "documentUrl" + ], + "type": "object" + } + ] + }, + "minItems": 1, + "type": "array" + } +] - added
Input schema / properties / messages / items / properties / content / descriptionAdded value: +"Message text or an ordered list of text, image, and document content parts." - added
Input schema / properties / messages / items / properties / role / descriptionAdded value: +"Message author: system for instructions, user for requests, or assistant for prior replies." - removed
Input schema / properties / model / enumRemoved value: -[ - "pixtral-large-latest", - "pixtral-12b-latest", - "mistral-large-latest", - "mistral-medium-latest", - "mistral-small-latest" -] - added
Input schema / properties / model / maxLengthAdded value: +200 - added
Input schema / properties / model / minLengthAdded value: +1 - added
Input schema / properties / seed / maximumAdded value: +9007199254740991 - added
Input schema / properties / seed / minimumAdded value: +-9007199254740991 - added
Input schema / properties / temperature / descriptionAdded value: +"Sampling temperature: higher values make output more random; lower values make it more focused. Prefer adjusting this or top_p, not both." - added
Input schema / properties / top_p / descriptionAdded value: +"Nucleus sampling probability mass: 0.1 considers tokens in the top 10% of probability mass. Prefer adjusting this or temperature, not both." - changed
Output schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Added
process_document - Changed
voxtral_transcribe13 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false - changed
Input schema / properties / audio / anyOfPrevious value: -[ - { - "additionalProperties": false, - "properties": { - "fileUrl": { - "description": "HTTPS URL to an audio file (mp3/wav/flac/ogg/webm/m4a).", - "type": "string" - }, - "type": { - "const": "file_url", - "type": "string" - } - }, - "required": [ - "type", - "fileUrl" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "fileId": { - "description": "ID of an audio file previously uploaded via the Files API (purpose=audio).", - "type": "string" - }, - "type": { - "const": "file", - "type": "string" - } - }, - "required": [ - "type", - "fileId" - ], - "type": "object" - } -]New value: +[ + { + "properties": { + "fileUrl": { + "description": "HTTPS URL to an audio file (mp3/wav/flac/ogg/webm/m4a).", + "type": "string" + }, + "type": { + "const": "file_url", + "description": "Transcribe audio from the URL in fileUrl.", + "type": "string" + } + }, + "required": [ + "type", + "fileUrl" + ], + "type": "object" + }, + { + "properties": { + "fileId": { + "description": "ID of an audio file previously uploaded via the Files API (purpose=audio).", + "type": "string" + }, + "type": { + "const": "file", + "description": "Transcribe an uploaded audio file identified by fileId.", + "type": "string" + } + }, + "required": [ + "type", + "fileId" + ], + "type": "object" + } +] - added
Input schema / properties / audio / descriptionAdded value: +"Audio to transcribe, supplied as a public URL or an uploaded file ID." - added
Input schema / properties / contextBias / descriptionAdded value: +"Words or phrases to favor when decoding the audio." - added
Input schema / properties / diarize / descriptionAdded value: +"Identify speakers in the returned transcription segments. Defaults to false." - removed
Input schema / properties / model / enumRemoved value: -[ - "voxtral-mini-latest", - "voxtral-small-latest" -] - added
Input schema / properties / model / maxLengthAdded value: +200 - added
Input schema / properties / model / minLengthAdded value: +1 - added
Input schema / properties / temperature / descriptionAdded value: +"Sampling temperature for transcription." - changed
Output schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - added
Output schema / properties / language / anyOfAdded value: +[ + { + "type": "string" + }, + { + "type": "null" + } +] - removed
Output schema / properties / language / typeRemoved value: -[ - "string", - "null" -]
- Removed
workflow_execute - Removed
workflow_interact - Removed
workflow_status
24 tool updates
v0.7.0- Removed
batch_cancel - Removed
batch_create - Removed
batch_get - Removed
batch_list - Changed
codestral_fim1 field changed- added
Input schema / properties / seedAdded value: +{ + "description": "Random seed for deterministic sampling. Maps to Mistral's `random_seed`.", + "type": "integer" +}
- Removed
files_delete - Removed
files_get - Removed
files_list - Removed
files_signed_url - Removed
files_upload - Removed
mcp_sample - Removed
mistral_agent - Changed
mistral_chat4 fields changed- added
Input schema / properties / reasoning_effortAdded value: +{ + "description": "Controls reasoning depth for Magistral models. 'high' enables full chain-of-thought; 'none' disables it. Ignored on non-reasoning models.", + "enum": [ + "none", + "high" + ], + "type": "string" +} - added
Input schema / properties / response_formatAdded value: +{ + "anyOf": [ + { + "additionalProperties": false, + "properties": { + "type": { + "const": "text", + "type": "string" + } + }, + "required": [ + "type" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "type": { + "const": "json_object", + "type": "string" + } + }, + "required": [ + "type" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "json_schema": { + "additionalProperties": false, + "properties": { + "description": { + "type": "string" + }, + "name": { + "description": "Identifier for the schema; surfaced in API errors.", + "maxLength": 64, + "minLength": 1, + "type": "string" + }, + "schema": { + "additionalProperties": {}, + "description": "JSON Schema object the response must conform to.", + "type": "object" + }, + "strict": { + "description": "If true, the API rejects responses that do not strictly match the schema.", + "type": "boolean" + } + }, + "required": [ + "name", + "schema" + ], + "type": "object" + }, + "type": { + "const": "json_schema", + "type": "string" + } + }, + "required": [ + "type", + "json_schema" + ], + "type": "object" + } + ], + "description": "Force a structured output: `{type:\"json_object\"}` for JSON mode, `{type:\"json_schema\", json_schema:{...}}` for strict schema mode." +} - added
Input schema / properties / seedAdded value: +{ + "description": "Random seed for deterministic sampling. Maps to Mistral's `random_seed`.", + "type": "integer" +} - added
Output schema / properties / reasoning_contentAdded value: +{ + "description": "Reasoning trace returned by Magistral models. Absent for non-reasoning models.", + "type": "string" +}
- Removed
mistral_chat_stream - Removed
mistral_classify - Removed
mistral_embed - Removed
mistral_moderate - Changed
mistral_ocr7 fields changed- added
Input schema / properties / bbox_annotation_formatAdded value: +{ + "additionalProperties": false, + "properties": { + "json_schema": { + "additionalProperties": false, + "properties": { + "description": { + "type": "string" + }, + "name": { + "minLength": 1, + "type": "string" + }, + "schema": { + "additionalProperties": {}, + "type": "object" + }, + "strict": { + "type": "boolean" + } + }, + "required": [ + "name", + "schema" + ], + "type": "object" + }, + "type": { + "const": "json_schema", + "description": "Only json_schema is accepted by OCR annotation formats.", + "type": "string" + } + }, + "required": [ + "type", + "json_schema" + ], + "type": "object" +} - added
Input schema / properties / confidence_scores_granularityAdded value: +{ + "enum": [ + "page", + "word" + ], + "type": "string" +} - added
Input schema / properties / document_annotation_formatAdded value: +{ + "$ref": "#/properties/bbox_annotation_format" +} - added
Input schema / properties / document_annotation_promptAdded value: +{ + "type": "string" +} - added
Output schema / properties / annotationsAdded value: +{ + "additionalProperties": false, + "properties": { + "document_annotation": { + "type": "string" + }, + "image_annotations": { + "items": { + "additionalProperties": false, + "properties": { + "annotation": { + "type": "string" + }, + "bbox": { + "additionalProperties": false, + "properties": { + "bottom_right_x": { + "type": "number" + }, + "bottom_right_y": { + "type": "number" + }, + "top_left_x": { + "type": "number" + }, + "top_left_y": { + "type": "number" + } + }, + "type": "object" + }, + "image_id": { + "type": "string" + }, + "page_index": { + "type": "integer" + } + }, + "required": [ + "page_index", + "annotation" + ], + "type": "object" + }, + "type": "array" + } + }, + "type": "object" +} - added
Output schema / properties / pages / items / properties / confidence_scoresAdded value: +{ + "additionalProperties": false, + "properties": { + "average_page_confidence_score": { + "type": "number" + }, + "minimum_page_confidence_score": { + "type": "number" + }, + "word_confidence_scores": { + "items": { + "additionalProperties": false, + "properties": { + "confidence": { + "type": "number" + }, + "start_index": { + "type": "integer" + }, + "text": { + "type": "string" + } + }, + "required": [ + "text", + "confidence", + "start_index" + ], + "type": "object" + }, + "type": "array" + } + }, + "required": [ + "average_page_confidence_score", + "minimum_page_confidence_score" + ], + "type": "object" +} - added
Output schema / properties / pages / items / properties / images / items / properties / image_annotationAdded value: +{ + "type": "string" +}
- Removed
mistral_tool_call - Changed
mistral_vision1 field changed- added
Input schema / properties / seedAdded value: +{ + "description": "Random seed for deterministic sampling. Maps to Mistral's `random_seed`.", + "type": "integer" +}
- Removed
voxtral_speak - Added
workflow_execute - Added
workflow_interact - Added
workflow_status
22 tool updates
v0.4.0- First observed
batch_cancel - First observed
batch_create - First observed
batch_get - First observed
batch_list - First observed
codestral_fim - First observed
files_delete - First observed
files_get - First observed
files_list - First observed
files_signed_url - First observed
files_upload - First observed
mcp_sample - First observed
mistral_agent - First observed
mistral_chat - First observed
mistral_chat_stream - First observed
mistral_classify - First observed
mistral_embed - First observed
mistral_moderate - First observed
mistral_ocr - First observed
mistral_tool_call - First observed
mistral_vision - First observed
voxtral_speak - First observed
voxtral_transcribe
TDQS
Scored across 6 tools
Each tool targets a distinct modality/action: code FIM, text chat, OCR, vision chat, document pipeline, and audio transcription. There is minor overlap between mistral_chat and mistral_vision (both chat) and between process_document and mistral_ocr, but the descriptions explicitly steer selection, so confusion is limited.
Names mix provider-prefixed conventions (codestral_fim, mistral_chat, mistral_ocr, mistral_vision, voxtral_transcribe) with a plain action name (process_document). The pattern is somewhat readable but not a predictable verb_noun scheme and the provider prefixes fragment consistency.
Six tools is well-scoped for a multimodal Mistral wrapper, with each tool covering a distinct capability rather than redundancy. Slightly broad for a server named 'Document Extraction' since it also includes code completion and chat, but not excessive.
Covers OCR, vision, chat, transcription, and a combined pipeline, but several tools reference a Files API (fileId) that is not exposed as a tool, creating a dead-end for uploading files. No model-listing or file-management operations leave workflows partially blocked.
Maintenance
Related MCP Connectors
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…
MCP server for progressive tool usage at any scale (see https://klavis.ai)
- QuallaaOAuthcom.quallaa
Talk to your public-facing AI from any MCP client — Claude, ChatGPT, Cursor, Cline, Windsurf.
Nifty's MCP server — exposes tasks, projects, messages, and files as tools for AI agents.
Related MCP Servers
- -licenseNot gradedqualityDmaintenanceA TypeScript implementation of a Model Context Protocol server and client that enables interaction with language models (specifically Mistral running on Ollama).-
- -licenseNot gradedqualityNot gradedmaintenanceA production-ready TypeScript MCP server providing basic tools (add, echo, timestamp), resources (server info, greetings, data access), and prompt templates (analyze, code-review, summarize). Serves as a foundation for building custom MCP servers with extensible architecture.398 npm-
- AlicenseBqualityCmaintenanceA TypeScript-based MCP server that provides tools to interact with local Codex and Gemini CLIs via stdio transport. It enables users to execute prompts through the ask_codex and ask_gemini tools, supporting custom models and timeout configurations.110 npm6MIT
- FlicenseAqualityDmaintenanceA TypeScript MCP server template with Zod validation, dual transport (stdio/HTTP), and modular architecture for building MCP-compatible tools, resources, and prompts.11-