yandex-vision-ocr-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| YANDEX_API_KEY | No | API key (recommended for long-lived usage). One of YANDEX_API_KEY or YANDEX_IAM_TOKEN is required. | |
| YANDEX_FOLDER_ID | No | Yandex Cloud folder ID. Optional; only sent as x-folder-id when set. | |
| YANDEX_IAM_TOKEN | No | Short-lived IAM token (~12h). Use instead of an API key. |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| recognize_textA | Recognize text in an image (JPEG/PNG) or a single-page PDF using Yandex Vision OCR (synchronous). Picks the recognition model via 'model' (default 'page' for printed text; 'handwritten', 'table', 'markdown', etc.). Provide a local file 'path' or 'base64' content. For multi-page/large PDFs, use recognize_pdf instead. |
| recognize_pdfA | Recognize text in a PDF document (single or multi-page) using Yandex Vision OCR via asynchronous recognition (recognizeTextAsync + getRecognition polling). Use this for PDFs and large files. Provide a local file 'path' or 'base64' content, and pick a 'model' (default 'page'). |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 2 tools
The two tools have clearly distinct purposes: one for images and single-page PDFs (synchronous), the other for multi-page/large PDFs (asynchronous). No overlap in functionality.
Both tools follow the 'recognize_' prefix pattern with clear suffixes ('pdf' and 'text'), making them predictable and self-explanatory.
Two tools is minimal but appropriate for a focused OCR server. It covers the core distinction between image/single-page and multi-page PDF processing without unnecessary complexity.
The tool surface covers the essential OCR use cases: image and PDF text recognition. The model parameter handles handwritten, table, and markdown, so no major gaps.