Orion Vision MCP Server
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Orion Vision MCP Serverextract data from this invoice: https://example.com/invoice.pdf"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Orion Vision MCP Server π
π Compatible with Cline, Cursor, Claude Desktop, and any other MCP Clients!
Orion Vision MCP is also compatible with any MCP client
The Model Context Protocol (MCP) is an open standard that enables AI systems to interact seamlessly with various data sources and tools, facilitating secure, two-way connections.
The Orion Vision MCP server provides:
Seamless integration with Azure Form Recognizer / Document Intelligence
Document analysis and form data extraction capabilities
Support for various document types (receipts, invoices, ID documents, etc.)
Type-safe operations with TypeScript
Prerequisites π§
Before you begin, ensure you have:
Azure Form Recognizer / Document Intelligence endpoint and key
Claude Desktop or Cursor
Node.js (v20 or higher)
Git installed (only needed if using Git installation method)
Related MCP server: MCP PDF
Orion Vision MCP server installation β‘
Running with NPX
npx -y orion-vision-mcp@latestInstalling via Smithery
To install Orion Vision MCP Server for Claude Desktop automatically via Smithery:
npx -y @smithery/cli install @orion-vision/mcp --client claudeConfiguring MCP Clients βοΈ
Configuring Cline π€
The easiest way to set up the Orion Vision MCP server in Cline is through the marketplace with a single click:
Open Cline in VS Code
Click on the Cline icon in the sidebar
Navigate to the "MCP Servers" tab (4 squares)
Search "Orion Vision" and click "install"
When prompted, enter your Azure Form Recognizer credentials
Alternatively, you can manually set up the Orion Vision MCP server in Cline:
Open the Cline MCP settings file:
# For macOS:
code ~/Library/Application\ Support/Code/User/globalStorage/saoudrizwan.claude-dev/settings/cline_mcp_settings.json
# For Windows:
code %APPDATA%\Code\User\globalStorage\saoudrizwan.claude-dev\settings\cline_mcp_settings.jsonAdd the Orion Vision server configuration to the file:
{
"mcpServers": {
"orion-vision-mcp": {
"command": "npx",
"args": ["-y", "orion-vision-mcp@latest"],
"env": {
"AZURE_FORM_RECOGNIZER_ENDPOINT": "your-endpoint-here",
"AZURE_FORM_RECOGNIZER_KEY": "your-key-here"
},
"disabled": false,
"autoApprove": []
}
}
}Save the file and restart Cline if it's already running.
Configuring Cursor π₯οΈ
Note: Requires Cursor version 0.45.6 or higher
To set up the Orion Vision MCP server in Cursor:
Open Cursor Settings
Navigate to Features > MCP Servers
Click on the "+ Add New MCP Server" button
Fill out the following information:
Name: Enter a nickname for the server (e.g., "orion-vision-mcp")
Type: Select "command" as the type
Command: Enter the command to run the server:
env AZURE_FORM_RECOGNIZER_ENDPOINT=your-endpoint AZURE_FORM_RECOGNIZER_KEY=your-key npx -y orion-vision-mcp@latestImportant: Replace
your-endpointandyour-keywith your Azure Form Recognizer credentials
Configuring the Claude Desktop app π₯οΈ
For macOS:
# Create the config file if it doesn't exist
touch "$HOME/Library/Application Support/Claude/claude_desktop_config.json"
# Opens the config file in TextEdit
open -e "$HOME/Library/Application Support/Claude/claude_desktop_config.json"For Windows:
code %APPDATA%\Claude\claude_desktop_config.jsonAdd the Orion Vision server configuration:
{
"mcpServers": {
"orion-vision-mcp": {
"command": "npx",
"args": ["-y", "orion-vision-mcp@latest"],
"env": {
"AZURE_FORM_RECOGNIZER_ENDPOINT": "your-endpoint-here",
"AZURE_FORM_RECOGNIZER_KEY": "your-key-here"
}
}
}
}Usage in Claude Desktop App π―
Once the installation is complete, and the Claude desktop app is configured, you must completely close and re-open the Claude desktop app to see the orion-vision-mcp server. You should see a hammer icon in the bottom left of the app, indicating available MCP tools.
Orion Vision Examples
Analyze a Document:
Analyze the document at "https://example.com/document.pdf" using Azure Form Recognizer.Extract Form Data:
Extract data from the invoice at "https://example.com/invoice.pdf".Process ID Document:
Process the ID document at "https://example.com/id.pdf" and extract relevant information.Troubleshooting π οΈ
Common Issues
Server Not Found
Verify the npm installation by running
npm --versionCheck Claude Desktop configuration syntax
Ensure Node.js is properly installed by running
node --version
Azure Form Recognizer Credentials Issues
Confirm your Azure Form Recognizer endpoint and key are valid
Check the credentials are correctly set in the config
Verify no spaces or quotes around the credentials
Document Processing Issues
Verify the document URL is accessible
Check the document format is supported
Ensure the document is not corrupted or password-protected
Acknowledgments β¨
Model Context Protocol for the MCP specification
Anthropic for Claude Desktop
Microsoft Azure for Form Recognizer / Document Intelligence
Available Tools
2 toolsanalyze-documentC
Analyzes a document using Azure Form Recognizer and returns structured data
| Name | Required | Description | Default |
|---|---|---|---|
| modelId | No | Optional model ID for custom models | |
| url | Yes | URL of the document to analyze |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the analysis method and output type but omits critical details like authentication needs, rate limits, processing time, error handling, or what 'structured data' entails, leaving significant gaps for a tool performing external API calls.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with zero waste, clearly front-loading the core functionality. Every word contributes to understanding the tool's purpose without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of document analysis with an external service, no annotations, and no output schema, the description is incomplete. It lacks details on behavioral traits, output format, error cases, and usage context, which are essential for effective tool invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters (modelId and url). The description adds no additional parameter semantics beyond what the schema provides, such as examples or constraints, meeting the baseline for high coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('analyzes') and resource ('a document') using Azure Form Recognizer, with the outcome of returning structured data. It distinguishes from the sibling 'extract-form-data' by specifying the analysis method, though not explicitly contrasting their use cases.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus the sibling 'extract-form-data' or other alternatives. The description implies usage for document analysis but lacks context on prerequisites, constraints, or comparative scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract-form-dataC
Extracts structured data from forms using Azure Form Recognizer
| Name | Required | Description | Default |
|---|---|---|---|
| formType | Yes | Type of form to analyze | |
| url | Yes | URL of the form document to analyze |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. While it mentions the technology (Azure Form Recognizer), it doesn't describe what happens during extraction - whether it's a read-only operation, if it modifies data, authentication requirements, rate limits, or error handling. For a tool with no annotation coverage, this leaves significant behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise - a single sentence that directly states the tool's purpose without any unnecessary words. It's front-loaded with the core functionality and doesn't waste space on redundant information. Every word earns its place in this minimal description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that there's no output schema and no annotations, the description should provide more context about what the tool returns and how it behaves. For a data extraction tool with 2 required parameters, the description is too minimal - it doesn't explain the extraction results format, error conditions, or practical usage examples. The completeness is inadequate for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters thoroughly. The description doesn't add any meaningful parameter semantics beyond what's in the schema - it doesn't explain how 'formType' affects extraction results or provide examples of valid URLs. With complete schema coverage, the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Extracts structured data from forms using Azure Form Recognizer'. It specifies the action (extracts), resource (structured data from forms), and technology (Azure Form Recognizer). However, it doesn't explicitly differentiate from its sibling 'analyze-document', which might have overlapping functionality.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus its sibling 'analyze-document' or other alternatives. It doesn't mention prerequisites, limitations, or specific scenarios where this tool is preferred. The only implied context is that it works with forms, but no explicit usage instructions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
2 tool updates
v1.0.0- First observed
analyze-document - First observed
extract-form-data
TDQS
The two tools have nearly identical purposes: both use Azure Form Recognizer to extract structured data from documents/forms. 'analyze-document' and 'extract-form-data' are functionally indistinguishable, with no clear boundary between them. This high ambiguity will cause agents to misselect between tools.
Both tools use kebab-case naming, which is consistent. However, the verb choices ('analyze' vs 'extract') are different despite similar functionality, creating minor inconsistency. The naming pattern is readable but not perfectly aligned in purpose.
With only 2 tools, the server feels thin for a vision/document processing domain. A typical MCP server for this scope would include more operations like text extraction, image analysis, or OCR configuration. The minimal tool count limits functionality and suggests incomplete coverage.
For a vision/document processing server, there are significant gaps: no image analysis tools, no OCR configuration, no batch processing, and no support for different document types beyond forms. The surface is severely limited, focusing only on form data extraction without broader vision capabilities.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
DocForge turns documents into structured data. Upload a PDF, image, or Office file and get fielded JSON back with per-field confidence scores. 95 templates (invoices, receipts, bank statements, ID docs), custom JSON Schema mode, auto-detect, natural-language instructions. Keyless demo tool included. Free 7-day trial.
- DatanemOAuthcom.datanem
Turn PDFs, scans and photos into a queryable database. Invoices, CVs, receipts, in bulk.
Turn any PDF into structured JSON via AI + OCR: invoices, bank statements, contracts.
Invoice and receipt extractor: reads PDF and image invoices/receipts with AI, pulling dateβ¦
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables AI agents and users to process documents through natural language, supporting PDF operations like text extraction, redaction, splitting, form filling, annotations, and content search.27561MIT
- AlicenseNot gradedqualityFmaintenanceEnables AI-powered extraction and analysis of PDF documents with 40+ specialized tools for text, tables, images, layout analysis, security assessment, and document intelligence. Supports both text-based and scanned PDFs with OCR capabilities.10MIT
- AlicenseNot gradedqualityFmaintenanceEnables PDF document processing including text, image, and table extraction, as well as intelligent classification and similarity analysis across multiple languages.49MIT
- AlicenseAqualityCmaintenanceEnables AI assistants to extract and structure content from documents (PDFs, images, Office files) using Upstage AI's document digitization and information extraction APIs.23MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/Cognitive-Stack/orion-vision-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server