Skip to main content
Glama
Cognitive-Stack

Orion Vision MCP Server

Orion Vision MCP Server πŸš€

πŸ”Œ Compatible with Cline, Cursor, Claude Desktop, and any other MCP Clients!

Orion Vision MCP is also compatible with any MCP client

The Model Context Protocol (MCP) is an open standard that enables AI systems to interact seamlessly with various data sources and tools, facilitating secure, two-way connections.

The Orion Vision MCP server provides:

  • Seamless integration with Azure Form Recognizer / Document Intelligence

  • Document analysis and form data extraction capabilities

  • Support for various document types (receipts, invoices, ID documents, etc.)

  • Type-safe operations with TypeScript

Prerequisites πŸ”§

Before you begin, ensure you have:

  • Azure Form Recognizer / Document Intelligence endpoint and key

  • Claude Desktop or Cursor

  • Node.js (v20 or higher)

  • Git installed (only needed if using Git installation method)

Related MCP server: MCP PDF

Orion Vision MCP server installation ⚑

Running with NPX

npx -y orion-vision-mcp@latest

Installing via Smithery

To install Orion Vision MCP Server for Claude Desktop automatically via Smithery:

npx -y @smithery/cli install @orion-vision/mcp --client claude

Configuring MCP Clients βš™οΈ

Configuring Cline πŸ€–

The easiest way to set up the Orion Vision MCP server in Cline is through the marketplace with a single click:

  1. Open Cline in VS Code

  2. Click on the Cline icon in the sidebar

  3. Navigate to the "MCP Servers" tab (4 squares)

  4. Search "Orion Vision" and click "install"

  5. When prompted, enter your Azure Form Recognizer credentials

Alternatively, you can manually set up the Orion Vision MCP server in Cline:

  1. Open the Cline MCP settings file:

# For macOS:
code ~/Library/Application\ Support/Code/User/globalStorage/saoudrizwan.claude-dev/settings/cline_mcp_settings.json

# For Windows:
code %APPDATA%\Code\User\globalStorage\saoudrizwan.claude-dev\settings\cline_mcp_settings.json
  1. Add the Orion Vision server configuration to the file:

{
  "mcpServers": {
    "orion-vision-mcp": {
      "command": "npx",
      "args": ["-y", "orion-vision-mcp@latest"],
      "env": {
        "AZURE_FORM_RECOGNIZER_ENDPOINT": "your-endpoint-here",
        "AZURE_FORM_RECOGNIZER_KEY": "your-key-here"
      },
      "disabled": false,
      "autoApprove": []
    }
  }
}
  1. Save the file and restart Cline if it's already running.

Configuring Cursor πŸ–₯️

Note: Requires Cursor version 0.45.6 or higher

To set up the Orion Vision MCP server in Cursor:

  1. Open Cursor Settings

  2. Navigate to Features > MCP Servers

  3. Click on the "+ Add New MCP Server" button

  4. Fill out the following information:

    • Name: Enter a nickname for the server (e.g., "orion-vision-mcp")

    • Type: Select "command" as the type

    • Command: Enter the command to run the server:

    env AZURE_FORM_RECOGNIZER_ENDPOINT=your-endpoint AZURE_FORM_RECOGNIZER_KEY=your-key npx -y orion-vision-mcp@latest

    Important: Replace your-endpoint and your-key with your Azure Form Recognizer credentials

Configuring the Claude Desktop app πŸ–₯️

For macOS:

# Create the config file if it doesn't exist
touch "$HOME/Library/Application Support/Claude/claude_desktop_config.json"

# Opens the config file in TextEdit
open -e "$HOME/Library/Application Support/Claude/claude_desktop_config.json"

For Windows:

code %APPDATA%\Claude\claude_desktop_config.json

Add the Orion Vision server configuration:

{
  "mcpServers": {
    "orion-vision-mcp": {
      "command": "npx",
      "args": ["-y", "orion-vision-mcp@latest"],
      "env": {
        "AZURE_FORM_RECOGNIZER_ENDPOINT": "your-endpoint-here",
        "AZURE_FORM_RECOGNIZER_KEY": "your-key-here"
      }
    }
  }
}

Usage in Claude Desktop App 🎯

Once the installation is complete, and the Claude desktop app is configured, you must completely close and re-open the Claude desktop app to see the orion-vision-mcp server. You should see a hammer icon in the bottom left of the app, indicating available MCP tools.

Orion Vision Examples

  1. Analyze a Document:

Analyze the document at "https://example.com/document.pdf" using Azure Form Recognizer.
  1. Extract Form Data:

Extract data from the invoice at "https://example.com/invoice.pdf".
  1. Process ID Document:

Process the ID document at "https://example.com/id.pdf" and extract relevant information.

Troubleshooting πŸ› οΈ

Common Issues

  1. Server Not Found

    • Verify the npm installation by running npm --version

    • Check Claude Desktop configuration syntax

    • Ensure Node.js is properly installed by running node --version

  2. Azure Form Recognizer Credentials Issues

    • Confirm your Azure Form Recognizer endpoint and key are valid

    • Check the credentials are correctly set in the config

    • Verify no spaces or quotes around the credentials

  3. Document Processing Issues

    • Verify the document URL is accessible

    • Check the document format is supported

    • Ensure the document is not corrupted or password-protected

Acknowledgments ✨

  • Model Context Protocol for the MCP specification

  • Anthropic for Claude Desktop

  • Microsoft Azure for Form Recognizer / Document Intelligence

Available Tools

2 tools
analyze-documentC

Analyzes a document using Azure Form Recognizer and returns structured data

ParametersJSON Schema
NameRequiredDescriptionDefault
modelIdNoOptional model ID for custom models
urlYesURL of the document to analyze

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the analysis method and output type but omits critical details like authentication needs, rate limits, processing time, error handling, or what 'structured data' entails, leaving significant gaps for a tool performing external API calls.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with zero waste, clearly front-loading the core functionality. Every word contributes to understanding the tool's purpose without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of document analysis with an external service, no annotations, and no output schema, the description is incomplete. It lacks details on behavioral traits, output format, error cases, and usage context, which are essential for effective tool invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters (modelId and url). The description adds no additional parameter semantics beyond what the schema provides, such as examples or constraints, meeting the baseline for high coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('analyzes') and resource ('a document') using Azure Form Recognizer, with the outcome of returning structured data. It distinguishes from the sibling 'extract-form-data' by specifying the analysis method, though not explicitly contrasting their use cases.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus the sibling 'extract-form-data' or other alternatives. The description implies usage for document analysis but lacks context on prerequisites, constraints, or comparative scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

extract-form-dataC

Extracts structured data from forms using Azure Form Recognizer

ParametersJSON Schema
NameRequiredDescriptionDefault
formTypeYesType of form to analyze
urlYesURL of the form document to analyze

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. While it mentions the technology (Azure Form Recognizer), it doesn't describe what happens during extraction - whether it's a read-only operation, if it modifies data, authentication requirements, rate limits, or error handling. For a tool with no annotation coverage, this leaves significant behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise - a single sentence that directly states the tool's purpose without any unnecessary words. It's front-loaded with the core functionality and doesn't waste space on redundant information. Every word earns its place in this minimal description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that there's no output schema and no annotations, the description should provide more context about what the tool returns and how it behaves. For a data extraction tool with 2 required parameters, the description is too minimal - it doesn't explain the extraction results format, error conditions, or practical usage examples. The completeness is inadequate for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters thoroughly. The description doesn't add any meaningful parameter semantics beyond what's in the schema - it doesn't explain how 'formType' affects extraction results or provide examples of valid URLs. With complete schema coverage, the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Extracts structured data from forms using Azure Form Recognizer'. It specifies the action (extracts), resource (structured data from forms), and technology (Azure Form Recognizer). However, it doesn't explicitly differentiate from its sibling 'analyze-document', which might have overlapping functionality.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus its sibling 'analyze-document' or other alternatives. It doesn't mention prerequisites, limitations, or specific scenarios where this tool is preferred. The only implied context is that it works with forms, but no explicit usage instructions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 2 tool updatesv1.0.0
    • First observedanalyze-document
    • First observedextract-form-data

TDQS

C2.6/5.0
Disambiguation1/5

The two tools have nearly identical purposes: both use Azure Form Recognizer to extract structured data from documents/forms. 'analyze-document' and 'extract-form-data' are functionally indistinguishable, with no clear boundary between them. This high ambiguity will cause agents to misselect between tools.

Naming Consistency3/5

Both tools use kebab-case naming, which is consistent. However, the verb choices ('analyze' vs 'extract') are different despite similar functionality, creating minor inconsistency. The naming pattern is readable but not perfectly aligned in purpose.

Tool Count2/5

With only 2 tools, the server feels thin for a vision/document processing domain. A typical MCP server for this scope would include more operations like text extraction, image analysis, or OCR configuration. The minimal tool count limits functionality and suggests incomplete coverage.

Completeness2/5

For a vision/document processing server, there are significant gaps: no image analysis tools, no OCR configuration, no batch processing, and no support for different document types beyond forms. The surface is severely limited, focusing only on form data extraction without broader vision capabilities.

Maintenance

ActivityInactive
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents and users to process documents through natural language, supporting PDF operations like text extraction, redaction, splitting, form filling, annotations, and content search.
    275
    61
    MIT
  • A
    license
    Not graded
    quality
    F
    maintenance
    Enables AI-powered extraction and analysis of PDF documents with 40+ specialized tools for text, tables, images, layout analysis, security assessment, and document intelligence. Supports both text-based and scanned PDFs with OCR capabilities.
    10
    MIT
  • A
    license
    Not graded
    quality
    F
    maintenance
    Enables PDF document processing including text, image, and table extraction, as well as intelligent classification and similarity analysis across multiple languages.
    49
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/Cognitive-Stack/orion-vision-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server