Skip to main content
Glama

Vision MCP Server

MCP server for image processing via Ollama vision models (Gemma 4, Gemma 3, LLaVA...).
Enables LLM clients without vision capability (DeepSeek, Qwen, etc.) to process images by delegating to a local vision model through Ollama.

Features

Tool

Description

describe_image

Describe image content (brief / detailed / exhaustive)

ocr_image

Extract text from image with language hints (vi, en, ja, zh, ko)

ask_image

Ask any question about an image with a custom prompt

process_clipboard_image

Read image directly from macOS clipboard — no file path needed

Related MCP server: mcp-vision

Requirements

  • macOS (clipboard tool uses osascript)

  • Python 3.12+

  • uv — Python package manager

  • Ollama — local LLM runtime

Installation

1. Clone the repo

git clone https://github.com/nguyenduc/vision-mcp-server.git
cd vision-mcp-server

2. Install dependencies

uv sync

uv sync creates .venv/ and installs all packages from uv.lock. No need for pip install or uv init.

3. Pull a vision model

ollama pull gemma4

Other compatible vision models: gemma3, llava, llava-llama3, moondream.

4. Make sure Ollama is running

ollama serve

Verify:

curl http://127.0.0.1:11434/api/tags

5. Test the server

uv run server.py

The server runs over stdio — press Ctrl+C to stop.

MCP Client Configuration

OpenCode

Add to .opencode.json (project-level or ~/.opencode.json):

{
  "mcpServers": {
    "vision": {
      "enabled": true,
      "type": "local",
      "command": ["uv", "run", "server.py"],
      "cwd": "/absolute/path/to/vision-mcp-server",
      "env": ["OLLAMA_BASE_URL=http://127.0.0.1:11434", "VISION_MODEL=gemma4"]
    }
  }
}

Note: In OpenCode, command is an array and env is an array of "KEY=VALUE" strings, not an object.

Claude Desktop

Add to ~/Library/Application Support/Claude/claude_desktop_config.json:

{
  "mcpServers": {
    "vision": {
      "command": "uv",
      "args": ["run", "server.py"],
      "cwd": "/absolute/path/to/vision-mcp-server"
    }
  }
}

Cursor / Windsurf / Cline

{
  "mcpServers": {
    "vision": {
      "command": "uv",
      "args": ["run", "server.py"],
      "cwd": "/absolute/path/to/vision-mcp-server",
      "env": {
        "OLLAMA_BASE_URL": "http://127.0.0.1:11434",
        "VISION_MODEL": "gemma4"
      }
    }
  }
}

Environment Variables

Variable

Default

Description

OLLAMA_BASE_URL

http://127.0.0.1:11434

Ollama API endpoint

VISION_MODEL

gemma4

Model name in Ollama (must have vision capability)

How It Works

┌─────────────┐     ┌───────────────────┐     ┌─────────────┐
│  LLM Client │────▶│  Vision MCP Server │────▶│   Ollama    │
│ (DeepSeek)  │◀────│   (stdio/MCP)      │◀────│  (Gemma 4)  │
└─────────────┘     └───────────────────┘     └─────────────┘
      │                       │
      │ [Image 1] + prompt    │ osascript: clipboard → PNG
      │                       │ base64 → /v1/chat/completions
      ▼                       ▼
  Receives text           Returns vision
  description/OCR         analysis result

Clipboard flow: User pastes image → LLM calls process_clipboard_image → server grabs image from macOS clipboard via osascript → encodes to base64 → sends to Ollama → returns text.

File path flow: User provides path → LLM calls describe_image / ocr_image / ask_image with path → server reads file → encodes → sends to Ollama → returns text.

Troubleshooting

Error

Cause

Fix

404 Not Found

Model doesn't exist in Ollama

ollama pull gemma4

Connection refused

Ollama is not running

ollama serve

No image found in clipboard

Clipboard is empty or not an image

Copy an image to clipboard first

Timeout

Model too large for hardware

Switch to a smaller model: moondream

License

MIT

Install Server
F
license - not found
A
quality
C
maintenance

Maintenance

Maintainers
Response time
Release cycle
Releases (12mo)
Commit activity

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Servers

  • A
    license
    -
    quality
    C
    maintenance
    An MCP server that enables any LLM to describe images from file paths, URLs, or base64 data by forwarding them to a supported vision provider such as OpenAI, Anthropic, or local Ollama models.
    714
    9
    MIT
  • A
    license
    -
    quality
    C
    maintenance
    MCP server that provides a 'borrowed eye' for text-only LLMs, enabling them to identify and describe local images via the Qwen VL vision model, including face recognition, scene description, OCR, and targeted visual questioning.
    1
    Apache 2.0
  • A
    license
    -
    quality
    B
    maintenance
    MCP server for local Ollama vision analysis, enabling text-only agents like Claude Code to inspect images via a single tool. Processes images locally with Ollama, keeping image bytes on the machine and returning text reports.
    2
    MIT

View all related MCP servers

Related MCP Connectors

  • OCR, transcription, file extraction, and image generation for AI agents via MCP.

  • MCP server for Clipkit — gives AI agents a video toolbox via the Clipkit schema.

  • A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…

View all MCP Connectors

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/ducnm9/ollama-vision-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server