Skip to main content
Glama
97,619 servers. Updated

Matching MCP tools:

Matching MCP Connectors:

  • A
    license
    A
    quality
    A
    maintenance
    A minimal local vision bridge for text-only VS Code Copilot, enabling it to see images by sending them to a local Ollama vision model and returning text descriptions. It provides MCP tools for image description, listing, OCR, and status checks.
    4
    1
    Apache 2.0
  • A
    license
    A
    quality
    A
    maintenance
    A Python MCP server that gives vision capabilities to text-only LLMs by exposing an analyze_image tool that sends local images to a vision-capable Ollama model and returns textual descriptions.
    1
    MIT
  • F
    license
    A
    quality
    D
    maintenance
    MCP server enabling LLM clients without vision capability to process images by delegating to local Ollama vision models. Supports describing images, OCR, asking questions, and processing clipboard images.
    4
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables local image analysis and document processing through MCP workers (Vision and Document) using Ollama models, with cloud orchestration for mixed hierarchical tasks. Supports tools for image understanding, requirement extraction, test case generation, and GPU memory lifecycle management.
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    Enables AI agents to capture and analyze screenshots of macOS applications, windows, or the entire screen using local (Ollama) or cloud-based AI vision models, with non-intrusive, fast screen capture via Apple's ScreenCaptureKit.
    3
    12 npm
    2
    MIT
  • F
    license
    A
    quality
    D
    maintenance
    An MCP server that captures screenshots of URLs or local app windows and analyzes them with a local Ollama vision model, enabling Claude to visually inspect web pages and desktop applications without sending image data externally.
    3
    -
  • A
    license
    A
    quality
    B
    maintenance
    Provides local vision and audio perception for MCP-compatible agents, enabling them to read images, transcribe text from visual media, analyze videos, and convert speech to text entirely on-device. It is privacy-focused with no cloud upload or API keys by default, using Ollama for inference.
    6
    1
    MIT
  • A
    license
    B
    quality
    C
    maintenance
    Enables Claude Code to call local generative AI services on a LAN, letting users list and chat with Ollama models and list checkpoints, LoRAs, and samplers or generate and save images via Stable Diffusion WebUI Forge Neo. Requests can be made in natural language, with per-call options for model, image size, and Hires.fix settings.
    1
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables AI assistants to control Adobe Photoshop through natural-language requests for tasks like background removal, portrait retouching, web and social export, carousel creation, and batch watermarking. This privacy-focused fork supports a local, free Ollama provider so prompts and images can be processed without an API key or cloud account.
    14,938 npm
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    An MCP server that enables any LLM to describe images from file paths, URLs, or base64 data by forwarding them to a supported vision provider such as OpenAI, Anthropic, or local Ollama models.
    1,150 npm
    9
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    MCP server for local Ollama vision analysis, enabling text-only agents like Claude Code to inspect images via a single tool. Processes images locally with Ollama, keeping image bytes on the machine and returning text reports.
    2
    MIT