Skip to main content
Glama
452,745 tools. Updated 2026-08-13 12:35

"A server for image recognition and OCR" matching MCP tools:

  • Extract text from local image files using OCR. Converts image-based content into structured JSON and a text summary for downstream processing.
    MIT
  • Capture desktop, window, or region screenshots. Select detail level: metadata for window positions, UI elements for interactive controls, or OCR for text recognition.
    MIT
  • Retrieve OCR text from a Joplin image or scanned PDF without downloading the attachment. Use when you need the text content of a resource after OCR has finished processing.
    MIT

Matching MCP Servers

Matching MCP Connectors

  • OCR for images and Korean ID documents

  • Arabic-first OCR, translation and document extraction. First call mints a free trial key.

  • Capture screenshots and extract text with OCR in one step, returning both the image and text with word coordinates for Hyprland desktop automation.
    MIT
  • Extract visual and textual context from video, audio, or image files locally—frames, transcripts, OCR, and scene changes—for AI analysis.
    Apache 2.0
  • Check which OCR engines are available on your machine before selecting an engine for text recognition.
    MIT
  • Navigate to a URL and extract page content as DOM elements and OCR text, defaulting to zero image tokens to reduce costs. Optionally capture a visual screenshot.
    MIT
  • Extract text from images (JPEG/PNG) or single-page PDFs using Yandex Vision OCR. Supports printed, handwritten, table, and markdown models with language selection.
    MIT
  • Extract visible text from local images using OCR. Ideal for code screenshots, terminal output, and error dialogs.
    MIT
  • Resolve local or OpenCode image paths, analyze via multiple vision APIs and OCR, and return JSON context for non-vision models.
    MIT
  • Evaluate OCR output by comparing extracted text or an image/PDF (OCR'd automatically) against a ground truth file. Returns CER, WER, and accuracy metrics.
    MIT
  • Extract text from PDF documents (single or multi-page) using Yandex Vision OCR with asynchronous recognition. Supports models for page, table, handwritten, and markdown output.
    MIT
  • Extract and process images from file paths for visual content analysis, OCR text extraction, and object recognition. Supports screenshots, photos, diagrams, and documents in PNG, JPG, GIF, and WebP formats.
    MIT
  • Execute a performance benchmark on OCR and YOLO detection using a provided base64 image or a generated synthetic image.
    MIT
  • Extract text from images using OCR. Provide a public image URL or base64-encoded image data to recognize text in multiple output formats, with cost of $0.05 USDC on Base.
    MIT