Skip to main content
Glama
448,306 tools. Updated 2026-08-12 11:30

"OCR agent" matching MCP tools:

  • Extract visible text from your screen using OCR. Returns text grouped by detected windows with bounding boxes. Use when you need code, terminal, chat, or document content without visual layout. Avoids sending data to the cloud.
    AGPL 3.0
  • Extract transcript, key frames, OCR text, and annotated timeline from video URLs. Supports Loom, YouTube, direct files, and more.
    MIT
  • Specify a screen region with device pixel coordinates to extract text via OCR, returning JSON with text, confidence, and bounding box for targeted analysis.
    MIT
  • Verify identity documents by submitting front and optional back images. Get structured OCR data and authenticity checks for fraud prevention.
    MIT
  • Convert scanned PDFs into searchable, editable documents using OCR. Choose quality, language, and skip OCR when text is already digital.
    MIT

Matching MCP Servers

Matching MCP Connectors

  • OCR for images and Korean ID documents

  • Scan any URL for AI agent readability — Vercel Spec, llmstxt.org, and agent-protocol manifests.

  • Capture desktop, window, or region screenshots. Select detail level: metadata for window positions, UI elements for interactive controls, or OCR for text recognition.
    MIT
  • Capture any window, region, or hidden background and get metadata, UI elements, or OCR text with click coordinates to automate desktop interactions.
    MIT
  • Extract text from PDFs located in Notion, shared mounts, local paths, or Drive without moving bytes over MCP. Supports OCR for scanned/image-only PDFs.
    MIT
  • Retrieve OCR text from a Joplin image or scanned PDF without downloading the attachment. Use when you need the text content of a resource after OCR has finished processing.
    MIT
  • Check which OCR engines are available on your machine before selecting an engine for text recognition.
    MIT
  • Run OCR on scanned PDFs to add a searchable text layer using Tesseract, making the text copyable and searchable.
    MIT
  • Navigate to a URL and extract page content as DOM elements and OCR text, defaulting to zero image tokens to reduce costs. Optionally capture a visual screenshot.
    MIT
  • Load a saved game from Civilization VI's main menu using OCR-guided clicking. Specify a save name or load the most recent autosave.
    MIT