Business Card Watchdog is an installable, user-scoped service for processing large batches of business card images from watched folders. It exposes CLI, API, MCP, and watched-folder surfaces for deterministic orchestration and routing to Google Contacts or Odoo.
Provides image recognition capabilities to Codex/Claude Code by routing 'look at screen/screenshot' requests to a multimodal model via an API relay. Enables UI automation agents to analyze local screenshots and return descriptions or structured JSON coordinates.
Enables AI clients like Codex, Claude Code, and Kimi Code to analyze public video URLs by downloading media, uploading it to Gemini, and returning timestamped production breakdowns covering shots, visual design, animation, motion, narration, music, sound effects, and editing, with follow-up Q&A and session management tools.
Enables intelligent document processing by extracting text, classifying document types, and generating structured summaries from PDFs and images using vision LLMs.
A streamlined MCP server for XMP metadata embedding with beautiful formatting and smart filename indicators, enabling metadata embedding, reading, validation, and report generation for lifestyle, product, and orbit schemas.
This server enables interaction with Google's Video Intelligence API for advanced video analysis, auto-generated using AG2's MCP builder to provide a standardized multi-agent interface.
MCP server for MarkItUp's AI image-annotation pipeline. Generate polished marketing-visual variations of any screenshot, regenerate, AI outpaint, and remove backgrounds —
powered by Claude analysis + Gemini rendering.
MCP server for FreezeText — OCR anything on your Mac screen from Claude, Cursor, or any MCP client. Freeze the screen and extract text via Apple Vision (videos, popups, protected PDFs), OCR a region or a base64 image, and manage a searchable capture history. 12 tools. Bridge open-source (MIT), FreezeText app is free.
Enables text-only agents to process images by accepting image files, base64 data, or URLs, sending them to multimodal models, and returning structured text results via MCP.
Enables natural language interaction with complex computer vision workflows such as auto-labeling, class mapping, and embedding selection through an LLM-agnostic MCP orchestration layer.
An MCP server that recreates Microsoft Comic Chat (1996) to generate comic strips from conversation summaries, using original character art and deterministic layout, with no LLM calls.
Enables AI image generation, editing, and composition using Google's Gemini image models (Nano Banana Pro and Nano Banana). Supports text-to-image generation, multi-image composition, flexible aspect ratios, high-resolution output up to 4K, and real-time information grounding.
MCP server that gives AI agents visual intelligence — search Pinterest, analyze images with LLM vision, build a semantic reference library, and retrieve by style or mood.
MCP server for ScanToBill, an AI-powered OCR platform that extracts structured JSON from invoices and business documents in Arabic and English. It enables Claude Desktop to extract invoice data from URLs, list past extractions, and retrieve detailed results.
An MCP server that integrates with Siril astronomical image processing software to process Seestar telescope images, offering tools for binary detection, version checking, mosaic processing, and project analysis.