Provides screen capture, OCR text extraction, and visual language model scene understanding capabilities with continuous monitoring and automatic memory storage integration.
Enables medical image analysis, structured medical report generation, and medical Q\&A through the Lingshu medical AI model. Provides healthcare professionals and developers with AI-powered medical assistance capabilities via a FastMCP server interface.
MCP server that turns articles, transcripts, and markdown into LinkedIn carousel PDFs, Instagram PNGs, and Threads PNGs. Content in, slides out. No web UI, no cloud service.
Enables coding agents to turn videos from social platforms or local files into a small set of distinct frames plus a manifest, so they can answer visual questions by reading image paths. Also provides metadata lookup for captions and authors without downloading the video.
Local stdio MCP service that calls GPT Image API to generate images from text, edit images with references or masks, and save images locally while returning absolute paths and file URIs.
Enables image and video generation across GPT-Image, Gemini, Grok, and Jimeng with file-based outputs, multi-reference support, and model capability lookup.
MCP server for MarkItUp's AI image-annotation pipeline. Generate polished marketing-visual variations of any screenshot, regenerate, AI outpaint, and remove backgrounds —
powered by Claude analysis + Gemini rendering.
Enables AI assistants to generate images from text prompts and transform existing images using Google Gemini's nano banana model through the Nanana AI service. Supports both text-to-image generation and image-to-image transformation capabilities.
MCP server for AI image generation supporting text-to-image and image-to-image editing via any OpenAI-compatible service, with configurable models, aspect ratios, and sizes.
Exposes local Umi-OCR v2 capabilities to AI agents via MCP, enabling image text extraction, batch OCR, PDF OCR, and status checks without manually starting the service.
A self-hostable MCP server that exposes Roshan AI's Alefba OCR service as tools, enabling document reading, page extraction, status polling, and export to PDF/Word/Excel for Persian, Arabic, and English text.
Enables AI agents to use Kakao's Local, Search, and Vision APIs for address conversion, place search, web/blog/search, and OCR, requiring only a REST API key.
Enables creating and managing HappyHorse video generation tasks (edit, image-to-video, text-to-video) via RunAPI, with optional polling for completion and pricing lookup.
Combines screen vision, RuneLite HTTP API, click coordinates, OSRS Wiki lookups, and Grand Exchange prices to give AI agents eyes and knowledge for playing Old School RuneScape. Enables game state analysis, skill tracking, item/NPC lookup, and interactive automation.
Provides tools for generating optimized images via Google's Gemini model and fetching weather forecasts and alerts from the National Weather Service. It enables users to create visual content and retrieve environmental data seamlessly within MCP-compatible clients.