Cross-platform desktop automation MCP server that lets AI agents capture screenshots, run OCR with UI-element classification, control mouse/keyboard, and launch programs on Linux, macOS, and Windows.
An MCP server that enables automated dataset creation and custom object detection model training through natural language interactions. It integrates foundation models like GroundedSAM for auto-labeling and supports training specialized YOLOv8 models using local or Unsplash images.
A Model Context Protocol server that provides screenshot capabilities for AI assistants, including window, region, and long screenshots across Windows, macOS, and Linux.
Exposes geospatial foundation model inference as MCP tools for LLM agents, enabling building extraction, water segmentation, object detection, change detection, zero-shot scene classification, and embeddings on CPU.
An MCP server that enables AI agents to capture targeted screenshots of specific application windows on Windows and Linux, with smart window state restoration and focus management.
Enterprise-grade screenshot capture server for AI agents with multi-format support, PII masking, multi-monitor support, and security controls for capturing full screens, specific windows, or custom regions across Linux, macOS, and Windows.
Enables scanner capture via SANE on Linux, providing tools for device discovery, scan jobs with ADF/duplex support, document batching, and multipage assembly through a local MCP server.
A reliable MCP server for taking screenshots with customizable timestamp overlays, region captures, and file management capabilities across macOS, Windows, and Linux.
Enables AI agents to segment 2D and 3D biomedical images using Microsoft's MedImageParse foundation models via free-text prompts, returning masks and statistics. Supports both local stdio MCP server and a hosted Azure API Management gateway for remote agents like Microsoft Foundry.