Provides screen capture and optical character recognition (OCR) capabilities for entire displays or specific application windows. It enables users to list running applications, take screenshots, and extract text from images using multi-language support.
A basic MCP server that provides arithmetic calculation tools and system information retrieval capabilities, including platform, CPU, and memory details.
Enables Windows GUI automation through the Model Context Protocol, providing tools, workflows, and generative capabilities for natural language-driven interaction with desktop applications.
Enables an MCP-capable agent to operate a heterogeneous cluster of Debian/Armbian Linux nodes over SSH from a single inventory, inspecting status and logs, managing packages, services, files, storage and NFS, and deploying new capabilities to nodes through installable modules. Built-in cluster tools fan out to nodes by explicit names, roles or tags with bounded concurrency, while modules can be installed on compatible nodes and run on demand or as systemd services.
Provides comprehensive computer control capabilities including mouse and keyboard automation, screen capture, OCR text recognition, and window management through MCP protocol.
Enables LLMs to capture and analyze screenshots of your screen, windows, or regions with smart detection capabilities. Features natural language queries, automatic window targeting, and text enhancement for UI debugging and visual inspection.
kira-mcp is a local MCP server that gives any MCP-compatible agent host full computer-use capabilities on the host machine, including screen perception with OmniParser and desktop automation.
Enables desktop agents such as Claude Code and Codex to control an Android phone as a long-lived action endpoint, pairing over outbound WebSocket and exposing phone capabilities as MCP tools with local pause and revocation.
Enables AI assistants to automate Wayland desktop environments through screenshot analysis, mouse control, and keyboard input simulation. It supports visual context via VLM providers like Gemini and OpenRouter to perform complex, multi-step desktop actions.
An MCP server that bridges AI agents with GUI automation capabilities, allowing them to control mouse, keyboard, windows, and take screenshots to interact with desktop applications.
MCP server that provides computer control capabilities including mouse movements, keyboard actions, screenshot capture with OCR, and window management through a unified API.
Enables read-only MCP access to Tactical RMM, allowing agents to query devices, clients, sites, audit logs, software, and pending actions without any mutating capabilities.
This server provides:
* Fast file search capabilities using Everything SDK
* Windows-specific implementation
* Complements existing filesystem servers with specialized search functionality
A modified JetBrains MCP Server that adds WebSocket monitoring capabilities, allowing users to monitor MCP tool calls in real-time while maintaining compatibility with the original implementation.
Lets ChatGPT Web act as the reasoning layer for an agent that runs shell commands, reads and writes files inside a constrained workspace, captures screenshots as image content, and controls the mouse and keyboard. It also keeps persistent Markdown memory and identity/skill context files that load automatically into the conversation.
Provides tools for executing shell commands both synchronously and asynchronously with real-time output streaming and process management capabilities. It enables users to start background tasks, monitor progress, and manage long-running processes via Stdio or HTTP transports.
A Model Context Protocol server that provides desktop automation capabilities using RobotJS and screenshot capabilities, enabling LLMs to control mouse movements, keyboard inputs, and capture screenshots of the desktop environment.