An open-source MCP server that lets any coding agent operate your computer like a person does—reading the UI through accessibility trees, clicking and typing in the background, showing an agent pointer, zooming into regions, and driving your signed-in Chrome on macOS, Windows, and Linux.
Enables AI agents to see and drive Kasm Workspaces desktops by starting and stopping sessions, taking screenshots, and performing clicks, typing, key presses, and scrolling through MCP clients or as a native Hermes Agent plugin.
Enables Claude to see and control a real Android phone via MCP, reading the accessibility tree to tap, type, navigate, launch apps, and diagnose connectivity over ADB or cellular.
Exposes Anthropic's computer-use action surface (screenshot, click, move, keyboard, clipboard, batch) against a persistent desktop display via MCP stdio protocol. Enables AI agents to control a virtual desktop environment through natural language instructions.
Enables safe automation of Chrome browser through a local MCP server and Chrome extension, allowing LLMs to control browser tabs, pages, and computer-use actions with permission controls.
Enables AI agents to control a persistent Chrome browser through a local daemon, with Jev performing routine actions and the host agent handling reasoning, text, and verification.
Enables AI agents to autonomously operate web browsers through the Model Context Protocol, including navigating pages, clicking, typing, filling forms, and extracting structured data.
Enables MCP agents to run, continue, observe, check status of, and abort bounded GUI automation tasks through a shared local runtime with visual detection and deterministic input.
Enables MCP-capable agents to run TypeSafe's Jev judgment model as typed yes/no, choice, and score tools, with calibrated probabilities, confidence thresholds, escalation for uncertain or non-judgment tasks, and an optional action gate that fails open.
Enables local Chrome DOM access through MCP for reading visible page text, operating DOM controls, and managing tabs without screenshots or screen sharing.
A framework-agnostic computer-use MCP server that exposes core desktop operations (screen capture, mouse, keyboard, and file access) as standard MCP tools, enabling any MCP-compatible agent to drive a computer.
Enables MCP clients to discover desktop applications with structured metadata and receive native accessibility trees and screenshot content through Open Computer Use, preserving upstream tool results and image blocks.
Enables Windows desktop automation via MCP, allowing AI agents to control mouse, keyboard, and screen capture with the same interface as Anthropic's computer-use tool.
Standalone MCP server that gives AI agents full GUI control over macOS — screenshots, mouse, keyboard, apps, clipboard, and multi-display — with zero private dependencies.
Enables AI agents to see, locate UI elements, and operate any Windows desktop app through natural language, using accessibility-tree matching with optional vision-model fallback, plus an autonomous visual loop with introspection and meta-learning.