desktop-touch-mcp
Related Servers
Alternatives to desktop-touch-mcp
- AlicenseAqualityBmaintenanceAllow AI agents to see and control a real Windows PC you own: observe (UIA + screenshots), click/type/drag/scroll, launch apps, owner Live View. BYOH — your machine, your key.16Apache 2.0
Related Servers
- AlicenseNot gradedqualityBmaintenanceEnables AI clients on Windows to control the local mouse, keyboard, and screen understanding via MCP stdio, allowing automated workflows like viewing the screen, locating elements, clicking, typing, and verifying results.1Apache 2.0
- AlicenseNot gradedqualityDmaintenanceEnables AI clients to automate Windows desktop applications through window manipulation, image recognition, OCR, keyboard/mouse simulation, and memory operations via the MCP protocol.MIT
- AlicenseAqualityBmaintenanceEnables any AI agent to perceive and control a Windows desktop through live screen capture, OCR, and safety-gated mouse/keyboard actions via a local HTTP API and MCP.162MIT
- AlicenseNot gradedqualityAmaintenanceGives AI agents full control of a Windows computer over MCP through 63 native tools — screenshots and UI trees for vision, mouse/keyboard input, Chrome control via CDP, plus files, terminal, processes, and SSH/SFTP. The agent can look at the screen, click, type, launch apps, run commands, and verify each step with a fresh screenshot.29MIT
- AlicenseAqualityDmaintenanceEnables LLM agents to capture screenshots, control mouse/keyboard, and manage windows on desktop platforms, primarily Windows, via an MCP server.161MIT
- FlicenseNot gradedqualityBmaintenanceEnables AI agents to seamlessly integrate with the Windows operating system, performing tasks such as file navigation, application control, UI interaction, and QA testing via the MCP protocol.-
TDQS
Scored across 30 tools
Multiple clusters overlap in function: three click tools (browser_click, click_element, mouse_click), four browser-discovery tools (overview/search/locate/form), and two shell tools (shell_session, terminal). The descriptions' consistent 'Prefer X over Y' guidance resolves most ambiguity, but an agent must track CDP-vs-UIA-vs-coordinates distinctions across several near-synonym tools.
The browser_* prefix (9 tools) plus mouse_*, workspace_*, and screenshot_* clusters are highly consistent, but the remaining tools mix conventions: verb_noun (click_element, focus_window, run_macro), noun_verb (window_dock, notification_show), and bare nouns (keyboard, scroll, excel). All lowercase snake_case keeps it readable, but action placement is unpredictable across the set.
At 30 tools this exceeds the 25+ threshold and creates a heavy context and selection burden for agents. Several tools are peripheral or mergeable (screenshot_query/screenshot_gc are both cache management, server_status is diagnostics, notification_show is a one-shot utility), though the genuinely broad desktop-automation scope softens the excess somewhat.
The browser lifecycle is well covered (open, navigate, click, fill, observe, wait), but the native-desktop workflow has dead ends: desktop_discover and desktop_act are referenced as required precursors in multiple descriptions yet are absent from the tool set, so agents following documented guidance will hit tool-not-found errors. Missing browser_close/tab management and first-class element enumeration are notable gaps.