VisionCLI
Related Servers
Alternatives to VisionCLI
No user-submitted related servers found.
Related Servers
- AlicenseNot gradedqualityCmaintenanceEnables Claude to see and interact with any macOS application using natural language commands. Perfect for testing Mac applications, UI automation, and app development with AI assistance.33MIT
- AlicenseNot gradedqualityAmaintenanceEnables driving macOS apps from Claude Code — screenshots, accessibility trees, clicks, typing and keys — by borrowing the Computer Use eyes and hands from OpenAI's ChatGPT/Codex desktop app, with persistent JavaScript scripting, text-targeted batching, diff output and macros that collapse many actions into one call without consuming Codex quota or calling a model.MIT
- AlicenseAqualityDmaintenanceEnables AI agents to capture and analyze screenshots of macOS applications, windows, or the entire screen using local (Ollama) or cloud-based AI vision models, with non-intrusive, fast screen capture via Apple's ScreenCaptureKit.312 npm2MIT
- AlicenseNot gradedqualityCmaintenanceProvides AI coding assistants with privacy-first, synchronized screen context including cursor position, active window metadata, and visual verification to aid in building and debugging graphical interfaces.MIT
- AlicenseNot gradedqualityBmaintenanceScreenshot and diagram tool for AI agents. Capture and annotate screenshots to show Claude what you mean — or let the agent render Mermaid diagrams and open them for visual review. Approve, annotate, or request changes with text feedback. Built-in review mode with structured responses. CLI and MCP server for Claude Code, Cursor, Windsurf, Cline. macOS, open source, free.12 npm327MIT
- AlicenseAqualityBmaintenanceMCP server that lets Claude use your Mac — native apps and the browser — by remembering what it has seen instead of screenshotting every step.121MIT
TDQS
Scored across 17 tools
Tools cluster into clear domains (browser control, screen/device capture, database, voice), and the capture tools each target a distinct source (screen, simulator/device, clipboard, user selection). Minor overlap between browser_open (new tab) and browser_navigate (active tab), and between browser_switch_tab and browser_tabs, but descriptions disambiguate them.
Browser tools share a consistent browser_ prefix and db tools a db_ prefix, mostly verb_noun. The vision tools deviate slightly with a bare verb ('look') alongside verb_noun forms (list_windows, select_region, device_screenshot), but all remain snake_case and readable.
17 tools is slightly heavy but justified: browser control legitimately needs ~8 tools, capture needs several sources, database needs schema+query, plus voice. No tool feels redundant or out of scope.
Covers browser navigation, page reading, element interaction, multiple capture sources, DB introspection, and voice. Gaps are minor and mostly by design (read-only DB, no explicit back/forward history navigation), so agents can work around them.