VisionCLI
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
| prompts | {
"listChanged": true
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| lookA | Take a screenshot of what the user is looking at and return it as an image. By default captures the front-most window that is not the terminal (target "window"). Use target "screen" for a whole display. Optionally pick a window by app name/title or window id. |
| list_windowsA | List on-screen app windows (front-most first) with ids, app names and titles. Use before |
| select_regionA | Show the native crosshair so the user can drag a rectangle around the exact part of the screen they mean (press Space to pick a window instead, Esc to cancel). Blocks until they finish. Tell the user to make the selection. |
| clipboard_imageB | Return the image currently on the user's clipboard (e.g. a screenshot taken with Cmd+Ctrl+Shift+4). |
| device_screenshotA | Full-resolution screenshot of the booted iOS Simulator (platform ios) or a connected Android device/emulator (platform android), without window chrome. Use it to explain or debug app UI, then find the view in code. |
| db_schemaA | Tables, columns and foreign keys of the project's database (connection read from the project's .env: Laravel DB_* or DATABASE_URL; MySQL, Postgres, SQLite). Pass a table to narrow it. Read-only. |
| db_queryA | Run one SELECT/SHOW/DESCRIBE/EXPLAIN/WITH query against the project's database inside a read-only transaction. Results are capped at 50 rows; use LIMIT. |
| browser_tabsA | Open tabs in the front-most Brave/Chrome (ids like t1.3 = window 1, tab 3), marking the active one. |
| browser_openB | Open a URL in the browser, in a new tab (default) or the current one. |
| browser_switch_tabB | Bring a tab (id from browser_tabs) to the front. |
| browser_navigateC | Navigate the active tab. |
| browser_readA | Title, URL, selected text and readable text of the active tab (capped). Faster and more exact than a screenshot. |
| browser_elementsA | Visible links, buttons and form fields on the active page with ids (e1, e2, …) for browser_click / browser_type. |
| browser_clickB | Click an element by id from browser_elements (or by its visible text). |
| browser_typeB | Set the value of a form field (id from browser_elements), optionally submitting the form. |
| browser_scrollC | Scroll the active page. |
| start_voiceA | Launch the voice overlay for this project: the user holds Option+Space, asks a question out loud, and the answer appears in a bubble above their mouse pointer. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
| see | Capture what you're looking at and ask about it |
| point | Drag a box around something on screen and ask about it |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 17 tools
Tools cluster into clear domains (browser control, screen/device capture, database, voice), and the capture tools each target a distinct source (screen, simulator/device, clipboard, user selection). Minor overlap between browser_open (new tab) and browser_navigate (active tab), and between browser_switch_tab and browser_tabs, but descriptions disambiguate them.
Browser tools share a consistent browser_ prefix and db tools a db_ prefix, mostly verb_noun. The vision tools deviate slightly with a bare verb ('look') alongside verb_noun forms (list_windows, select_region, device_screenshot), but all remain snake_case and readable.
17 tools is slightly heavy but justified: browser control legitimately needs ~8 tools, capture needs several sources, database needs schema+query, plus voice. No tool feels redundant or out of scope.
Covers browser navigation, page reading, element interaction, multiple capture sources, DB introspection, and voice. Gaps are minor and mostly by design (read-only DB, no explicit back/forward history navigation), so agents can work around them.