GhostDesk
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| TZ | No | IANA timezone (POSIX standard, e.g. 'Europe/Paris'). | America/New_York |
| LANG | No | POSIX locale (e.g. 'fr_FR.UTF-8'). | en_US.UTF-8 |
| GHOSTDESK_PORT | No | MCP server listening port. | 3000 |
| GHOSTDESK_TLS_KEY | No | Path to the TLS private key (matching GHOSTDESK_TLS_CERT). | /etc/ghostdesk/tls/server.key |
| GHOSTDESK_TLS_CERT | No | Path to the TLS certificate. When the file exists, 'websockify' and the MCP server auto-switch to 'wss://' / 'https://'. | /etc/ghostdesk/tls/server.crt |
| GHOSTDESK_AUTH_TOKEN | No | Bearer token required on every MCP request. Generate with 'openssl rand -hex 32'. | |
| GHOSTDESK_SCREEN_WIDTH | No | Virtual screen width in pixels. | 1280 |
| GHOSTDESK_VNC_PASSWORD | No | Password for wayvnc (username is 'agent' in the prod image). Generate with 'openssl rand -hex 16'. | |
| GHOSTDESK_SCREEN_HEIGHT | No | Virtual screen height in pixels. | 1024 |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| screen_shotA | Capture the screen, optionally cropped to a region. Args: region: Area to capture (full screen if omitted). format: "webp" (default, smaller payload) or "png" (lossless). stabilize: Wait for the page to stop moving before capturing (max 2.5 s). Useful right after navigation. |
| mouse_clickA | Click at screen coordinates. Use coordinates from screen_shot() or inspect(). Returns a dict with:
|
| mouse_double_clickA | Double-click at screen coordinates. Use for opening files or selecting words. Returns a dict with:
|
| mouse_dragA | Drag from one position to another. Use for selecting text, moving items, or resizing. Returns a dict with:
|
| mouse_scrollA | Scroll at a position. direction: up/down/left/right. amount: number of scroll steps (max 5). Returns a dict with:
|
| key_typeA | Type text. Handles Unicode, newlines, and tabs. Returns the standard |
| key_pressA | Press a key or key combination. Friendly names accepted: Examples: Returns the standard |
| app_listA | Return installed GUI apps. Scans Returns a list of dicts, each with:
|
| app_launchA | Launch a desktop GUI application and return its PID and log file path. Only applications listed by Returns a dict with:
On failure, returns a dict with a single |
| app_statusA | Check whether a launched app is still running and read its logs. Only PIDs returned by Args:
pid: Process ID returned by Returns a dict with:
On failure, returns a dict with a single |
| clipboard_getA | Read the current clipboard text. |
| clipboard_setA | Write text to the clipboard. Use with key_press("ctrl+v") to paste. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 12 tools
Each tool has a clearly distinct purpose with no overlap: app_launch, app_list, and app_status form a coherent app management group; clipboard_get/set handle clipboard operations; key_press and key_type cover keyboard input; mouse_click, mouse_double_click, mouse_drag, and mouse_scroll provide distinct mouse actions; and screen_shot handles screen capture. The descriptions clearly differentiate their functions, preventing misselection.
The naming is mostly consistent with a verb_noun pattern (e.g., app_launch, clipboard_get, mouse_click), but there are minor deviations: key_press and key_type use 'key' instead of 'keyboard', and screen_shot uses 'shot' instead of 'capture'. These deviations are minor and do not significantly hinder readability or predictability.
With 12 tools, the count is well-scoped for a desktop automation server covering app management, clipboard, keyboard, mouse, and screen operations. Each tool earns its place, providing a comprehensive yet manageable set for the domain without being overly sparse or bloated.
The tool set offers complete coverage for desktop automation: app management (launch, list, status), clipboard operations (get/set), keyboard input (press/type), mouse actions (click, double-click, drag, scroll), and screen capture. There are no obvious gaps; agents can perform full workflows from launching apps to interacting with them via input and monitoring via screenshots.