Skip to main content
Glama

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
DISPLAYNoDetected automatically when the host doesn't pass it, preferring the display your desktop session uses
LCU_GLIDENo0: jump instead of the curved spring motion1
XAUTHORITYNoDetected automatically when the host doesn't pass it, preferring the display your desktop session uses
LCU_OVERLAYNo0: don't draw the agent cursor1
LCU_IME_BYPASSNo0: don't switch ibus to a plain engine while typing1
LCU_CURSOR_ICONNoYour own PNG/SVG iconCodex glyph
LCU_CURSOR_SIZENoHeight of a custom icon, in px28
LCU_CURSOR_COLORNoFixed cursor colour, #rrggbbwallpaper
LCU_CURSOR_LABELNoName tag next to the cursor; auto uses the client name (Claude/Codex)none
LCU_CURSOR_SCALENoCursor scale (1.0 = 14 px)1.0
LCU_PLAIN_ENGINENoibus engine used while typingxkb:us::eng
LCU_REMAP_SETTLENoSeconds to wait after rebinding keys for non-layout characters (Hangul, emoji). Raise it if the first such character is dropped or wrong0.12
LCU_MAX_LONG_EDGENoLong edge of screenshots, in px1280
LCU_CURSOR_HOTSPOTNoClick point inside that icon, in px0,0
LCU_SUPERVISOR_LOGNoFile for supervisor restart logsstderr
LCU_VIRTUAL_POINTERNo0: share the user's pointer (no own pointer, no overlay)1
DBUS_SESSION_BUS_ADDRESSNoDetected automatically when the host doesn't pass it, preferring the display your desktop session uses

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}
prompts
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}
experimental
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
screenshotA

Capture the screen. Optionally pass x/y/width/height (screenshot coords) to zoom into a region for reading small text; coordinates you act on must still be full-screen screenshot coordinates.

screen_infoB

Screen size (real and screenshot space), scale, and pointer position.

active_windowA

Where your keystrokes will go (keys_go_to: the window under your pointer when you have your own pointer) and which window the user has focused. Check this before typing.

cursor_positionA

Current mouse pointer position in screenshot coordinates.

clickB

Click at (x, y). count=2 for double-click, 3 for triple-click. modifiers e.g. ["ctrl"] or ["shift"] are held during the click.

mouse_moveB

Move the pointer to (x, y) without clicking (e.g. to reveal hover menus).

dragC

Press at start, move smoothly to end, release.

mouse_downA

Press and hold a mouse button at the current pointer position.

mouse_upC

Release a mouse button.

scrollC

Scroll the wheel at (x, y). amount = number of wheel clicks.

type_textA

Type text into the focused widget. Any Unicode works (Korean, emoji...). Newlines press Return. expect_window: substring of the focused window's title or class; if it doesn't match, nothing is typed.

keyC

Press a key or chord: "Return", "ctrl+c", "ctrl+shift+t", "alt+F4", "super", "Escape", "Page_Down". X keysym names are accepted. expect_window works as in type_text.

hold_keyB

Hold key(s) down for seconds (max 10), e.g. for games or key repeat.

waitB

Wait for the UI (max 30s), then by default return a screenshot.

list_windowsB

Top-level windows known to AT-SPI: app, title, active/visible, box.

ui_treeA

List visible widgets as [id] role "name" @(x,y) with center coordinates in screenshot space. Filter by app/window name substring, or active_window_only=True. Ids are valid until the next ui_tree call. Apps only appear if accessibility is enabled (Chrome/Electron need --force-renderer-accessibility).

click_elementB

Activate a ui_tree element. auto = AT-SPI action if available, else a real mouse click at its center.

set_textB

Replace the contents of an editable ui_tree element. Falls back to focusing it, select-all and typing when the widget isn't directly editable.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

A3.5/5.0

Scored across 18 tools

Disambiguation4/5

Most tools target clearly distinct primitives (coordinate click vs. semantic click_element, type_text vs. set_text, key vs. hold_key, drag vs. mouse_down/up). Minor overlap: screen_info already reports the pointer position that cursor_position returns, and active_window/list_windows/ui_tree all relate to window/widget state and could be briefly confusing. Descriptions largely resolve these ambiguities.

Naming Consistency4/5

All names use consistent snake_case, which reads cleanly across the set. There is some variance between verb-based (mouse_move, list_windows, click_element, set_text) and noun-based (screenshot, cursor_position, active_window, ui_tree) names, but the vocabulary is predictable for a low-level GUI primitive API.

Tool Count4/5

18 tools is slightly above the typical 3-15 range but justified for full GUI automation, since each covers a genuinely distinct primitive (mouse buttons, keys, accessibility, screen capture). No redundant tools inflate the count.

Completeness5/5

The surface covers screen capture and geometry, full mouse control (click, move, drag, down/up, scroll), keyboard input (type, chords, holds), accessibility inspection and activation, plus waiting and focus verification. This is a complete lifecycle for computer-use tasks with no obvious dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues