Skip to main content
Glama

English | 简体中文

License: MIT TypeScript MCP Node Sandbox Status

Why · How it works · Quickstart · MCP tools · Support · Roadmap


Poltr is a local MCP server that gives any agent (Claude, Gemini, or your own) an isolated desktop, a demonstration recorder, and a skill engine.

🧠 The agent reasons. 🤖 Poltr stays deterministic: recording, validation, execution, verification and safety.

💡 Why

Computer-use agents re-reason every step of every run. That is slow, expensive and flaky.

Without Poltr

With Poltr

🐌 The agent re-plans the whole task each run

⚡ Learn once, replay a stored semantic skill

🎲 Clicks depend on screenshots and guesses

🎯 Targets are resolved semantically, never by replaying coordinates

🤷 "It probably worked"

✅ Every step is deterministically verified

💥 A UI change breaks everything

🩹 Failures produce evidence, the agent proposes a repair, Poltr tests it and saves v2

Related MCP server: windows-gui-mcp

🔄 How it works

Principle

What it means

🔌

Model-agnostic

Recording, validation, storage and execution need no model API or key

🧾

Evidence-grounded

Every skill step must reference real recording steps and screenshots

🙅

Honest failures

Unsupported targets, actions and checks return structured errors, never silent success

🛡️

Safe by default

Bounds checks, blocked hotkeys, typing and scroll limits, app launch limited to approved sandbox apps (no generic shell)

🌱

Versioned healing

A repair creates skill.v2 from v1. Originals are never modified and runtime parameter values are never stored

🚀 Quickstart

Requires Node 18+. The sandbox needs Linux with Xvfb, Openbox, xdotool, dbus-daemon and at-spi2-core (at-spi-bus-launcher). GTK/ATK applications such as Mousepad must expose accessibility through their toolkit bridge.

git clone https://github.com/janghotan/Poltr && cd Poltr
npm install
npm run build

node dist/src/cli/index.js sandbox start
node dist/src/cli/index.js computer probe --computer sandbox

🔗 Connect an agent over stdio

{
  "mcpServers": {
    "poltr": {
      "command": "node",
      "args": ["dist/src/cli/index.js", "mcp"]
    }
  }
}

▶️ Run a skill from the CLI

node dist/src/cli/index.js skill list
node dist/src/cli/index.js run <skill-name> --dry-run
node dist/src/cli/index.js run <skill-name> --param name=value

💡 No display handy? Add --computer mock to try everything without a desktop.

mcp · sandbox start|status|stop|reset · computer probe|info · skill list|show · run · recording list|show|delete · healing list|show

Option

Meaning

--computer local|sandbox|mock|placeholder

Backend (default sandbox)

--display :99

Target display

--version N

Skill version (default latest)

--dry-run

Validate and plan without firing actions

--param name=value

Runtime parameter

🧰 MCP tools

Group

Tools

🖱️

Computer control

poltr_computer_screenshot _click _double_click _move _type _key _hotkey _scroll _drag _wait _screen_size _active_window _windows _get_ui_tree · poltr_inspect_computer

🎥

Recording

poltr_start_recording poltr_stop_recording poltr_list_recordings poltr_show_recording poltr_get_recording_steps poltr_get_recording_screenshot

🧩

Skills

poltr_save_skill poltr_list_skills poltr_show_skill poltr_execute_skill (supports dryRun and parameters)

🩹

Healing

poltr_get_healing_context poltr_test_skill_repair

The agent learns progressively: list recordings → show one → fetch steps → fetch individual screenshots → submit a skill.

🧬 Skill format

A skill has typed parameters (string number boolean path command application, substituted as {{name}} at runtime), preconditions, steps and postconditions. Each step carries an intent, application, semantic target, action, expected result, verification and evidence references. See skills-library/.

📊 Current support

✅ Supported

🚧 Not yet

🎯 Targets

application window · ui-element / text (only with a real UI tree)

region workspace custom

⚡ Actions

launch focus type key hotkey scroll wait drag click double_click

run_command navigate inspect custom

🔍 Verification

application-state window-state process

terminal-output visual custom

♿ The X11 sandbox now exposes a real AT-SPI2 accessibility tree through a sandbox-private D-Bus session. ui-element and text targets are resolved from current application/window/role/name/text/state data. When a target is resolved, its current screen bounds are used for the computer action; this is not coordinate replay. If dbus-daemon or at-spi-bus-launcher is unavailable, Poltr returns an explicit unsupported-capability result and does not inspect the host accessibility bus.

AT-SPI support is toolkit-dependent. GTK/ATK applications such as Mousepad are the primary tested case; other Linux toolkits may expose only part of their accessibility tree or require their own accessibility configuration. LocalComputer remains unchanged and does not silently fall back to the sandbox tree.

poltr skill compile is a deterministic prototype. Real skill authoring happens through the agent and poltr_save_skill.

🗺️ Roadmap

  • Isolated Xvfb + Openbox desktop sandbox

  • Demonstration recorder with screenshot evidence

  • Skill schema, evidence provenance validation, immutable versions

  • Semantic executor with deterministic verification

  • Evidence-driven healing with sandboxed repair tests

  • AT-SPI accessibility tree for the sandbox

  • Visual and terminal-output verification

  • Runnable example skills

  • Reference agent that learns and runs skills end to end

  • Demo recording

🛠️ Development

npm run build
npm test
npm run lint
POLTR_RUN_SANDBOX_TEST=1 npm test   # opt-in real X11 sandbox test

📐 Architecture details: docs/architecture.md


MIT © Jangho Tan · Built for agents that should not have to relearn the desktop every time

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    F
    maintenance
    Enables creation of reusable browser automation skills through demonstration by recording user actions in a browser while narrating, then converting those workflows into executable skills that can be invoked through natural language.
    1
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables AI coding agents to automate Windows desktop applications through semantic UI Automation instead of brittle coordinate clicks, with tools for discovering windows, finding controls by stable identifiers, and verifying actions.
    39 PyPI
    2
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables an AI agent to operate Windows applications through vision-driven UI Automation and record polished demo videos with pre-click camera zoom, narration, and cinematic effects.
    MIT