Poltr
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Poltrreplay the invoice upload skill and tell me if any step failed"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
English | 简体中文
Why · How it works · Quickstart · MCP tools · Support · Roadmap
Poltr is a local MCP server that gives any agent (Claude, Gemini, or your own) an isolated desktop, a demonstration recorder, and a skill engine.
🧠 The agent reasons. 🤖 Poltr stays deterministic: recording, validation, execution, verification and safety.
💡 Why
Computer-use agents re-reason every step of every run. That is slow, expensive and flaky.
Without Poltr | With Poltr |
🐌 The agent re-plans the whole task each run | ⚡ Learn once, replay a stored semantic skill |
🎲 Clicks depend on screenshots and guesses | 🎯 Targets are resolved semantically, never by replaying coordinates |
🤷 "It probably worked" | ✅ Every step is deterministically verified |
💥 A UI change breaks everything | 🩹 Failures produce evidence, the agent proposes a repair, Poltr tests it and saves |
Related MCP server: windows-gui-mcp
🔄 How it works
Principle | What it means | |
🔌 | Model-agnostic | Recording, validation, storage and execution need no model API or key |
🧾 | Evidence-grounded | Every skill step must reference real recording steps and screenshots |
🙅 | Honest failures | Unsupported targets, actions and checks return structured errors, never silent success |
🛡️ | Safe by default | Bounds checks, blocked hotkeys, typing and scroll limits, app launch limited to approved sandbox apps (no generic shell) |
🌱 | Versioned healing | A repair creates |
🚀 Quickstart
Requires Node 18+. The sandbox needs Linux with Xvfb, Openbox, xdotool,
dbus-daemonandat-spi2-core(at-spi-bus-launcher). GTK/ATK applications such as Mousepad must expose accessibility through their toolkit bridge.
git clone https://github.com/janghotan/Poltr && cd Poltr
npm install
npm run build
node dist/src/cli/index.js sandbox start
node dist/src/cli/index.js computer probe --computer sandbox🔗 Connect an agent over stdio
{
"mcpServers": {
"poltr": {
"command": "node",
"args": ["dist/src/cli/index.js", "mcp"]
}
}
}▶️ Run a skill from the CLI
node dist/src/cli/index.js skill list
node dist/src/cli/index.js run <skill-name> --dry-run
node dist/src/cli/index.js run <skill-name> --param name=value💡 No display handy? Add
--computer mockto try everything without a desktop.
mcp · sandbox start|status|stop|reset · computer probe|info · skill list|show · run · recording list|show|delete · healing list|show
Option | Meaning |
| Backend (default |
| Target display |
| Skill version (default latest) |
| Validate and plan without firing actions |
| Runtime parameter |
🧰 MCP tools
Group | Tools | |
🖱️ | Computer control |
|
🎥 | Recording |
|
🧩 | Skills |
|
🩹 | Healing |
|
The agent learns progressively: list recordings → show one → fetch steps → fetch individual screenshots → submit a skill.
🧬 Skill format
A skill has typed parameters (string number boolean path command application, substituted as {{name}} at runtime), preconditions, steps and postconditions. Each step carries an intent, application, semantic target, action, expected result, verification and evidence references. See skills-library/.
📊 Current support
✅ Supported | 🚧 Not yet | |
🎯 Targets |
|
|
⚡ Actions |
|
|
🔍 Verification |
|
|
♿ The X11 sandbox now exposes a real AT-SPI2 accessibility tree through a sandbox-private D-Bus session.
ui-elementandtexttargets are resolved from current application/window/role/name/text/state data. When a target is resolved, its current screen bounds are used for the computer action; this is not coordinate replay. Ifdbus-daemonorat-spi-bus-launcheris unavailable, Poltr returns an explicit unsupported-capability result and does not inspect the host accessibility bus.AT-SPI support is toolkit-dependent. GTK/ATK applications such as Mousepad are the primary tested case; other Linux toolkits may expose only part of their accessibility tree or require their own accessibility configuration.
LocalComputerremains unchanged and does not silently fall back to the sandbox tree.
poltr skill compileis a deterministic prototype. Real skill authoring happens through the agent andpoltr_save_skill.
🗺️ Roadmap
Isolated Xvfb + Openbox desktop sandbox
Demonstration recorder with screenshot evidence
Skill schema, evidence provenance validation, immutable versions
Semantic executor with deterministic verification
Evidence-driven healing with sandboxed repair tests
AT-SPI accessibility tree for the sandbox
Visual and terminal-output verification
Runnable example skills
Reference agent that learns and runs skills end to end
Demo recording
🛠️ Development
npm run build
npm test
npm run lint
POLTR_RUN_SANDBOX_TEST=1 npm test # opt-in real X11 sandbox test📐 Architecture details: docs/architecture.md
MIT © Jangho Tan · Built for agents that should not have to relearn the desktop every time
This server cannot be deployed
Maintenance
Related MCP Connectors
- openhelmOAuthai.openhelm
Autonomous cloud agent tasks: real browser + your tools, structured evidence-backed results.
- SkilderOAuthai.skilder
One place to build, share, and govern the skills and tools your AI agents use at work.
Operate Linux, macOS and Windows from your LLM. Every action runs through an auditable allowlist.
Loom for agents: AI agents record narrated product demos, plus transcripts, summaries, and search.
Related MCP Servers
- FlicenseNot gradedqualityFmaintenanceEnables creation of reusable browser automation skills through demonstration by recording user actions in a browser while narrating, then converting those workflows into executable skills that can be invoked through natural language.1-
- AlicenseNot gradedqualityCmaintenanceEnables AI coding agents to automate Windows desktop applications through semantic UI Automation instead of brittle coordinate clicks, with tools for discovering windows, finding controls by stable identifiers, and verifying actions.39 PyPI2MIT
- AlicenseNot gradedqualityBmaintenanceEnables users to record Windows desktop and browser workflows once and generate replayable Cursor Skills, using Playwright MCP and Windows Computer MCP for semantic, non-coordinate replay.1MIT
- AlicenseNot gradedqualityAmaintenanceEnables an AI agent to operate Windows applications through vision-driven UI Automation and record polished demo videos with pre-click camera zoom, narration, and cinematic effects.MIT