Skip to main content
Glama
README.md
<div align="center">

<img src="docs/assets/banner.svg" alt="Poltr" width="100%"/>

**English** | [简体中文](README.zh-CN.md)

[![License: MIT](https://img.shields.io/badge/license-MIT-22d3ee?style=for-the-badge)](LICENSE)
[![TypeScript](https://img.shields.io/badge/TypeScript-5.5-3178c6?style=for-the-badge&logo=typescript&logoColor=white)](tsconfig.json)
[![MCP](https://img.shields.io/badge/MCP-server-a78bfa?style=for-the-badge)](https://modelcontextprotocol.io)
[![Node](https://img.shields.io/badge/node-18%2B-339933?style=for-the-badge&logo=nodedotjs&logoColor=white)](package.json)
[![Sandbox](https://img.shields.io/badge/sandbox-Xvfb%20%2B%20Openbox-f472b6?style=for-the-badge&logo=linux&logoColor=white)](#-quickstart)
[![Status](https://img.shields.io/badge/status-early%20alpha-fbbf24?style=for-the-badge)](#-current-support)

[**Why**](#-why) · [**How it works**](#-how-it-works) · [**Quickstart**](#-quickstart) · [**MCP tools**](#-mcp-tools) · [**Support**](#-current-support) · [**Roadmap**](#-roadmap)

</div>

---

**Poltr** is a local [MCP](https://modelcontextprotocol.io) server that gives any agent (Claude, Gemini, or your own) an isolated desktop, a demonstration recorder, and a skill engine.

> 🧠 **The agent reasons.** 🤖 **Poltr stays deterministic**: recording, validation, execution, verification and safety.

<!-- TODO: replace with a real demo GIF: ![demo](docs/assets/demo.gif) -->

## 💡 Why

Computer-use agents re-reason every step of every run. That is slow, expensive and flaky.

| Without Poltr | With Poltr |
|---|---|
| 🐌 The agent re-plans the whole task each run | ⚡ Learn once, replay a stored **semantic skill** |
| 🎲 Clicks depend on screenshots and guesses | 🎯 Targets are resolved semantically, **never by replaying coordinates** |
| 🤷 "It probably worked" | ✅ Every step is **deterministically verified** |
| 💥 A UI change breaks everything | 🩹 Failures produce evidence, the agent proposes a repair, Poltr tests it and saves **`v2`** |

## 🔄 How it works

<div align="center">
<img src="docs/assets/pipeline.svg" alt="Poltr pipeline" width="100%"/>
</div>

<br/>

| | Principle | What it means |
|---|---|---|
| 🔌 | **Model-agnostic** | Recording, validation, storage and execution need no model API or key |
| 🧾 | **Evidence-grounded** | Every skill step must reference real recording steps and screenshots |
| 🙅 | **Honest failures** | Unsupported targets, actions and checks return structured errors, never silent success |
| 🛡️ | **Safe by default** | Bounds checks, blocked hotkeys, typing and scroll limits, app launch limited to approved sandbox apps (no generic shell) |
| 🌱 | **Versioned healing** | A repair creates `skill.v2` from `v1`. Originals are never modified and runtime parameter values are never stored |

## 🚀 Quickstart

> Requires **Node 18+**. The sandbox needs **Linux** with Xvfb, Openbox, xdotool, `dbus-daemon` and `at-spi2-core` (`at-spi-bus-launcher`). GTK/ATK applications such as Mousepad must expose accessibility through their toolkit bridge.

```bash
git clone https://github.com/janghotan/Poltr && cd Poltr
npm install
npm run build

node dist/src/cli/index.js sandbox start
node dist/src/cli/index.js computer probe --computer sandbox
```

**🔗 Connect an agent over stdio**

```json
{
  "mcpServers": {
    "poltr": {
      "command": "node",
      "args": ["dist/src/cli/index.js", "mcp"]
    }
  }
}
```

**▶️ Run a skill from the CLI**

```bash
node dist/src/cli/index.js skill list
node dist/src/cli/index.js run <skill-name> --dry-run
node dist/src/cli/index.js run <skill-name> --param name=value
```

> 💡 No display handy? Add `--computer mock` to try everything without a desktop.

<details>
<summary><b>⌨️ CLI reference</b></summary>

`mcp` · `sandbox start|status|stop|reset` · `computer probe|info` · `skill list|show` · `run` · `recording list|show|delete` · `healing list|show`

| Option | Meaning |
|---|---|
| `--computer local\|sandbox\|mock\|placeholder` | Backend (default `sandbox`) |
| `--display :99` | Target display |
| `--version N` | Skill version (default latest) |
| `--dry-run` | Validate and plan without firing actions |
| `--param name=value` | Runtime parameter |

</details>

## 🧰 MCP tools

| | Group | Tools |
|---|---|---|
| 🖱️ | **Computer control** | `poltr_computer_screenshot` `_click` `_double_click` `_move` `_type` `_key` `_hotkey` `_scroll` `_drag` `_wait` `_screen_size` `_active_window` `_windows` `_get_ui_tree` · `poltr_inspect_computer` |
| 🎥 | **Recording** | `poltr_start_recording` `poltr_stop_recording` `poltr_list_recordings` `poltr_show_recording` `poltr_get_recording_steps` `poltr_get_recording_screenshot` |
| 🧩 | **Skills** | `poltr_save_skill` `poltr_list_skills` `poltr_show_skill` `poltr_execute_skill` (supports `dryRun` and `parameters`) |
| 🩹 | **Healing** | `poltr_get_healing_context` `poltr_test_skill_repair` |

The agent learns progressively: list recordings → show one → fetch steps → fetch individual screenshots → submit a skill.

## 🧬 Skill format

A skill has typed parameters (`string` `number` `boolean` `path` `command` `application`, substituted as `{{name}}` at runtime), preconditions, steps and postconditions. Each step carries an intent, application, semantic target, action, expected result, verification and evidence references. See [`skills-library/`](skills-library).

## 📊 Current support

| | ✅ Supported | 🚧 Not yet |
|---|---|---|
| 🎯 **Targets** | `application` `window` · `ui-element` / `text` (only with a real UI tree) | `region` `workspace` `custom` |
| ⚡ **Actions** | `launch` `focus` `type` `key` `hotkey` `scroll` `wait` `drag` `click` `double_click` | `run_command` `navigate` `inspect` `custom` |
| 🔍 **Verification** | `application-state` `window-state` `process` | `terminal-output` `visual` `custom` |

> ♿ The X11 sandbox now exposes a real AT-SPI2 accessibility tree through a sandbox-private D-Bus session. `ui-element` and `text` targets are resolved from current application/window/role/name/text/state data. When a target is resolved, its current screen bounds are used for the computer action; this is **not coordinate replay**. If `dbus-daemon` or `at-spi-bus-launcher` is unavailable, Poltr returns an explicit unsupported-capability result and does not inspect the host accessibility bus.
>
> AT-SPI support is toolkit-dependent. GTK/ATK applications such as Mousepad are the primary tested case; other Linux toolkits may expose only part of their accessibility tree or require their own accessibility configuration. `LocalComputer` remains unchanged and does not silently fall back to the sandbox tree.
>
> `poltr skill compile` is a deterministic prototype. Real skill authoring happens through the agent and `poltr_save_skill`.

## 🗺️ Roadmap

- [x] Isolated Xvfb + Openbox desktop sandbox
- [x] Demonstration recorder with screenshot evidence
- [x] Skill schema, evidence provenance validation, immutable versions
- [x] Semantic executor with deterministic verification
- [x] Evidence-driven healing with sandboxed repair tests
- [x] AT-SPI accessibility tree for the sandbox
- [ ] Visual and terminal-output verification
- [ ] Runnable example skills
- [ ] Reference agent that learns and runs skills end to end
- [ ] Demo recording

## 🛠️ Development

```bash
npm run build
npm test
npm run lint
POLTR_RUN_SANDBOX_TEST=1 npm test   # opt-in real X11 sandbox test
```

📐 Architecture details: [docs/architecture.md](docs/architecture.md)

---

<div align="center">

MIT © Jangho Tan · Built for agents that should not have to relearn the desktop every time

</div>