mcp-dotpals
Integrates with OpenAI's Codex coding agent to monitor its coding sessions, check test results, and provide MCP tools like check_my_work and ready_to_merge so it can query what really happened.
dotpals
See what your coding agent actually did.
A small floating pal that watches Claude Code, Codex or any agent and checks its work: which files changed, which commands ran, and whether the tests really passed, read from their output, not the agent's word. When Claude breaks the tests, dotpals sends it back to fix them. And any AI tool can ask dotpals what really happened.
npx dotpals@latest setupOne command on Windows, macOS or Linux. Free, open source, and everything stays on your computer. Or straight from GitHub: npx --allow-git=all github:rikinshah787/dotpals setup
🔊 Watch the launch video with sound (47 s, 1080p) · square cut
Install · Make your own pal · Plug in any agent · What's new
⭐ Star dotpals if your agent ever said "Done!" and you weren't sure · Tell us what's confusing
New in 0.10
dotpals used to show you what your agent did. Now it checks the work, and every AI tool can ask it.
Make agents fix failing tests. When Claude Code's tests fail, it's told at once, with why ("expected 3, got -1 (test/math.test.js:5)"), and sent back if it tries to finish, commit or push anyway. On by default; one switch in Settings, or the gear in the notch, turns it off.
Catches the faked pass. Green only because a test lost an assertion, got a skip or now expects the wrong thing? Claude is sent back to put the test back and fix the code, and you're told.
Ask dotpals from any agent (MCP). Claude Code, Codex, Cursor and other assistants ask
check_my_work,ready_to_merge,todayand more, and get the evidence instead of the agent's word.
All of it in the changelog.
Related MCP server: GitWarren
Why
Coding agents do a lot in a single request. They read dozens of files, edit a handful, run tests, retry and search. The chat scrolls by and the diff is spread across files. dotpals keeps a live, plain-language record next to your editor, so at any moment you can answer:
What did it change? Every file edited, created or deleted, with the diff one click away.
What did it run, and did it work? Every command, with its output, duration and ✓ or ✕. A step that failed and was retried says fixed on try 2 or still failing after 3 tries.
Did it fix what it broke? When Claude Code's tests fail, it's sent back to fix them before it finishes or commits, and a test changed just to make it pass is caught.
Was the code as it is now tested? "Changed 2 files after the tests passed: not tested since" is impossible to miss, so an old green result doesn't pass for a check of the latest edits.
What is it doing right now? The pal thinks, works, asks for your OK and celebrates, live.
What did I get done today? A running tally, and one click copies it as Markdown for a standup or PR.
Features
Make agents fix failing tests (Claude Code, on by default): when a test run fails, Claude is told right away. When it tries to finish, commit or push while the tests fail or weren't run after its last change, it's sent back to fix them, at most twice per request; then it may stop and the pal tells you. A request that changed no code ("run the tests and tell me") just reports the failure, and a test that was already failing before Claude changed anything is reported, not forced on it. A result nobody can read (output cut by
| tail): Claude runs the tests again; if it still can't be read, the pal asks you to look. And if the tests only went green because Claude changed them (took out an assertion, added a skip, changed what one expects), dotpals catches it: Claude is sent back and you're told. Projects without tests are never held up. Turn it off in Settings, or from the gear in the notch.Ask dotpals from any agent:
dotpals mcplets Claude Code, Codex, Cursor and other assistants ask what your agents really did, with the evidence:check_my_workbefore an agent says "done",ready_to_merge,today,recap,risky_stepsand more. See below.The story, not the log: each request reads as a few chapters, such as Changed 5 files +42 −7 · Tests failed twice, then passed · Committed and pushed, instead of hundreds of tool calls. Anything worth a second look is flagged:
.envchanged, a force-push, the same command failing 3 times, two agents editing the same file, or code changed without testing it.Two agents, one file: when an agent is about to change a file another agent changed in the last few minutes, Claude Code asks you first ("Codex (api) changed billing.ts 2 minutes ago. Edit anyway?"), or tells Claude to re-read it. Other agents can't be stopped beforehand: the notch and the pal tell you as it happens.
Hand-off: Continue in ▾ hands a session to Codex, Claude Code or Gemini CLI, in a new terminal in the same project, with a note on what was asked, what was done, how the tests stand and what's left. Or copy the note and paste it anywhere.
Chat with your agents (optional, off by default): type to an agent from the pal, and answer its permission requests and questions right there. OpenCode for now; Claude Code and others are coming. Turn it on in Dashboard → Settings.
Retries: when a step fails and the agent tries the same thing again, the tries are linked: fixed on try 2, still failing after 3 tries. Click a try to jump to it. On the dashboard, paste a step's ID (
toolu_…) to open it.Was it tested? One line per session says Tests passed · 48 passed · 7:08 PM, after the last change, Tests passed at 7:08 PM · 3 files changed since or No tests run by the agent, and whether the last commit was tested. Each result says where it came from: the test output's own summary (jest, vitest, mocha, node:test, pytest, go, cargo, dotnet, Maven, Gradle, PHPUnit, RSpec and more), or (exit code only). Zero tests or only skipped ones read as Tests unclear: no tests actually ran, never as passed. Optionally, an unclear result can be double-checked by Laya on your computer (one click sets it up: Settings → Set up Laya, or
dotpals laya; needs Python 3.10+) or TypeSafe's Jev in the cloud (off by default). Only tests the agent ran count.Simple or Detailed: Simple (the default) sums up each request in one plain sentence, such as Changed billing.ts, the tests passed after one retry, and committed and pushed., plus only the warnings that matter. Detailed shows every chapter, with small steps folded away. Switch in the pal, the notch or the dashboard.
Setup asks, in the terminal: your pal, the notch, Simple or Detailed, approvals, test double-checks (Off, Local Laya or Cloud Jev, with your key typed hidden) and more. Press Enter for the defaults, or run
setup --yes.The notch: an island at the top of your screen with every agent, a live diff of the file it's editing, its plan ("2/4 · Detecting the system setting"), its context window and your Claude and Codex usage limits. It opens by itself when an agent needs you, and you can allow or deny from the keyboard. Hide the pal and the notch takes over; – minimizes it, so nothing sits at the top while agents work.
Context and limits: the pal gets worried as a session's context window fills up and cheers after it compacts. Usage bars show your 5-hour and weekly limits with reset times (for Claude, run
dotpals statuslineonce).Summary: one card per request, with what you asked, what the agent said it did, and a tally such as Changed 3 files · Ran 5 commands, 1 failed · Used 1 skill. Show steps lists every step as a short sentence.
Tools: every tool call as it happens. Click one to see the exact command and output, or the lines an edit changed.
Files: every file read, changed, created or deleted, with diffs. Click to open it in VS Code.
Today: requests, files changed, commands run and time the agent spent working. Copy today gives you a ready-made standup note.
Copy recap: copy any request as Markdown for a PR description or commit message. It keeps ✅ ran successfully, ❌ failed, ❔ unclear and ⚪ not run apart, and every claim carries its evidence: the command, the result it was read from ("48 passed", or the exit code) and the step's ID, which the dashboard's search opens. Quick look-ups like
greparen't counted as failures.One tab per session: Claude Code and Codex sessions never mix, and the window follows whichever is active.
Dashboard: every session with its requests, files and full log, with search and export to Markdown or JSON. It also shows requests per day, time by project and a live view of which agents are connected.
Settings: choose your pal, turn sounds and notifications on or off, and decide how long to keep history (or clear it). Settings are shared by the pal and the dashboard.
History: survives restarts, kept on your computer in
~/.dotpals/history.json(7 days by default).Notifications and sounds: a ping when the agent needs your OK, a chime when it's done, and a desktop notification if you've looked away.
Every session, every agent: all your Claude Code sessions show up, even ones started before dotpals was installed, next to Codex and anything else you plug in. In small mode each agent gets its own pal, with a round bar above them naming each one.
Make your own pal: pick a body, eyes, something on top, a color and a name. See below.
A pal with personality: eight ready-made characters that think, work, talk, wait, celebrate and sulk. Drag it anywhere; it stays on top, and clicks on the empty space around it go through to your editor.
The notch
A small island that hangs from the top of your screen. It has four sizes:
Hidden when nothing is running, or you've been away for 3 minutes: just a thin, invisible strip at the top edge. Hover it and a small island peeks out; rest there a moment and it opens.
Bar while agents work: a mini pal for each agent, the current step, the plan step ("2/4") and a ring for your highest usage limit. Hover it for about 200 ms, or click, to open. Don't want it there? – in the open notch minimizes it (remembered): it stays hidden while agents work, still opens when one needs you, and the top edge still peeks.
Open (640 px): the agent in focus as a big pal on the left, one card on the right, a column of mini pals for the other agents, and two tabs:
Now: a live diff of the file it's editing (or a checklist of its steps), its plan, context window, helpers and your usage limits. When it needs your OK, an approval card with Deny and Allow, or Ctrl+Alt+N and Ctrl+Alt+Y (⌘⌥N and ⌘⌥Y on macOS), which work only while the card is showing. When it's done or fails, a short card says what happened.
Story: today's totals with Copy today, whether the code was tested since its last change, the plan, helpers, the context window with Copy /compact, what it's been using, a note when two agents changed the same file, and the last few requests as chapters you can expand.
Alerts open it by themselves, one at a time. One that needs you shows even if you've been away, and stays until you answer. Done and error cards close after about 5 and 8 seconds. When you open it yourself, it closes 8 seconds after the pointer leaves (a shrinking line shows the last seconds), or after a quiet minute with the pointer resting on it. Esc closes it while the pointer is over it. Its window lets clicks through everywhere except the island, and the peek never takes a click, so it doesn't get in the way of your browser tabs.
By default it appears when you hide the pal. You can keep it on always or never show it, from the tray or with dotpals notch --auto | --off.
Claude Code shares its usage limits only with a status line command, so run dotpals statusline once to see them. If you already have a status line, it keeps showing yours; dotpals statusline --off puts everything back. Codex's limits come straight from its logs.
Make your own pal
Open the dashboard (▦ on the pal, or dotpals dashboard), go to Settings → Make your own pal, and mix:
Body: round, boxy, fluffy, pointy, heart or frog
Eyes: dots, button, googly, pixel, visor or shades
On top: cat ears, horns, antenna, sprout, sparkle, bow, crown or beret
Color: any color, fluffy or smooth
Name: yours to pick
Try it thinking, working and celebrating right there, then Use this pal. The floating pal switches straight away. Surprise me rolls a random one. There are thousands of combinations.
In your own app it's one call: registerCustom({ name: 'Pip', shape: 'bean', eyes: 'googly', top: 'crown', color: '#16c6ae' }), then <dot-pal character="custom">.
Install
One command
npx dotpals@latest setup(Needs Node 20 or newer. To install straight from GitHub instead: npx --allow-git=all github:rikinshah787/dotpals setup.)
That's all. It:
installs the desktop pal in
~/.dotpals, and downloads its runtime (Electron, about 100 MB, once),adds the Claude Code plugin, if Claude Code is installed,
picks up Codex automatically, if it's installed,
starts the pal, turns on open when I log in, and opens the dashboard.
Options: --no-claude (skip the plugin), --no-login (don't start at login) and --no-start. Run it again any time to update.
Only the Claude Code plugin
/plugin marketplace add rikinshah787/dotpals
/plugin install dotpals@dotpalsRestart Claude Code, then run /dotpals:pals. The first time, it offers to download the desktop window's runtime. After that, the pal opens by itself whenever a Claude Code session starts.
Codex
There's nothing to install on the Codex side. dotpals follows Codex's session logs (~/.codex/sessions), so the Codex CLI, IDE extension and app all show up while the pal is running. The one-command setup starts it at login.
Cursor, Gemini CLI, OpenCode and GitHub Copilot CLI
Open the dashboard's Agents page and press Connect on the agent you use. Each card shows whether the agent is installed, whether it's connected, and when its last event arrived.
Connect adds one small command, or a plugin for OpenCode, to that agent's own config. It backs up the original first, merges instead of overwriting, and leaves a file it can't read untouched.
Disconnect takes out only what dotpals added.
Send a test event runs the real command. If it reaches dotpals, a pal says hello.
The switch on each card turns an agent off without disconnecting it.
The Connect button changes these files:
Agent | What Connect changes |
Cursor |
|
Gemini CLI |
|
OpenCode | adds |
GitHub Copilot CLI | adds |
Any other agent
Send JSON to the local bridge from your agent loop, a hook script or a wrapper. See Plug in any agent.
Using it
Ctrl+Alt+P (⌘⌥P on macOS) | Show or hide the pal from anywhere |
Drag the pal | Move the window; it remembers where you put it |
▦ | Open the dashboard: sessions, logs, stats and settings |
⤡ | Switch between just the pal and the full view |
× | Hide to the tray. The tray menu has Dashboard, Just the pal, Notifications, Open when I log in and Quit |
🔊 | Sounds on or off |
From a terminal, after setup (or with npx dotpals <command>):
dotpals start # open the floating pal
dotpals dashboard # open the dashboard
dotpals status # what's running and connected
dotpals doctor # no pal or notch? checks and fixes it (setup runs the same check at the end)
dotpals bridge # only the bridge, e.g. on a machine without a desktop; dashboard at http://127.0.0.1:5175/dashboard
dotpals mcp # an MCP server, so any assistant can ask dotpals what your agents did (see below)Ask dotpals from any agent (MCP)
dotpals mcp is an MCP server: Claude Code, Codex, Cursor and other assistants can ask dotpals what your agents really did, and get the evidence instead of the agent's word. It only reads, from the pal running on this computer, so keep dotpals running.
claude mcp add dotpals -- npx dotpals mcpCodex, in ~/.codex/config.toml:
[mcp_servers.dotpals]
command = "npx"
args = ["dotpals", "mcp"]Cursor and others: command npx, arguments dotpals mcp.
Tool | Answers |
| Which agents worked in the last 2 hours, on what, whether they're still at it, and how their tests stand |
| Is it really tested? Passed or failed (read from the output), why the last run failed ("expected 3, got -1 (test/math.test.js:5)"), and whether code changed since |
| The checks a reviewer would make: tests ran after the last change and passed (not because the tests were changed), nothing left failing, nothing risky, and the work is committed. Mid-request too, as it stands now. Pass |
| What each agent was asked and did, request by request, with the files it changed, why its tests failed, and tests that passed only after they were changed (a project or a session, |
| Today's work per agent, for a standup: requests, files changed and test runs, then each request in a sentence |
| Force pushes, recursive deletes, changes to |
| A note for another agent to pick up a session's work: the ask, what was done, the files, how the tests stand, what's left |
| For the agent itself, before it says "done": "Looks done", or what to fix first (with why the tests fail, or "put the test back and fix the code"). Tests that were failing before the session changed anything are its to report, not to fix |
Ask things like "What are my agents doing?", "Is this really tested?", "Is feat/login ready to merge?" or "What did Codex change in billing this morning?". The project is the folder the assistant runs in, unless you name another. dotpals matches it by the folder each agent really works in, so two folders with the same name, a monorepo's packages and git worktrees stay apart (a project's name alone matches by name). It notes the branch when an agent finishes a request in a git repository. Only tests an agent ran count: dotpals can't see the ones you run yourself, or CI.
Privacy
Everything stays on your machine. The bridge listens only on 127.0.0.1. It reads Claude Code hook events and transcripts and Codex's session logs locally, and it sends nothing anywhere. The one exception is opt-in: if you choose Cloud (Jev) under Settings → Double-check unclear test results, the end of an unclear test run's output is sent to TypeSafe, after removing anything that looks like a password, key, email or IP address. Make agents fix failing tests (on by default) adds short notes about Claude's own test runs to its context, which Claude Code sends to its model like anything else there; turn it off in Settings. History is a plain JSON file in ~/.dotpals. Set DOTPALS_HISTORY=0 to turn it off, or DOTPALS_CODEX=0 to stop following Codex. See SECURITY.md.
How accurate is it?
We check. test/accuracy/cases holds 36 real requests (cleaned of paths, keys and prompts): 21 from building dotpals itself, and 15 from a small practice project with a planted bug, where Claude was asked to run the tests, add code, and fake a passing test. Every claim dotpals makes about them was checked by hand against what really happened: 255 claims.
Claim | Claims | Right | Wrong | Unsure |
Test results (passed, failed) | 99 | 91 | 1 | 7 |
Risky steps | 14 | 14 | 0 | 0 |
Retries ("fixed on try 2") | 15 | 14 | 1 | 0 |
Ready to merge? | 36 | 35 | 1 | 0 |
Why it stopped | 36 | 35 | 1 | 0 |
What changed (git) | 10 | 9 | 1 | 0 |
Tests changed to pass | 20 | 20 | 0 | 0 |
Why tests failed | 25 | 25 | 0 | 0 |
All | 255 | 243 (95.3%) | 5 | 7 |
Tests changed to pass is the faked pass: green only because an edit to a test took out an assertion, turned a test off or changed what it expects (3 fakes caught; rightly quiet on 17 others: honest test edits while building dotpals, a skip taken out again before a real fix, fixes that didn't touch the tests). Why tests failed is the reason dotpals reads from a failed run's output and tells the agent ("expected 3, got -1").
"Unsure" is a missing answer, not a false one: a test run dotpals called unclear though a person could tell (that's what the optional Jev or Laya check is for), or a failed run it gave no reason for though the output shows one. The 5 wrong claims are listed by npm run accuracy. One test result and one "Ready to merge?" are wrong because the bridge stored a long command cut off (fixed for new requests). A retry and a "Why it stopped" are wrong because a render that failed twice worked later with different arguments, and dotpals didn't link the two. Git can't tell who changed a file, so a file you reset during a request counts as changed by it. The test suite fails if a claim that's right turns wrong. Add your own cases with node scripts/accuracy-capture.mjs.
How it works
Claude Code ── hooks + transcripts ┐
Codex ──────── session logs ───────┼──▶ bridge (127.0.0.1:5175) ──▶ floating pal (Summary · Tools · Files)
your agent ─── POST /event ────────┘ one activity model or any browser tabYour agents report what they do. Claude Code sends hook events as it works, and dotpals also reads each session's transcript in
~/.claude/projects, so every session shows up, including ones that started before dotpals was installed. Codex writes session logs to~/.codex/sessions, which dotpals follows. Nothing to set up on the Codex side. Any other agent can POST JSON.A small local server (the bridge) turns that into one activity feed. It runs on
127.0.0.1:5175, only answers your own computer, and keeps history in~/.dotpals. Nothing is sent anywhere.The pal shows it. The desktop app (Electron, always on top) and the dashboard read the feed live: the pal's mood, the Summary, Tools and Files tabs, stats and history.
Each agent connects through an adapter in bridge/adapters/, and every adapter produces the same activity entries (bridge/activity.js). The pal itself is a dependency-free Web Component that you can also drop into your own app (see below).
Platforms: Windows, macOS and Linux (Node 20+). On macOS the pal lives in the menu bar instead of the Dock. On Linux the small window can't pass clicks through its empty space, because Linux doesn't support it.
Plug in any agent
The bridge is harness-agnostic. Each agent tool connects through an adapter in bridge/adapters/, and every adapter feeds the same activity model (bridge/activity.js).
Harness | How it connects | Setup |
Claude Code | Hooks for live state (including permission prompts), plus the session transcript, so the history is complete even if the pal opened late | |
Codex (CLI, IDE extension, app) | Follows Codex's session logs in | None. Keep the pal running (tray: Open when I log in). |
Cursor (editor and CLI) | Hooks in | Connect on the dashboard's Agents page |
Gemini CLI | Hooks in | Connect on the Agents page |
OpenCode | A plugin in | Connect on the Agents page, then restart OpenCode |
GitHub Copilot CLI | Hooks in | Connect on the Agents page |
Anything else | POST JSON to | A few lines in your agent loop, a hook script or a wrapper. The Agents page has copy-paste snippets for curl, PowerShell, Node, Python and the shell |
Every integration can be switched off on the Agents page, or in ~/.dotpals/config.json with { "agents": { "cursor": false } }. The ids are claude, codex, cursor, gemini, opencode, copilot and generic.
The hook-based integrations run node ~/.dotpals/app/bridge/hook.js <agent>, so Node has to be on your PATH. The command posts to POST /hook?agent=<id>, and the matching module in bridge/adapters/ turns the events into activity. To add another agent, write a module there with the same shape (id, name, detect(), connect(), disconnect(), apply(); see bridge/adapters/index.js) and list it in index.js.
The event format
Send the pal's state, activity rows, or both. Rows with the same id are merged, so you can send a tool call when it starts and again when it finishes:
# a tool call starts…
curl -s localhost:5175/event -d '{
"session": "run-42", "harness": "my-agent", "label": "my-project",
"state": "working", "text": "Running tests",
"activity": { "id": "call-1", "kind": "run", "tool": "shell", "title": "Run the tests",
"status": "running", "body": { "command": "npm test" } }
}'
# …and finishes
curl -s localhost:5175/event -d '{
"session": "run-42", "harness": "my-agent", "state": "thinking",
"activity": { "id": "call-1", "status": "ok", "ms": 5120, "body": { "output": "42 passing" } }
}'
# a file edit, with its diff
curl -s localhost:5175/event -d '{
"session": "run-42", "harness": "my-agent",
"activity": { "id": "call-2", "kind": "edit", "tool": "write_file", "title": "src/app.js", "status": "ok",
"files": [{ "path": "/abs/path/src/app.js", "change": "edit" }],
"body": { "patch": "-const a = 1;\n+const a = 2;" } }
}'Field | Values |
| any id; each session gets its own pal and tab |
| shown on the tab, e.g. "My-agent · my-project" |
| the pal's state (see Agent states) and bubble text |
|
|
|
|
| `[{ path, change: "read" |
|
|
Any event the pal already understands (Anthropic, OpenAI or Agent SDK stream events, or { "state", "text" }) works here too. To add a first-class adapter, see bridge/adapters/codex.js. It's a good template for any harness that writes a session log.
Use the pal in your own app
The pal is a dependency-free Web Component, <dot-pal>, for chat UIs, IDE panels and dashboards. Send it your agent's state and it shows thinking dots while the model reasons, a progress bubble while tools run, a talking mouth while text streams, a question bubble when it needs approval, a jump when it's done and a frown when something fails. It works in plain HTML, React, Vue, Svelte, Angular, Electron and VS Code webviews.
Characters
Id | Pal | Click action |
| Blu, a blue cloud in a beret | jump |
| Hop, a green frog | jump |
| Sunny, a yellow gumdrop in glasses | wiggle |
| Lovi, a pink heart in sunglasses | love |
| Muse, a violet flame with sparkles | spin |
| Grok, a slate bot with a glowing visor | nod |
| Nova, an orange bot with a light-bulb antenna | jump |
| Byte, a teal cat with pixel eyes | wiggle |
Quick start
<script type="module" src="https://unpkg.com/dotpals"></script>
<dot-pal id="agent" character="grok"></dot-pal>
<script type="module">
const pal = document.getElementById('agent');
pal.setState('thinking');
pal.setState('working', { text: 'Running tests…' });
pal.setState('done', { text: 'All green!' });
</script>Or from npm:
npm install dotpalsimport 'dotpals';Agent states
State | What the pal does |
| breathes, blinks and follows the cursor |
| leans in with wide eyes, for while the user is typing |
| looks up, shows a bubble with bouncing dots |
| busy bob, eyes down, shows a progress bar or your |
| mouth moves, for while tokens stream in |
| hops, then keeps bouncing with wide eyes, shows a |
| jumps with a burst of sparkles and happy eyes, then settles back to calm |
| jitters, then looks sad and desaturated, with × eyes |
| eyes closed, floating zs |
A soft glow behind the pal follows the state (amber while waiting, red on errors, green when done), and moods and moves blend into each other instead of snapping.
You can set a state three ways:
<dot-pal character="muse" state="thinking"></dot-pal>pal.state = 'speaking';
pal.setState('working', { text: 'web_search' });Plug into your harness
1. Stream events straight in
connectAgent accepts an EventSource, a WebSocket, any EventTarget, or an async iterable (such as an SDK stream). It maps each event to a state automatically.
import { connectAgent } from 'dotpals';
// Server-Sent Events from your backend
connectAgent(pal, new EventSource('/agent/events'));
// WebSocket
connectAgent(pal, new WebSocket('wss://my-harness/agent'));
// An SDK stream (async iterable), e.g. the Anthropic TypeScript SDK
const stream = client.messages.stream({ model, max_tokens, messages, tools });
connectAgent(pal, stream);It returns a function that disconnects.
2. Call it from your own event loop
import { agentHandler } from 'dotpals';
const onEvent = agentHandler(pal);
for await (const event of myAgent.run(prompt)) {
onEvent(event); // unknown events are ignored
render(event);
}Events it understands
Source | Events | State |
Anthropic Messages API (streaming) |
| thinking |
| thinking | |
| working, with the tool name | |
| speaking | |
| done | |
Claude Agent SDK |
| thinking |
| working, with the tool name | |
| speaking | |
| done, or error if it failed | |
OpenAI Responses API (streaming) |
| thinking |
| working | |
| speaking | |
| done | |
| error | |
Generic |
| the matching state |
Your own |
| exactly what you send |
Plain strings work too: 'thinking', or a JSON string of any of the above.
Custom mapping
connectAgent(pal, source, {
map: (e) => {
if (e.kind === 'plan') return { state: 'thinking', text: 'Planning…' };
if (e.kind === 'shell') return { state: 'working', text: `$ ${e.cmd}` };
return toAgentState(e); // fall back to the built-in mapping
},
});Runnable example
npm run example:agent # opens a Server-Sent Events harness on http://localhost:5174See examples/sse-harness. The server side is about 20 lines. Replace the fake runAgent with your real loop.
More ways to use a pal
// Loading feedback for any promise: thinking, then happy or sad
const data = await pal.during(fetch('/api/save'), { successText: 'Saved!' });
// A form companion: follows the caret, covers its eyes on passwords,
// frowns at invalid fields and cheers on submit
const stop = pal.watch('#login-form');
// Speech bubble
pal.say('Hi! Ask me anything.');
// Show a mood for a moment
pal.flash('surprised', 1500);
// One-shot actions: jump · squish · wiggle · shake · nod · spin · love · hop · jitter · hello · dizzy
await pal.play('love');
// Say hello: rise up from below, squint happily, hop and blink twice
await pal.greet();
// A face for a moment: happy · love · star · wide · closed · dizzy · oops · hey · sweat
await pal.emote('love', 1600);
// Throw particles: heart · sparkle · star · sweat · z, or any text or emoji
pal.burst('sparkle', 8);Faces and reactions
Expression eyes: pals swap in happy arcs, closed lids, wide eyes, × ("oops"), spinning spirals, hearts and sparkle-stars to match their mood (happy, sleepy, surprised or waiting, and the error state) or an
emote().Reactions: hover and it blinks; rest the mouse on it for 2 seconds and it gets heart eyes; click and it plays its tap action with a "hey" face; click 3 times quickly and it gets dizzy. Each click fires
dotpal-pokewith{ count }.staticturns these off.Tiny pals: under 48 px a pal becomes an avatar (the
tinyattribute and the read-onlypal.tinyproperty): no fur, bigger eyes, no glow and no particles.Pointing from outside the page:
DotPal.pointAt(x, y)tells every pal where the cursor is (viewport CSS px), for apps that track it themselves.DotPal.emoteslists every emote.
Attributes
Attribute | Values | Default |
| any id from the table above, or a registered name |
|
|
|
|
|
|
|
| number (px) or any CSS length |
|
| any CSS color | the character's color |
|
|
|
|
|
|
|
| leans a little |
| boolean: turns off the hover and click reactions | – |
| accessible name | the character's name |
| set by the pal itself while it's smaller than 48 px | – |
A state is the agent lifecycle; each state sets a mood. Use mood directly if you aren't driving an agent.
Events
pal.addEventListener('dotpal-state', (e) => e.detail); // { state, text }
pal.addEventListener('dotpal-mood', (e) => e.detail); // { mood }
pal.addEventListener('dotpal-action', (e) => e.detail); // { action }
pal.addEventListener('dotpal-poke', (e) => e.detail); // { count }: quick clicks in a rowStyling
dot-pal {
--dp-size: 200px; /* same as the size attribute */
--dp-color: hotpink; /* same as the color attribute */
--dp-glow: transparent; /* turn off the glow behind the pal */
}
dot-pal::part(bubble) { background: #111; color: #fff; }
dot-pal::part(svg) { filter: drop-shadow(0 10px 20px rgb(0 0 0 / .4)); }The parts you can style are root, idle, actor, svg and bubble.
Frameworks
React 19+: import 'dotpals', then <dot-pal character="grok" state={agentState} />.
Vue: set compilerOptions.isCustomElement = (tag) => tag === 'dot-pal'.
TypeScript: types are included, and document.querySelector('dot-pal') is typed as DotPal.
SSR: importing on the server is safe. The element renders once it reaches the browser.
Add your own character
Characters are plain SVG drawn in a 200×200 viewBox. They sit on the bottom edge and "peek" up over it.
import { registerCharacter } from 'dotpals';
registerCharacter('ghost', {
label: 'Ghost',
color: '#e8e8ff',
tap: 'spin',
look: 6, // how far the eyes follow the cursor
mouth: [100, 170], // where mood mouths are drawn
cheek: 34, // blush distance from the mouth
eyes: { at: [[80, 130], [120, 130]], r: 9 }, // where expression eyes go
render: ({ body }) => ({
body: `<rect fill="${body}" x="30" y="50" width="140" height="220" rx="70"/>`,
face: `
<g class="dp-look">
<g class="dp-blink"><circle cx="80" cy="130" r="9"/></g>
<g class="dp-blink"><circle cx="120" cy="130" r="9"/></g>
</g>`,
}),
});The body is automatically covered in fur and shaded.
Put
class="dp-blink"on each eye so it blinks and reacts to moods.Put
class="dp-look"on anything that should follow the cursor.eyes(optional) says where the eyes are, so the pal can swap in expression eyes:at(the two centres),r(their size), and optionallyink(their color),glow(trueor a color) andown(expressions your eyes already do well, e.g.['wide']). While they show, the parts markedclass="dp-eyes"hide (or the.dp-blinkparts). Withouteyes, the eyes just squint for moods.Let bodies run below
y=200, so jumping reveals more body instead of a flat edge.
You can add actions too, with registerAction('pop', { keyframes, duration, particles }). particles is a shape (heart, sparkle, star, sweat or z, drawn as SVG) or any text or emoji.
Accessibility
Each pal has
role="img"and anaria-labelthat includes its current mood, for example "Grok (working)".With
prefers-reduced-motion: reduce, the pal keeps its faces, blinks and state changes, but skips the big moves: idle loops, eye wandering, leaning, particles, the floating zs, state entry moves and the hover, click and dizzy moves.Speech bubbles are decorative. Keep your own visible status text for screen-reader users.
Documentation
The full guide is at rikinshah787.github.io/dotpals/guide: getting started, every feature, each agent integration, the CLI, configuration and environment variables, the bridge's HTTP API, the <dot-pal> component, privacy and security, and troubleshooting. Its source is in site/guide/ in this repository, so it's also published wherever the site is hosted.
For contributors, docs/ARCHITECTURE.md explains how the pieces fit together: adapters, the bridge, the activity model, the story engine, the desktop app, and how to add an adapter.
Roadmap
"It's stuck" alerts: a gentle ping when an agent goes in circles (no progress, the same file back and forth, a test that won't pass).
Morning brief and weekly recap: what your agents did, what's unfinished and what's failing, per project.
Token use per request, from the agents' own logs.
More agents: Windsurf, Cline, Aider and others, as each gets a documented way in.
Signed installers for Windows and macOS, so Node isn't needed.
Ideas and pull requests are welcome. Open an issue to discuss.
Contributing
npm install # dev only: Electron for the desktop window
npm test # node --test, no dependencies needed
npm run float # the desktop pal
npm run dashboard # the dashboard
npm run dev # the web component playground on http://localhost:5173See CONTRIBUTING.md. To support a new agent, add an adapter next to bridge/adapters/codex.js, which is a good template for any agent that writes a session log.
Credits
Reading test results and double-checking unclear ones builds on claude-referee by Ismail Dasci (MIT): its test-output parsers, redaction rules and "done" question are adapted in
bridge/ui/testout.js,bridge/redact.jsandbridge/checker.js. See THIRD_PARTY_NOTICES.The optional checkers are Laya by Convai Innovations (runs on your computer; not bundled, dotpals installs it from PyPI when you click Set up Laya) and TypeSafe's Jev, through its MIT-licensed SDK
@typesafe-ai/sdk.
Trademarks
Character names are playful nicknames. dotpals is not affiliated with or endorsed by Anthropic, OpenAI or any other AI company, and the characters are original artwork, not logos.
License
Available Tools
8 toolsagents_nowA
What the coding agents on this computer (Claude Code, Codex, Cursor and others dotpals watches) are doing, or did in the last 2 hours, newest first: the project, whether each is working, waiting for you, finished or stopped, what it was asked, its latest step or result, and how its tests stand. dotpals reads this from what the agents actually did (their tool calls and the test output), not from what they said. Use it for "what are my agents doing?" and to find a session ID for the other tools.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does so well: it discloses data provenance (reads actual tool calls and test output, not agent self-reports), a time window (last 2 hours), and ordering (newest first). It does not cover pagination, limits, or the shape of the session ID returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core capability is front-loaded in the first sentence, followed by the provenance note and the usage hint. The first sentence is long and list-heavy, but every clause adds distinct information about returned fields rather than restating the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description appropriately explains the returned data (fields, statuses, test standing) and the data source. It is nearly complete for a zero-parameter read tool, lacking only edge-case coverage such as what an empty result means or how the session ID should be consumed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters (empty object schema), so there is no parameter meaning to convey. Per the rubric this is the baseline 4 for a parameterless tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource (what coding agents on this computer are doing) with concrete output fields (project, working/waiting/finished/stopped status, latest step, test standing) and a scope (last 2 hours, newest first). It is clear but does not explicitly differentiate itself from overlapping siblings like recap or today, which likely also summarize agent activity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit usage triggers ("what are my agents doing?") and points to the downstream use case of finding a session ID for other tools. There is no "when-not-to-use" or direct naming of an alternative sibling, which keeps it short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_my_workA
Call this before you tell the user your work is done. dotpals checks your own session in this project the way a reviewer would: did the tests run after your last change, did they pass (and why not), did you change the tests to make them pass, did anything fail and stay failing, did you do anything risky. It answers "Looks done" or says what to fix first. The evidence is what you actually ran, not what you remember.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Only if dotpals can’t find your session: your project's folder path or name. Default: the folder this server runs in. | |
| session | No | Only if dotpals picks the wrong session: your session ID (or its first characters). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does reasonably well: it enumerates the checks performed (tests run after last change, pass/fail and causes, test edits used to force passes, lingering failures, risky actions) and the shape of the answer ("Looks done" versus what to fix first). It does not state that the tool is read-only/side-effect free, nor any cost, auth or timing traits, which is the remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the imperative call-to-action, then explains the checks and the output in three compact sentences. Slightly chatty with phrases like "the way a reviewer would," but every sentence is doing work.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-optional-param, no-annotation, no-output-schema tool, the description covers what is inspected and what the response means well enough to invoke it correctly. It omits how it relates to the other audit/status tools, which is the only real completeness gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both params are optional with clear "only if dotpals picks the wrong session/project" semantics already in the schema. The description itself adds no parameter meaning, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: it checks the agent's own session in the current project, described concretely as a reviewer-style audit of tests, failures, test tampering and risky moves. It is clear what the tool does, but it does not distinguish itself from siblings like test_status and risky_steps, whose territory it partly claims.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger: "Call this before you tell the user your work is done." That is a clear when-to-use signal. It stops short of naming alternatives (e.g. test_status, ready_to_merge, risky_steps) or saying when not to call it, so the routing guidance is incomplete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
handoff_noteA
A hand-off note, so another agent (or a new session) can pick up where a session left off: the original ask and the last request, what was done with the evidence, the files it changed, how the tests stand (with the failure), what's left on its plan, what's risky and its last message. Markdown, ready to paste. Default: the newest session in this project.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | The project: its folder's path (best: dotpals matches where each agent works, so same-named folders, a monorepo's packages and worktrees stay apart) or its name. Default: the folder this server runs in. | |
| session | No | A session ID from agents_now (or its first characters). Default: the newest session in the project. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the output format ('Markdown, ready to paste'), the exact content sections, and the default behavior ('the newest session in this project'). It stops short of stating whether the operation is read-only or what permissions are needed, but for a note-generation tool the disclosure is strong.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but front-loads the core purpose and then lists the components of the note in a compact sequence. Every clause contributes information, though the long extended sentence could be structured with clearer delimiters for easier scanning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter note-generation tool with no output schema and no annotations, the description is nearly sufficient: it explains the output format, the content, and the default session. It could be more complete by noting how the session ID is obtained or by distinguishing the tool from adjacent recap/status tools, but the essential information is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema already explains both parameters, including the project folder default and the session ID-from-agents_now default. The description repeats the session default ('the newest session in this project') but adds no new syntax, format, or constraint beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific deliverable: a hand-off note that lets another agent or session resume, and enumerates the exact contents (original ask, last request, evidence, files, test failures, plan, risks, last message). It does not, however, differentiate itself from siblings like recap or check_my_work.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The scenario is implied — 'so another agent (or a new session) can pick up where a session left off' — which gives an agent enough context to understand the intended use case. But there is no explicit when-to-use guidance, no exclusions, and no mention of alternatives such as recap or check_my_work.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ready_to_mergeA
Is an agent session's work ready to merge? The checks a reviewer would make, each ✓ or ✗: tests ran, after the last change, and passed (and not because the tests were changed); no failure left behind; nothing risky (force pushes, .env changes, recursive deletes…); the commit was tested. Pass branch to ask about a git branch ("is feat/login ready?"). Default: the newest session in this project.
| Name | Required | Description | Default |
|---|---|---|---|
| branch | No | A git branch ("feat/login"): the newest session whose work ended on it, this project's first. | |
| project | No | The project: its folder's path (best: dotpals matches where each agent works, so same-named folders, a monorepo's packages and worktrees stay apart) or its name. Default: the folder this server runs in. | |
| session | No | A session ID from agents_now (or its first characters). Default: the newest session in the project. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the exact review criteria, notable edge cases ("not because the tests were changed"), and expected output form (each check ✓ or ✗). It still does not state that the operation is read-only/side-effect free or what is returned when no session exists, so a gap remains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core question, then the check list, then parameter guidance — an efficient ordering. Slightly dense with a parenthetical aside ("force pushes, .env changes, recursive deletes…") but every sentence contributes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by describing the ✓/✗ check-by-check return structure, which is what an agent needs to interpret results. No annotations means safety/non-mutation behavior is unaddressed, keeping this just short of fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters are already documented with defaults and resolution rules. The description restates the `branch` use case and the newest-session default but adds no syntax or precedence detail beyond what the schema provides, matching the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific question it answers ("Is an agent session's work ready to merge?") and enumerates the concrete checks performed (tests after last change, no leftover failures, no risky operations, tested commit). It is clearly a readiness-verification tool, but it never distinguishes itself from close siblings like check_my_work or test_status, leaving the agent to infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear invocation context: pass `branch` to query a git branch ("is feat/login ready?"), otherwise default to the newest session in the project. It does not, however, say when to prefer this over check_my_work or test_status, so alternative selection is only implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recapA
What agents actually did, request by request: what each was asked, one plain sentence on what it did (files changed, how the tests went, what it shipped), warnings worth a look (tests that passed only after they were changed, too), why its tests failed, and the files it changed. For a project (default: this one, the last 24 hours) or one session (all of it).
| Name | Required | Description | Default |
|---|---|---|---|
| since | No | Only work since then: minutes ago ("90") or a time ("2026-10-04T09:00"). | |
| project | No | The project: its folder's path (best: dotpals matches where each agent works, so same-named folders, a monorepo's packages and worktrees stay apart) or its name. Default: the folder this server runs in. | |
| session | No | A session ID from agents_now (or its first characters). Default: the newest session in the project. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose meaningful behavior: the default window is the last 24 hours for a project versus the full history for a session, and it flags warnings such as tests that passed only after being edited. It does not explicitly state the operation is read-only, but the nature of a 'recap' makes that clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single long sentence densely packed with nested parentheticals and em-dash asides, which makes it hard to parse on first read. Front-loading is decent ('What agents actually did'), and every clause carries content, but the run-on structure undercuts readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-output-schema, no-annotation, three-optional-parameter read tool, the description sufficiently covers what content is returned and the scope defaults. It stops short of fully spelling out how warnings or session defaults interact, but nothing critical to calling it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters and their defaults, making 3 the baseline. The description only echoes the project/session scoping and adds the 24-hour default window, adding little beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource (agent activity) and enumerates exactly what the recap contains: what each agent was asked, one-sentence summaries, warnings, test-failure reasons, and changed files. It is clear what the tool produces, but it never names or contrasts any sibling (today, agents_now, check_my_work), so an agent must infer the distinction itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives scoping context: 'For a project (default: this one, the last 24 hours) or one session (all of it).' That implies when each parameter applies, but there is no explicit when-to-use/when-not guidance and no routing to the alternatives among the siblings, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
risky_stepsA
Risky things agents did, each with when, which agent and the command: force pushes, recursive deletes, git changes thrown away, dropped database tables, scripts piped from the internet, sudo and permission changes, force-stopped programs, changes to .env and other secret files, and the same command failing again and again. Use it for "did my agents do anything dangerous?" or before trusting their work. For a project (default: this one, the last 24 hours) or one session (all of it).
| Name | Required | Description | Default |
|---|---|---|---|
| since | No | Only work since then: minutes ago ("90") or a time ("2026-10-04T09:00"). | |
| project | No | The project: its folder's path (best: dotpals matches where each agent works, so same-named folders, a monorepo's packages and worktrees stay apart) or its name. Default: the folder this server runs in. | |
| session | No | A session ID from agents_now (or its first characters). Default: the newest session in the project. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does disclose useful traits beyond the schema: the detection categories and the scoping defaults (project = this one, last 24 hours; session = whole session). It still says nothing about whether it's read-only, result volume, or performance, which keeps it below a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The long enumeration in the first sentence is dense but each item earns its place by defining the tool's detection scope, and the usage conditions follow immediately. It is slightly run-on but front-loaded and waste-free.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description compensates by stating what each entry contains ("when, which agent and the command") and by covering scope and time defaults. An agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds a default time window ("the last 24 hours") that the schema's `since` parameter does not state, and clarifies session scope means "all of it." That is genuine added meaning over the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It names a specific concept (risky agent actions) and enumerates the concrete verbs/events that qualify: force pushes, recursive deletes, dropped tables, sudo, .env edits, repeated failures. An agent can immediately tell this is a risk-reporting tool and distinguish it from siblings like test_status or ready_to_merge.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit triggering questions: "did my agents do anything dangerous?" or "before trusting their work." That is clear when-to-use guidance, but it names no sibling alternative and gives no when-not-to-use condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
test_statusA
Is the code really tested? How the tests stand in an agent session, from the test runs dotpals saw: passed or failed (read from the test output, with the counts, and why the last run failed), when, and whether code changed after the last run (stale) or was never tested. Only tests an agent ran count. Default: the newest session in this project.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | The project: its folder's path (best: dotpals matches where each agent works, so same-named folders, a monorepo's packages and worktrees stay apart) or its name. Default: the folder this server runs in. | |
| session | No | A session ID from agents_now (or its first characters). Default: the newest session in the project. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it does meaningful work: it discloses what is measured (only agent-run tests), what the output conveys (pass/fail, counts, last-failure reason, timestamp) and a non-obvious derived trait (stale when code changed after the last run, or never tested). It omits whether the call is read-only, its cost, or error behavior, so it is solid but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It front-loads with a rhetorical question ('Is the code really tested?') that delays the actual operation, then packs return-value, scoping, and staleness detail into a single dense run-on sentence. The content mostly earns its place, but the structure is cluttered rather than cleanly segmented.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must describe the return payload, and it does: pass/fail, counts, why the last run failed, when it ran, and stale/never-tested state. Combined with complete schema coverage for both parameters, an agent has enough to call the tool and interpret results, with minor gaps only around read-only/error behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both 'project' and 'session' are already documented in the schema, including their defaults. The description only restates the project-default scoping and the session default ('newest session in this project'), adding no syntax or format detail beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource and outcome: the test status of code within an agent session (passed/failed, counts, failure reason, timing, staleness). It is clearly a status-inspection tool distinct from siblings like ready_to_merge or check_my_work, though it never names an alternative to route against. Purpose is clear but not sharply differentiated from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implicit guidance is present — 'Only tests an agent ran count' and 'Default: the newest session in this project' tell the agent when the tool is meaningful and which session it targets. However, there is no explicit statement of when to use this versus siblings (e.g., agents_now for sessions, check_my_work for review), so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
todayA
What the agents did today, for a standup or an end-of-day note: per agent, how many requests, files changed and test runs, then each request with what it was asked and one plain sentence on what it did, warnings worth a look (tests that passed only after they were changed, too) and why its tests failed. For a project (default: this one).
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | The project: its folder's path (best: dotpals matches where each agent works, so same-named folders, a monorepo's packages and worktrees stay apart) or its name. Default: the folder this server runs in. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does surprisingly well: it discloses the exact contents surfaced, including subtle signals like tests that passed only after being changed and the reasons tests failed. It does not state read-only safety explicitly or any auth/rate behavior, but the output-shape disclosure is strong.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The lead clause 'What the agents did today' is well front-loaded, and every fragment is informative, but the whole thing is one sprawling sentence with nested parentheticals that is harder to parse than it needs to be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly takes on the job of explaining return contents in detail, and it also states the default scope. For a single-optional-param read tool it is close to complete, missing only tool-selection routing against its siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the single 'project' parameter is already fully documented. The description's 'For a project (default: this one)' merely restates the schema default, adding no syntax or matching nuance beyond it. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly conveys the resource being produced: a per-agent daily activity digest covering request counts, files changed, test runs, per-request detail, warnings, and test failures. It is specific about content and scope, though it never names the closely related siblings (recap, handoff_note) to help an agent distinguish the right reporting tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It offers a use-case hook ('for a standup or an end-of-day note'), which implies when to reach for it. However, there are no explicit when-not conditions and no named alternatives to route against the several other summary/status siblings, leaving selection to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.10.0- First observed
agents_now - First observed
check_my_work - First observed
handoff_note - First observed
ready_to_merge - First observed
recap - First observed
risky_steps - First observed
test_status - First observed
today
TDQS
Scored across 8 tools
Most tools are distinguishable by domain facet (test status, activity, risk, handoff), but several overlap heavily: recap vs today both summarize what agents did, and ready_to_merge vs check_my_work both run near-identical reviewer checks (differing mainly in whose session is the default). test_status also overlaps with the test facets embedded in ready_to_merge and check_my_work, so an agent can plausibly misselect among these.
All names use snake_case, which is consistent, but the convention itself is mixed: some are question-style phrases (agents_now, ready_to_merge, check_my_work), some are bare nouns (recap, today), and others are noun_phrases (test_status, risky_steps, handoff_note). Readable but no single predictable verb_noun pattern.
Eight tools is well within a reasonable range for an agent-observability server and each maps to a plausible use case. Slightly heavy given that at least two pairs (recap/today, ready_to_merge/check_my_work) are near-duplicates, but overall the count is sensible.
The surface covers the domain well: per-session and per-project activity, test status, merge readiness, risk detection, handoff, and self-review. Coverage of the coding-agent lifecycle is broad; minor gaps (e.g. no explicit session-listing or time-range customization beyond defaults) are workable.
Maintenance
Related MCP Connectors
Read-only AI coding tools for change verification, release readiness, capacity, and guidance.
Cross-agent artifact workspace with provenance across Claude Code, Codex, Cursor, LangGraph.
Change-aware CI validation and affected-test guidance for coding agents.
Deterministic AI code review, with an audit record. Governance inside the agent loop.
Related MCP Servers
- AlicenseAqualityAmaintenanceProvides a workspace-safe, read-only bridge between browser-based AI planning/review and local coding agents, enabling structured plan, execution summary, and review handoffs without granting shell, file write, or Git push access.11MIT
- AlicenseAqualityAmaintenanceLocal code review for people and their coding agents: an agent opens a review of its changes and hands over a link, the person comments on lines, and the agent reads the threads, replies and resolves what it fixed. Reads the git worktree directly, uncommitted work included; nothing leaves the machine.21776 npm18GPL 3.0
- AlicenseNot gradedqualityAmaintenanceObservation-only CI/CD validation for AI coding agents. Analyze a change, compare safe test commands, and report evidence without altering required CI.1,694 npmAGPL 3.0
- AlicenseAqualityCmaintenanceEnables read-only inspection of local coding-agent sessions across Claude Code, OpenCode, Cursor, and Codex, including running status, topics, todos, and transcripts, without writing or sending messages.4169 npm1MIT