Skip to main content
Glama

open-mcp-apps

CI npm license node MCP Registry

English | 简体中文

Give your AI a persistent, reusable UI. It builds the app once — you keep it forever.

open-mcp-apps is an open engine built on MCP Apps (ui://, io.modelcontextprotocol/ui) — an extension to the core Model Context Protocol specification, and the first official one, live since 26 January 2026. It gives any MCP-Apps-capable host (Claude Desktop, claude.ai, Codex, ChatGPT, …) three things the extension itself doesn't provide:

  1. An app registry the AI can write to. Ask for a UI that doesn't exist — the AI reads the authoring guide, writes a single-file HTML app against a tiny window.oma API, and saves it. From that moment you can open it by name, in this chat and every future one.

  2. Persistent, versioned data — separate from the UI. Apps bind to generic collections of items backed by SQLite plus an append-only change_event ledger. Every mutation is an idempotent domain command (command_id) with optimistic concurrency (expected_version). The AI and the human edit the same store — the widget is just a view.

  3. A shell runtime so AI-written apps actually work. Serving ui://, the engine wraps the app with the official MCP App bridge, host theming (Claude's design tokens, light/dark), and the window.oma data API. What you write is a view; the protocol, persistence, idempotency and theming are the engine's problem.

Which of those hosts this engine reaches. It runs on your machine and binds 127.0.0.1, so it serves the hosts on that same machine: Claude Desktop, Claude Code, Codex, plus its own browser viewer. A browser host cannot reach a loopback server on your laptop, so claude.ai and ChatGPT web need a remote deployment — today that means the hosted openmcp.app, which runs this same engine for you. Running that remote shape yourself is on the roadmap and not done.

Version

0.7.2 (CHANGELOG.md)

License

MIT, whole repository (LICENSE · LICENSING.md)

npm

@2nd1st/open-mcp-apps — scoped; the unscoped name is an unrelated package

Command

npx -y @2nd1st/open-mcp-apps — the line your host's MCP config runs; a stdio server, not something to run by hand (typed into a terminal it just waits, and says so)

Requires

Node 22 or newer. git too, on the installer path

Surface

33 tools · a built-in App Store · 3 system apps seeded

Platforms

macOS · Windows · Linux

Hosts

Claude Desktop · Claude Code · Codex · ChatGPT web — see Host support

Hosted

openmcp.app — the remote shape, and the way to reach browser hosts (claude.ai, ChatGPT)

Install

open-mcp-apps runs as a local MCP server. First get it connected to your host (below); then onboarding happens inside the host, separately — that's where the AI builds your first app.

From npm — nothing to clone

If you are comfortable editing your host's config file, point it at the published package and let npx fetch the engine. This path needs only Node 22 — no git, and no checkout for you to keep updated. Paste this into your host's MCP server config:

{
  "mcpServers": {
    "open-mcp-apps": {
      "command": "npx",
      "args": ["-y", "@2nd1st/open-mcp-apps"]
    }
  }
}

Two things the installer below does that this path does not: it registers the server into every host it finds, and it pre-seeds the built-in system apps (settings, dashboard, App Store) into your store — so on the npx path your registry starts empty and your AI installs what it needs from the App Store on demand, which is fully available either way. Your data lives in the same fixed per-user store, so you can move between an npx server and a cloned one without migrating anything.

A note on npm: this project publishes under the scoped name @2nd1st/open-mcp-apps. The unscoped open-mcp-apps on the registry is not this project — that name is held by an unrelated package. Check for the @2nd1st/ prefix; the scope is the only thing telling the two apart.

With the installer — one command

Installing needs a shell, so the chat apps (Claude Desktop, Codex) can't install themselves — use one of these instead:

curl -fsSL https://raw.githubusercontent.com/2nd1st/open-mcp-apps/main/install.sh | sh

It opens a short picker to choose which hosts to register into — Claude Desktop, Claude Code, Codex — plus your permission preference. Skip it with -s -- --yes, or target one host with -s -- --host codex.

Where it puts the clone, before you pipe anything into a shell: ~/open-mcp-apps. Set OMA_DIR to put it somewhere else — curl -fsSL <url> | OMA_DIR=~/src/oma sh. Re-running the one-liner updates that same clone in place instead of making a second one. Your apps and data are not in it (they live in the per-user store under Configuration), so the folder is safe to move or delete — and node uninstall.mjs does not delete it for you.

With a coding agent (Claude Code, Codex CLI — they have a shell), paste:

Read https://raw.githubusercontent.com/2nd1st/open-mcp-apps/main/install.md and follow it.

Either way, install.mjs registers the server into each host you pick, idempotently — it never clobbers your other servers, pins a stable node launcher (native SQLite ABI), reports what changed, and cleans up a pre-rename entry if one lingers. Your data lives in a fixed per-user store (not inside the clone), so every host shares the same apps and data.

From a clone — for development

git clone https://github.com/2nd1st/open-mcp-apps && cd open-mcp-apps
npm install
node install.mjs        # same picker as the one-liner above

To wire a clone into a host by hand instead, point it at the checkout — this is the shape install.mjs writes:

{
  "mcpServers": {
    "open-mcp-apps": {
      "command": "node",
      "args": ["/absolute/path/to/open-mcp-apps/src/server.mjs"]
    }
  }
}

Hosted

openmcp.app runs the engine for you. The engine in this repository binds 127.0.0.1 by design, so a self-hosted remote deployment is not a supported shape yet — see Status and roadmap.

Uninstall

node uninstall.mjs unregisters the server from every host it finds — but keeps your data: the shared store stays put, so re-installing later restores every app and all data. It also leaves the installer's clone (~/open-mcp-apps, or wherever OMA_DIR pointed) on disk — nothing here ever deletes that folder, so remove it yourself when you want the checkout gone.

node uninstall.mjs           # unregister from all detected hosts — keeps your data
node uninstall.mjs --purge   # also delete the shared store (apps + data), irreversible
node uninstall.mjs --check   # read-only: show what's registered and what would change

Related MCP server: AI PC Assistant MCP Server

Requirements

  • Node 22 or newer, on macOS, Windows or Linux.

  • git — only on the installer path. The npx path above needs neither git nor a checkout. The installer checks for both and stops with a message rather than half-installing if either is missing.

  • A host that renders ui:// if you want widgets in the conversation. Terminal hosts (Claude Code in a terminal, codex CLI) drive the same data by design and put the UI on a browser screen beside the terminal instead — one that can follow along, showing whatever the AI just opened. The per-host detail is in Host support.

  • After installing or updating, fully quit and reopen the host (Cmd-Q, not just closing the window) — it keeps its old server process on the old data until fully quit.

Configuration

Every setting is an environment variable, set in the env block of your host's MCP server entry:

{
  "mcpServers": {
    "open-mcp-apps": {
      "command": "npx",
      "args": ["-y", "@2nd1st/open-mcp-apps"],
      "env": {
        "OMA_VIEWER": "1",
        "PORT": "8787",
        "OMA_DYNAMIC_TOOLS": "0"
      }
    }
  }
}

Variable

Default

What it does

OMA_VIEWER

1

The browser viewer on loopback. 0 doesn't start it at all.

PORT

8787

Where the viewer listens.

OMA_DYNAMIC_TOOLS

0

1 also publishes one open_<name> tool per saved app. Off by default because it costs prompt cache — and one approval prompt per app.

OMA_DB

per-user store

Path to the SQLite store. Set it to isolate a store.

Where your data lives. The whole store is one SQLite file, open-mcp-apps.db, in ~/Library/Application Support/open-mcp-apps/ (macOS), %APPDATA%\open-mcp-apps\ (Windows), or $XDG_DATA_HOME else ~/.local/share/open-mcp-apps/ (Linux). It is outside any clone, which is why every host shares the same apps and data.

First-run permissions. The first few tool calls each show an approval dialog — pick "Always allow". The tool set is small and stable on purpose: read-only tools generally skip approval, and the single open_app tool covers opening every app (including ones the AI creates later) behind that one grant, so nothing new asks again — on every host the installer registers, with no exceptions any more. From 2026-07-28 to 2026-08-16 there were two: Claude Desktop and Claude Code were registered with OMA_DYNAMIC_TOOLS=1, which routed around a chat-surface bridge regression by giving every app its own open_<name> tool, at one approval prompt per app. Re-measured on Desktop 1.30096.5, that symptom is gone, so the installer no longer sets the flag for anybody — KNOWN-ISSUES.md carries both readings. If you installed during that window, your entry still has the flag: node install.mjs --check reports it as stale, and re-running the installer removes that one key while leaving every other env value you have set exactly where it is. You can also batch approvals in Settings → Connectors → open-mcp-apps → Tool permissions.

Usage

Start in your host. Restart it after installing. New here? The engine ships one MCP prompt, get_started. A host that surfaces prompts lists it as Get started with open-mcp-apps; hosts that render prompts as slash commands spell it /mcp__open-mcp-apps__get_started. Picking it hands the AI the whole opening move. Not every host surfaces prompts — where yours doesn't, nothing is lost, because the prompt is just a sentence you can say yourself: "I just installed open-mcp-apps — show me how to use it with a couple of examples, and suggest a few apps that fit how I work." Either way it looks at what you already have and what the App Store already offers, asks you a couple of questions, and sets up a first app tailored to you. This step is separate from install and lives in the host. Or just ask directly:

  • "make me a board for what I'm juggling right now" → the AI writes it, seeds it, and opens it (persistent)

  • "make me a habit tracker" → watch it read the guide, write the app, save it, open it

  • close the app, reopen, ask again → everything is still there

The loop

"make me a kanban"
      │
      ▼
list_apps ── exists? ──► open_app {app: "kanban"}     (reuse, instant)
      │ no
      ▼
get_app_guide ──► AI writes HTML ──► save_app
      │
      ▼
open_app {app: "kanban"}  →  rendered inline, themed, persistent — reusable in every future chat

Apps accumulate. Each one is single-purpose and independent — a board, a tracker, a splitter — minted for the task in front of you and kept for the next time you need it.

What it looks like

Apps render inline, in the chat you were already having. Ask for one and the AI writes it:

Codex — asking for a reading tracker; the AI writes it and it renders inline, already holding the three books

Come back in another chat — or another host — and it's still there, with your data in it:

Claude — a new chat opens the same reading list, now eight books long

The built-in App Store — rebuilt in 0.5.0 as a real storefront — ships 22 ready-made apps, with working previews and one-click install:

The App Store — live previews of ready-made apps

Companion — an AI character with shared memory

Family Week — dinners, chores rotation, shopping and weekend plans

Study Cards — spaced repetition with review heatmap and deck shelf

Knowledge Cards — a visual library of saved answers

Every app above is a single HTML file bound to plain data collections — written with the same window.oma API and authoring guide your AI will use for the apps it builds you.

Multiple widgets in one conversation work fine (habit-streaks + meal-planner side by side).

The browser viewer, and the port it binds

Every install runs a small local web server on http://127.0.0.1:8787. It is how you see your apps outside a chat window — one page per app, the same data your AI is reading — and in a terminal host it is the only way to see them at all, so the AI hands you the link when it builds or opens something.

It starts on its own; OMA_VIEWER and PORT above change that. If the port is already taken by another open-mcp-apps process, that one is already serving the same data and this one just shares its address; if it is taken by something else, you get no viewer and no links rather than a link into a stranger's server.

There is no password on it, and that is deliberate. The listener is hard-wired to 127.0.0.1, so there is no setting that makes it answer from another machine. Any program on your computer that could reach the port can already open the SQLite file directly — a password would be a lock beside an open wall. The one way this reaches the internet is a tunnel you start yourself, which is its own deliberate decision; while a tunnel is up, treat its URL as a secret, because it is currently the only thing standing between the internet and your data.

Host support

Live-tested 2026-07-22; ChatGPT web row updated 2026-07-28. Both readings predate 0.5.0 — the largest change so far, and later than either date. Apart from the cells that carry their own 2026-08-16 date, nothing in this table has been re-tested on 0.5.0 or newer; a date says when that row was true, not that it was checked again since.

Host

Renders widgets

Human clicks widget

AI operates data

Same store

Claude Desktop (local stdio)

✅ — re-checked 2026-08-16 on 1.30096.5: the universal open_app renders in chat without the OMA_DYNAMIC_TOOLS workaround that shipped for 1.24012.9 (see KNOWN-ISSUES)

✅ full loop incl. sendMessage reply

✅

✅

Browser viewer (/view/<name>)

✅

✅ (no chat attached — sendMessage degrades to a notice)

via CLI AI

✅

Codex desktop (ChatGPT app, enable_mcp_apps flag) — tested against a local engine; remote not established

✅ experimental

◐ updates/toggles from widget clicks work; adds were blocked host-side. The umbrella request openai/codex#28912 (an enhancement: "make MCP apps work end-to-end in the Codex GUI") closed as completed on 2026-08-05 — but #30092, the bug matching this exact failure and reproduced there by a third party, was still open on 2026-08-16. Not re-tested here either way, so the cell stays ◐. See KNOWN-ISSUES

✅

✅

Claude Code — in a terminal (claude mcp)

— in the chat, by design (text fallback) — but see a screen beside the terminal

— in the chat

✅

✅

Claude Code — the Code surface inside the Claude app

✅ live-tested 2026-08-16: an app opened with the universal open_app renders inline, the same shape the chat surface gives

not measured on this surface

✅

✅

codex CLI / IDE

— in the chat, by design (text fallback) — but see a screen beside the terminal

— in the chat

✅

✅

ChatGPT web (Work mode)

✅ live-tested 2026-07-28 (remote HTTPS) — renders at full height, no clamping; a widget loses its data after a page refresh (mitigation shipped, awaiting live re-test here — see KNOWN-ISSUES)

✅ a widget button added a row and it stuck

✅

✅

Everything rides the MCP Apps bridge, so host fixes upstream (e.g. #28912) benefit this project with zero changes.

On Claude Code specifically: it is one product with two surfaces, and only one of them can draw — which is why it takes two rows. In a terminal there is no inline widget surface at all: that is architecture rather than a gap, and it is exactly what a screen beside the terminal is for. The Code surface inside the Claude app has a UI and renders inline; the 2026-08-16 reading there came through the universal open_app.

On Codex specifically: plugins are registered on the web side, so a locally-installed engine is reached as an MCP server, not as a plugin — which is the right path for a self-hosted install anyway. Widget rendering in the ChatGPT desktop app also appears to depend on how you are signed in (we have seen it work under an account sign-in; not yet established under an API key).

A screen beside the terminal

Those two — cells say the chat shows text. They do not say there is no UI. Since 0.5.1 the engine remembers which app was opened last and pushes that pointer to the viewer on the /events frame, and an app can place a region — oma.embed("@live", {into}) — that mounts whatever the AI opened last and swaps itself when the AI opens another. The App Store ships one: install live, open http://127.0.0.1:8787/view/live in a window you then leave alone, and the terminal keeps the conversation while that screen shows the app. The AI opens or writes from the CLI, the screen follows, and you can click, edit and type into it there. Headless CLI use is what it was built for.

The cheapest form of "that screen" is a browser pane in the same tiled workspace — one column over from the agent, in the window you are already working in. Same machine, no tunnel, no second device; most modern terminal setups can put a browser next to a shell, and that is all this needs. A second monitor is the same idea with more desk, and a screen on another device is the same idea again with the caveat below.

The costs are real and worth stating plainly. There is still no widget in the transcript. sendMessage degrades to a notice on a standalone page, exactly as in the Browser viewer row — clicks change data, they do not talk back to the chat. The viewer has to be running (OMA_VIEWER, on by default) with a browser pointed at it; this is not zero-config. And the listener is bound to 127.0.0.1, so "a spare tablet on the wall" means this machine's screen unless you put up the tunnel described above and accept what that section says about it. Inside a chat host the same region deliberately draws a placeholder instead of following anything.

Described from the code as built, not measured on a host — the live-test dates above cover the table, not this section.

Writing an app yourself

The AI is the usual author, but it isn't the only one — its context window shouldn't be the ceiling on what an app can be. Build one in your own editor, with your own bundler, and install it:

node install-app.mjs ./my-app.html              # yours, full trust — same as an AI-authored app
node install-app.mjs ./my-app.html --sandboxed  # untrusted: runs behind the runner, no capabilities
node install-app.mjs --list                     # what's installed, and under whose provenance

# a build pipeline's output: a readable template + its bundle, plus the declaration as its own file
node install-app.mjs ./ui.html --name my-app --manifest ./manifest.json \
  --asset ./dist/app.js --asset ./dist/app.css --update

Two shapes are accepted. One self-contained HTML document — no size cap (keep it lean: data lives in the collection, source is read in windows) — the engine injects the kit CSS, the host's design tokens and window.oma. Or a template plus a bundle: the HTML is a readable mount point that references its own build output (<script type="module" src="oma-asset:app.js">, <link rel="stylesheet" href="oma-asset:app.css">), --asset pushes those files into the app's file plane, and the engine inlines them the moment the document leaves the store — a widget's CSP allows no external subresource, and a host iframe could not reach this machine anyway.

The trade: the AI can no longer iterate on it — your files are the source of truth, you rebuild and re-install. It can still read a single-document app's source; for a template + bundle app it reads the template and the bundle stays yours (edit_app and the AI's save_app refuse with built_outside; re-running the command above is the edit). Either way the app shares your data like any other. Provenance is not overwritable in either direction, so an app installed --sandboxed stays sandboxed until you delete it. What a widget may connect to is the app's own declaration (manifest.csp, relayed to the host per the MCP Apps spec — nothing by default), and an app's functions run engine-side and may fetch — see RUNTIME.md §5.1 and §6.1, and the two type files for what an app sees — @2nd1st/open-mcp-apps/types/window-oma (the window.oma API) and @2nd1st/open-mcp-apps/types/oma-function (the args/api a function body gets) — both resolve through the package exports.

RUNTIME.md is the contract — the window.oma API in both modes, what a sandboxed app can still do, and the traps that only bite authors who aren't the AI. It carries a version (oma.contract) and test/runtime-contract.mjs pins it to the two runtimes' real surfaces, so it can't drift from them silently.

Security model

Trust is tiered by where an app came from. Locally-authored and system apps run in direct mode. The engine also ships a runner — a sandboxed srcdoc iframe with a CSP-first document and a minimal read-scoped bridge — as the mandatory execution mode for any app that isn't locally trusted, plus reserved security:* / policy:* config keys that generic data writes can't touch and an out-of-band privileged writer.

Honest status: everything in the OSS version — your apps, AI-built apps, and the built-in App Store apps (all first-party) — runs locally in direct mode with full trust; there is nothing third-party to sandbox yet. The runner is built and tested but dormant: it is the ready seam for shared/published apps later, where review + sandboxing arrive together. See SECURITY.md for the full threat model and trust tiers.

Design positions (why it's built this way)

  • UI and data persist separately, both versioned. Apps are views; collections are truth; the ledger is history. Swap either without losing the other.

  • The AI talks domain commands, never SQL, never raw state. That's what makes human+AI concurrent editing safe (idempotency + optimistic concurrency at the command layer).

  • Extension-first. Everything rides the MCP Apps bridge — no host-private APIs. One codebase should serve every host that renders ui://.

  • Single-purpose, not composite. Each app owns one scenario and its own collection; the engine mints a new one rather than cramming features into an old one. System apps (settings, dashboard) are the deliberate exception — engine-owned, privileged, allowed to see across collections.

Troubleshooting

Symptom

What it is

Updated, but the host still shows the old behaviour

The host keeps its old server process on the old data until fully quit (Cmd-Q, not just closing the window).

Approval dialogs came back after a Claude Desktop auto-update

A Desktop auto-update occasionally resets these decisions (upstream #56954, closed 2026-06-23 as not planned) — no fix is coming from that issue, so just re-allow.

One approval prompt per app

OMA_DYNAMIC_TOOLS=1 is in your host entry — either you put it there, or you installed between 2026-07-28 and 2026-08-16, when the installer set it for Claude Desktop and Claude Code as a workaround. node install.mjs --check calls such an entry stale; re-running the installer removes that one key and keeps the rest of your env. See Configuration.

No viewer link, or the viewer is somebody else's

The port is taken by a non-open-mcp-apps process. Set PORT to something free.

A widget loses its data after a page refresh (ChatGPT web)

Known, mitigation shipped, live re-test pending — KNOWN-ISSUES.md.

Widget clicks can update but not add (Codex desktop)

Blocked host-side. The umbrella request openai/codex#28912 closed as completed on 2026-08-05, but that one is an enhancement, not this defect: the matching bug, #30092, was still open on 2026-08-16. Update Codex and try, but expect it to still bite.

You want to start completely clean

Fully quit your host(s), delete open-mcp-apps.db (plus its -wal/-shm siblings) from the store directory under Configuration. All apps and data gone, irreversibly, while staying installed.

pnpm install exits 1 with ERR_PNPM_IGNORED_BUILDS

pnpm 11 refuses third-party build scripts until you decide about them, and calls that an error. Nothing here needs building — better-sqlite3 loads a prebuilt binary it ships, and esbuild's binary comes from its platform package — so the tree it leaves behind is complete and working. Answer pnpm approve-builds however you like, or use npm. We do not declare those scripts as allowed, because that would make pnpm compile better-sqlite3 on machines with no toolchain (a container, most CI) and fail there for nothing.

Development

src/server.mjs

stdio MCP server; single open_app path (per-app open_<name> tools off unless OMA_DYNAMIC_TOOLS=1)

src/http.mjs

/mcp (stateless Streamable HTTP) + /view/<name> browser viewer, bound to 127.0.0.1

src/store.mjs

SQLite: items + app registry + change_event ledger (idempotent, OCC)

src/shell-runtime.js

browser runtime injected into every app (window.oma)

src/shell.mjs

wraps stored HTML with runtime + design-token fallbacks at serve time

src/guide.mjs

the authoring contract the AI reads before generating an app

install-app.mjs

install an app you wrote yourself, from a file — the one door into the registry that doesn't go through the AI

components/

3 system apps installed on seed (settings, dashboard, app-store) + 22 App Store apps — not auto-installed; browse the app-store app for live previews with sample data and one-click install

npm test                     # every suite below, plus the static invariants and budget checks
node test/server-smoke.mjs   # 453 assertions over real stdio — incl. runtime app creation
node test/http-smoke.mjs     #  81 assertions over the HTTP transport (incl. SSE /events, viewer)
node test/provenance.mjs     #  39 assertions that an app's author — its trust tier — is not overwritable
node test/seed-smoke.mjs     #  22 assertions on the seed / design-kit pipeline
node test/files-smoke.mjs    #  41 assertions on the per-app file store (chunked uploads, GC races)

Contributions need nothing signed — MIT in, MIT out (CONTRIBUTING.md).

Status and roadmap

Early v0 — proven end-to-end on Claude Desktop; cross-vendor render + shared store proven on Codex desktop and the browser viewer.

What 0.5.0 changed (breaking, and the largest change so far — CHANGELOG.md has the full account):

  • An app's declaration is a first-class object. save_app takes ui and manifest as two slots instead of a manifest block buried in the document, and every revision snapshots both, so restoring brings back the pair.

  • An app can expose a function — a data→data closure the AI calls with call_function, run by the engine against that app's own collections. The seat is opt-in at createEngine and absent by default, so a hosted deployment cannot inherit it.

  • Deleting a row is confirmed by the engine, inside the store transaction every path passes through. App authors no longer write confirmation UI; the apps that carried their own arm-then-delete had it removed.

  • promote_app turns a one-off visual into a kept app in one atomic step, and edit_app takes a hash-checked {offset, length} range, so a model that has read a window can edit it without sending an anchor back up.

  • Settings and the App Store were rebuilt — rail navigation, in-place detail pages, and the storefront pictured above.

  • Underneath: SDK v1 → v2, 2026-07-28 in the supported protocol versions, and a tool surface audited down to 33 tools. Renamed and removed tools mean hosts will ask you to approve the tools once more after upgrading.

Where it stands:

  • engine: registry + shell + generic data commands + ledger

  • system apps installed (settings, dashboard, app-store); 22 App Store apps with live previews, one-click install

  • AI app creation loop (guide → save → open)

  • in-context onboarding (ask how to use it → the AI reads your history/memory and builds a tailored starter set)

  • security foundation: trust tiers + sandboxed runner + reserved config keys

  • multi-host discovery installer (Claude Desktop · Claude Code · Codex) + shared per-user store

  • npx one-command install (@2nd1st/open-mcp-apps on npm)

  • self-hosted remote (Streamable HTTP) as a supported shape → claude.ai / ChatGPT / mobile off an engine you run — the transport exists (src/http.mjs) and has been live-tested over HTTPS; what's missing is the hosted story, since the engine binds 127.0.0.1 by design. Those browser hosts already work against the hosted openmcp.app; this box is about doing it yourself

  • one-click install with no shell

  • app export/import → sharing → community App Store (review + runner sandbox activate here)

License

MIT, for the whole repository — the engine and the apps in components/ alike (LICENSE · LICENSING.md). Use it, fork it, modify it, embed it, run a modified version as a hosted service; keep the copyright notice with substantial portions you redistribute. That is the whole obligation. Up to v0.5.2 the engine was AGPL-3.0-only under a directory split — see LICENSING.md for what changed and why.

The names open-mcp-apps, openmcp.app, SecondFirst, and 2nd1st, and their logos, are not granted by the license — see TRADEMARKS.md. Fork the code freely; give your fork its own name.

Copyright © 2026 2nd1st.

© 2026 2nd1st

Available Tools

33 tools
apply_data_writesMany writes, one transactionA
DestructiveIdempotent

Apply up to 200 writes in ONE transaction — for applying many known rows in one go instead of one call per row. Each command is exactly what you would send to data_add_item / data_update_item / data_move_item, as {type, ...args}: type is add_item | update_item | move_item. Deletion is not one of them: a delete needs its own confirmation, which an all-or-nothing batch cannot give per row, so it goes through data_delete_item. All or nothing: the first failure rolls back everything and names which command failed. The reply is one line per command ({id, seq}) — not the rows, which you already have.

ParametersJSON Schema
NameRequiredDescriptionDefault
actorNo
commandsYes[{type: "add_item", collection, fields, group?}, {type: "update_item", id, fields}, …] — same arguments as the single-write tools
command_idYesidempotency key — generate a fresh uuid per action

Output Schema

ParametersJSON Schema
NameRequiredDescription
okYes
eotNo
seqNothe batch's last ledger position — pass it to data_changes as `since`
noteNo
countNo
reasonNo
resultsNo
failed_indexNo

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, so the safety profile is covered. The description adds genuine value beyond that: all-or-nothing rollback on first failure, failure attribution naming the offending command, and the one-line-per-command reply shape. It does not mention auth/permission requirements, so it falls short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and scale limit, then rules, then return shape. Every sentence carries distinct information (scale, mapping, exclusion, transaction semantics, output), with no repetition of the schema or annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a batch mutation tool with nested command objects and an existing output schema, the description supplies everything an agent needs: batching threshold, command vocabulary, argument provenance, rollback behavior, and error attribution. Return values are additionally summarized despite the output schema, which is a bonus rather than a gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67% and the description meaningfully augments it: it defines the {type, ...args} shape of commands, enumerates the three valid type values, and points to the single-write tools as the source of truth for the per-command arguments. The actor parameter is left undocumented in both schema and description, keeping this off a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (apply writes), a hard scale limit (up to 200), and the batching rationale (many known rows in one go instead of one call per row). It explicitly distinguishes itself from the single-row siblings data_add_item / data_update_item / data_move_item by naming them and by naming data_delete_item as the excluded case.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when (many known rows in one go), an explicit when-not (deletion is not one of the supported command types), and routes the excluded case to the correct alternative, data_delete_item, with the reason (a batch cannot give per-row confirmation). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

app_store_listList App Store appsA
Read-onlyIdempotent

Browse the built-in App Store: ready-made, high-quality apps shipped with the engine that the user can install into their registry. Shows install state. Renders no UI — open_app {app: "app-store"} shows the browsable App Store app.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
entriesYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds genuinely useful context beyond them: it is a non-UI call, it reports install state, and the catalog is engine-shipped rather than user-created content.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the resource and scope, and the sibling pointer arrives last where it is most actionable. The phrase 'high-quality' is mild marketing filler that does not help selection, keeping this below a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be detailed, and the description still flags 'shows install state' as the key payload. For a zero-parameter read tool with full annotation coverage this is close to complete; only the absence of explicit install/preview routing leaves a small gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing for the description to disambiguate, and it correctly adds no parameter narrative.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (list the built-in App Store catalog) and scopes it as engine-shipped apps installable into the user's registry, which separates it from list_apps (registry contents). It also explicitly distinguishes itself from open_app, so an agent can tell the two apart without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit alternative and the condition that selects it: this tool renders no UI, whereas open_app {app: "app-store"} shows the browsable App Store app. It does not address when to prefer this over install_from_app_store or preview_app_store_entry, so routing guidance is clear but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

call_functionCall an app functionA
DestructiveIdempotent

Run a function an app declares (manifest.functions) — data in, data out, no UI needed. The body is code the app's author wrote and it may make outbound network requests (a host may route them through an allowlisted gateway). Args are checked against the declared params; failures return the declared schema so the retry needs no extra read. The reply carries the return value plus a receipt per write.

ParametersJSON Schema
NameRequiredDescriptionDefault
appYes
argsNo
actorNo
functionYes
command_idYesidempotency key — generate a fresh uuid per action

Output Schema

ParametersJSON Schema
NameRequiredDescription
okYes
eotNo
noteNo
reasonNo
resultNo
writesNo
availableNo
violationsNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=true, openWorldHint=true, and idempotentHint=true, so the description's additional context about outbound network requests, allowlisted gateways, argument validation, and receipt-per-write adds meaningful behavioral detail. It doesn't contradict annotations. However, it could more explicitly address the destructiveHint implications (e.g., what writes occur and their effects).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense sentence that front-loads the core action and packs in key behavioral details without waste. It could be slightly more structured for readability, but every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description needn't explain return values, but it does mention the reply carries the return value plus a receipt per write, which is helpful. For a tool with 5 parameters, nested objects, and complex annotations, the description covers most essential aspects: function execution, network behavior, argument validation, and failure handling. It could be more complete by addressing the actor parameter or edge cases, but it's largely sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 20%, so the schema documents only command_id's idempotency purpose. The description adds some meaning by noting that args are checked against declared params and that failures return the declared schema, but it doesn't explain the roles of app, function, or actor parameters beyond what their names imply. With low coverage, more parameter detail would be expected.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states what the tool does: run an app-declared function with data in/data out. It distinguishes itself from siblings like open_app or get_app_html by emphasizing 'no UI needed.' However, it doesn't explicitly name or contrast with any specific sibling tool, leaving some ambiguity about when to choose this over data operations or file operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for running declared functions without UI, but doesn't explicitly state when to use this vs alternatives like data_add_item or open_app. No exclusions or prerequisites are mentioned, leaving the agent to infer context from the purpose statement alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

data_add_itemAdd itemB
Idempotent

Add an item to a collection. group is the app-defined lane/section (e.g. a kanban column); fields is a JSON object (e.g. {title, done, notes…}).

ParametersJSON Schema
NameRequiredDescriptionDefault
actorNo
groupNo
fieldsYes
positionNo
collectionYes
command_idYesidempotency key — generate a fresh uuid per action

Output Schema

ParametersJSON Schema
NameRequiredDescription
idNo
okYes
eotNo
seqNo
itemNo
noteNo
reasonNo
deletedNo
collectionNo
violationsNo
expected_versionNo
prev_collection_seqNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=true, and destructiveHint=false, so safety and retry behavior are covered structurally. The description adds useful semantics for `group` (an app-defined lane/section) and `fields` (a JSON object shape), but says nothing about persistence, whether the write is immediately visible, or what happens on duplicate calls beyond what the schema's command_id note already states.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core action and then the two most ambiguous parameters. Efficient, though the parameter glosses read more like schema notes than a usage-oriented description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. But for a 6-parameter mutation tool with nested object fields and low schema coverage, the description omits the relationship to apply_data_writes and the meaning of `position`, leaving the agent to infer ordering and persistence behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 17% across 6 parameters, so the description carries real weight, and it does explain two key params: `group` as an app-defined lane and `fields` as a JSON object with an example. It leaves `actor`, `position`, and `collection` undefined beyond their names, so it only partially compensates for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Add) and resource (item to a collection), which cleanly separates it from data_update_item, data_move_item, and data_delete_item in the sibling list. It does not, however, explicitly name those siblings or contrast itself with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus data_update_item, data_move_item, or the two-phase apply_data_writes sibling. The presence of apply_data_writes in the sibling set suggests staged writes, but the description never explains whether this call persists immediately or requires a follow-up.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

data_changesWhat changed since you last lookedA
Read-onlyIdempotent

What happened in a collection after a ledger position YOU hold — including edits the USER made in the widget, which never pass through you. Any mark you already have works as since: a write ack's seq, data_list's version, or the last call's next_since. Returns the contiguous run of events right after it (oldest first, whole events: actor, item id, fields) plus next_since to continue; omit since to just learn the current position.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNomax events to return (default 50, max 500)
sinceNoyour held mark — events strictly after this seq; omit to just learn latest_seq
collectionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
eotNo
noteNo
sinceNo
totalNo
eventsYes
droppedNo
returnedNo
collectionYes
latest_seqYes
next_sinceYes

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), and the description adds genuinely non-redundant behavior: events form a contiguous run, are returned oldest-first as whole events (actor, item id, fields), include user edits made in the widget that never pass through the agent, and carry `next_since` for continuation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, and the paragraph packs purpose, mark sources, return shape, and continuation without filler. It is dense enough that the continuation detail sits mid-sentence rather than being broken out, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema and rich annotations present, the description only needs to add mechanism, and it does: mark provenance, ordering, event contents, and pagination continuation. It omits failure/edge behavior (e.g. what happens if `since` is older than retained history), which is the one gap for a paging read tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67% (`collection` is undocumented in the schema), and the description compensates for the most ambiguous parameter by explaining that `since` accepts several mark formats and what omitting it means. `limit` adds nothing beyond the schema, and `collection` is still only implied.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+scope: report what changed in a collection after a ledger position the caller holds, including widget-side edits. It distinguishes itself from generic reads by naming the ledger-position mechanism, but it never names the closest sibling (get_data_version), which is the tool an agent would most likely confuse it with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete triggering context: pass any mark you hold (write ack's `seq`, data_list's `version`, last call's `next_since`) as `since`, and omit `since` to just learn the current position. It does not state when to prefer get_data_version instead, so one routing question is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

data_delete_itemDelete itemB
DestructiveIdempotent

Delete an item permanently. May return reason:"confirmation_required" with a request_state — show the user what is named in note, then re-send the same call with request_state attached.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
actorNo
command_idYesidempotency key — generate a fresh uuid per action
request_stateNo
expected_versionNo
require_confirmationNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
idNo
okYes
eotNo
seqNo
itemNo
noteNo
reasonNo
deletedNo
previewNo
collectionNo
expires_atNo
violationsNo
request_stateNo
expected_versionNo
prev_collection_seqNo

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=true, so the safety profile is covered. The description adds real value beyond that by disclosing the confirmation_required response case and the request_state re-send pattern, which annotations cannot express. It still omits whether confirmation is mandatory or gated by require_confirmation/actor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action and then the edge-case flow. Efficient overall, though the inline quoting of reason:"confirmation_required" is a bit awkward and the field name `note` is used without definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, 6-parameter tool with only 17% schema coverage and a required idempotency key and expected_version, the description is too thin. The output schema removes the need to explain returns, but the parameter surface and the confirmation gating remain undocumented.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 17% across 6 parameters, so the description carries the burden of compensating and largely fails. It mentions request_state only in the context of the re-send flow, and never explains id, actor, expected_version, or require_confirmation; 'note' is referenced but is not even a declared parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Delete an item permanently'), and the 'permanently' qualifier signals irreversibility. It does not explicitly differentiate itself from siblings like file_delete or delete_app, but the 'item' resource distinguishing it from those is reasonably inferable from the name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides useful procedural guidance for the confirmation_required flow (show the note, re-send with request_state), which is genuinely actionable. However, it says nothing about when to prefer this over data_update_item or apply_data_writes, and gives no prerequisites or conditions for invoking it in the first place.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

data_listList collection dataA
Read-onlyIdempotent

Read items in full — every field, plus the item id you need to update or delete it. No UI (use open_app for that). Returns a PAGE: up to limit (default 100) matching items, with total and — when more exist — next_cursor; returned/total make a short delivery self-evident. match filters: a bare value means equals; an object is operators {ne, lt, lte, gt, gte, contains, prefix, exists} (numeric filters compare numerically, strings lexicographically — ISO dates work). Paging is a live keyset walk; items moved mid-page can be skipped or repeated.

ParametersJSON Schema
NameRequiredDescriptionDefault
groupNoonly items in this group/lane
limitNopage size (1-500, default 100)
matchNofilter on top-level fields, e.g. {"done": false} or {"amount": {"gte": 100}}
cursorNoopaque cursor from the previous page's next_cursor
collectionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
okNo
eotNo
noteNo
groupNo
itemsYes
totalNohow many items exist (the collection/group, pre-filter) — compare it with the rows you actually received
versionYes
returnedNo
collectionYes
next_cursorNopass back as `cursor` for the next page; null = no more
files_versionNo
settings_versionNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already carry the safety profile (readOnly, idempotent, non-destructive), so the bar is lower. The description adds genuinely new behavior: keyset pagination is a live walk where items moved mid-page can be skipped or repeated, and it warns that returned vs total can reveal short deliveries.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose and the open_app exclusion, then layers pagination and filter semantics efficiently. It is dense and slightly packed with parenthetical asides, but every sentence carries information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given five parameters, nested match objects, and a live-keyset paging model, the description covers what to expect from a call including page shape and cursor chaining. An output schema exists, so it is not obligated to detail return values further.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 80%, putting the baseline at 3, and the description exceeds it by defining `match` semantics the schema only illustrates: a bare value means equals, object values are operators {ne, lt, lte, gt, gte, contains, prefix, exists}, and numeric vs lexicographic (ISO date) comparison rules.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Read items in full") plus scope (every field, plus the item id). It distinguishes itself from open_app explicitly, but does not name the closest data siblings (data_changes, data_update_item), so an agent still has to infer boundaries from naming alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes UI needs to open_app, which is a real exclusion. However it gives no conditions for choosing between this and sibling data tools like data_changes or list_data_collections, so the when-to-use is clear only for the UI alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

data_move_itemMove itemC
Idempotent

Move an item to another group and/or position (e.g. kanban column).

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
actorNo
groupNo
positionNo
command_idYesidempotency key — generate a fresh uuid per action
expected_versionNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
idNo
okYes
eotNo
seqNo
itemNo
noteNo
reasonNo
deletedNo
collectionNo
violationsNo
expected_versionNo
prev_collection_seqNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare non-destructive, idempotent, non-open-world behavior, so the safety profile is covered. The description adds the domain notion of a group/kanban column, but omits concurrency behavior (expected_version), what happens to surrounding items' positions, and whether group can be cleared — context the bare annotations do not supply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the example is useful rather than wasteful. It is efficient, though arguably too terse for a six-parameter mutation tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with low schema coverage and an output schema (so returns need not be explained), the description leaves significant gaps: optimistic locking via expected_version, the actor parameter, and idempotency rationale are all unaddressed. It is not sufficient for confident, correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 17% (just command_id), so the description must compensate. It loosely characterizes group and position (with the kanban example) but says nothing about actor, expected_version (optimistic-concurrency semantics), or id, leaving four of six parameters undocumented in both description and schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ("Move an item") and the target dimensions ("another group and/or position"), including a concrete example ("kanban column"). It is clear what the tool does, though it does not distinguish itself from the sibling data_update_item, which plausibly could also change group/position.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus alternatives like data_update_item or apply_data_writes, nor any prerequisite or context. Usage must be inferred entirely from the name and the parenthetical example.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

data_update_itemUpdate item fieldsA
DestructiveIdempotent

Shallow-merge fields into an item (set a key to null to remove it). Uses optimistic concurrency.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
actorNo
fieldsYes
command_idYesidempotency key — generate a fresh uuid per action
expected_versionNothe item version you last saw

Output Schema

ParametersJSON Schema
NameRequiredDescription
idNo
okYes
eotNo
seqNo
itemNo
noteNo
reasonNo
deletedNo
collectionNo
violationsNo
expected_versionNo
prev_collection_seqNo

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint and idempotentHint, but the description adds genuinely new behavior: shallow-merge semantics, null-to-remove deletion, and optimistic concurrency. It stops short of saying what happens on a version mismatch (error vs retry) or how command_id relates to idempotency, which would complete the picture.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loading the merge behavior and deletion rule before the concurrency note. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. However, for a destructive mutation with a required idempotency key and version parameter, the description omits conflict/error behavior and does not compensate for the 40% schema coverage on id/actor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 40%; command_id and expected_version are described in the schema, but id and actor are undocumented anywhere. The description does add real semantics for the fields parameter (null removes a key), partially compensating, but does not explain expected_version's role in the optimistic-concurrency flow.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('shallow-merge fields into an item') and clarifies scope with the null-removal rule. It distinguishes itself from data_add_item/data_delete_item by implication, but never names a sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus the closely related siblings apply_data_writes or data_changes, which also mutate data. The concurrency and idempotency mechanics are mentioned but no context is given for choosing this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_appDelete appA
DestructiveIdempotent

Delete an app from the registry. Default data:"keep" is a tombstone: its data, files and history are KEPT and restore_app can bring the app back. data:"cascade" ALSO permanently deletes the data provably only this app used (plus its settings) — it always returns a disposition plan first: read the plan to the user (what will be deleted AND what will be kept and why), then re-send the same call with request_state. Cascade is NOT undoable. Shared or unprovable collections are always kept.

ParametersJSON Schema
NameRequiredDescriptionDefault
dataNo"keep" (default): tombstone, data survives. "cascade": delete the app AND the data only it used — two-step confirmed, permanent
nameYes
actorNo
command_idYesidempotency key — generate a fresh uuid per action
request_stateNothe state from a prior confirmation_required answer — re-send the identical call with it attached

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint and idempotentHint, but the description goes well beyond them: what is deleted vs. kept, that shared/unprovable collections are always kept, that cascade is not undoable, and that it requires a two-step confirmation. This is exactly the behavioral context a destructive tool needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the mode distinction and the irreversibility warning, and every sentence carries load. Slightly heavy: the data mode semantics are restated in both the description and the schema's data field, creating minor redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, yet the description tells the agent how to interpret the response (a disposition plan listing what will be deleted and what will be kept and why) and what follow-up call to make. Nothing needed to call this destructive tool safely is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 60% schema coverage, the description usefully expands the data enum (tombstone vs. permanent two-step delete) and explains the request_state round-trip. The actor parameter is left undocumented in both description and schema, and command_id semantics live only in the schema, so it doesn't fully compensate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb and resource (delete an app from the registry) and immediately distinguishes the two deletion modes by name. It also names the sibling restore_app as the counterpart, so an agent can place it in the toolset without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states the default (data:"keep" is a tombstone) and when to use the alternative (data:"cascade"). It also prescribes the exact workflow: read the returned disposition plan to the user, then re-send the identical call with request_state — which is precisely the when/how guidance an agent needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit_appEdit app sourceA
Idempotent

Surgical edits to an app WITHOUT round-tripping the whole source. Two edit forms, mixable: RANGE {offset, length, expect_hash, new_string} replaces a span you read with get_app (cheapest — echo the window's offset/returned/hash, no anchor text travels); STRING {old_string, new_string} replaces an exact-once match (or set replace_all). Range offsets always address the expected_version document and must not overlap; string edits apply after ranges, in order. All edits apply together, or nothing applies. The #oma-manifest block is re-read on save.

ParametersJSON Schema
NameRequiredDescriptionDefault
appYesapp name (this tool says `app`; save_app and get_app say `name`)
actorNo
editsYeseach item is RANGE (offset+length+expect_hash) or STRING (old_string)
command_idYesidempotency key — generate a fresh uuid per action
expected_versionYesREQUIRED — the version the edits were authored against (from get_app)

Output Schema

ParametersJSON Schema
NameRequiredDescription
okYes
eotNo
nameNo
noteNo
sizeNo
reasonNo
appliedNo
createdNo
versionNo
prev_sizeNo
manifest_actionNo
expected_versionNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: discloses all-or-nothing atomicity ('All edits apply together, or nothing applies'), ordering rules (string edits after ranges, in order), non-overlap constraint on ranges, version anchoring to expected_version, and that the #oma-manifest block is re-read on save. This is exactly the extra context annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core value proposition, then tight clauses covering each form and the atomicity rule. Dense but every sentence carries information; the parenthetical asides are purposeful rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need no explanation, and the description covers everything else an agent needs: form selection, ordering, atomicity, version binding, and the manifest side effect. Nothing material is left to inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Description supplies real meaning beyond the 80%-covered schema: it explains that RANGE fields echo get_app's window output (no anchor text travels), that old_string must match exactly once unless replace_all is set, and that offsets target the expected_version document. command_id idempotency is left to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource and a differentiating constraint: 'Surgical edits to an app WITHOUT round-tripping the whole source.' An agent immediately understands this is a partial-edit tool that stands apart from full-source siblings like save_app.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly states which of the two edit forms to choose (RANGE is 'cheapest', STRING for exact-once matches) and how they compose. It references get_app as the source of offset/hash, but never explicitly says when to prefer this tool over save_app for a given workflow.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

file_deleteDelete an app fileC
DestructiveIdempotent

Permanently delete one file an app has stored.

ParametersJSON Schema
NameRequiredDescriptionDefault
appYes
pathYes
actorNo
command_idYesidempotency key (uuid)
request_stateNo
expected_versionNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
okNo
appYes
noteNo
pathYes
reasonNo
deletedNo
previewNo
expires_atNo
files_versionNo
request_stateNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true and readOnlyHint=false, so the safety profile is covered. The description adds 'Permanently', reinforcing irreversibility beyond the hint, but says nothing about required actor identity, the concurrency role of expected_version, or what happens to the app if the path is wrong.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the key word 'Permanently' front-loaded and zero filler. It is appropriately sized, though borderline under-specified rather than merely concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, but for a destructive, idempotency-keyed mutation with 6 parameters and 17% schema coverage, the description omits the concurrency and identity semantics an agent needs to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 17% (just command_id), and the description adds no meaning for app, path, actor, request_state or expected_version. With 6 params and 5 undocumented in both places, the description fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (delete), resource (a file), scope (one file, app-stored), and adds 'Permanently'. An agent can distinguish it from delete_app and data_delete_item by the file/app scoping, though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of the adjacent siblings an agent might confuse it with (file_write/file_read for the same files, delete_app for the whole app). Usage must be inferred from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

file_listList app filesA
Read-onlyIdempotent

List the files an app (app) has stored — a PAGE of {path, size, mime, version} plus usage totals; limit/cursor page through, prefix narrows. These are opaque user files (attachments, exports) the app keeps — separate from its structured data collection. Renders no UI.

ParametersJSON Schema
NameRequiredDescriptionDefault
appYesthe app whose files to list
limitNopage size (default 200)
cursorNoopaque cursor from the previous page's next_cursor
prefixNoonly paths starting with this prefix

Output Schema

ParametersJSON Schema
NameRequiredDescription
appYes
eotNo
noteNo
filesYes
totalNo
usageYes
returnedNo
next_cursorNo
files_versionYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this a read-only, idempotent, non-destructive, closed-world operation, so the safety profile is covered. The description adds genuinely new behavioral context: results are paginated ('a PAGE of...'), outputs include usage totals, files are opaque, and it 'renders no UI.' That is value beyond the annotations, though nothing is said about limits or failures.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and then progressively adds scope, page shape, and the data_list distinction — each clause earns its place. Some density (the parenthetical '(app)', the stacked clauses) leaves minor room for tightening, so not a full 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present the description needn't explain returns, yet it still outlines the page contents; annotations cover safety and idempotency; all four parameters are covered at 100% schema coverage plus description-level paging semantics. Nothing an agent needs to invoke this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds how the parameters interact: 'limit/cursor page through, prefix narrows.' This explains the paging contract and prefix filtering behavior beyond the per-parameter schema text, lifting it above the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('List the files an app has stored') and even sketches the returned page shape. It explicitly distinguishes this from the sibling data_list by calling out that these are user files 'separate from its structured data collection.' An agent can tell it apart from data_list and file_read without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear context for when this is the right tool: enumerating opaque user files (attachments, exports) rather than structured data, with paging via limit/cursor and narrowing via prefix. What is missing is an explicit contrast with file_read or data_list by name, so it stops short of full when/when-not/alternatives guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

file_readRead an app fileA
Read-onlyIdempotent

Read one file an app has stored, as a WINDOW of its bytes: offset/length select it, data_base64 carries exactly that window, next_offset continues (same window grammar as get_app, and for the same reason). Reassemble by concatenating decoded windows; sha256 is the WHOLE file's hash, so reassembly is checkable.

ParametersJSON Schema
NameRequiredDescriptionDefault
appYes
pathYesthe file's logical name
lengthNomax bytes for this window (default fits the result budget)
offsetNobyte offset to read from (default 0)

Output Schema

ParametersJSON Schema
NameRequiredDescription
appYes
eotNo
mimeYes
pathYes
sizeYes
totalNo
offsetNo
sha256Yes
versionYes
returnedNo
data_base64No
next_offsetNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so safety is covered. The description adds non-obvious behavior the annotations cannot convey: reads are windows rather than whole-file, the default length fits a result budget (i.e. results may be truncated), and sha256 covers the whole file so reassembly is verifiable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence with the core action front-loaded and the reassembly/verification contract at the end. Every clause earns its place, though the parenthetical asides make it slightly harder to parse than necessary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be spelled out, yet the description still explains the window/next_offset/sha256 contract that an agent must understand to read a file correctly. Nothing needed to invoke or complete a chunked read is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75% and the schema documents offset/length defaults, but the description adds meaning beyond it: it explains the relationship between offset/length and the returned window, and names the response fields (data_base64, next_offset) that drive continued reads. This goes beyond the schema's per-parameter text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read one file an app has stored') with clear scope ('one file', 'an app has stored'), which separates it from file_list, file_write and get_app. It even cross-references the sibling whose window grammar it reuses, so an agent can place it without opening other schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage through the windowing model (call with offset/length, follow next_offset until done), but never states when to choose this tool over alternatives such as file_list or get_app, nor any exclusions. Usage is inferable rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

file_writeStore an app fileA
DestructiveIdempotent

Store a file for an app (create or overwrite by path). data_base64 carries the file bytes, base64-encoded. Overwriting an existing path bumps its version. Single-call writes are limited to a few MiB. Files persist and are the app's own, reusable across chats.

ParametersJSON Schema
NameRequiredDescriptionDefault
appYesthe app this file belongs to
mimeNocontent type, e.g. 'image/png' (default application/octet-stream)
pathYeslogical file name, e.g. 'receipt.pdf' or 'exports/2026-q1.csv'
command_idYesidempotency key — a fresh uuid per write
data_base64Yesfile bytes, base64-encoded
expected_versionNothe version you last saw, for optimistic concurrency (optional)

Output Schema

ParametersJSON Schema
NameRequiredDescription
appYes
mimeYes
pathYes
sizeYes
sha256Yes
versionYes
files_versionYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover safety (destructive=true, idempotent=true), and the description meaningfully adds version-bump-on-overwrite semantics, the per-call size ceiling, and that files persist and are reusable across chats. It stops short of describing what exactly an overwrite destroys or any auth prerequisites, but it goes beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four compact sentences, purpose front-loaded first, then encoding, versioning, and limits/persistence. No filler and every sentence carries a distinct fact.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, and the description covers purpose, encoding, versioning, size, and persistence. It lacks explicit pointer to the chunked-write fallback, which is the only real gap for a 6-param tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description clarifies that data_base64 carries base64-encoded bytes and that overwriting bumps the version (relevant to expected_version), but most parameter meaning still lives in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource with scope: "Store a file for an app (create or overwrite by path)." An agent immediately knows this is a single-call write. It does not name the sibling chunked-write tools (file_write_begin/chunk/commit), so differentiation is only implicit via the size limit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Single-call writes are limited to a few MiB" implies that larger files need the chunked siblings, but this is never stated as a routing rule and the alternatives are not named. No explicit when-to-use vs when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

file_write_beginBegin a chunked file writeA

Start a chunked upload for a file too big for file_write's single call. Returns an upload_id; send the bytes in order with file_write_chunk (each chunk up to ~5 MiB of raw bytes), then file_write_commit names the file. Uploads expire after 30 idle minutes; per-file ceiling 250 MiB.

ParametersJSON Schema
NameRequiredDescriptionDefault
appYesthe app this file will belong to

Output Schema

ParametersJSON Schema
NameRequiredDescription
upload_idYes
file_limit_bytesYes
chunk_limit_bytesYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover the safety profile (not read-only, not idempotent, not destructive). The description adds operational context the annotations cannot: the returned upload_id, the ~5 MiB chunk size, the 30-idle-minute expiry, and the 250 MiB per-file ceiling — exactly the lifecycle facts an agent needs to avoid a failed upload.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: the trigger, the call sequence with return value, and the limits. The scoping condition is front-loaded before the workflow details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema already defines the return shape, the description need not restate it — and it doesn't, instead covering the multi-call lifecycle, chunk sizing, and expiry/ceiling limits. Nothing an agent needs to complete a chunked upload correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single parameter with 100% schema description coverage ('the app this file will belong to'), so the schema already carries the parameter meaning. The description adds no parameter-level detail, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('start a chunked upload for a file') and explicitly scopes it against the sibling file_write ('too big for file_write's single call'). An agent can distinguish this from file_write, file_write_chunk, and file_write_commit without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the exact condition that selects this tool over file_write (file too large for a single call) and lays out the full workflow sequence with file_write_chunk then file_write_commit. Usage is unambiguous rather than inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

file_write_chunkAppend a chunk to an uploadA

Append the next chunk of bytes (base64) to an upload started with file_write_begin. Send chunks strictly in order, one at a time. Pass seq (0-based chunk index) so a resend after a lost response is acknowledged instead of double-appended.

ParametersJSON Schema
NameRequiredDescriptionDefault
seqNo0-based index of this chunk — a resend of an already-staged index is a safe no-op
upload_idYes
data_base64Yesthis chunk's bytes, base64-encoded

Output Schema

ParametersJSON Schema
NameRequiredDescription
bytesYestotal bytes staged so far
chunksNo
duplicateNo
upload_idYes

TDQS

A4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description usefully discloses ordering and resend semantics (same `seq` is acknowledged rather than double-appended), but this directly conflicts with the annotation idempotentHint=false: an identical re-call with the same seq and data is described as a safe no-op, which is idempotent behavior. Aside from the contradiction, no chunk-size limits or error behavior on out-of-order chunks are given.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences: what it does, how to sequence calls, and why `seq` matters. Front-loaded and every sentence carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and the description covers the critical ordering/resend contract for a stateful streaming write. It could still state a maximum chunk size or what happens if `seq` arrives out of order.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%; `seq` and `data_base64` are already documented in the schema, but the description adds the origin of `upload_id` (from file_write_begin) and reinforces the 0-based ordering contract, adding meaning beyond the raw fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Append the next chunk of bytes ... to an upload') and ties itself explicitly to the sibling that must precede it ('started with file_write_begin'), so an agent can distinguish it from file_write_begin/file_write_commit without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context: it is used only on an upload begun by file_write_begin, chunks must be sent strictly in order and one at a time. It stops short of naming file_write_commit as the terminating step, so the full lifecycle isn't spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

file_write_commitCommit a chunked file writeA
DestructiveIdempotent

Finalize an upload as an app file (create or overwrite by path) — the chunked equivalent of file_write. The upload is consumed either way; on failure, restart from file_write_begin.

ParametersJSON Schema
NameRequiredDescriptionDefault
mimeNo
pathYeslogical file name, e.g. 'video.mp4' or 'exports/backup.zip'
upload_idYes
command_idNoidempotency key (uuid); auto-generated if omitted. A retried commit with the same id returns the original receipt instead of demanding a re-upload
expected_versionNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
appYes
mimeYes
pathYes
sizeYes
sha256Yes
versionYes
files_versionYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructive, idempotent, and non-read-only, so the safety profile is covered. The description adds genuinely non-obvious behavior: the upload is consumed either way, and a failure forces you back to file_write_begin. It does not mention auth requirements or what the receipt contains, but the consumption/recovery detail is real added value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with what the tool does, followed by the recovery caveat. Every clause carries information; nothing is padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need no explanation, and the failure/consumption behavior is well disclosed. However, for a 5-parameter mutation tool at 40% schema coverage, the description leaves mime and especially expected_version unexplained, so an agent cannot fully reason about the call from the definition alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 40% (mime, upload_id, and expected_version are undocumented), so the description carries the burden of compensating — and it does not. It gestures at 'path' and 'an upload' implicitly but never names or explains upload_id, mime, or expected_version, the last of which implies optimistic concurrency that an agent must understand.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Finalize an upload as an app file') plus the create-or-overwrite semantics, and explicitly differentiates itself from the sibling file_write by calling itself 'the chunked equivalent of file_write'. An agent can place it in the write flow without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It positions itself relative to file_write (the non-chunked alternative) and names file_write_begin as the restart path on failure. It stops short of explicitly stating the trigger condition ('call after file_write_chunk completes') or when_not_to_use, so it is clear context rather than full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_appGet app sourceA
Read-onlyIdempotent

Read an app's ui source as a WINDOW — offset/length select it, next_offset continues, total is the full length. Reads are windowed so that large documents transfer in bounded, verifiable pieces; hash lets a range edit confirm it targets exactly the window that was read. Carries version — the expected_version for edit_app / save_app — and hash, the expect_hash for a range edit of exactly this window. node jumps the window to the element marked data-oma-node="". slot:"manifest" returns the declaration object instead (no window mechanics).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
nodeNoread the element marked data-oma-node="<node>" — the window (and its hash) covers exactly that element (ui slot only)
slotNodefault ui; manifest returns {manifest: object|null} whole
lengthNomax characters for this window (default fits the result budget; ui slot only)
offsetNocharacter offset to read from (default 0; ui slot only)

Output Schema

ParametersJSON Schema
NameRequiredDescription
eotNo
hashNo
nameYes
textNo
totalNo
offsetNo
versionYes
manifestNo
returnedNo
next_offsetNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so the bar is lower, and the description still adds real context: reads are windowed and verifiable, hash/version are returned to be fed to edit_app/save_app (expected_version, expect_hash). It stops short of stating limits on window size beyond the default budget or any failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The first sentence front-loads the window model well, but the middle portion is dense and repetitive: hash is explained twice ('hash lets a range edit confirm it targets exactly the window that was read' vs 'hash, the expect_hash for a range edit of exactly this window'). Several clauses could be tightened without losing meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers the windowing model, the node jump, the manifest slot, and how version/hash feed downstream edits. Only minor gaps remain, such as whether windows can be requested concurrently or any error behavior for out-of-range offsets.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 80%, so the schema already documents name, node, slot, length and offset; the description mostly restates node/slot semantics and explains the returned window fields (next_offset, total, hash) rather than adding parameter syntax. With high schema coverage this is the expected baseline of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Read an app's ui source') and immediately qualifies the retrieval model ('as a WINDOW'), which distinguishes it from the sibling get_app_html (rendered html) and get_app_guide (documentation). The slot enum and node behavior are named, so an agent can tell what this tool returns without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear usage context: windowed reads are for transferring large documents in bounded pieces, node jumps the window to a marked element, and slot:"manifest" switches to declaration reading. It does not explicitly say when to prefer get_app_html or get_app_guide instead, so there are no stated exclusions or named alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_app_guideApp authoring guideA
Read-onlyIdempotent

The authoring contract for creating or editing an app: the window.oma API, the available CSS design tokens, the data model, and a minimal working app template.

ParametersJSON Schema
NameRequiredDescriptionDefault
topicNowhich chapter (default basics: the contract + template). Each chapter stands alone

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered without the description. The description adds that this is a static written contract with fixed chapters ('each chapter stands alone'), which is useful context about the return content, but says nothing about size or auth.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that lists the guide's contents with zero filler. It is appropriately sized for a one-line documentation lookup, though the enumeration is dense and could have flagged the chapter default.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, one-optional-param documentation fetch with annotations covering safety and no output schema, the description adequately tells the agent what it will receive. The only minor gap is that the optional 'topic' chapters are not mentioned, but the schema carries them.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the enum values and the default behavior are already fully documented in the schema. The description does not touch the 'topic' parameter at all, so it adds no meaning beyond what the schema provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource (the authoring contract) and enumerates exactly what it contains: the window.oma API, CSS design tokens, the data model, and a working template. It is clearly a documentation retrieval tool, distinguishable from mutation siblings like save_app or edit_app, though it never names or contrasts a sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for creating or editing an app' implies the context in which to reach for this guide, and the schema note 'default basics: the contract + template' hints at entry points. But there is no explicit when-to-use vs an alternative like get_app_html, and no stated preconditions, so usage is only loosely implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_app_htmlApp HTML (internal)A
Idempotent

Internal: returns raw app HTML plus its trust tier and capability grants for the universal loader widget. Not useful to call directly — use get_app to read source.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
mountNothe caller is mounting this app on screen now (records the live pointer); refetches and framing bricks leave it unset

Output Schema

ParametersJSON Schema
NameRequiredDescription
cspNo
capsYes
hostNolabel for the client this widget is running under; usually empty
htmlYes
nameYes
tierYes
authorYes
lockedYesa fixed system app (settings renders these read-only)
versionYes
collectionYesthe collection this app opens on
declarationNo

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=true and destructiveHint=false, so the safety profile is partially covered. The description discloses the returned payload (trust tier, capability grants) but does not explain why a getter is not read-only (the live-pointer side effect noted in the mount param) or any auth requirements, leaving that inference to the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, no waste, with the internal-only scoping and the redirect to get_app front-loaded so the agent sees the routing advice immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be enumerated; the description still summarizes the payload (HTML, trust tier, capability grants). The remaining gap is the undocumented required 'name' parameter, which is not covered anywhere.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%: 'mount' is documented in the schema but 'name' (the sole required parameter) is undocumented in both places. The description does not compensate by clarifying what 'name' refers to or the effect of the mount flag on returned HTML.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('returns raw app HTML plus its trust tier and capability grants') and explicitly scopes it to the universal loader widget. It also names the sibling it is not ('use get_app to read source'), so an agent can route correctly without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when-not to call it ('Not useful to call directly') and names the alternative (get_app) for the normal read path. This is exactly the routing guidance an internal tool needs to prevent misuse.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_data_versionData change probeA
Read-onlyIdempotent

The cheapest possible change check: returns the store's global change counter (seq) plus settings/files sub-counters. An unchanged counter means nothing in the store changed. Widgets read it for adaptive polling.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
seqYes
files_versionYes
schema_versionYes
settings_versionYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and idempotentHint, so safety needs no restating. The description adds genuine value beyond them: precisely what is returned (global seq plus settings/files sub-counters), the cost claim, and the interpretation rule that an unchanged counter means no change occurred.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the cost/semantics claim. The trailing 'Widgets read it for adaptive polling' is mildly consumer-oriented rather than agent-oriented, but it does convey the intended use pattern without bloat.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be documented, yet the description still summarizes the return shape and interpretation. With zero params and full annotation coverage of the safety profile, nothing an agent needs to invoke this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4; the description correctly spends no words on parameter syntax. There is nothing for it to compensate for.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: it returns the store's global change counter (seq) plus settings/files sub-counters. The phrase 'the cheapest possible change check' implicitly distinguishes it from heavier change-listing siblings like data_changes, but it never names that alternative, so sibling differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: 'An unchanged counter means nothing in the store changed' and 'Widgets read it for adaptive polling' suggest polling-before-expensive-work, but there is no explicit when-to-use versus data_changes or any exclusion. Adequate context, no routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_ui_preference_schemaShared preference catalogA
Read-onlyIdempotent

The engine-owned catalog of SHARED preferences (key, type, label, default, options) that the settings app renders. Apps read effective values via oma.pref(); this tool only describes what exists. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
sharedYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/idempotentHint/destructiveHint, so safety is covered. The description adds that the catalog is engine-owned and that the tool only describes existence rather than returning effective values, which is genuine behavioral context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short clauses front-load the resource and its fields, then immediately disambiguate from oma.pref(). No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Zero-param read-only catalog with an output schema present, so return shape is already covered. The description supplies the one thing an agent needs to not confuse this with value retrieval, plus the field list.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. No parameter description is needed and none is misleading.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (engine-owned catalog of SHARED preferences) with the exact fields it returns (key, type, label, default, options). It also contrasts itself with the sibling-ish mechanism oma.pref(), making clear this tool describes rather than reads values.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Apps read effective values via oma.pref(); this tool only describes what exists,' which tells the agent this is not the value-fetch path. It doesn't enumerate explicit when-not-to-use cases beyond that, but the routing cue to oma.pref() is clear context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

install_from_app_storeInstall an App Store appA
Idempotent

Install (or update) a ready-made app from the built-in App Store into the user's registry. Once installed it behaves like any app of the user's own; its UI updates come from the App Store (newer store version), not from edits. Use app_store_list to see what's available.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesApp Store entry name (see app_store_list)
command_idNoidempotency key (uuid); auto-generated if omitted

Output Schema

ParametersJSON Schema
NameRequiredDescription
nameYes
tierYes
outcomeYesinstalled = first install; updated = the store had a newer version; current = already installed and unchanged
updatedNo
versionYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (idempotentHint=true, destructiveHint=false, openWorldHint=false), so the bar is lower. The description adds genuine post-call behavior the annotations cannot convey: the installed app behaves like the user's own but receives UI updates from the store rather than local edits. It does not mention permissions or failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the action and scope, then the post-install behavior, then the routing hint. Every sentence adds information an agent would otherwise lack.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be described, and annotations cover safety. The description covers source, effect, and update semantics, though it omits prerequisites such as whether the user needs prior store access or how conflicts with an existing app are handled.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both name and command_id are already documented in the schema. The description only points to app_store_list for valid names and adds no format or idempotency-key guidance beyond what the schema states, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Install/update), the resource (a ready-made app from the built-in App Store), and the destination (the user's registry). It also distinguishes itself from the sibling app_store_list by naming it as the discovery tool, so an agent can tell the browse and install operations apart.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent to app_store_list to find what's installable, which is the main pre-call step. It implies but does not state when to prefer update over install or any exclusion conditions, so it stops short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_app_checkpointsApp history (checkpoints)A
Read-onlyIdempotent

List an app's checkpoints as {checkpoint, ts, ui_size} — metadata only, NEVER the source (keeps context small; use get_app for the current source). Checkpoint 1 is the oldest; restore_app takes that number. History survives delete_app (tombstone). Each checkpoint snapshots BOTH slots (ui + manifest); restore brings back the pair.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
nameYes
historyYes

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover safety (readOnly/idempotent/non-destructive); the description adds substantial non-obvious behavior: metadata-only returns that never include source, that history survives delete_app as a tombstone, and that each checkpoint snapshots BOTH ui and manifest slots which restore_app restores as a pair.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and return shape, then packs the scoping constraint, numbering, tombstone, and snapshot semantics into tight parenthetical clauses. Every sentence carries a distinct fact with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists so return-value explanation is optional, yet the description still names the fields. Combined with the tombstone and dual-slot snapshot notes, an agent has everything needed to call this and understand its relationship to restore_app and delete_app.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single 'name' parameter has 0% schema description coverage, and the description never explicitly says name is the app identifier it lists checkpoints for — it is only inferable from 'an app's checkpoints'. It does clarify the semantics of the checkpoint number returned (oldest = 1, consumed by restore_app), which is useful but pertains to output, not input.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('List an app's checkpoints') and even gives the return shape {checkpoint, ts, ui_size}. It explicitly distinguishes itself from sibling get_app ('use get_app for the current source'), so an agent can separate them without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It routes the agent to get_app for current source and names restore_app as the consumer of the checkpoint number, which implicitly defines when this tool is the right choice. It never states an explicit 'use this when' condition or a when-not, so it falls short of 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_appsList appsA
Read-onlyIdempotent

List UI apps in the registry (reusable across all chats). Lists the user's openable apps by default — pass name to look one up, or widen with kind/visibility.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNodefault app: what a person opens and reuses
nameNoexact app name — the fastest way to answer "open my X"
visibilityNodefault featured+listed; archived/unlisted are retired or long-tail

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, closed-world, so the safety profile is covered. The description adds the meaningful context that the registry is reusable across all chats and that the default scope is the user's openable apps. It doesn't discuss result limits or ordering.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences: purpose and cross-chat scope first, then the two usage modes. No filler, and the default-scope constraint is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Adequate for a zero-required-param read tool whose annotations and schema are rich. The default behavior and the three widening levers are all surfaced; only minor gaps like default result count remain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both enums are documented with defaults (kind default app, visibility default featured+listed). The description restates the lookup/widen behavior but adds no syntax or format detail beyond the schema, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (UI apps in the registry), and adds scope that distinguishes it from siblings: 'reusable across all chats.' Clearly separates it from get_app (single retrieve) and app_store_list (store catalog).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete routing: pass name to look one up, or widen with kind/visibility, and the schema notes 'exact app name — the fastest way to answer open my X.' It implies the default-openable-apps behavior but doesn't explicitly name when a sibling like get_app or app_store_list would be preferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_data_collectionsList collectionsA
Read-onlyIdempotent

List every data collection that exists (name, item count, last activity). Use when unsure where data lives, what boards the user has, or which collection to bind an app to. Renders no UI.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
noteNo
collectionsYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds 'Renders no UI', a genuine behavioral trait not expressed in the annotations, though it says nothing about pagination, permissions, or result limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, both earning their place: the first defines scope and return shape, the second gives usage triggers. The scope statement is front-loaded and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only listing tool with an output schema present, nothing an agent needs to invoke it correctly is missing. The description covers purpose, when to reach for it, and the no-UI behavior, and it need not explain return values since the output schema exists.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no parameter semantics to document and the baseline is 4. The description correctly adds nothing about inputs and instead describes outputs, which is appropriate here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List every data collection') and even enumerates the returned fields (name, item count, last activity), so the agent knows exactly what it gets. It never names or contrasts a sibling such as data_list or list_apps, so differentiation is left to inference from the domain rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The sentence 'Use when unsure where data lives, what boards the user has, or which collection to bind an app to' gives concrete triggering situations, which is stronger than implied usage. However it lists no alternatives or exclusions, so the agent must infer when a different listing tool is preferable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

open_appOpen any appA
Read-onlyIdempotent

Open ANY app from the registry by name. Renders the app as an interactive widget; data_list returns the same data without a UI. Works IMMEDIATELY for apps saved moments ago in this same chat (the dedicated open_ tools may take a while to appear).

ParametersJSON Schema
NameRequiredDescriptionDefault
appYesapp name in the registry (see list_apps)
collectionNodata collection to bind (default: the one the app declares, else its own name)

Output Schema

ParametersJSON Schema
NameRequiredDescription
okNo
appYes
eotNo
noteNo
itemsYes
totalNo
versionYes
returnedNo
collectionYes
files_versionNo
settings_versionNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only, idempotent, non-destructive behavior, and the description adds genuinely new context: rendering as an interactive widget versus data_list's plain data, and the timing nuance that it works immediately for freshly saved apps. It does not describe loading/failure behavior, but the added context is meaningful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with zero padding; the core purpose and the widget-vs-data_list distinction are front-loaded, and the timing caveat earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read tool with an output schema and full annotation coverage, the description supplies everything an agent needs: registry source, widget output, the data_list alternative, and the fresh-app timing edge case.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the schema already documents both 'app' and 'collection' including the collection default rule, so the description adds only 'by name' at the margin. Baseline 3 is appropriate when the schema carries the parameter burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Open ANY app from the registry by name') and distinguishes the tool from its near-siblings by naming the dedicated open_<name> tools and data_list. An agent can tell what it does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly indicates the context in which this tool wins (any registry app, including ones saved moments ago) and names data_list as the no-UI alternative. It stops short of saying when to prefer get_app or the dedicated open_<name> tools explicitly, so it is clear but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

preview_app_store_entryApp Store live-preview payload (internal)A
Read-onlyIdempotent

Internal: returns an App Store entry's ui and manifest plus its mock-data fixtures, so the App Store app can render a LIVE sandboxed preview card (srcdoc + stub oma + mock snapshot). Read-only; nothing is installed. Not useful to call directly.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesApp Store entry name (see app_store_list)

Output Schema

ParametersJSON Schema
NameRequiredDescription
htmlYesthe entry's ui slot (field named for its format — the preview iframe srcdocs it)
nameYes
titleYes
categoryYes
fixturesYesembedded mock snapshot items, or null if the entry ships none
manifestYes
descriptionYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, non-destructive, closed-world behavior, and the description reinforces this with 'nothing is installed' and 'sandboxed'. It also adds real behavioral context the annotations cannot convey: the payload is srcdoc + stub oma + mock snapshot. No rate limits or error behavior are described, but that is minor here.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One dense sentence plus two short qualifiers, with the important context ('Internal:', 'Read-only', 'Not useful to call directly') front-loaded and easy to scan. Slightly jargon-heavy (srcdoc, stub oma) but not padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no prose, and the single-parameter surface is fully covered by the schema. For an internal read-only tool the description supplies enough to understand what it does and to avoid misusing it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single parameter ('name') is already documented in the schema with a pointer to app_store_list; the description adds no syntax, format, or naming detail beyond that. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (returns) and enumerates the exact payload: an entry's ui, manifest, and mock-data fixtures, plus the end goal (render a live sandboxed preview card). It clearly is not install_from_app_store or get_app, though it never explicitly names a sibling to compare against.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear negative directive ('Not useful to call directly') and explains why (internal, consumed by the App Store app to render a preview), which effectively tells an agent to stay away. It does not, however, state the positive condition under which a direct call is warranted.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

promote_appPromote visual to appA
Idempotent

Upgrade a kind:"visual" app to a full app in ONE atomic step: the engine flips kind in the stored manifest, keeping every other declared key, and saves a new version (OCC-guarded, history kept). Already an app is a no-op; downgrades are refused — demoting is an author edit (save_app with the manifest), not a lifecycle verb.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesan existing app with kind "visual" (list_apps {kind:"visual"} shows them)
command_idNoidempotency key (uuid); auto-generated if omitted

Output Schema

ParametersJSON Schema
NameRequiredDescription
okYes
eotNo
wasNo
kindNo
nameNo
noteNo
reasonNo
versionNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: it discloses that the operation is atomic, that it preserves every other manifest key, that it creates a new version, and that the write is OCC-guarded with history retained. This is rich behavioral context an agent cannot get from idempotentHint/destructiveHint alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and packs the preconditions, no-op case, and refusal case into a compact span with no filler. The dense jargon-laden clause structure is slightly heavy, but every clause carries meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the description fully covers the state change, side effects, and edge cases an agent needs. Complete for this lifecycle verb.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented in the schema (name must reference an existing visual app, command_id is an idempotency key). The description adds no parameter-level detail beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('upgrade a kind:"visual" app to a full app') with a concrete mechanism ('flips `kind` in the stored manifest'). It clearly distinguishes itself from siblings like save_app, edit_app, and list_apps by naming the exact state transition it performs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly covers when it applies (visual app), when it is a no-op ('Already an app is a no-op'), when it refuses ('downgrades are refused'), and names the correct alternative for the reverse case (save_app with the manifest). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

restore_appRestore an app versionA
Idempotent

Roll an app back to one of its earlier checkpoints: re-saves that checkpoint's HTML as a NEW current one (nothing is lost — history is preserved and you can roll forward again). Use when a newer edit broke the UI. Get the checkpoint number from list_app_checkpoints; after restoring, open_app to view it.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
checkpointYeswhich checkpoint to restore, 1 = the oldest (see list_app_checkpoints)
command_idNoidempotency key (uuid); auto-generated if omitted

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds real context beyond annotations: the operation is non-destructive because history is preserved and the agent can roll forward again, and the restore creates a new current version rather than mutating history. Annotations already declare destructiveHint=false and idempotentHint=true, so the description's main added value is explaining the mechanism and reversibility, though it doesn't mention permissions or effects on associated app data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences: purpose/mechanism first, then the trigger, then the prerequisite and follow-up. Every clause carries information and nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the description covers the mechanism, reversibility, trigger, prerequisite, and next step. It is nearly complete, though it says nothing about whether app data (versus HTML) is affected or any permission requirements.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%: 'checkpoint' and 'command_id' are already documented in the schema, and the description's 'Get the checkpoint number from list_app_checkpoints' largely repeats the schema note. The 'name' parameter is undocumented in both places. With mid-level coverage, the description neither compensates nor detracts, so baseline 3 fits.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('roll an app back') and resource ('one of its earlier checkpoints'), and immediately explains the mechanism (re-saves that checkpoint's HTML as a NEW current one). This clearly separates it from siblings like save_app, edit_app, and promote_app.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the triggering condition ('Use when a newer edit broke the UI'), the prerequisite source ('Get the checkpoint number from list_app_checkpoints'), and the follow-up action ('after restoring, open_app to view it'). The when-to-use and routing are fully specified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_appSave appA
Idempotent

Create or update a UI app in the persistent registry. Two slots, each optional on update (an omitted slot keeps its current value): ui — the complete self-contained HTML document (contract in get_app_guide; window.oma, no external resources, NO embedded manifest block) — and manifest, the app's declaration as a JSON object (kind, collections, settings, scene; keys in get_app_guide). manifest: null clears the declaration. Creating needs ui. Every save snapshots both slots as one new version (history kept). After saving, open it IMMEDIATELY with open_app.

ParametersJSON Schema
NameRequiredDescriptionDefault
uiNocomplete self-contained HTML document using window.oma; omit on update to keep the current one
nameYesapp name, ^[a-z][a-z0-9-]{0,31}$ (e.g. 'kanban', 'habit-tracker')
manifestNodeclaration object (whole-value replace), null to clear, omit to keep
command_idNoidempotency key (uuid); auto-generated if omitted
descriptionNoone line: what this app shows and what data fields it uses
expected_versionNoREQUIRED when overwriting an existing app: the version you read (get_app). Creating a new name needs none

Output Schema

ParametersJSON Schema
NameRequiredDescription
okYes
eotNo
nameNo
noteNo
sizeNo
reasonNo
appliedNo
createdNo
versionNo
prev_sizeNo
manifest_actionNo
expected_versionNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=true, and destructiveHint=false, so safety is covered. The description adds genuinely useful behavior beyond that: every save snapshots both slots as one new version with history kept, and manifest: null clears the declaration. It doesn't restate annotation content, which is a plus.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then slot semantics, then the critical next action (open_app). Dense but each clause carries information. Em-dash nesting makes it slightly heavy to parse, but nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers the mutation semantics, versioning, and follow-up action. It omits mention of expected_version's overwrite requirement and command_id idempotency, though both are documented in the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning on top: the omit-to-keep vs null-to-clear semantics for each slot, and the requirement that creating needs ui. The contract pointer (get_app_guide) for ui/manifest contents is useful context the schema does not carry.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: 'Create or update a UI app in the persistent registry,' which immediately tells the agent this is the write path for app definitions. It further clarifies the two-slot model (ui + manifest). It does not, however, explicitly differentiate itself from the sibling edit_app, leaving the create-vs-edit boundary to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear operational rule: 'Creating needs ui' and 'After saving, open it IMMEDIATELY with open_app,' which routes the agent to the correct follow-up tool. It also explains omitted-slot behavior on update. It stops short of naming alternatives (e.g., when to prefer edit_app or restore_app over this tool).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

security_setSet a security policy keyA
Idempotent

Privileged writer for reserved settings keys (security:* / policy:) — the ONLY tool that can write them; the generic data_ tools refuse reserved keys. Upserts one key/value in the settings collection.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYesa reserved key, e.g. security:kanban:send_message (cap suffixes are snake_case — the caps field names)
valueYesthe policy value: allow | deny (true/false too); delete_items also takes confirm; call_tools takes "*", a JSON array or a comma-separated tool list; policy:csp:<app> (or policy:csp:*) takes a JSON object of origin arrays, e.g. {"connectDomains":["https://api.example.com"]}. An unknown value is REFUSED, never stored
command_idNoidempotency key (uuid); auto-generated if omitted

Output Schema

ParametersJSON Schema
NameRequiredDescription
idNo
okYes
eotNo
seqNo
itemNo
noteNo
reasonNo
deletedNo
collectionNo
violationsNo
expected_versionNo
prev_collection_seqNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint=false, idempotentHint=true, destructiveHint=false, openWorldHint=false), so the bar is lower, and the description adds real context beyond them: privileged status, exclusivity for reserved keys, and upsert (overwrite) semantics. It does not describe permissions/auth prerequisites for a 'privileged' writer, leaving one meaningful gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero filler, and the differentiating facts (privileged, reserved keys, exclusivity) are front-loaded before the generic upsert statement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With annotations declaring safety/idempotency and a full output schema plus 100% parameter documentation, the description carries little remaining burden and covers purpose, exclusivity, and mutation semantics adequately. The only shortfall is not flagging the privileged access requirement it asserts.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents key, value, and command_id in significant detail (accepted value forms, refusal of unknown values, idempotency). The description adds only the generic 'one key/value' framing, which is the expected baseline when the schema carries the burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Upserts one key/value in the settings collection') and adds a precise scope constraint ('reserved settings keys security:* / policy:*'). It also distinguishes itself from siblings by noting the generic data_* tools refuse reserved keys, so an agent can route without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear selection condition (reserved security:*/policy:* keys) and rules out the alternative family ('the generic data_* tools refuse reserved keys'), which functions as a when-not. It stops short of an explicit directive naming the tool to use for non-reserved writes, so the guidance is strong but not fully enumerated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 18 tool updatesv0.7.2
    • Removedapp_history
    • Removedapp_html
    • Removedapp_store_preview
    • Addedapply_data_writes
    • Removeddata_batch
    • Removeddata_collections
    • Changeddata_list1 field changed
      • removedOutput schema / properties / host
        Removed value: -{
        -  "type": "string"
        -}
    • Removeddata_version
    • Addedget_app_html
    • Addedget_data_version
    • Addedget_ui_preference_schema
    • Changedinstall_from_app_store2 fields changed
      • addedOutput schema / properties / outcome
        Added value: +{
        +  "description": "installed = first install; updated = the store had a newer version; current = already installed and unchanged",
        +  "enum": [
        +    "installed",
        +    "updated",
        +    "current"
        +  ],
        +  "type": "string"
        +}
      • changedOutput schema / required
        Previous value: -[
        -  "name",
        -  "version",
        -  "tier"
        -]New value: +[
        +  "name",
        +  "version",
        +  "tier",
        +  "outcome"
        +]
    • Addedlist_app_checkpoints
    • Addedlist_data_collections
    • Changedopen_app1 field changed
      • removedOutput schema / properties / host
        Removed value: -{
        -  "type": "string"
        -}
    • Addedpreview_app_store_entry
    • Changedrestore_app1 field changed
      • changedInput schema / properties / checkpoint / description
        Previous value: -"which checkpoint to restore, 1 = the oldest (see app_history)"New value: +"which checkpoint to restore, 1 = the oldest (see list_app_checkpoints)"
    • Removedui_prefs_schema
  2. 10 tool updatesv0.6.0
    • Changedapp_html1 field changed
      • addedOutput schema / properties / csp
        Added value: +{
        +  "additionalProperties": {
        +    "items": {
        +      "type": "string"
        +    },
        +    "type": "array"
        +  },
        +  "propertyNames": {
        +    "type": "string"
        +  },
        +  "type": "object"
        +}
    • Changeddata_collections2 fields changed
      • removedOutput schema / properties / collections / items / properties / last_activity / anyOf
        Removed value: -[
        -  {
        -    "type": "string"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]
      • addedOutput schema / properties / collections / items / properties / last_activity / type
        Added value: +[
        +  "string",
        +  "null"
        +]
    • Changeddata_list2 fields changed
      • removedOutput schema / properties / next_cursor / anyOf
        Removed value: -[
        -  {
        -    "type": "string"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]
      • addedOutput schema / properties / next_cursor / type
        Added value: +[
        +  "string",
        +  "null"
        +]
    • Changededit_app2 fields changed
      • removedOutput schema / properties / prev_size / anyOf
        Removed value: -[
        -  {
        -    "type": "number"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]
      • addedOutput schema / properties / prev_size / type
        Added value: +[
        +  "number",
        +  "null"
        +]
    • Changedfile_list2 fields changed
      • removedOutput schema / properties / next_cursor / anyOf
        Removed value: -[
        -  {
        -    "type": "string"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]
      • addedOutput schema / properties / next_cursor / type
        Added value: +[
        +  "string",
        +  "null"
        +]
    • Changedfile_read2 fields changed
      • removedOutput schema / properties / next_offset / anyOf
        Removed value: -[
        -  {
        -    "type": "number"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]
      • addedOutput schema / properties / next_offset / type
        Added value: +[
        +  "number",
        +  "null"
        +]
    • Changedget_app2 fields changed
      • removedOutput schema / properties / next_offset / anyOf
        Removed value: -[
        -  {
        -    "type": "number"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]
      • addedOutput schema / properties / next_offset / type
        Added value: +[
        +  "number",
        +  "null"
        +]
    • Changedsave_app2 fields changed
      • removedOutput schema / properties / prev_size / anyOf
        Removed value: -[
        -  {
        -    "type": "number"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]
      • addedOutput schema / properties / prev_size / type
        Added value: +[
        +  "number",
        +  "null"
        +]
    • Changedsecurity_set1 field changed
      • changedInput schema / properties / value / description
        Previous value: -"the policy value: allow | deny (true/false too); delete_items also takes confirm; call_tools takes \"*\", a JSON array or a comma-separated tool list. An unknown value is REFUSED, never stored"New value: +"the policy value: allow | deny (true/false too); delete_items also takes confirm; call_tools takes \"*\", a JSON array or a comma-separated tool list; policy:csp:<app> (or policy:csp:*) takes a JSON object of origin arrays, e.g. {\"connectDomains\":[\"https://api.example.com\"]}. An unknown value is REFUSED, never stored"
    • Changedui_prefs_schema2 fields changed
      • removedOutput schema / properties / shared / items / properties / default / anyOf
        Removed value: -[
        -  {
        -    "type": "string"
        -  },
        -  {
        -    "type": "number"
        -  },
        -  {
        -    "type": "boolean"
        -  }
        -]
      • addedOutput schema / properties / shared / items / properties / default / type
        Added value: +[
        +  "string",
        +  "number",
        +  "boolean"
        +]

TDQS

A3.6/5.0

Scored across 33 tools

Disambiguation4/5

Most tools map to clearly distinct resource+action targets, and the verbose descriptions actively distinguish near-neighbors (save_app vs edit_app, data_list vs get_data_version vs data_changes, file_write vs the chunked trio). The two tools explicitly labeled 'Internal... not useful to call directly' (get_app_html, preview_app_store_entry) add minor noise, but boundaries are otherwise well drawn.

Naming Consistency3/5

Names are uniformly snake_case, but the ordering convention is genuinely mixed: action_resource (get_app, list_apps, save_app, edit_app, restore_app) coexists with resource_action (data_add_item, data_list, file_write, app_store_list, security_set). Still readable and clusterable by prefix, but no single predictable pattern.

Tool Count3/5

33 tools is on the heavy side, though the surface spans several genuinely large sub-domains (app lifecycle, data CRUD, files, app store, settings) so most tools earn their place. It sits above an ideal scoped set and the two internal helpers could arguably be folded away, making it borderline rather than clean.

Completeness4/5

Coverage is strong: apps have full create/read/edit/delete plus checkpoints, restore, and promotion; data has CRUD, batch writes, a change feed, version counter, and collection listing; files have single and chunked writes, read, list, and delete; app store browse/install and settings are present. Only minor gaps (e.g. app duplication/rename, explicit search) remain.

Maintenance

ActivityActive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables interacting with YouTube Data API v3 through MCP tools (get-video, get-channel, get-latest-video) and bundled UI apps for video and channel profiles.
    7 npm
    4
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Federates UI component registries, enabling AI agents to fetch exact component code, dependencies, and setup prerequisites directly into the workspace. Includes sandboxed previews, anti-slop layout auditing, and offline-to-cloud telemetry sync.
    7 npm
    1
    MIT