Finitact
Enables computer-use control of Blender, a native Windows application, through screen observation, candidate selection, and synthetic input.
Enables computer-use control of Unity, a native Windows application, through screen observation, candidate selection, and synthetic input.
Enables browser-based computer-use interactions with Wikipedia, including multi-step navigation and search tasks.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@FinitactIn Chrome, estimate 2 t3.micro EC2 instances for one month on AWS Pricing Calculator"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Finitact
Finitact is an MCP server that lets coding agents (Claude Code, Codex, any MCP client) do computer use by leaning on a one-decision model such as Jev. Finitact observes the screen, reduces it to a finite list of candidates, has a small decision model (currently TypeSafe's Jev) pick one, sends the input, and judges the result.
Finitact started as an independently maintained copy ofbrowser-use/jev-ultrafast. It is a personal project, not an
official release of Browser Use or TypeSafe, and it is not affiliated with windows-mcp. The original MIT
copyright stays in LICENSE. NOTICE records the provenance.
Intent
jev-ultrafast and windows-mcp are both excellent prior work, but honestly I wanted a bit more capability and cost-efficiency.
Use cases. jev-ultrafast is a fast browser-driving demo, and a great one. But it could not touch native Windows apps (including Blender and Unity, which expose almost no accessibility information), and it got nowhere on the AWS Pricing Calculator, one of the websites that annoys me most.
Cost. windows-mcp has everything computer use needs, but it burns tokens. On nine short Windows tasks the calling agent used 1.4× to 6.5× more tokens with windows-mcp than with Finitact (results), because every observation returns a desktop snapshot that stays in its context.
So I tried an implementation where the expensive agent decides what to do in a few words, and a cheap, fast model decides which of the observed controls that means, many times per task.
Related MCP server: screen-use
Honest status: Jev is hard to use
My feed is full of posts praising Jev. Honestly, though, Jev is hard to use. Current pain points:
Jev only knows the words on the screen. Jev picks among candidate words on something like intuition, so it keeps making choices that anyone with knowledge of the app would rule out. Finitact therefore stops with
provider_uncertainand hands the top five candidates back to the calling agent, which costs a round trip of several seconds (ADR-0019, ADR-0020, ADR-0022).DONEis not proof. Its progress signals cannot be trusted yet. Finitact decides success with a rule over the observed screen change, combined with Jev's judgment of whether that change satisfies the goal. The judgment is conservative: on held-out samples it confirmed only 44–56% of real successes, and it never confirms a fill that replaces a one-line document. A miss becomes "unverified", not a false success (ADR-0030).Safety cannot be delegated to the model. Jev judged correctly only 75–82% of the time whether a region is an input field, or whether a popup is transient. Finitact guarantees that a candidate is executable from observed evidence or checks just before input. Jev only chooses among executable candidates.
Agents do not use the interface you design for them. Calling agents ignored or confused the first two ways to pick a candidate directly. Instead of tuning tool descriptions, I collapsed them into one field on the goal (
ref) (ADR-0041–ADR-0043).Decision quality: hoping for more. On a fixed set of 13 decision cases × 10 repetitions, Jev chose the expected action 89.2% of the time. For reference, a local Qwen2.5-14B scored 84.6%. Neither reached the preregistered 90% bar, and their failures took different shapes. Laya 0.3.21 scored lower still, at 53.8% (decision-model comparison).
I wanted to avoid having the agent itself do so-called prompt engineering to steer Jev, but in places it ended up doing some.
How the work is split
flowchart LR
U([User task]) --> A
subgraph A[Calling agent: Claude Code, Codex, ...]
A1[Split the task into short goals]
A2[Choose the target tab or window]
A3[Read outcomes, then rephrase or pick by ref]
end
A -- "run_browser / run_windows<br/>goals, fill values, optional ref" --> F
subgraph F[Finitact MCP server]
F1[Observe: DOM, or UIA + OCR]
F2[Build finite candidates]
F4[Re-check the target and send input<br/>after input-lock, idle and hit-test checks]
F5[Verify success from the observed change]
end
F1 --> F2
F2 -- "goal + candidates + history" --> J
J[Jev: pick one candidate id,<br/>or WAIT / DONE / BLOCKED,<br/>with probabilities] -- choice --> F4
F4 --> F5
F5 -- "observed change" --> J2[Jev: does this change<br/>satisfy the goal?]
J2 --> F5
F -- "per-goal outcome,<br/>or top-5 candidates when uncertain" --> ARole | Decides | Never does |
Calling agent | What the goals are, which window or tab, the exact text to enter, what to do after a failure | See every intermediate screen (unless it asks with |
Finitact | What is observable and executable, when to stop, whether input may be sent, whether the goal was reached | Invent selectors or coordinates, or resend input whose delivery is uncertain |
Jev | Which observed candidate matches the goal, and whether an observed change fits the goal | Produce coordinates, selectors, or text, or guarantee safety |
A small text model writes a value only when a browser goal needs text and the agent did not supply fill_values.
One run, step by step
sequenceDiagram
participant A as Calling agent
participant F as Finitact
participant J as Jev
participant W as Window / tab
A->>F: run_windows(target, goals)
loop each goal
F->>W: observe
alt first goal carries a ref from observe_window
F->>F: resolve the ref (skip Jev)
else
F->>J: goal + candidates + history
J-->>F: choice + confidence
opt low confidence (below 0.4, or BLOCKED below 0.6)
F-->>A: provider_uncertain + top-5 candidates
Note over A,F: agent picks one by ref or rephrases, then calls again
end
end
F->>W: re-check, then send input (stop if the screen changed)
F->>W: observe the change
F->>J: observed change: does it satisfy the goal?
J-->>F: yes / no
F-->>F: verified_success, or unverified and continue
end
F-->>A: per-goal outcomeWhen a run fails
Finitact has no automatic fallback. It never re-sends an input on its own and never switches routes on its own. Each result tells the calling agent why the run stopped, and the agent decides how to recover.
In the result | What the agent can do |
| Pick one of |
| Check the state with |
| Call |
| Send a new run with only the remaining goals. |
| Call again with a larger budget or a later deadline. |
No response | Call again with the same |
| Do not retry under a new |
get_run_journal returns the record of a run, and cancel_run stops a run. Finitact does not include a second
computer-use method. If Finitact cannot do a task, the agent must use its own tools, for example a shell or another MCP server.
Results
Measured on 2026-10-04 with Claude Code and claude-sonnet-5-5 as the calling agent, one Finitact commit, and the
same prompt for both systems. Details and limits are in docs/report/.
9 short Windows tasks (Notepad, Calculator, VS Code, Tk, Unity, Blender), 10 trials each: windows-mcp succeeded 10/10 on every task, Finitact 10/10 on 8 and 9/10 on one (Blender, a menu left open). Finitact was faster on 7 of 9, and the calling agent used 1.4× to 6.5× fewer tokens with it on all 9.
5 multi-step tasks, 5 trials each: Finitact succeeded 5/5 on all. windows-mcp succeeded 5/5 on four and 4/5 on the Wikipedia task (the search route could not be shown). The calling agent used fewer tokens with Finitact on four (1.4× to 2.6×) and more on one (1.4×).
Finitact's own calls to Jev are not in these token counts. Counting them at list prices, a trial still cost less with Finitact on 13 of 14 tasks (1.3× to 6.3× short, 1.4× to 1.9× multi-step) and 1.1× more on the Wikipedia task (see the report).
Setup
Requirements
What | Needed for | Notes |
Python 3.12+ and uv | everything | |
TypeSafe API key ( | every run | Jev chooses each action. Get a key from TypeSafe |
OpenAI-compatible text model key ( | fill goals without | Composes the text to type. The default endpoint is OpenRouter ( |
Chrome started with |
| |
Windows 10/11 x64, Windows-native Python 3.12 with the |
| OCR is RapidOCR (ONNX Runtime / OpenVINO), installed by pip. RapidOCR downloads its models once. You do not need Tesseract or other system OCR. UI Automation uses |
The server reads the checkout's .env at start-up (variables set by the MCP client take precedence), so the keys
do not have to be passed through the client configuration.
git clone https://github.com/yo-xe/finitact.git
cd finitact
uv sync
cp .env.example .env # set TYPESAFE_API_KEY (and TEXT_MODEL_API_KEY for generated text)Browser
Start Chrome with remote debugging (we recommend a disposable profile), then register the server with your MCP client:
claude mcp add finitact -- uv --directory /path/to/finitact run finitact-mcpSet BU_CDP_URL=http://127.0.0.1:9222 if Chrome listens on a non-default endpoint.
Prerequisites for browser operation
run_browser needs a Chrome reachable over CDP: started with --remote-debugging-port (default 9222, otherwise
BU_CDP_URL). run_windows hands a goal on a Chrome window to the browser path only when all of these conditions hold.
Otherwise it stays on screen input:
enabled by
routing: "browser_if_singleton"on the call, or byFINITACT_WINDOW_BROWSER_ROUTE=1for calls that omitrouting.BU_CDP_URLis a localhttpendpoint (127.0.0.1,localhost,::1) andBU_CDP_WSis unset.the CDP browser process owns exactly one visible, non-minimized window, the target, and that window has exactly one tab. Finitact ignores the browser's own bubbles, such as the download bubble.
the tab's title prefixes the window title, its page is
http(s)and visible, and the window does not change while Finitact checks this.the call has one goal and no
ref, drop target, operation limits, label constraints or selection policies.
An everyday Chrome with several tabs therefore stays on screen input. Use a dedicated profile with a single tab.
Windows
Only Windows-native Python sends mouse and keyboard input. Install into a Windows venv (x64) with the screen extra and download the OCR models once:
py -3.12 -m venv .venv; .venv\Scripts\python.exe -m pip install -e ".[screen]"
.venv\Scripts\python.exe -c "from finitact.stage1_extractors import default_ocr_extractor; default_ocr_extractor()"Register that interpreter with a Windows MCP client:
claude mcp add finitact -- C:\path\to\finitact\.venv\Scripts\python.exe -m finitact.mcp_serverFrom a client running in WSL, set FINITACT_WINDOWS_PYTHON in .env to that python.exe and register
scripts/finitact-mcp-windows.sh. The script forwards the keys through WSLENV.
Tools
Tool | Purpose |
| Targets as |
| Read-only candidates with |
| Execute ordered goals and return the per-goal |
| Inspect or stop a run |
Treat confirmed as "input was sent", not "goal succeeded". Read outcome. Replaying the same run_id with the
same input returns the cached result without repeating input.
Safety and limits
Mouse and keyboard input starts only after one second without user input and while an on-screen indicator is visible. Each run is bounded by a deadline, an action budget and a budget of decision-model calls. Ending the server process stops a run.
The indicator is a glowing orb in a screen corner. A small amber spark beside it lights while Finitact waits on the decision model, and a smaller orb marks roughly where each input lands. Size, place, colors and motion can be set with
FINITACT_INDICATOR_*in.env(see.env.example). The indicator itself cannot be turned off.Input is never resent automatically. If it is unclear whether input arrived, the server refuses further input until it restarts.
Screen
fillreplaces the clipboard and does not restore it.Tested on the cases in the report only. Electron menus through UIA, long tasks across many windows, and other operating systems are untested.
Development
uv run ruff check .
uv run pytest # offline
uv buildLive evaluation scripts under scripts/ drive real windows and make paid API calls.
Design records (docs/adr/, docs/bugs/) and the specification are in Japanese. This README, the evaluation report,
decisions.md and known-issues.md are
in English.
Available Tools
7 toolscancel_runAIdempotent
Request cancellation of an active run. Cancellation is checked before each provider call or mutation.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds genuinely non-obvious semantics: cancellation is cooperative and only takes effect at the next provider call or mutation, meaning the agent should not expect or poll for immediate effect. It still omits what the response looks like or how to confirm the run actually stopped.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences with the purpose front-loaded and the behavioral caveat immediately after. Nothing is redundant and nothing extraneous remains.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, non-destructive cancellation call the definition covers purpose and the key timing behavior. However, with no output schema present, the description says nothing about the return value or how the agent learns whether cancellation was accepted, leaving a modest but real gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter (run_id) with 0% schema description coverage, so the description technically has room to compensate. It only indirectly conveys that run_id must reference an active run; no format or sourcing guidance is given, but the parameter is a trivially understood identifier.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Request cancellation of an active run'), which is unambiguous and clearly distinct from sibling tools like observe_window or list_windows. It does not explicitly name or contrast itself with siblings, but the action is self-evident.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Active run' implies the prerequisite that the run must still be in progress, but there is no explicit when-to-use/when-not guidance, no mention of alternatives, and no statement of what to do if the run has already completed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_run_journalARead-onlyIdempotent
Read the redacted event journal for a run. Journals omit page text, credentials, and field values.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish the safety profile (readOnlyHint=true, idempotentHint=true, destructiveHint=false, openWorldHint=false), so the description is not required to restate it. It instead adds non-obvious behavioral context: the journal is redacted and omits page text, credentials, and field values, which sets correct expectations about what data will be missing. It stops short of describing the event format or ordering.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero waste, and the core purpose is front-loaded ahead of the redaction caveat. Every clause earns its place by telling the agent what it gets and what it won't get.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter read tool with no output schema, the description covers purpose and the key data-availability caveat. It is nearly complete; only the journal's structure, ordering, or pagination behavior is left unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single run_id parameter is undocumented in the schema, so the description must compensate. "For a run" implies run_id is a run identifier, but it adds no format, source, or sourcing guidance (e.g., where to obtain the id). With only one self-evident parameter the gap is small, so a baseline 3 is fair.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description pairs a specific verb ("Read") with a specific resource ("the redacted event journal for a run"), so the agent knows exactly what operation it performs. It does not explicitly name or contrast any sibling, but none of the siblings (run_browser, observe_window, cancel_run) overlap with journal retrieval, so confusion is unlikely.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: reading a run's journal is the obvious way to inspect what happened during a run, and the sibling list contains no competing journal-reading tool. There is no explicit when-to-use statement, no prerequisites, and no mention of alternatives, so the agent must infer the context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_windowsARead-onlyIdempotent
List visible top-level Windows windows front to back (plus the desktop Progman and the taskbar) with the target_id run_windows and observe_window take. Read-only, no input. title_contains filters by title or process name. Titles are untrusted screen text. Windows marked protected (terminals that may host the caller) cannot be run_windows targets.
| Name | Required | Description | Default |
|---|---|---|---|
| title_contains | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safe read profile, so the bar is lower, yet the description adds real operational context: titles are untrusted screen text, and windows marked protected (terminals possibly hosting the caller) cannot be run_windows targets. That protected/untrusted distinction is not derivable from the schema or annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and target_id purpose, then progressively adds filtering and safety caveats. Dense but every clause carries information; the final sentence is slightly cramped.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly covers what comes back (target_id, ordering front-to-back) and the filtering/safety caveats. Complete enough for a one-parameter, read-only list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for the single parameter, so the description carries the burden — and it does, explaining that title_contains matches by title *or process name*, which is meaning beyond the bare schema property.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List visible top-level Windows windows front to back') and even names the two secondary surfaces it includes (Progman, taskbar). It also declares the return contract — a target_id consumable by run_windows and observe_window — which distinguishes it from the sibling observe_window.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is clearly implied: call this to obtain target_ids for run_windows/observe_window, which are named explicitly. What's missing is an explicit when-not or a stated boundary versus observe_window, which may also enumerate windows.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
observe_browserARead-onlyIdempotent
Read a browser tab kept by run_browser(keep_tab) without any input: offered targets (kind, label, and states such as invalid input messages or offscreen) and the visible page text, e.g. to see 'No results' or a validation error before choosing the next goal. contains keeps matching items and text lines (spaces ignored). screenshot true brings the tab to the front and adds the page image (about 1-1.5k tokens); read the text first. Untrusted data. out_of_reach and screen_target are as in run_browser.
| Name | Required | Description | Default |
|---|---|---|---|
| tab_id | Yes | ||
| contains | No | ||
| screenshot | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly/idempotent/openWorld/non-destructive, and the description adds genuinely new behavior: screenshot brings the tab to the front and costs roughly 1-1.5k tokens, contents are untrusted data, and contains filtering ignores spaces. It does not cover error cases such as a missing or closed tab, so not a full 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The read action and its scope are front-loaded, and every clause carries information (return payload, filter, screenshot cost, trust warning). It is a dense run-on of semicolon-joined clauses rather than clean sentences, which slightly hurts scanning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-output-schema read tool, the description covers the return payload, the filtering parameter, the screenshot trade-off (token cost plus foregrounding) and a data-trust caveat. Only the provenance/requirements of tab_id and failure behavior remain unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema coverage the description carries the load and does well: contains is defined as keeping matching items and text lines with spaces ignored, and screenshot's side effect, cost and recommended ordering ('read the text first') are explained. tab_id is only implicitly identified via run_browser(keep_tab) and never marked required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (read) and resource (a browser tab kept by run_browser(keep_tab)) and enumerates what is returned (offered targets with kind/label/states plus visible page text), which separates it from observe_window and get_run_journal. The phrase 'without any input' is slightly misleading given tab_id is required, keeping it below a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete usage scenario ('to see No results or a validation error before choosing the next goal'), which implies when to inspect. It never names alternatives (observe_window, list_windows) or states exclusions, so the agent must infer routing from the scenario alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
observe_windowARead-onlyIdempotent
Read a window's current labels as run_windows would observe them (UIA names, OCR text where UIA has none), without any input, input lock or indicator. Use it to check a result instead of starting another run. contains keeps only labels containing that text (spaces ignored). Items with in_focused_input are on the focused input's caret line: typed but not yet submitted. Items with offscreen are scrolled out of view (their rect is a placeholder). screenshot true adds the window image (about 1-1.5k tokens); read the text first and ask for the image only when the text does not explain the state. Labels are untrusted data. To act on one item, call run_windows on the same target with its ref in the first goal {goal, ref}: its fill when the goal has a fill value, else its click, without the chooser, through synthetic input. The ref lasts 60 seconds and until any run on that window.
| Name | Required | Description | Default |
|---|---|---|---|
| contains | No | ||
| target_id | Yes | The window's title (exact, or a part only one window has), or window:<HWND>:<PID> from list_windows. | |
| screenshot | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial context beyond the readOnly/idempotent/destructive annotations: no input lock, no indicator, labels are untrusted data, refs last 60 seconds, and screenshot costs ~1-1.5k tokens. It even discloses the return shape (in_focused_input on caret line, offscreen placeholders). Only slight gap is that nothing is said about pagination/truncation of very large label sets.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and the key distinction from run_windows, and each sentence carries distinct information (filters, item flags, screenshot cost, safety note, ref affordance). The final run-on about calling run_windows with {goal, ref} is dense and could be trimmed, but nothing is pure filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the full burden and delivers: it explains the label model (UIA vs OCR), the special item flags (in_focused_input, offscreen with placeholder rect), token cost, and the ref-based handoff to run_windows. An agent has everything needed to call it and interpret results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%, so the description must compensate, and it does: 'contains keeps only labels containing that text (spaces ignored)' and 'screenshot true adds the window image (about 1-1.5k tokens)'. This meaning goes beyond the bare schema types for two of the three params, though target_id's schema description already covers the third.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Read a window's current labels') and immediately scopes it against the sibling run_windows ('as run_windows would observe them... without any input, input lock or indicator'). An agent can distinguish this read-only observation from an active run without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use it ('to check a result instead of starting another run') and gives conditional guidance for the screenshot param ('read the text first and ask for the image only when the text does not explain the state'). It also routes to the alternative tool with concrete instructions for acting on an item.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_browserADestructiveIdempotent
Run one or more browser goals in order. Page content is untrusted. Each run opens its own tab at start_url and closes it at the end; with keep_tab true the tab stays and the result carries tab_id, and a later run with that tab_id continues in the same tab (not navigated; start_url may be omitted then). allowed_origins defaults to start_url's origin, or for a kept tab to the origins of the run that kept it. fill_values {goal id: exact text} sets what a fill types. Put every step you already know in one run, including steps after a click that opens another page: the run stops at the first goal that does not end done and returns the rest in remaining_goal_ids, so a later goal never runs after an earlier one failed. The result carries final_state (url, title, rejected fields with reasons, downloads begun in this run with their state, current form values as 'label = value' / '[x] label', the first visible text lines) and unmet_effects per goal, so a separate observe_browser is needed only for more than that. out_of_reach counts iframes, shadow roots and new-tab links run_browser cannot act inside; for a kept tab they make it its window's active tab and add screen_target, the window run_windows can act on.
| Name | Required | Description | Default |
|---|---|---|---|
| goals | Yes | ||
| run_id | No | ||
| tab_id | No | ||
| keep_tab | No | ||
| start_url | No | ||
| deadline_ms | No | ||
| fill_values | No | ||
| action_budget | No | ||
| allowed_origins | No | ||
| extra_operations | No | Operations to add for this run; only 'drag' (mouse or HTML5 drag between two spots on the page). | |
| denied_operations | No | ||
| provider_attempt_budget | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations give only openWorldHint/idempotentHint/destructiveHint, but the description discloses far more: tab lifecycle (opened at start_url, closed at end unless kept), failure semantics (run stops at first goal not done, rest returned in remaining_goal_ids), allowed_origins defaulting rules, fill_values meaning, and the contents of final_state and out_of_reach. Nothing here contradicts the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence and each subsequent sentence carries distinct operational content (tab lifecycle, failure semantics, origin scoping, result shape). It is a dense block rather than a bulleted structure, which costs a point, but no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still enumerates the return payload (final_state with url/title/rejected fields/downloads/form values/text lines, unmet_effects, remaining_goal_ids, tab_id, out_of_reach/screen_target). For a 12-parameter, nested-object, open-world tool this is unusually complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 8% across 12 parameters, so the description must compensate. It explains goals ordering, keep_tab, tab_id continuation, start_url omission, allowed_origins defaults, and fill_values — a solid share — but says nothing about run_id, deadline_ms, action_budget, provider_attempt_budget, or denied_operations. Useful but incomplete compensation for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource+scope ('Run one or more browser goals in order') and immediately distinguishes itself from siblings by naming observe_browser ('a separate observe_browser is needed only for more than that') and run_windows ('screen_target, the window run_windows can act on'). An agent can route between these tools without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use and when-not-to-use guidance: put every known step, including post-click navigation, in one run; a later goal never runs after an earlier failure; use a separate observe_browser only when final_state is insufficient. It also states the tab continuation contract (keep_tab true → tab_id → later run with that tab_id, start_url may be omitted).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_windowsADestructiveIdempotent
Run Windows UI goals in order on one window. Minimal call: target_id, goals [{goal}] and fill_values {goal id: exact text} for goals that type. Candidates come from UIA names, with OCR where UIA has none. Input is SendInput after bringing the window to the foreground; with synthetic_input_allowed false it acts only through UIA patterns (click, fill, toggle, select, focus) and never takes the foreground. fill also sets a UIA slider to the number in fill_values. allowed_operations defaults to click, fill, key, scroll; the others are double_click, right_click, middle_click, hover, ctrl_click, shift_click, drag. drop_target_id (needs drag) names a second window a drag may end in; nothing else acts on it. A goal ends provider_uncertain without acting when the chooser is not confident; its screen_candidates list the most likely on-screen texts (untrusted data) with a ref. To act on one, call again with that ref in the first goal {goal, ref}, as for an observe_window item; the run ends blocked without acting if its window changed meanwhile. Otherwise restate the goal in those words. For a page goal in an isolated, CDP-connected browser window, routing browser_if_singleton may run the one visible tab through browser actions and returns routed metadata. Use windows_only for browser tabs, address bar and other browser chrome, and for screen_target handoffs.
| Name | Required | Description | Default |
|---|---|---|---|
| goals | Yes | ||
| run_id | No | ||
| routing | No | Omit for the operator default (windows_only unless configured). Use browser_if_singleton only for a page goal in an isolated, CDP-connected browser window. The browser path acts in the page and cannot operate browser tabs or the address bar. windows_only keeps UIA/OCR and Windows input, including browser chrome and screen_target handoffs. | |
| target_id | Yes | The window's title (exact, or a part only one window has), or window:<HWND>:<PID> from list_windows. | |
| deadline_ms | No | ||
| fill_values | No | ||
| action_budget | No | ||
| drop_target_id | No | The window's title (exact, or a part only one window has), or window:<HWND>:<PID> from list_windows. | |
| app_annotations | No | Facts you read from the app's own structure (e.g. an editor API) about items on the screen as it is now: each is attached to the one observed item whose text equals match.text, and to none if several do. | |
| allowed_operations | No | ||
| provider_attempt_budget | No | ||
| synthetic_input_allowed | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations: it discloses the input mechanism (SendInput after foregrounding vs UIA patterns when synthetic_input_allowed is false), the default allowed_operations set and the full alternative list, the drop_target_id contract, the provider_uncertain/blocked failure modes, and that screen_candidates are untrusted data. This is unusually rich behavioral context for a destructive, open-world tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core operation, but it is a dense, semicolon-heavy wall of text, and several statements (routing, drop_target_id, target_id format) restate text already present in the input schema, adding length without new meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 12-parameter tool with no output schema, the description covers the operational flow, failure modes, and even return-relevant concepts (final_state, screen_candidates). The remaining gap is that the numeric budgets and run_id are never explained, so an agent cannot reason about timeouts or run continuation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%, so the description must compensate, and it does add real meaning for fill_values (goal id: exact text), allowed_operations (defaults plus the full alternate list), drop_target_id, routing, and synthetic_input_allowed. However, run_id, deadline_ms, action_budget, provider_attempt_budget, and app_annotations are left completely opaque, so the gaps roughly match the covered parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb and resource with scope: 'Run Windows UI goals in order on one window.' It also distinguishes the Windows path from the browser path via routing, but the sibling differentiation (vs run_browser/observe_window) is embedded in later routing detail rather than stated crisply up front.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use conditions are given for the key decision points: when to use windows_only vs browser_if_singleton, when to pass a ref instead of restating a goal, and what to do when a goal ends provider_uncertain or blocked. Alternatives and their selecting conditions are named, not implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v0.1.1- First observed
cancel_run - First observed
get_run_journal - First observed
list_windows - First observed
observe_browser - First observed
observe_window - First observed
run_browser - First observed
run_windows
TDQS
Scored across 7 tools
Each tool targets a clearly distinct resource and operation: list Windows targets, run browser goals, run Windows UI goals, observe browser pages, observe Windows labels, cancel a run, and read a run journal. The browser/window split is consistently reflected in both run and observe tool names. No two tools appear interchangeable.
All tool names use snake_case with a verb-first pattern: list_windows, run_browser, run_windows, observe_browser, observe_window, cancel_run, get_run_journal. The convention is predictable and readable throughout.
Seven tools is well-scoped for a UI and browser automation control server. Each tool covers a necessary capability without obvious redundancy or missing core surface.
The set covers the core automation lifecycle: target discovery, running browser and Windows goals, observing results, cancelling runs, and reading journals. Minor gaps remain, such as no explicit way to enumerate existing browser tabs or manage/close kept tabs beyond the run_browser keep_tab flow, and no list-active-runs tool. These are workable limitations rather than severe omissions.
Maintenance
Related MCP Connectors
- openhelmOAuthai.openhelm
Autonomous cloud agent tasks: real browser + your tools, structured evidence-backed results.
Deterministic contextual decision arbitration and action routing for autonomous software. Takes current state, context, or intent plus caller-supplied candidate actions, state transitions, routes, refusals, escalations, tools, or models and returns a deterministic ordered candidate field. Also provides persistent machine representations for memory, retrieval, indexing, and downstream coherence measurement.
Operate Linux, macOS and Windows from your LLM. Every action runs through an auditable allowlist.
Eyes and hands on real Windows PCs — observe, click, type via Glasswarp API.
Related MCP Servers
- AlicenseBqualityDmaintenanceEnables AI agents to control Windows desktop applications by wrapping Codex's Computer Use capability.1311MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to see, locate UI elements, and operate any Windows desktop app through natural language, using accessibility-tree matching with optional vision-model fallback, plus an autonomous visual loop with introspection and meta-learning.3MIT
- AlicenseNot gradedqualityAmaintenanceEnables safe Windows desktop automation and computer use through natural language, including window observation, UI Automation, and execution of verified actions like clicking, typing, and scrolling.2MIT
- AlicenseAqualityAmaintenanceEnables users to drive a selected Windows application toward a natural-language goal by reading the UI with local OCR/UIA and deciding next actions.661MIT