clearcote-jet-mcp
OfficialSummary: This MCP server lets an assistant run an entire web task in a Clearcote browser and read back what the page says, then close the tabs it left open.
browse(goal, url?, tab_id?)— carry out a whole task on a website (e.g. "Find the date of the next bank holiday in England and Wales") and get the page's answer as markdown. Every value that must be typed has to be written into the goal.Start a task from a
url, or continue an existing tab withtab_id(a follow-up step, or resuming afterneeds_input/blocked).Interpret the returned status:
done,blocked,needs_input(retry with the sametab_idand a goal containing the missing value),budget(step limit hit) orerror.Leave the browser window open for the user to watch; the browser shuts itself down after a few idle minutes, closing its tabs.
close_tab(tab_id)— close a tab thatbrowseleft open.Scope note: this schema exposes only these two tools; the README's
snapshot,act,list_skillsandforget_skillare not available here, so inspection and manual steps aren't possible through this server.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@clearcote-jet-mcpfind the next UK bank holiday on GOV.UK"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Clearcote Jet
Tell it what you want done on a website. Jet reads the page, picks the next step in about a quarter of a second, does it with a real mouse and keyboard in Clearcote, and hands you the answer.

Three live runs, recorded in real time; the captions are added afterwards from each run's log. Every step shows what Jet did, how sure it was, and how long it took to decide. Watch the MP4.
Why use it
Fast. Each decision took 235 to 473 ms in our runs. Most of a run is the browser, not the thinking.
Cheap. A finished task averaged 17,973 input tokens: about $0.0008, or $0.76 per 1,000 tasks.
Learned once, then free. A task that worked is saved as a skill: the same goal from the same page then runs again with no model calls (24 of 24 replays in our runs), and where the page has changed, the agent takes over for that part and saves it again.
You can see why. Every step comes with the probability behind it and the options it passed over, so a shaky step stands out.
It uses the page like a person. Curved mouse paths, key-by-key typing, dropdowns picked from the open list, and a cookie banner is just another thing to click (it declines when it can).
One loop for every site. There is no per-site code: the same loop looks at the page again after every step.
Related MCP server: Cloudflare Playwright MCP
Try it in two minutes
git clone https://github.com/clearcotelabs/clearcote-jet && cd clearcote-jet
pip install -e .
cp .env.example .env # add TYPESAFE_API_KEY
python examples/run_task.py # watch it find the next UK bank holiday, cursor drawnJust the library: pip install git+https://github.com/clearcotelabs/clearcote-jet.
What it did in our runs
Live, in a visible Clearcote window, 29 September 2026:
Task | Result | Steps | Decisions | Time | Input tokens | Cost |
GOV.UK: find the next bank holiday in England and Wales | done: 25 December | 4 | 5 | 14.8 s | 13,654 | $0.0006 |
Python docs: search | done | 3 | 4 | 10.9 s | 29,048 | $0.0012 |
Python docs: switch the page to version 3.12 | done | 1 | 2 | 4.4 s | 19,349 | $0.0008 |
Hacker News: open the top story's comments | done | 1 | 2 | 2.7 s | 29,877 | $0.0013 |
RFC 9110: jump to the section on 404 Not Found | stopped at the 60-step cap | 60 | 60 | 92.3 s | 368,958 | $0.0155 |
The last row is a failure we kept on purpose: on a very long document it scrolled instead of using the table of contents. Jet now jumps to a section the goal names: on the same page it goes straight to "404 Not Found" in one step (13,793 input tokens, $0.0006, in 3 of 3 runs on 2 October 2026). For a goal that names no place it still scrolls, and stops after 12 scrolls in a row.
A GOV.UK run, step by step:
1. click "Reject additional cookies" p=0.85 decided in 361 ms
2. type "next bank holiday" into "Search" p=1.00
3. click "Search GOV.UK" p=0.93
4. click "UK bank holidays" p=0.99
done · 4 actions · 5 decisions · 14.8 s · 13,654 input tokens ≈ $0.0006What it costs
Decision model: $0.042 per million input tokens; output tokens are free. That is the published rate on 29 September 2026; set
CLEARCOTE_JET_USD_PER_MTOKif yours differs.Measured (seven everyday tasks, three runs each, 2 October 2026): 10,372 to 30,031 input tokens per finished task, so $0.0004 to $0.0013, about a quarter fewer than before that day's request trimming. A replayed skill: 0.
Not included: your Clearcote licence, and an optional text model, which bills with its own provider.
Your own runs: every result carries
usage(tokens, requests,estimated_usd), andpython examples/estimate_cost.pytotals everything inruns/and projects the cost per 1,000 tasks.
Learned once, replayed for free
The first time a goal finishes from a start page, Jet saves how it did it as a skill: each step with how to find its control again (role, label, and its place among controls labelled alike; numbers in a label may change, so "87 comments" is found again as "213 comments"), how the run ended, and how the page's list of results is read. The next time the same goal is run from the same page, Jet replays the skill with no model calls. A cookie banner that doesn't show again is skipped. If the page has changed (a button renamed, a step that is gone, an ending that doesn't match), the agent takes over from that point, decides only what changed, and saves the skill again.
If the results came from a JSON request, the skill keeps that request too, and a replay reads the rows with it from inside the page, without clicking at all.
2 October 2026, 8 tasks × 3 runs | Worked out step by step | Replayed |
Passed | 21 of 24 | 24 of 24 |
Model requests per task | 5.6 | 0 |
Input tokens per task | 17,450 | 0 |
Replays are not much faster: most of a run is the browser and its human-paced input, not the model.
Skills are JSON files in ~/.clearcote-jet/skills (CLEARCOTE_JET_SKILLS). On the command line, --no-skill works a
task out step by step and saves nothing; over MCP, browse(..., reuse=False) does the same, and list_skills /
forget_skill manage them. A skill is for one goal from one page: another search term is another skill.
Results without a model
When a run finishes on a page with a list of results (search hits, products, listings), items holds its rows, read
from the page's layout with no model: a title, a link, a price (the sale price, not a struck-out one) and an image,
each found with a selector that works for most rows. The navigation, header and footer are left out.
Use it
Python
import asyncio
from clearcote_jet import Session, run
async def main():
session = await Session.launch(headless=False)
page = await session.new_tab("https://www.gov.uk/")
result = await run(session, "Find the date of the next bank holiday in England and Wales.", page=page)
print(result["status"], result["url"], result["usage"]["estimated_usd"])
for step in result["trace"]:
print(step["step"], step["kind"], step["action"], step["probability"], step["alternatives"])
print(result["markdown"])
await session.close()
asyncio.run(main())Command line
clearcote-jet run --url https://docs.python.org/3/ --goal "Find the entry for asyncio.gather" --show-cursor
# keep one browser open and send it tasks
clearcote-jet browser --port 9222
clearcote-jet run --cdp http://127.0.0.1:9222 --url https://news.ycombinator.com/ --goal "Open the top story's comments"Flags: --show-cursor (draw the mouse), --keep-open, --headless, --no-humanize, --profile DIR,
--confirm (stop before a click that can't be taken back), --no-skill (don't learn or replay; see below).
As a tool for your AI assistant (MCP)
{
"mcpServers": {
"clearcote-jet": {
"command": "clearcote-jet-mcp",
"env": { "TYPESAFE_API_KEY": "...", "CLEARCOTE_LICENSE_KEY": "..." }
}
}
}Your assistant gets these tools:
Tool | What it does |
| runs a whole task and returns every step, the rows of any list of results, and the page content. A task started from a |
| shows a tab exactly as Jet sees it: the numbered element table ( |
| one step by hand: |
| the tasks it has learned; forget one so it is worked out again |
| closes a tab |
When browse stops (blocked, needs_input), the assistant can look with snapshot, do a step with act, and hand
back to browse(goal, tab_id=...). act returns the new snapshot, so steps chain; if what a step depends on changed
since the snapshot, it does nothing and answers stale. Tabs stay open between calls, several tasks can run at once in
their own tabs (a second call on a busy tab answers busy), and the browser shuts itself down after a few idle minutes.
Screenshots are taken without touching the page's DOM.
Settings
Put them in .env (see .env.example) or the environment:
Variable | Default | What it does |
| required | key for the decision model (from console.typesafe.ai) |
|
| your Clearcote licence; without one the open build runs |
| unset | optional OpenAI-compatible model that writes typed values; without it Jet types words taken from your goal |
|
| browser profile; cookies and consent choices persist |
| unset | route through a proxy; timezone, language and location follow its exit IP |
| off |
|
| on |
|
| unset | attach to a browser that is already running |
|
| MCP server: close the browser after this long without calls |
|
| where learned tasks are kept, one JSON file each |
|
| rate used for the cost estimate |
Reading a result
Field | Meaning |
|
|
| every step: what was done, the text typed, its probability, the runner-up options and how long the decision took (replayed steps are marked |
| steps that were decided again because the page changed before Jet could act |
| decision-model tokens and requests, |
| the parts of the final page that answer the goal; for a run that did not finish, what was on screen |
| the rows of the final page's list of results (title, link, price, image), read without a model; empty if there is none |
| with skills: |
| with |
How it works
Look. Jet lists the controls a person could use right now: visible, enabled, not hidden behind a pop-up, including those inside embedded frames (iframes), same-site or cross-site. Password fields can be filled but are never read: Jet only sees whether one is empty, never what it holds. When it gets stuck on a page (the model finds no way forward there, or it has scrolled three times in a row), it also lists the links inside the page's closed menus: navigation dropdowns, menu panels and collapsed sections, each named by its menu (
Docs › Recommended settings). To click one, it opens the menu the way a person does, pointing at it (or clicking it, if pointing opens nothing), moves down inside the menu and clicks the link. It lists, too, what is scrolled out of view inside a box that scrolls on its own (a phone menu's panel, a dialog, a side list, a row of cards), which scrolling the page does not reach; it wheels over that box until the control is in view.Decide. One request to the decision model picks the kind of step and the control together, with a probability for every option. The model only chooses from that list; it never writes code or coordinates.
Act. Jet checks that the page has not changed, then moves the mouse along a curved path, clicks, and types key by key. Page reads run in a separate, isolated context, so the page's own scripts don't see them. On a long page it can also jump to words from your goal, through the page's own anchor when there is one.
Answer. When the goal is met, the page is turned into markdown, and the sections around where the run ended are scored against the goal. A list of results on the page is read row by row, from its layout, with no model.
Remember. A finished task is saved as a skill (see below), so the next run of it needs no model.
Examples
File | What it shows |
the smallest program: one goal, the steps, the answer | |
any goal in a visible window with the cursor drawn; the run is saved to | |
four tasks on four kinds of site, with a summary | |
tokens and dollars for your saved runs, projected per 1,000 tasks | |
records live runs and turns them into a captioned MP4 and GIF (needs ffmpeg) |
Known limits
On a long page it jumps to a section the goal names; when the goal names none, it may scroll rather than use a table of contents, and stops after 12 scrolls in a row with
blocked.Links that open a new tab are not followed, and closed shadow roots, canvas apps, file uploads and CAPTCHAs are not handled. Controls inside iframes are; scrolling inside an iframe is not.
Links inside closed menus, and controls scrolled away inside a scroll box, are offered only once a run is stuck on a page, and one level deep: a submenu inside a menu is not opened, nor a box scrolled away inside another box.
doneis the model's judgement. Check what matters.Jet only types words that appear in your goal (unless you configure a text model), so write the values into it.
Built for privacy, testing, research and lawful automation.
Development
pip install -e . pytest ruff
ruff check . && pytest tests -q # offline: no browser, no model calls
python tests/e2e/check_browser_layer.py # visible browser, scripted steps, a local test page
python tests/e2e/check_full_loop.py # the whole loop with a stand-in for the model
python tests/e2e/check_frames.py # controls inside same-site and cross-site iframes
python tests/e2e/check_mcp_tools.py # snapshot and act on a local form (one live step if a key is set)
python tests/e2e/check_mcp_stdio.py # the MCP server over stdio, as an assistant runs it
python tests/e2e/check_skills.py # learn, replay, repair, lists, the request shortcut, confirm and findLicense
MIT, see LICENSE.
Available Tools
2 toolsbrowseA
Carry out a task on a website in a Clearcote browser and return what the page says about it, as markdown.
Args: goal: What you want done, e.g. "Find the date of the next bank holiday in England and Wales" or "Search the Python documentation for asyncio.gather and open its entry". Write every value that has to be typed into the goal: typed text is taken from it. url: Start page. Required unless continuing an existing tab. tab_id: Continue in a tab returned earlier (a follow-up step, or after needs_input / blocked).
Status values: done, blocked, needs_input (the goal lacks a value a field requires: call again with the same tab_id and a goal that includes it), budget (step limit reached), error. The tab stays open so the user can see it; its tab_id is returned. The browser closes itself after a few idle minutes, which also closes its tabs.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | ||
| goal | Yes | ||
| tab_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does substantial work: it enumerates the status values (done, blocked, needs_input, budget, error), explains that the tab stays open and returns a tab_id, and that the browser self-closes after idle minutes, taking its tabs with it. It does not cover authentication, permissions, or rate/cost limits, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The content is front-loaded, with the core action in the first sentence and the arguments and status semantics following in scannable blocks. It is dense rather than padded, though the prose sections are slightly long for what remains after the status list.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is agentic and multi-turn, and the description covers the multi-turn mechanics (tab_id reuse, needs_input re-calls), the terminal states, and tab lifecycle. An output schema exists, and the description correctly avoids over-explaining the return payload while still noting it is markdown, so nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate and it does: goal is explained with concrete examples plus a critical operational detail that typed text is drawn from the goal; url is described as the start page required unless continuing a tab; tab_id is described as continuing a previously returned tab. Each parameter gains meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Carry out a task on a website'), the resource (Clearcote browser), and the return format ('return what the page says about it, as markdown'), which is far more than a restatement of the name. An agent can distinguish this from the sibling close_tab immediately, since browse performs work while close_tab only tears down a tab.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly states when to pass tab_id (a follow-up step, or after needs_input / blocked) and when url is required ('Required unless continuing an existing tab'), which maps usage onto the tool's own lifecycle. It does not explicitly name close_tab as the alternative for ending a session, so the routing guidance is strong but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
close_tabB
Close a tab that browse left open.
| Name | Required | Description | Default |
|---|---|---|---|
| tab_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It only says 'close a tab' and omits whether the action is destructive, whether it requires the tab to exist, and what happens on failure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. The only weakness is the grammatically informal 'that browse left open', which slightly muddies the object reference.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the tool is simple with one required parameter. However, the absence of annotations and any mention of tab_id or failure behavior leaves gaps for an agent invoking a mutation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must define the single required parameter, but it never mentions tab_id or its expected format/source. The parameter remains undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Close a tab') and ties it to the sibling tool browse, so the agent can roughly distinguish it from browsing. The phrasing 'that browse left open' is slightly awkward but conveys the intended object.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: use this to close a tab previously left open by browse. There is no explicit when-not-to-use guidance or mention of error conditions when the tab is already closed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.1.0- First observed
browse - First observed
close_tab
TDQS
Scored across 2 tools
browse is a high-level agentic web task tool; close_tab is a cleanup utility for the tabs it creates. Their purposes do not overlap, so an agent can easily select the correct tool.
Both names are imperative snake_case verbs, but browse is a bare verb while close_tab follows a verb_noun pattern. This is a minor inconsistency, though overall naming is predictable.
Two tools is thin for a browser automation server, even one with a high-level browse design. The count sits at the borderline where the interface may feel under-specified.
The surface covers starting a task, handling follow-ups, and closing tabs. Minor gaps exist, such as no way to list open tabs or inspect browser state, but the agentic browse tool mitigates most needs.
Maintenance
Related MCP Connectors
Run multi-step tasks in a real Chrome browser: persistent environments, live view, human takeover.
Web search, browser automation, scraping, crawling and CAPTCHA solving for AI agents.
- openhelmOAuthai.openhelm
Autonomous cloud agent tasks: real browser + your tools, structured evidence-backed results.
AI-powered browser automation — navigate, click, fill forms, and extract data from any website.
Related MCP Servers
- FlicenseNot gradedqualityCmaintenanceEnables AI assistants to control a browser through a set of tools, allowing them to perform web automation tasks like navigation, typing, clicking, and taking screenshots.-
- FlicenseNot gradedqualityCmaintenanceEnables AI assistants to control a browser through a set of tools, allowing them to perform web automation tasks like navigation, typing, clicking, and taking screenshots.-
- FlicenseNot gradedqualityCmaintenanceEnables AI assistants to control a browser through a set of tools, allowing them to perform web automation tasks like navigation, typing, clicking, and taking screenshots.-
- FlicenseNot gradedqualityCmaintenanceEnables AI assistants to control a browser through a set of tools, allowing them to perform web automation tasks like navigation, typing, clicking, and taking screenshots.-