Cloak Agent
OfficialAllows navigating GitHub repositories, viewing repository details such as star count and latest release information.
Allows searching Google and extracting results from the search page, such as finding and opening links to specific websites.
Uses OpenAI-compatible chat models to generate values for form fields during browsing tasks, enabling filling out forms with contextually appropriate text.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Cloak AgentSearch Google for the best restaurants in San Francisco and list the top 3."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
CloakBrowser Agent: a Jev-powered stealth browser agent
Give it a goal in plain language. TypeSafe Jev decides every step in ~0.3 s, CloakBrowser carries it out like a human, and you get the result back as markdown.
browse(goal="Search Google for 'CloakBrowser GitHub', open the CloakHQ/CloakBrowser repository on GitHub
from the results, and find how many stars it has and what the latest release is.",
url="https://www.google.com")
status: done
tab_id: t1 (still open)
url: https://github.com/CloakHQ/cloakbrowser/releases
title: Releases · CloakHQ/CloakBrowser
steps: 3 actions, 5 decisions, 12918 ms
actions taken (p = Jev's probability for the chosen target; runner-ups in brackets):
1. fill 'Išči' = 'CloakBrowser GitHub' p=1.0
2. click 'cloakbrowser github' p=0.59 ['Iskanje Google' p=0.25, 'cloakhq cloakbrowser github' p=0.13]
3. click 'CloakHQ/CloakBrowser' p=0.96 ['Open Išči' p=0.04]
...
<untrusted_page_content>
[Star 31.8k](...)
# Releases: CloakHQ/CloakBrowser
## Chromium v152.0.7977.82.1 — ... [Latest](https://github.com/CloakHQ/CloakBrowser/releases/latest)
...A real run, trimmed. Google's labels are in Slovenian because of where the test machine is.
It works as an MCP server (Claude Code, Cursor, Claude Desktop, any MCP client), a CLI, or a Python library. It runs on CloakBrowser, a stealth Chromium with human-like mouse and keyboard input.
Why Jev
Most browser agents ask a large language model to write the next action, which costs seconds per step. Jev is TypeSafe's first System One model. It doesn't generate text: you give it the current state plus typed questions, and it returns an answer with calibrated probabilities. A browser step is exactly that kind of question: which of these controls, doing what?
One request per step. Jev answers "which operation" and "which element" together in one request (speculative fan-out: a target question for each possible operation, and only the chosen one is used).
Fast. In our runs, Jev decisions averaged 0.28–0.43 s each, about 1 s for a whole search task. A text model is called only when a field needs typing.
Probabilities, not prose. Every decision comes with a distribution and a confidence, so the code can see when the model is unsure.
Nothing to parse. Answers are typed choices from options we built, so there's no free-form output to go wrong.
Jev ranks the result too. When the task is done, Jev scores every section of the final page against the goal, and only the relevant sections come back.
Related MCP server: Cloudflare Playwright MCP
How it works
goal ─► OBSERVE read the page → numbered table of the controls a user can actually reach
▲
│ DECIDE one Jev request: which operation (CLICK, TYPE_TEXT, SELECT, SCROLL, WAIT, DONE, BLOCKED)
│ and which element, answered together
│ TYPE_TEXT → a small text model writes only the value for that one field
│
└─ ACT human-like click / typing, after checking the page did not change
DONE → the page is turned into markdown, Jev scores each section against the goal, the best sections are returnedNo site-specific code. The same loop runs on every site and re-plans from the current page after every action. Cookie walls, popups and changed layouts are just more elements to choose from.
Model output never becomes code. The model picks from indices we built. It never writes selectors, coordinates or scripts.
Nothing is injected into the page. The page is read without adding anything to it that the site's scripts can see.
Only reachable controls are offered. Hidden, disabled, and covered elements (for example, behind a modal) aren't in the table.
A DONE answer isn't taken as proof. Check results that matter.
Requirements
Python 3.10+
A TypeSafe API key (Jev)
A key for any OpenAI-compatible chat model (OpenRouter, OpenAI, DeepSeek, …). It is used only to write field values.
A CloakBrowser license key for the latest stealth build. A free key takes one GitHub sign-in: run
cloakbrowser loginor go to cloakbrowser.dev/free. A free key allows one browser session at a time, and a paid key raises that limit. Without any key, the older build is used.
The browser binary downloads automatically on first use. Node is not needed.
Install
pip install cloakbrowser-agentThis installs the cloak-agent and cloak-agent-mcp commands, plus cloakbrowser (for cloakbrowser login).
For MCP clients you don't even need to install it: the configs below use uvx, which fetches and runs it on demand.
From source: git clone https://github.com/CloakHQ/CloakBrowser-Agent && cd CloakBrowser-Agent && pip install -e .
Configure
Variable | Required | Meaning |
| yes | Jev decisions |
| no | default |
| yes | OpenAI-compatible base URL, e.g. |
| yes | key for that endpoint |
| yes | model id, e.g. a small fast model |
| no | sent as |
| no | JSON object of extra request headers, if your provider needs any |
| recommended | CloakBrowser license key ( |
Use as an MCP server
Claude Code
claude mcp add cloak-agent --scope user \
-e TYPESAFE_API_KEY=... \
-e TEXT_MODEL_BASE_URL=https://openrouter.ai/api/v1 -e TEXT_MODEL=... -e TEXT_MODEL_API_KEY=... \
-e CLOAKBROWSER_LICENSE_KEY=cb_... \
-- uvx --from cloakbrowser-agent cloak-agent-mcpCursor / Claude Desktop (mcpServers in the client's config)
{
"mcpServers": {
"cloak-agent": {
"command": "uvx",
"args": ["--from", "cloakbrowser-agent", "cloak-agent-mcp"],
"env": {
"TYPESAFE_API_KEY": "...",
"TEXT_MODEL_BASE_URL": "https://openrouter.ai/api/v1",
"TEXT_MODEL": "...",
"TEXT_MODEL_API_KEY": "...",
"CLOAKBROWSER_LICENSE_KEY": "cb_..."
}
}
}
}Tools
browse(goal, url?, tab_id?) runs one whole task and returns:
status:done,blocked(no control can make progress),needs_input(the goal lacks a value a field needs),budget(step limit reached), orerrortab_id(the tab stays open), the final URL and title, and the number of actions, decisions and millisecondsevery step taken: what was clicked or typed, Jev's probability
pfor it, and the top runner-ups in brackets. A lowp, or a runner-up close behind, shows where the agent was unsure.stale retries, if any: steps re-decided because the page changed before acting, with what changed
the relevant page content as markdown, fenced as
<untrusted_page_content>
While a call runs, each step is also sent as a live progress notification (clients that show MCP progress display it).
close_tab(tab_id) closes a tab.
Working with tabs
Call | What happens |
| New tab, opens |
| Continues on the page tab |
| Navigates tab |
| Error: give a |
A tab_id stays valid until you close_tab it, close it in the browser, or the browser idles out. After that, browse answers status: error (tab 't1' is gone; open tabs: ...).
Multi-call examples
# 1. A task that needs values the goal did not include
browse(goal="Fill in the pizza order form and submit it.", url="https://httpbin.org/forms/post")
→ status: needs_input (No value in the goal for field: Customer name:) tab_id: t1
browse(goal="Fill in the pizza order form with customer name Jane Doe, telephone 555-0100, "
"email jane@example.com, size medium, and submit it.", tab_id="t1")
→ status: done (fills all four fields on the same form, clicks 'Submit order')
# 2. A follow-up step on the page the last task ended on
browse(goal="Search Google for 'CloakBrowser GitHub' and open the CloakHQ/CloakBrowser repository.",
url="https://www.google.com")
→ status: done tab_id: t2
browse(goal="Open the Issues tab of this repository and list the titles of the newest issues.", tab_id="t2")
→ status: done (1 action: click 'Issues')Tips:
Put every value the task needs into the goal (search terms, form values), because fields are filled from the goal.
Several
browsecalls can run at once. Each gets its own tab in the same browser.
Browser lifecycle
The browser starts on the first call, headed by default so you can watch it.
Tabs stay open after a task so you can see the result.
After
CLOAK_AGENT_IDLE_MINUTESwithout calls (default 5), the browser closes itself along with its tabs, and it relaunches on the next call. It also closes when the MCP server stops.An open browser holds one CloakBrowser session. On a free key (one session), close it or let it idle out before running another CloakBrowser script.
Variable | Default | Meaning |
|
| close the browser after this long without calls |
| off |
|
| on |
|
|
| persistent browser profile (cookies and consent choices survive) |
| unset | attach to an already running browser, e.g. |
A profile can only be open in one browser at a time. Give each instance its own CLOAK_AGENT_PROFILE.
Use from the command line
export TYPESAFE_API_KEY=... TEXT_MODEL_BASE_URL=https://openrouter.ai/api/v1 TEXT_MODEL=... TEXT_MODEL_API_KEY=...
export CLOAKBROWSER_LICENSE_KEY=cb_... # or run `cloakbrowser login` once
# one task, own browser (closes at the end; add --keep-open to inspect)
cloak-agent run --url https://en.wikipedia.org --goal "Open the Wikipedia article about Gödel's incompleteness theorems."
# debugging: keep one headed browser running, then run tasks against it
cloak-agent browser --port 9222
cloak-agent run --cdp http://127.0.0.1:9222 --url https://www.google.com --goal "Search Google for 'CloakBrowser'"--no-humanize switches to instant input.
Use from Python
Uses the same environment variables as above, including CLOAKBROWSER_LICENSE_KEY (or the key saved by cloakbrowser login).
import asyncio
from cloak_agent import Session, run
async def main():
session = await Session.launch(headless=False) # or: await Session.connect("http://127.0.0.1:9222")
page = await session.new_tab("https://en.wikipedia.org")
result = await run(session, "Open the Wikipedia article about Gödel's incompleteness theorems.", page=page)
print(result["status"], result["url"])
print(result["markdown"])
await session.close()
asyncio.run(main())Speed
Most of a humanized run is spent moving the mouse and typing at a human pace. That's on purpose, because behavior is scored on protected sites.
Task (single runs, one machine) | Humanized |
|
Google search → results | 11.4 s | 3.4 s |
Wikipedia search → open article | ~12 s | — |
Jev decisions take ~0.3 s each, roughly 1 s per task. Filling a field with the text model adds 0.5–2 s depending on the model.
Limits
Frames, closed shadow roots, canvas apps, file uploads, links that open new tabs, and CAPTCHAs aren't handled.
Native
<select>values are set programmatically, not through a mouse-driven dropdown.Autocomplete: after typing, it may pick a close suggestion instead of submitting the exact text.
Password fields are never read or offered to the model.
A valid action can still be the wrong one, and
donemeans the model saw the goal as met. Check results that matter.
Security
Page content is untrusted. It's returned fenced as
<untrusted_page_content>so the calling agent doesn't treat it as instructions.API keys stay in the server's environment and are only sent to the endpoints you configure.
For Google result links, the real target is resolved with one direct request per link, since Google hides it.
Credits
The observe → choose → act loop, the element snapshot and the Jev instructions are adapted from browser-use/jev-ultrafast (MIT, see LICENSE-jev-ultrafast).
Available Tools
2 toolsbrowseA
Complete a web task in a stealth browser and return the relevant page content as markdown.
Args: goal: The task in plain language, e.g. "Search Google for X and show the results" or "Find the price of the iPhone 17 on idealo.de". Include every value the task needs (search terms, form values), because fields are filled from the goal. url: Start page. Required unless continuing an existing tab. tab_id: Continue in a tab returned earlier (a follow-up step, or after needs_input / blocked).
Status values: done, blocked, needs_input (the goal lacks a value a field requires: call again with the same tab_id and a goal that includes it), budget (step limit reached), error. The tab stays open so the user can see it; its tab_id is returned. The browser closes itself after a few idle minutes, which also closes its tabs.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | ||
| goal | Yes | ||
| tab_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden and largely meets it: it enumerates the terminal statuses (done, blocked, needs_input, budget, error), discloses that the tab stays open for the user, that the tab_id is returned, and that the browser self-closes after idle minutes. It omits auth/permission requirements and what triggers blocked vs budget, so it is strong but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core behavior is front-loaded in the first sentence, followed by a clean Args block and a compact status list. It is slightly longer than strictly necessary (the Google/idealo examples plus the repeated tab-lifecycle note), but every element is functional.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be re-explained, and the description fills the remaining gaps: parameter intent, the multi-turn continuation protocol, and the failure/status states. An agent can complete a session correctly from this alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does for all three params: goal is 'the task in plain language' with the critical rule that every needed value must be embedded because fields are filled from the goal; url is the start page; tab_id continues a prior tab. The example goals clarify the expected granularity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and outcome: 'Complete a web task in a stealth browser and return the relevant page content as markdown.' The agent immediately knows this is an autonomous browsing/retrieval tool, and it is clearly distinct from the only sibling, close_tab.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the operative conditions well: url is required unless continuing an existing tab, tab_id is for follow-ups or after needs_input/blocked, and the needs_input flow tells the agent to re-call with the same tab_id plus the missing value. It does not explicitly contrast with close_tab, but that routing is obvious enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
close_tabB
Close a tab that browse left open.
| Name | Required | Description | Default |
|---|---|---|---|
| tab_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It only says 'close a tab' and omits whether the action is destructive, whether it requires the tab to exist, and what happens on failure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. The only weakness is the grammatically informal 'that browse left open', which slightly muddies the object reference.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the tool is simple with one required parameter. However, the absence of annotations and any mention of tab_id or failure behavior leaves gaps for an agent invoking a mutation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must define the single required parameter, but it never mentions tab_id or its expected format/source. The parameter remains undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Close a tab') and ties it to the sibling tool browse, so the agent can roughly distinguish it from browsing. The phrasing 'that browse left open' is slightly awkward but conveys the intended object.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: use this to close a tab previously left open by browse. There is no explicit when-not-to-use guidance or mention of error conditions when the tab is already closed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.1.0- First observed
browse - First observed
close_tab
TDQS
Scored across 2 tools
The two tools have clearly distinct purposes: 'browse' handles all web tasks, while 'close_tab' is solely for cleanup. There is no overlap or ambiguity in their intent.
Both names use snake_case, but 'browse' is a bare verb while 'close_tab' follows a verb_noun pattern. This minor inconsistency is still readable and predictable.
With only two tools, the set feels thin for a web browsing agent, even though the 'browse' tool is designed to handle many tasks via natural language goals. A few more utilities (e.g., list tabs) could improve scope.
The surface covers the core lifecycle: start a browsing task and close a tab. Missing are minor conveniences like listing open tabs or closing all tabs, but agents can work around these via tab_ids.
Maintenance
Related MCP Connectors
Real Chrome for agents: start a browser, read pages as numbered markdown, click, type, hand off.
AI-powered browser automation — navigate, click, fill forms, and extract data from any website.
Web search, browser automation, scraping, crawling and CAPTCHA solving for AI agents.
Provides cloud browser automation capabilities using Stagehand and Browserbase, enabling LLMs to i…
Related MCP Servers
- FlicenseNot gradedqualityCmaintenanceEnables AI assistants to control a browser through a set of tools, allowing them to perform web automation tasks like navigation, typing, clicking, and taking screenshots.-
- FlicenseNot gradedqualityCmaintenanceEnables AI assistants to control a browser through a set of tools, allowing them to perform web automation tasks like navigation, typing, clicking, and taking screenshots.-
- FlicenseNot gradedqualityCmaintenanceEnables AI assistants to control a browser through a set of tools, allowing them to perform web automation tasks like navigation, typing, clicking, and taking screenshots.-
- FlicenseNot gradedqualityCmaintenanceEnables AI assistants to control a browser through a set of tools, allowing them to perform web automation tasks like navigation, typing, clicking, and taking screenshots.-