JevBrow
Enables local browser automation through a Brave session, providing tools to open pages, observe visible content and actions, and perform clicks, typing, selection, scrolling, and tab closure via the Chrome DevTools Protocol.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@JevBrowOpen google.com and search for MCP servers"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
JevBrow
Pre-alpha 0.1 · Local browser tools for Codex, backed by your own Brave or Chrome session.
JevBrow connects to a browser through the Chrome DevTools Protocol (CDP) and exposes a small set of MCP tools. Codex reads the visible page, chooses an action from JevBrow's observed list, and receives a new observation after each action. The MCP server does not call a second model or require a model API key.
This is an early release. It handles common HTML controls and still needs supervision on consequential actions.
What it does
Opens a background tab owned by JevBrow in a browser that is already running locally.
Reports visible text, supported controls, action IDs, and a page fingerprint.
Clicks, types, selects a native option, scrolls, or waits only through an action ID from the latest observation.
Refuses stale actions and consumes an observation before browser input, so a failed mutation is never replayed blindly.
Closes its own tab without closing your other tabs or browser process.
Page text and labels are untrusted website content. The MCP client sees that content to decide what to do. Websites you open may be remote services, even though the browser connection and MCP process run locally.
Related MCP server: Chrome DevTools MCP
Requirements
Python 3.12 or newer and uv
Brave or Chrome with remote debugging available at
http://127.0.0.1:9222An MCP client such as Codex
Use a browser profile you intend to automate. CDP can interact with pages and sessions in that profile; keep the debugging port bound to loopback and do not expose it to your network.
Install for Codex
git clone https://github.com/ZnOw01/JevBrow.git
cd JevBrow
uv sync
python scripts/codex_config.py installThe installer adds a marked jevbrow block to ~/.codex/config.toml, using the actual checkout path, and keeps recovery snapshots alongside that file. It leaves unrelated settings in place. Restart Codex so it discovers the new MCP server. If your CDP port differs, change BU_CDP_URL in the marked block to your loopback endpoint.
To remove the integration:
python scripts/codex_config.py uninstallYou can also add the MCP server manually:
[mcp_servers.jevbrow]
command = "uv"
args = ["run", "--directory", "/absolute/path/to/JevBrow", "jevbrow"]
startup_timeout_sec = 30
[mcp_servers.jevbrow.env]
BU_CDP_URL = "http://127.0.0.1:9222"MCP workflow
Call
jevbrow_open(url)for anhttporhttpspage, orjevbrow_observe()for the existing JevBrow tab.Choose an
action_idfrom the returnedactionsand pass itsfingerprinttojevbrow_click,jevbrow_type_text,jevbrow_select, orjevbrow_scroll_or_wait.Inspect the observation returned by that tool before another action. After an error or stale fingerprint, observe again.
Call
jevbrow_close()when finished.
The API accepts no model-generated selectors, JavaScript, shell commands, or coordinates. Text entry is limited to an observed editable field and does not submit the form by itself. JevBrow exposes at most 250 page actions per observation and reports how many were omitted.
Optional standalone mode
The original autonomous runner is separate from MCP. It uses a local OpenAI-compatible model endpoint (for example, CLIPRox) to choose actions and write field values. MCP mode never reads these model credentials.
cp .env.example .env
# Fill in JEVBROW_MODEL_API_KEY and adjust the local endpoint if needed.
uv run --env-file .env jevbrow-run --url https://example.com --goal 'Describe the page.'For the local inspector, run uv run --env-file .env jevbrow-demo and open the loopback URL it prints. .env, browser recordings, caches, and local agent settings are ignored by Git.
Development
uv sync --group dev
uv run pytest -q
uv run ruff check .
uv run python scripts/check_guards.pyThe first two checks run offline. check_guards.py uses your local CDP browser and a temporary JevBrow-owned tab; it makes no model calls.
Scope and provenance
This project keeps the original jev_ultrafast Python package name. It derives from the MIT-licensed browser-use/jev-ultrafast prototype; the original copyright notice is preserved in LICENSE. The MCP integration and local Codex setup are the focus of this fork.
JevBrow supports common visible HTML and ARIA controls. Shadow roots, frames, canvas controls, uploads, pop-ups, nested scrolling, and complex keyboard widgets can block progress. A visible action can still be the wrong action for a task. The older performance reports and demo document the upstream TypeSafe-based prototype; they are historical evidence, not benchmarks of this pre-alpha MCP integration.
Available Tools
7 toolsjevbrow_clickA
Click one observed action with its fingerprint, then inspect the new state.
| Name | Required | Description | Default |
|---|---|---|---|
| action_id | Yes | ||
| fingerprint | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the operation is not read-only, not idempotent, and not destructive. The description adds only 'then inspect the new state,' a mild postcondition. It does not explain side effects, the role of the fingerprint in preventing stale actions, or failure behavior, so it does not meaningfully go beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence conveys the action, the target, and the expected follow-up. Every word earns its place; there is no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with an output schema, the description is adequate but not complete. It omits the provenance and semantics of the fingerprint, which is a nontrivial parameter, and leaves the agent to infer that action_id/fingerprint come from a prior observe call. A sentence clarifying that would close the gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must supply parameter meaning. It maps fingerprint to the observed action and implies action_id identifies the action, but it never explains where these values come from, why the fingerprint is required, or any format constraints. This is insufficient compensation for the missing schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: 'Click one observed action.' The phrase 'with its fingerprint' and the contrast with sibling tools (observe, type_text, select, scroll) make clear this tool activates an already-observed action rather than opening, typing, or selecting. It is unambiguous what the tool is for.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool should be used after observation ('one observed action') and for click-style interactions, but it never explicitly states when to prefer it over siblings such as select or type_text, nor does it state when not to use it. Usage context is implied by the verb rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jevbrow_closeAIdempotent
Close only the background tab opened and owned by Jev; does not close the user's other browser tabs.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true and destructiveHint=false, so the safety profile is covered. The description adds the key behavioral constraint: it only closes the Jev-owned background tab, not the user's other tabs. This is valuable context beyond the annotations. It doesn't describe what happens if no Jev-owned tab exists, but the idempotent hint covers the no-op case implicitly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the action and scope, and the exclusion of user tabs is stated immediately. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with an output schema and annotations covering idempotency and destructiveness, the description is nearly complete. It could mention what happens when no Jev-owned tab exists, but the idempotentHint=true annotation implies a safe no-op. The description is sufficient for an agent to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there is no parameter semantics burden. The description correctly implies the tool takes no arguments and acts on the Jev-owned tab. Baseline 4 for zero params is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool closes a background tab opened and owned by Jev, and explicitly distinguishes it from closing the user's other tabs. The verb 'close' plus the resource 'background tab opened and owned by Jev' is specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool: when you need to close a tab that Jev opened, and it explicitly says it does not close the user's other tabs. It doesn't name alternative sibling tools, but the exclusion of user tabs provides clear context. A 4 is appropriate because it gives clear context but no explicit when-not-to-use or alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jevbrow_observeARead-onlyIdempotent
Read the current URL, visible text, and supported actions before acting.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds value by specifying exactly what is observed (URL, visible text, supported actions), which is context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the verb and object, with no filler. Every word contributes to purpose and usage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only observation tool with a rich output schema and comprehensive annotations, the description supplies all necessary context. It tells the agent what to read and when to read it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has no parameters, so there is nothing for the description to clarify. Baseline 4 applies because the description cannot add parameter-level meaning where none exist.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a concrete action ('Read') and specific resources: the current URL, visible text, and supported actions. This immediately distinguishes it from the action-oriented sibling tools like jevbrow_click and jevbrow_type_text.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'before acting' gives explicit timing guidance, positioning this tool as the prerequisite step before any other action tool. It does not explicitly name alternatives, but the usage context is unmistakable given the sibling set.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jevbrow_openA
Open an http(s) URL in a Jev-owned background tab in the local browser.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate this is not read-only, not idempotent, not destructive, and touches the outside world. The description adds useful behavioral detail beyond annotations: the tab opens in the background, is Jev-owned, and is in the local browser.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler or redundancy. Every word contributes meaning, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with an output schema and safety annotations, the description covers the essential context: what to open, where, and how. Minor omissions like loading behavior are acceptable for such a simple action.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the parameter documentation burden. It adds the http(s) scheme constraint and clarifies the parameter is a URL, though the parameter name already conveys much of this. It does not discuss formatting, required prefixes, or edge cases.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action (open), a resource (http(s) URL), and a precise location (Jev-owned background tab in the local browser). It is immediately distinguishable from the sibling tools like click, type, scroll, and close.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies its use case: navigate to or open a URL in the browser. It does not explicitly state when to prefer it over alternatives or mention any exclusions, but for a simple open action the context is fairly clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jevbrow_scroll_or_waitA
Execute one observed scroll or wait action by action_id, then return a fresh observation.
| Name | Required | Description | Default |
|---|---|---|---|
| action_id | Yes | ||
| fingerprint | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnly=false, openWorld=true, and destructive=false. The description adds that the tool executes exactly one action and then returns a fresh observation, which is useful. It does not detail side effects, timing, or prerequisites, but the core behavioral profile is adequately covered by the combination of annotations and description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence with no filler or redundancy. It front-loads the action ('Execute one observed scroll or wait action by action_id') before stating the result, making it efficient and immediately scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description does not need to explain return values in detail. However, the required fingerprint parameter is left unexplained, and the origin of action_id/fingerprint is only implied by the word 'observed'. For a low-complexity tool this is a minor but real completeness gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description needs to compensate. It partially does by explaining action_id as the identifier of the observed action to execute, but fingerprint, a required parameter, is never mentioned or explained. This leaves a meaningful gap in parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('execute'), a specific resource ('one observed scroll or wait action'), and the expected result ('return a fresh observation'). This distinguishes it from the sibling tools like jevbrow_click, jevbrow_type_text, and jevbrow_select without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: use this tool when you have an action_id for a scroll/wait action and want an updated observation. It does not explicitly list exclusions or alternatives such as 'for clicks use jevbrow_click', but the scroll/wait scoping makes the appropriate use case clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jevbrow_selectA
Select one observed native dropdown option by action_id; never accepts model-generated selectors or values.
| Name | Required | Description | Default |
|---|---|---|---|
| action_id | Yes | ||
| fingerprint | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations present (readOnlyHint=false, destructiveHint=false, idempotentHint=false), the description adds value by disclosing a key behavioral constraint: it only accepts observed options and never model-generated inputs. This is critical context not captured in annotations, but it does not describe side effects or prerequisites beyond the implication of 'observed'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tightly written sentence with the core action front-loaded and the key constraint appended. Every word contributes meaning, and there is no redundant or filler content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema, so return values are already covered. The description explains the action and a critical constraint, but it leaves a required parameter ('fingerprint') undefined and provides no guidance on how to obtain the action_id or fingerprint, making it incomplete for reliable invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains that 'action_id' identifies the observed dropdown option, but it completely neglects 'fingerprint', leaving its meaning and role unclear. An agent cannot correctly supply both parameters based on the description alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Select') and resource ('one observed native dropdown option'), clearly identifying what the tool does. It also distinguishes itself from siblings like click and type_text by focusing on native dropdown options and requiring prior observation, which makes its purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: selecting an observed native dropdown option. It also provides an explicit exclusion ('never accepts model-generated selectors or values'), which helps the agent avoid misuse, though it does not mention alternative sibling tools by name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jevbrow_type_textA
Type into one observed field. Inspect the new state before acting again.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| action_id | Yes | ||
| fingerprint | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate this is a non-read-only, non-idempotent action, and the description goes beyond them by warning that the state will change and should be re-inspected before further actions. This adds useful, non-redundant behavioral context, though it does not cover failure modes or validation requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely compact: two short sentences that front-load the core operation and then deliver a meaningful behavioral warning. There is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although an output schema exists and annotations provide some safety profile, the description leaves the required action_id and fingerprint semantics unaddressed. With zero schema-level parameter documentation, the description is not sufficient for reliable invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain any of the three required parameters. 'Type' only weakly hints at 'text'; 'action_id' and 'fingerprint' are entirely unexplained, so an agent cannot infer how to populate the arguments correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Type') and a specific target ('one observed field'), making the operation unambiguous. It also distinguishes this tool from siblings like click/select/scroll by focusing on text entry into an observed element.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool (enter text into an observed field) and adds a useful follow-up instruction ('Inspect the new state before acting again'). However, it does not explicitly state when not to use it or name alternatives, leaving some usage decisions to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v0.1.0- First observed
jevbrow_click - First observed
jevbrow_close - First observed
jevbrow_observe - First observed
jevbrow_open - First observed
jevbrow_scroll_or_wait - First observed
jevbrow_select - First observed
jevbrow_type_text
TDQS
Scored across 7 tools
Each tool maps to a distinct browser interaction: open, observe, click, type, select, scroll/wait, and close. Even though scroll_or_wait and observe both return observations, their triggering actions are clearly differentiated and never overlap.
All tools share a consistent jevbrow_ prefix and use clear imperative verb phrases like open, observe, click, type_text, select, and close. The naming pattern is predictable and easy to reason about as a unified set.
Seven tools is well-scoped for a browser automation toolset, covering the full lifecycle from opening a tab to observing state, performing actions, and closing the tab. There is no redundancy or bloated surface area.
The set covers the essential browser automation loop: open, observe, click, type, select, scroll/wait, and close. Minor gaps like explicit back/forward navigation or refresh are absent, but they are likely accessible through observed clickable actions.
Maintenance
Related MCP Connectors
- TabfleetOAuthcom.tabfleet
Launch, inspect, control, and share isolated cloud browsers for your agents.
Provides cloud browser automation capabilities using Stagehand and Browserbase, enabling LLMs to i…
Run multi-step tasks in a real Chrome browser: persistent environments, live view, human takeover.
Undetectable cloud browser sessions for AI agents and scrapers. Navigate, extract, click, captcha.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables AI coding assistants to control and inspect a live Chrome browser for automation, debugging, performance analysis, network monitoring, and DOM interaction through Chrome DevTools Protocol.3,138,727 npmApache 2.0
- -licenseNot gradedqualityNot gradedmaintenanceEnables AI coding assistants to control and inspect a live Chrome browser for automated debugging, performance analysis, and web interaction. It leverages Puppeteer and Chrome DevTools to provide capabilities like network monitoring, console logging, and automated browser actions.-
- AlicenseNot gradedqualityDmaintenanceLets coding agents control and inspect a live Chrome browser via the Model-Context-Protocol, providing advanced browser debugging, performance insights, and reliable automation.5 npmApache 2.0
- AlicenseAqualityCmaintenanceEnables Grok CLI to control your already-running Brave browser via Chrome DevTools Protocol, allowing tab management, clicking, typing, navigation, and screenshots using your existing session and logins.14MIT