jev-explorer
It runs a browser-using AI agent server (Jev Explorer) that lets an agent or user drive live browser sessions toward a goal, get evidence-backed answers, and manually intervene — without sending full page text back.
Open an isolated browser session for manual sign-in or initial page inspection (jev_open).
Delegate a complete browsing objective; the server navigates, chooses actions, and returns a compact handoff with verbatim source evidence (jev_explore).
Continue the same objective in the same live browser by adding missing values, facts, or supervisor notes (jev_continue).
Inspect the live session without any model call: get a readable handoff, page targets, and screenshots (jev_inspect).
Take explicit manual actions — click, type, press keys, select, check, scroll, navigate — using refs from inspection (jev_act).
Close the session at any time; local evidence files remain on disk (jev_close).
Receive clear statuses (answered, needs_value, needs_decision, blocked, spent) that tell you what to do next.
Rely on safety constraints: no arbitrary JavaScript, no model-written selectors, password-like values scrubbed from traces, but clicks can have real-world effects — guard via the goal.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@jev-explorerGo to example.com and find the minimum cancellation notice for video appointments."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Jev Explorer
Give a browser goal in plain words. Get back the answer, the exact text that proves it, and a browser that stays open.
An AI agent that browses a website spends its own memory on page text, clicks and retries. This server does the browsing somewhere else, with a small cheap model, and sends back only what matters.
The one rule
The code reads. Jev chooses.
Jev picks one option from a list. It cannot plan, and it cannot write text. So the code reads the page and builds the list, and Jev points at one item.
Every message to the model has the same shape:
state: the goal, and short facts the code collected
question: one fixed sentence, written by hand
options: a numbered list, built from the pageThere are three question sentences in the whole system, and they live in one
file, src/jev/questions.ts:
Which of these options moves toward the goal?
Which field should receive the value named X?
Which of these text blocks answers this question?
Nothing else is ever asked. A question about a calendar would force a question about every widget on earth, so there is none.
Related MCP server: jev-browser
What one step costs
One step of exploration sends one message: the question that chooses the action.
A test asserts this. If a change makes a step cost two messages, the test fails.
Findings
The model never writes the answer. It points at a block of text on the page, and the code copies that block word for word, with the page address. That is why the evidence can be trusted.
Safety
Nothing stops a click. The run presses whatever moves toward the goal, including save, send, buy, book and delete. The goal is the only brake, so write one that stops before the button you do not want pressed.
The model never supplies a selector, and no tool runs code on the page.
Values whose name looks like a password or a token are removed from the trace.
The model's judgement is evidence, not a security boundary. The boundary is the tool contract.
Install
Node 24 or later.
npm install
npx playwright install chromium
npm run build
cp .env.example .envPut your key in .env as TYPESAFE_API_KEY.
Connect it
{
"mcpServers": {
"jev-explorer": {
"command": "node",
"args": ["/absolute/path/to/jev-explorer/dist/mcp/main.js"],
"env": { "TYPESAFE_API_KEY": "..." }
}
}
}The four tools
Tool | What it does |
| Send a goal, questions and values. Send the same |
| Look at the live page. This never calls the model. |
| Do one step yourself, using a control reference from |
| Close the browser. The evidence files stay on disk. |
What comes back
Every call returns the same short report: a status, what is needed next, the findings with their source text, the last few steps, how much was used, and the paths to the evidence on disk. The page itself never goes back to the caller.
The status is one of five, and each one says what to do:
Status | What to do |
| Read the findings. |
| Send the missing value. |
| Take over: no offered step moves toward the goal. |
| Sign in, solve a challenge, or give up. |
| Raise the budget, or narrow the goal. |
The skill
skills/jev-explorer/SKILL.md teaches an agent how to use these four tools:
what to send, what the five statuses mean, and how to carry a session on.
Link it into your own skills folder:
ln -s "$PWD/skills/jev-explorer" ~/.claude/skills/jev-explorerTests
npm test # local pages, scripted answers, no key and no cost
npm run test:live # the real model against the same local pagesWhat this does not do
It does not handle calendars, autocomplete lists or any other named widget with its own code. A popup adds its items to the option list, and the same question picks one. This is simpler, and less reliable, on purpose.
It does not promise to finish any workflow on any website.
It does not measure itself against other browser agents.
Licence
Apache-2.0.
Available Tools
6 toolsjev_actADestructive
Take one explicit native browser action in the retained session, with no Jev call. Use a current ref from jev_inspect. For typing provide the text; passwords are not returned in the handoff. Then continue the exploration with jev_continue. This tool does not generate selectors or execute arbitrary JavaScript.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | ||
| sessionId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate destructive and non-read-only behavior, but the description adds critical context: it takes exactly one action, cannot generate selectors, and passwords are not returned. This goes beyond the annotations by explaining execution constraints. Minor gap: does not detail side effects of each action type, but annotations cover the destructive nature.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, each packed with essential information: the action scope, the ref requirement, the text instruction, the password caveat, and the follow-up with jev_continue. The negative constraints are front-loaded, and every sentence adds value without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of the action parameter with multiple command types, the description covers the common usage pattern and constraints clearly. The schema fully defines the action variants, so the description need not repeat them. A minor gap is not explaining when to use jev_act vs. jev_explore, but the description's focus on 'explicit' action and the follow-up with jev_continue implies the distinction.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and there are no parameter descriptions in the schema. The description compensates by explaining the 'action' parameter's structure (commands like typing, clicking) and the need for a current ref from jev_inspect. It also mentions the text field for typing. However, it does not enumerate all possible commands, relying on the schema's oneOf, but the description provides enough context for typical use.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states the tool performs one native browser action in the retained session, listing examples like typing and noting it does not generate selectors or execute JavaScript. This clearly distinguishes it from exploration tools like jev_inspect and jev_continue. The verb 'take' with resource 'native browser action' is specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit instructions: use a current ref from jev_inspect, provide text for typing, and then continue with jev_continue. It also warns that passwords are not returned in the handoff, guiding the agent on what to expect. This covers when to use and how to sequence with siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jev_closeADestructive
Close the browser session and retain its local evidence. Closing does not undo business effects. A closed session cannot be resumed by replaying its trace.
| Name | Required | Description | Default |
|---|---|---|---|
| sessionId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the operation as destructive and open-world. The description adds meaningful behavioral detail beyond those hints: local evidence is retained, business effects persist, and replaying the trace will not resume the session. This gives the agent a much fuller picture of consequences.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly scoped sentences with no filler. The core action is front-loaded, followed only by high-value consequences. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter tool with no output schema, the description adequately covers purpose, side effects, and irreversibility. It omits details like idempotency or error behavior, but these are not essential given the simplicity and the annotation coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions sessionId or how to identify the target session. The schema's property name and UUID format carry the meaning, but the description does nothing to compensate for the missing documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Close the browser session.' It adds distinctive constraints—evidence is retained, business effects are not undone, and the trace cannot be replayed—which clearly separates it from sibling tools like jev_open and jev_continue.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives practical context: closing is final and does not reverse side effects, and a closed session cannot be resumed via replay. This implies when not to use the tool, although it does not explicitly name alternative tools such as jev_continue.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jev_continueADestructive
Continue the same objective in the same live browser. Add missing values, a supervisor note or sourced facts. Keeps discoveries and prior effects and does not replay the whole workflow. Input values should be supplied data, not guessed answers.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | ||
| facts | No | ||
| values | No | ||
| sessionId | Yes | ||
| effectResolution | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as non-read-only, open-world, and destructive, so the description does not need to restate safety traits. It adds valuable behavioral nuance by explaining that discoveries and prior effects are preserved and that the workflow is not replayed. The instruction that input values should be supplied data, not guesses, is additional meaningful guidance. No contradiction with the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is composed of three short, purposeful sentences. The primary purpose is front-loaded, and each sentence contributes meaning: what the tool does, what state it preserves, and what kind of input is acceptable. There is no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with five parameters, nested objects, no output schema, and destructive annotations, the description covers the main purpose and persistence semantics but leaves gaps. sessionId and effectResolution are not semantically explained, and there is no guidance on expected results or how the continuation appears to the caller. The annotations handle the safety profile, making this adequate but not comprehensive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must carry the burden. It does explain the intended meaning of values, note, and facts, and clarifies that values should be real supplied data. However, it does not explain the required sessionId parameter or the meaning of effectResolution (confirmed/not_applied), leaving those under-specified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states a specific verb and resource: continue the same objective in the same live browser, adding missing values, a note, or sourced facts. The mention of not replaying the whole workflow distinguishes it from a fresh start, and it contrasts naturally with sibling tools like jev_open, jev_explore, or jev_act that perform new actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides useful context: use this when continuing an existing objective rather than starting over, and when you need to supply concrete data, notes, or facts. It even warns against guessed answers. However, it does not explicitly name sibling alternatives or state clear when-not-to-use conditions, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jev_exploreADestructive
Delegate a complete browser exploration objective to Jev. Supply questions whose answers are unknown; answers include verbatim source evidence. Returns a compact handoff, not the full DOM. The browser stays open for inspection and continuation. ready_for_review is evidence for the caller to review, not certified business correctness. Use only within the user-authorized scope. Default submission guards are conservative heuristics, not a read-only security boundary.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | ||
| headed | No | ||
| values | No | ||
| maxCalls | No | ||
| maxSteps | No | ||
| maxTokens | No | ||
| objective | Yes | ||
| questions | No | ||
| sessionId | No | ||
| timeoutMs | No | ||
| allowCommit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate destructive, open-world, non-read-only behavior, and the description honestly reinforces this with 'Default submission guards are conservative heuristics, not a read-only security boundary.' It also discloses return behavior, session continuation, and review semantics beyond the annotations. No contradiction exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but tightly packed, with the core purpose front-loaded in the first sentence. Each sentence contributes a distinct point: purpose, question semantics, output format, continuation, and safety caveats. It loses a point only because several important caveats are compressed into a single running paragraph.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 11 parameters, no output schema, nested objects, and destructive annotations, the description is only partially complete. It explains the high-level contract and safety posture but omits operational details such as how to configure session continuation, what `values` or `allowCommit` affect, and the exact structure of the returned handoff.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only loosely covers `objective` and `questions`. The other nine parameters, including `url`, `values`, `sessionId`, `timeoutMs`, `maxCalls`, `maxSteps`, `maxTokens`, `allowCommit`, and `headed`, receive no semantic explanation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly names the action: 'Delegate a complete browser exploration objective to Jev' and explains the core mechanism of supplying unknown questions with verbatim evidence. It also distinguishes itself from simpler browser tools by noting it returns 'a compact handoff, not the full DOM' and that the browser stays open. However, it does not explicitly name or differentiate sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use this tool: for complete exploration objectives and questions whose answers are unknown. It also warns to stay 'within the user-authorized scope' and clarifies that ready_for_review is evidence, not certified correctness. It lacks an explicit when-not-to-use or alternative selection compared to siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jev_inspectARead-only
Inspect a retained session with no model call. Default output stays compact. Request targets only when taking over; use nextOffset to page through targets. Request screenshot explicitly to receive the image. Refs become stale when the page changes.
| Name | Required | Description | Default |
|---|---|---|---|
| offset | No | ||
| targets | No | ||
| sessionId | Yes | ||
| screenshot | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint and destructiveHint annotations, the description discloses that no model call is made, that output is compact by default, and that refs become stale when the page changes. These are valuable behavioral traits that an agent needs to know and are not inferable from the schema or annotations alone. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each packed with information: purpose, default behavior, parameter usage, and a warning. No filler. The key differentiator ('no model call') is front-loaded, and the warning about stale refs is placed at the end as a natural caveat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only inspection tool with no output schema, the description covers the essential usage: what it does, how to get targets and screenshots, and the lifecycle caveat about refs. The schema handles parameter bounds (offset max 1000) and defaults. Nothing critical is missing for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must add meaning. It does so for three of the four parameters: 'offset' is implied by 'use nextOffset to page through targets' (though the name mismatch could confuse), 'targets' is explained as 'only when taking over', and 'screenshot' as 'Request screenshot explicitly'. sessionId is self-evident from the parameter name and format. This compensates for the missing schema descriptions, but the 'nextOffset' wording introduces a slight inconsistency with the actual parameter name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Inspect') and resource ('a retained session') and immediately clarifies a key distinction: 'no model call'. This differentiates it from sibling tools like jev_act or jev_continue. The phrase 'retained session' makes the target unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear conditions for optional parameters: 'Request targets only when taking over' and 'Request screenshot explicitly'. It also gives a pagination hint with 'use nextOffset to page through targets' and warns about stale refs. However, it does not explicitly name sibling tools or state when to prefer this over jev_explore or jev_continue, so the guidance is context-specific rather than comparative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jev_openADestructive
Open an isolated browser session without inference. Use for authentication or to inspect the initial page before delegating a goal. A headed session can be operated by the user. The session stays alive until closed or idle expiry.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| headed | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations (readOnlyHint=false, destructiveHint=true), the description adds meaningful behavior: 'The session stays alive until closed or idle expiry' and 'A headed session can be operated by the user.' It also clarifies it performs no inference, which is a key behavioral trait. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences, front-loaded with the core action and then usage and behavior. No filler; every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with no output schema, the description covers purpose, usage, and session lifecycle. It mentions idle expiry and user interaction, which are critical for correct invocation. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must compensate. It indirectly explains url by describing use cases involving a page, and explicitly explains headed: 'A headed session can be operated by the user.' While it doesn't detail url format (already provided by schema's uri format), it adds useful context for both parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action: 'Open an isolated browser session without inference,' with explicit use cases ('for authentication or to inspect the initial page'). It distinguishes from siblings by highlighting 'without inference,' which contrasts with likely inference-based tools like jev_inspect or jev_explore.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says when to use it: 'for authentication or to inspect the initial page before delegating a goal,' and implicitly excludes inference tasks via 'without inference.' However, it does not name sibling alternatives directly, leaving some inference about which tools to use instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.1.0- First observed
jev_act - First observed
jev_close - First observed
jev_continue - First observed
jev_explore - First observed
jev_inspect - First observed
jev_open
TDQS
Scored across 6 tools
Each tool has a clearly distinct role in the browser session lifecycle: open creates, inspect observes, explore delegates, continue extends, act manually intervenes, and close terminates. The descriptions reinforce these boundaries well, so an agent should have little trouble selecting the right tool.
All tool names follow the same 'jev_' prefix plus a single lowercase verb pattern. This is fully consistent and makes the action of each tool predictable from its name.
Six tools is well-scoped for a browser exploration server. Each tool covers a necessary part of the workflow without redundancy or bloat.
The tool set forms a complete lifecycle: open a session, inspect it, delegate exploration, continue with additional input, act manually when needed, and close the session. There are no obvious dead ends or missing operations for the stated purpose.
Maintenance
Related MCP Connectors
- openhelmOAuthai.openhelm
Autonomous cloud agent tasks: real browser + your tools, structured evidence-backed results.
Provides cloud browser automation capabilities using Stagehand and Browserbase, enabling LLMs to i…
Real Chrome for agents: start a browser, read pages as numbered markdown, click, type, hand off.
Stealth web browser for agents: search, fetch, click, download and type in persistent MCP sessions.
Related MCP Servers
- AlicenseAqualityBmaintenanceEnables AI agents to delegate complex web browsing goals to a real Chrome instance driven by Jev, completing tasks end-to-end in ~300ms per decision and returning only the final result.11830 npm5MIT
- FlicenseNot gradedqualityBmaintenanceEnables Claude Code and Claude Desktop to control your own Chrome browser, with Jev deciding each click, keystroke, and scroll.5-
- AlicenseAqualityBmaintenanceEnables coding agents to drive a real Chrome browser to complete natural-language web tasks and return evidence for verification.2MIT
- AlicenseAqualityBmaintenanceEnables Cursor to drive Google Chrome through the jev-browser-use click loop, with Cursor supplying tasks, typing text, and verifying results while Jev chooses the next click.51MIT