jev-browser-sidekick-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@jev-browser-sidekick-mcpSearch for 'MCP servers' on Google and click the first result"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
jev-browser-sidekick-mcp
An MCP server that drives a browser with Jev, TypeSafe's decision model. You write the steps. Jev chooses which control on the page carries out each one. The server does the clicking and the waiting, and it keeps the budgets.
The server has two tools. run_action does the browser work. use_jev_raw answers one typed question with no browser involved.
Jev named the package. Six candidates went into use_jev_raw with the download counts of every competing package and the trade-offs of each name written out. It picked this one at a probability of 0.66, against 0.17 for the runner-up, and it cost $0.000058 to ask. The maintainer had argued for a different name and lost.
Install agentic-playwright-mcp alongside it
Install agentic-playwright-mcp as well. This server runs at its best with that one beside it, and the two are built to be used together.
This server does not start its own browser. It attaches to the Chrome that agentic-playwright-mcp already runs, at http://127.0.0.1:9223. One Chrome, one profile, one set of cookies, shared by both servers and by your agent. A site you signed into through your Playwright tools is still signed in when a step runs here, and a cart this server filled is still there when you look at it yourself.
That shared session is what makes the pair stable. Pass a targetId from your Playwright tools into run_action, and pass the targetId values it returns back the other way, and both sides act on the same tab. When a step stops as blocked on a password or a captcha, the tab is one you already hold, so you can finish it in place and resume.
If no Chrome is listening, this server starts one of its own. It works, but nothing else can see that browser, so you lose the handoff and the sign-in you already had.
Related MCP server: MCP QA Demo Server
Install
Install both globally. agentic-playwright-mcp is a peer dependency, so npm pulls it in on its own, but naming it here also puts its commands on your PATH.
npm install -g agentic-playwright-mcp jev-browser-sidekick-mcpThat gives you four commands, a long form and a short form for each package.
Long | Short | What it runs |
|
| This server, and its |
|
| The browser this server attaches to |
npm writes a .cmd and a .ps1 shim next to each command on Windows, so all four work unchanged in PowerShell, in cmd, in Git Bash, and in WSL.
To skip the install and fetch on demand, use npx --yes with the full package name. npx resolves a package name rather than a command name, so npx --yes jev-bro does not work.
npx --yes jev-browser-sidekick-mcp doctorSave an API key
jev-bro setup --provider openrouter --api-key "$OPENROUTER_API_KEY"In PowerShell, use $env:OPENROUTER_API_KEY. In cmd, use %OPENROUTER_API_KEY%.
The command writes ~/.jev/.env. The MCP server, the CLI, and runAction all read that file.
To use a TypeSafe key instead, run this command.
jev-bro setup --provider official --api-key "$TYPESAFE_API_KEY"To add this server to ~/.cursor/mcp.json, pass --install-cursor. The command leaves the agentic-playwright-mcp entry alone.
To override ~/.jev inside this repo, copy .env.example to .env and set the same keys.
Check the setup
jev-bro doctorThe output includes jev decision: ok and a playwright: ok line with a targetId.
Add the servers
Put both in the host mcp.json. Neither needs a path, and npx fetches them on first run.
{
"mcpServers": {
"playwright": {
"command": "npx",
"args": ["--yes", "agentic-playwright-mcp"]
},
"jev": {
"command": "npx",
"args": ["--yes", "jev-browser-sidekick-mcp"]
}
}
}Start the playwright entry first, or just let the host start both. This server looks for that Chrome on every call, so the order only decides whether the first call attaches or opens its own browser.
Write the steps
Jev picks among labelled options and returns typed answers. It does not read plans and does not write text, so you write the plan and each step hands Jev one choice.
Step | What it does |
| Puts the words in the page's own search box |
| Chooses that entry out of a list |
| Reaches a place, such as the cart page |
| Presses the control carrying that label |
| Presses it until the page stops offering it |
| The same, for a delete control it finds itself |
| Hands the page's own words back to you |
| Answers "where am I" without the whole page |
Each step acts on one page. Use the words that appear on the screen, because Jev matches labels literally.
Adding one item to a cart is three steps, one per page.
["search colgate toothpaste", "open the best matching colgate toothpaste result", "click add to cart"]A single add colgate toothpaste to cart still runs, but it never leaves the results page, so it presses whatever on that page carries those words.
Each step starts a fresh Jev loop that sees only the live page, so one step never inherits another's page or history.
Repeat a press until the page stops offering it
clear cart is one use of a general rule. The server presses the same control until the page stops offering it. A step that names its own control keeps it, so keep clicking Load more works anywhere. A step that names a container instead, such as clear cart, falls back to whatever the page uses for delete or remove.
A repeat that gives up with controls still on the page returns partial, never completed. A repeat that presses nothing returns rejected with reason no_control, the same answer a single click gives, so an untouched cart never reads as a cleared one.
When the outcome already holds
A step that names a control the page no longer carries comes back rejected with reason no_control. For example, once an item is in the cart, a product page can replace "Add to cart" with "Go to cart", so click add to cart finds nothing and says so.
The server does not guess whether that means the work is already done. In testing, guessing it from the page marked steps finished that had never run. So a step that already looks satisfied still runs, and it still reports what happened. End the series with a read step and decide from the cart's own words.
Call run_action
Steps in tasks run in order on one tab.
Errands that do not depend on each other belong in groups, and they run at the same time, one tab each. Two sites, two accounts, or two separate searches are one call with two groups, never two calls. Use groups first, and fall back to a single series only when every step needs the page the step before it left. Series that share a targetId run one after another, because they share a tab.
{
"groups": [
{
"id": "cart",
"targetId": "TAB_A",
"expect": "subtotal",
"tasks": [
"open the cart page",
"clear cart",
"search colgate toothpaste",
"open the best matching colgate toothpaste from the results",
"click add to cart"
]
},
{
"id": "other",
"targetId": "TAB_B",
"tasks": ["open example.com"]
}
]
}Read the result
The result has is_finished, targetIds, groups, status, summary, and a handoff when something stopped. Each group has its own tasks, steps, and counts. snapshot appears only when you set returnSnapshot, and usage, elapsedMs, and per-step ms only when you set debug.
Every step has its own status.
Status | What it means |
| The step did what it said |
| It did some of the work and stopped with more to do |
| The page answered no. Read |
| The page wants something only you can give. Read |
| Every step ran, and |
| The budget or the clock ran out |
| An earlier step in the series stopped this one |
A group's status is worst-case across its steps, so a group where eight of ten steps completed still reads as rejected. Read counts for what actually happened.
{ "status": "rejected", "counts": { "completed": 12, "rejected": 1 } }Later steps depend on earlier ones, so a step that does not complete ends its series, and the rest come back as skipped naming the step that stopped them. Set noFail on a series whose steps are independent, and it runs them all.
Why a step was turned down
| Meaning |
| Nothing on that page does what the step named |
| The list held no entry matching what the step named |
| Out of stock, sold out, or not delivered here |
| The page offers a different route, such as other sellers |
| The page is not about the wanted thing |
| The page had not finished loading |
| Returned as |
Prove the run worked
A completed status is Jev's claim. Set expect to the text that proves it, and the server reads the final page for that text. Matching ignores case and spacing.
{ "tasks": ["open the cart page"], "expect": "subtotal (3 items)" }Pick text that only the finished state produces. Subtotal (3 items) works. A product name does not, because shops repeat product names in recommendation rails, so the text matches even on an empty cart. The result quotes the words either side of the match in proof, so you can see which it matched.
{ "verified": true, "proof": "…All Carts Subtotal (3 items): ₹509.00 Proceed to Buy…" }The check runs only when every step completed. A series that stopped reports proof: not checked, rather than claiming the text was missing from a page it never reached.
Pick up a stopped run
When a series stops early, the result includes a handoff.
{
"resumable": true,
"targetId": "TAB_A",
"url": "https://www.amazon.in/ap/signin",
"stoppedAt": "click add to cart",
"status": "blocked",
"reason": "credentials",
"remaining": ["open the cart page"],
"recent": [{ "step": 4, "operation": "CLICK", "detail": "clicked e42" }]
}The tab is still open and you already share it, so you can finish the step through your own Playwright tools, ask the user, or call run_action again with that targetId and the remaining steps.
Pages that block a step
On every step, the server asks Jev whether the task can happen on this page at all. When Jev says no, the server asks why, then stops the step with blocked and a reason.
| The page is |
| A sign-in wall |
| Asking for a password or a one-time code |
| Asking the user to prove they are a human |
Those three come back as blocked with a handoff, because you own the same tab and can still act. This server never fills a password, a one-time code, or a captcha itself. Type the value in the shared tab, hand it to the user, or stop. Pass ordinary strings such as an email or a postcode in values.
Other reasons, such as unavailable or wrong_page, come back as rejected. Those are the page's own answer, not something you can unblock.
Keep a call short
A call returns after its own timeout, which defaults to 90 seconds. On expiry you get a normal result holding the steps that finished and a handoff naming the rest, so you resume by calling again with that targetId and the remaining steps.
A call your MCP client drops is the case worth avoiding. The browser work still happens, you never see the result, and running the same steps again does them twice. Keep timeoutMs under your client's own transport timeout and resume instead of asking for one long call.
Token usage and timing
usage and elapsedMs appear only when you set debug, alongside tracePath. Every group and task then carries its own ms.
usage sums what each API response reported. Nothing in it is estimated.
Field | Source |
|
|
|
|
| The two above, added |
| Jev calls made |
| Calls to the text model that fills a field Jev cannot write |
| Present only when a provider returns a price |
TypeSafe returns tokens and no price, so costUsd is absent on a direct TypeSafe key and present through OpenRouter.
decisions: 0 is normal and not a failure. A search, a destination, and a control labelled exactly what the step said all resolve without a judgment call, so a run made only of those asks Jev nothing. A recent two-site run of 26 steps cost 19 decisions and about $0.001.
Ask Jev without a browser
use_jev_raw sends state and typed questions straight to Jev. Use it whenever a decision has more than one defensible answer and you are about to pick on instinct: which fix to do first, which name to ship when each has a real trade-off, whether a draft meets a bar you can write down, whether a step is risky enough to stop and ask the user.
You get a probability for every option and a confidence, so a close call reads as close. Ask every question you have in one call, because they share the state and answer in parallel. Three questions over a page of state run about a tenth of a cent.
{
"state": { "bug": "Checkout throws on an empty cart.", "fixes": { "guard": "...", "schema": "..." } },
"questions": {
"pick_fix": {
"type": "choice",
"instructions": "Which fix in `fixes` removes the cause of `bug` rather than hiding it?",
"criteria": { "guard": "Adds a runtime check.", "schema": "Makes the broken state unrepresentable." }
}
}
}The answer has the chosen option, a probability for every option, a confidence, and the tokens the API counted. The server publishes a jev://raw-decisions resource with the question types, the rules for writing criteria, and the size limits. Read it before the first call.
Nothing says the state has to be about code. If you are the sort of person who stands in the cereal aisle for ten minutes, your sidekick will happily take that one too. Give it the four job offers, the three flat listings, the cat names, or what to cook tonight, write down what you actually care about in criteria, and it tells you which one it likes and how sure it is. A tenth of a cent for a friend who never answers "I don't know, what do you want to do" is a fair trade. It already picked this package's name, and it was right.
What the server already handles
Leave these out of the plan. The server waits for loads, follows a link that opens its own tab, recovers element refs that went stale between the snapshot and the click, and skips invisible controls that carry real labels.
Finding a control and choosing it are separate. When a step names a control, the server collects every control on the page carrying those words, including ones drawn as plain text with no accessibility role, and Jev picks one or answers that none of them fits. When nothing fits, the step comes back rejected or blocked instead of pressing something at random.
Jev reads text only, so this server works from the page's own text and takes no screenshots. A control drawn without text is found by its DOM text instead.
A step that opens an entry is done only once the page shows that entry. A click that navigates, opens its own tab, or draws an overlay showing the entry all count. A click that leaves the page exactly as it was does not count, so the step tries something else rather than reporting success.
Traps on real sites
Each of these cost real runs in testing. Shopping sites found them, but nothing about them is particular to shopping.
One tab, two storefronts. A site can run more than one storefront in the same tab, and its search box keeps you in whichever one the tab is already in. Results in the wrong storefront open overlays rather than their own pages, so a click step finds nothing. Pass a startUrl that names the storefront you want. For example, Amazon Fresh sits beside the main Amazon store in one tab.
One site, two collections. A site can keep more than one cart, list, or queue, and show a count that spans all of them, so a collection you just emptied still reads as full. The page that lists everything usually shows the other collection read-only, with a link in place of the per-row controls, so a clear step there presses nothing and returns rejected with no_control. Open the collection that owns the items first. For example:
["open the cart page", "click Go to Fresh Cart", "clear cart"]A page that offers a different route. When the usual action is unavailable, a page often puts another control in its place. The step returns rejected with no_control or other_route. That is the page's answer, not a failure to retry. For example, a listing with no default offer shows "See All Buying Options" where "Add to cart" would be.
A near miss in place of a match. A search can answer with something close rather than the thing you named, most often when the real item sits behind a sign-in or in another storefront. Expect no_match, or check the substitution with a read step.
Debug a run
Set debug on a call to write every browser call and Jev decision to a JSONL file. The result names it in tracePath.
To trace every run whatever the caller passes, add --debug to the server command.
{
"mcpServers": {
"jev": {
"command": "npx",
"args": ["--yes", "jev-browser-sidekick-mcp", "--debug"]
}
}
}Traces land in ~/.jev/traces, one file per run.
Run one goal from the CLI
jev-bro run "Open example.com and click More information" --snapshotAdd --expect to check the final page, --task to write a series, and --groups-json for parallel groups.
Call runAction from code
npm install jev-browser-sidekick-mcpimport { runAction } from "jev-browser-sidekick-mcp";
const result = await runAction({
tasks: ["open the cart page", "clear cart"],
expect: "your cart is empty",
});
console.log(result.status, result.verified, result.usage.totalTokens);Available Tools
2 toolsrun_actionA
Drive a browser through steps you write, with Jev choosing on each page. Covers several independent errands in one call through groups, one tab each, so two sites never need two calls. Returns a status per step, a handoff when one stops, and the tokens the API reported. Read this server's instructions for the step shapes.
| Name | Required | Description | Default |
|---|---|---|---|
| goal | No | A single step. Use tasks for a series. | |
| debug | No | Record every browser call and Jev answer to the JSONL file named in tracePath. | |
| tasks | No | Steps in order on one tab, one primitive action each. | |
| expect | No | Text that proves the run worked. Read off the final page. | |
| groups | No | Independent errands, run at the same time, one tab each. Use this whenever the work splits across two sites, accounts, or searches instead of calling the tool twice. | |
| noFail | No | Run the rest of the series even after a step that did not complete. | |
| values | No | Text the steps may type, such as an email. Never a password or a one-time code. | |
| groupId | No | Tab group to open a new tab in. | |
| maxSteps | No | Page actions the whole call may spend. Default 40. | |
| startUrl | No | Address to open before the first step. | |
| targetId | No | Tab to work in. Omit to open one, which is returned in targetIds. | |
| timeoutMs | No | Ceiling for the whole call. Default 90000. On expiry the call returns normally with the steps that finished and a handoff naming the rest. Keep it under your own client's transport timeout, because a dropped call still leaves the browser work done. | |
| contextPaths | No | Files the steps may read. | |
| returnSnapshot | No | Include the final page's accessibility tree. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, but it does disclose key behavioral traits: it returns status per step, handoff on stop, and tokens reported by API. It also mentions the timeout behavior, which is significant. However, it does not elaborate on error handling, side effects, or what happens to browser state after the call. The description partially covers the safety/side-effect profile, but lacks depth on failure modes and browser state changes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is reasonably concise, with three sentences covering the main functionality, return values, and a pointer to more detailed instructions. It front-loads the purpose and then covers the grouping and return information. There is some redundancy with the schema, but the description is efficient and not bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (14 parameters, nested objects, no output schema), the description is insufficiently complete. It mentions return values (status, handoff, tokens) but does not detail the exact structure of those returns, which is critical for an agent to parse the results. It also relies on 'Read this server's instructions' for step shapes, which is an external reference and may not be available to the agent. The description needs more detail on return format and step syntax to be fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every parameter in detail. The description adds minimal extra meaning beyond what the schema provides, such as the concept of 'independent errands' and 'one tab each'. It does not add syntax or format details beyond the schema definitions. Baseline 3 is appropriate given high coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states that the tool drives a browser through user-written steps with Jev choosing actions, and covers multiple independent errands via groups and tabs. It distinguishes from the sibling 'use_jev_raw' by emphasizing the step-driven and group orchestration, though it does not explicitly name the sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides strong usage guidance: it recommends using groups for independent errands to avoid multiple calls, and mentions that steps should be single actions. It implicitly sets expectations for when to use the tool, but doesn't explicitly contrast with 'use_jev_raw' or state when not to use it. The guidance is clear for typical browser automation scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
use_jev_rawA
Ask Jev one typed question, or several, with no browser involved.
Use it whenever a decision has more than one defensible answer and you are about to pick on instinct. Jev gives a probability for every option and a confidence, so a close call reads as close, and a clear one reads as clear. Anything you would otherwise settle by coin flip and call judgment belongs here.
Decisions worth handing over: which of these fixes to do first, which name or design to ship when each has a real trade-off, whether this text meets a bar you can write down, which of two error messages a stranger understands faster, whether a step is risky enough to stop and ask the user, how to rank a list of candidates, which of two readings of an ambiguous request the user meant.
Ask every question you have in one call. They share the state, they answer in parallel, and each extra question costs its own tokens and almost no extra time. Three questions over a page of state run about a tenth of a cent, so the cost is rarely the reason to skip it.
The answer gives the chosen option, the probability of every option, a confidence, and the tokens the API counted. Read the jev://raw-decisions resource before the first call for the question shapes and the criteria rules.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Pin a Jev version. Default is the configured model. | |
| state | Yes | What every question is about. Plain text, or JSON whose fields the instructions name in backticks. | |
| questions | Yes | Questions keyed by an id you pick. Answers come back under the same ids. Ask them all in one call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and it delivers: no browser, parallel answering, shared state, per-question token cost, and a concrete output shape (chosen option, probabilities, confidence, API-counted tokens). It also points to a prerequisite resource to read before first use.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average but each block earns its place: purpose, when-to-use examples, batching behavior, cost, output, and prerequisite resource. It is front-loaded with the core purpose and organized into readable paragraphs, though a few example bullets could be trimmed without loss.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a nested 3-parameter tool with no output schema, the description is unusually complete: it covers the output fields, cost, parallelism, and a mandatory reference resource. An agent has enough information to invoke the tool correctly without needing additional context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents all three parameters, so the baseline is 3. The description adds value by explaining that questions in one call share the same state, run in parallel, and are billed per question, which clarifies how the 'questions' object should be used beyond schema wording.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific action and resource: 'Ask Jev one typed question, or several, with no browser involved.' It explains what Jev returns (probabilities, confidence) and gives concrete decision examples, so an agent understands exactly what the tool is for and can distinguish it from browser-action siblings like run_action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: 'Use it whenever a decision has more than one defensible answer' and provides a detailed list of worthwhile decisions. It does not spell out when not to use it or name an alternative tool, so it misses the when-not/alternatives part of the top band.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.1.2- First observed
run_action - First observed
use_jev_raw
TDQS
Scored across 2 tools
run_action is entirely about browser automation while use_jev_raw is explicitly a non-browser decision tool. Their purposes and outputs do not overlap, so there is no risk of an agent selecting the wrong tool.
Both names are snake_case and verb-first, which is mostly consistent. However, run_action follows a clean verb_noun pattern while use_jev_raw mixes a product name and an adjective, making it slightly less predictable.
Two tools is minimal and sits at the thin end of acceptable. Each tool covers a broad area, but the server still feels sparse for something positioned as a browser sidekick.
run_action appears to consolidate multi-step browser workflows and returns statuses and handoffs, while use_jev_raw covers decision support. The main gap is a lack of separate low-level browser inspection or control tools, but the step-based design likely lets agents work around it.
Maintenance
Related MCP Connectors
Run multi-step tasks in a real Chrome browser: persistent environments, live view, human takeover.
Hosted real Google Chrome MCP with per-user persistent state. Navigate, click, type, screenshot.
Automate cloud Chrome—navigate, click, type, screenshot, run code, record screen video
Headless browser primitives for AI agents when sites need real JS rendering.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceEnables plain-English browser automation via an MCP server, allowing agents to run objectives or test suites in a real browser without selectors or scripts.63 npm2Apache 2.0
- FlicenseNot gradedqualityCmaintenanceEnables natural language-driven browser automation by launching a real browser, performing interactions, and generating test artifacts like feature/step/page files.-
- AlicenseAqualityBmaintenanceEnables AI agents to delegate complex web browsing goals to a real Chrome instance driven by Jev, completing tasks end-to-end in ~300ms per decision and returning only the final result.11830 npm5MIT
- FlicenseNot gradedqualityBmaintenanceEnables Claude Code and Claude Desktop to control your own Chrome browser, with Jev deciding each click, keystroke, and scroll.5-