jev-browser-sidekick-mcp
This server lets you drive a browser through Jev-chosen steps and ask Jev typed questions without a browser.
Use
run_actionto search, open results/pages, click labelled controls, read pages, and execute ordered steps.Run independent errands in parallel
groupsacross tabs, withtargetId,startUrl, andgroupId.Use loop tasks to repeat steps until a page condition holds, with
maxRounds.Get per-step statuses (
completed,partial,rejected,blocked, etc.), reasons, counts, and handoffs for resumable runs.Verify final outcomes with
expect; inspect proof ornot checked.Share a Chrome session with
agentic-playwright-mcp; paused tabs can be resumed viatargetId.Control budgets via
maxSteps,timeoutMs,noFail; debug viadebug,returnSnapshot,contextPaths.Use
use_jev_rawto ask choice/score/noul questions over shared state, returning probabilities, confidence, and token usage.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@jev-browser-sidekick-mcpSearch for 'MCP servers' on Google and click the first result"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
jev-browser-sidekick-mcp
An MCP server that drives a browser with Jev, TypeSafe's decision model. You write the steps. Jev chooses which control on the page carries out each one. The server does the clicking and the waiting, and it keeps the budgets.
The server has two tools. run_action does the browser work. use_jev_raw answers one typed question with no browser involved.
Jev named the package. Six candidates went into use_jev_raw with the download counts of every competing package and the trade-offs of each name written out. It picked this one at a probability of 0.66, against 0.17 for the runner-up, and it cost $0.000058 to ask. The maintainer had argued for a different name and lost.
Install agentic-playwright-mcp alongside it
Install agentic-playwright-mcp as well. This server runs at its best with that one beside it, and the two are built to be used together.
This server does not start its own browser. It attaches to the Chrome that agentic-playwright-mcp already runs, at http://127.0.0.1:9223. One Chrome, one profile, one set of cookies, shared by both servers and by your agent. A site you signed into through your Playwright tools is still signed in when a step runs here, and a cart this server filled is still there when you look at it yourself.
That shared session is what makes the pair stable. Pass a targetId from your Playwright tools into run_action, and pass the targetId values it returns back the other way, and both sides act on the same tab. When a step stops as blocked on a password or a captcha, the tab is one you already hold, so you can finish it in place and resume.
If no Chrome is listening, this server starts one of its own. It works, but nothing else can see that browser, so you lose the handoff and the sign-in you already had.
Related MCP server: MCP QA Demo Server
Install
Install both globally. agentic-playwright-mcp is a peer dependency, so npm pulls it in on its own, but naming it here also puts its commands on your PATH.
npm install -g agentic-playwright-mcp jev-browser-sidekick-mcpThat gives you four commands, a long form and a short form for each package.
Long | Short | What it runs |
|
| This server, and its |
|
| The browser this server attaches to |
npm writes a .cmd and a .ps1 shim next to each command on Windows, so all four work unchanged in PowerShell, in cmd, in Git Bash, and in WSL.
To skip the install and fetch on demand, use npx --yes with the full package name. npx resolves a package name rather than a command name, so npx --yes jev-bro does not work.
npx --yes jev-browser-sidekick-mcp doctorSave an API key
jev-bro setup --provider openrouter --api-key "$OPENROUTER_API_KEY"In PowerShell, use $env:OPENROUTER_API_KEY. In cmd, use %OPENROUTER_API_KEY%.
The command writes ~/.jev/.env. The MCP server, the CLI, and runAction all read that file.
To use a TypeSafe key instead, run this command.
jev-bro setup --provider official --api-key "$TYPESAFE_API_KEY"To add this server to ~/.cursor/mcp.json, pass --install-cursor. The command leaves the agentic-playwright-mcp entry alone.
To override ~/.jev inside this repo, copy .env.example to .env and set the same keys.
Check the setup
jev-bro doctorThe output includes jev decision: ok and a playwright: ok line with a targetId.
Add the servers
Put both in the host mcp.json. Neither needs a path, and npx fetches them on first run.
{
"mcpServers": {
"playwright": {
"command": "npx",
"args": ["--yes", "agentic-playwright-mcp"]
},
"jev": {
"command": "npx",
"args": ["--yes", "jev-browser-sidekick-mcp"]
}
}
}Start the playwright entry first, or just let the host start both. This server looks for that Chrome on every call, so the order only decides whether the first call attaches or opens its own browser.
Write the steps
Jev picks among labelled options and returns typed answers. It does not read plans and does not write text, so you write the plan and each step hands Jev one choice.
Step | What it does |
| Puts the words in the page's own search box |
| Chooses that entry out of a list |
| Reaches a place, such as the cart page |
| Presses the control carrying that label |
| Runs that step once a round until the page shows those words |
| The same, with the condition read off the label |
| The same, for a delete control it finds itself |
| Hands the page's own words back to you |
| Answers "where am I" without the whole page |
Each step acts on one page. Use the words that appear on the screen, because Jev matches labels literally.
Adding one item to a cart is three steps, one per page.
["search colgate toothpaste", "open the best matching colgate toothpaste result", "click add to cart"]A single add colgate toothpaste to cart still runs, but it never leaves the results page, so it presses whatever on that page carries those words.
Each step runs against the live page. The decision also sees the series motive and a short steps_done line for each finished task, so Jev can refuse a step that earlier work already covered. Groups running in parallel share nothing, so each one sees only its own motive and its own finished steps.
Repeat steps until the page says to stop
A loop runs its body once a round, then Jev reads the live page and answers one question: does the condition hold yet? Clearing a cart is one use of it. Any page that hands back one item at a time needs the same shape.
["repeat click remove until the cart is empty"]The condition ends the loop, not a control disappearing. A site that redraws its list between rounds offers no controls for a moment, and reading that as "finished" reports an untouched cart as cleared. Write the condition as something the page shows, such as the cart is empty, rather than done.
A step that names its own control derives its condition, so keep clicking Load more and clear cart still work as written. Between rounds the server waits for the page's own scripts rather than a fixed pause, because a row that a site deletes over the network lands whenever its request comes back.
For a body of more than one step, pass a loop object in tasks.
{
"tasks": [
"open the cart page",
{ "loop": { "tasks": ["click remove", "click confirm"], "until": "the cart is empty", "maxRounds": 20 } },
"read the cart"
]
}No loop runs forever. Five things end one: the condition, the round ceiling (maxRounds, 12 by default and never above 50), the call's own deadline, the step budget, and two rounds that change nothing. The step reports completed when the page showed the condition, partial when a ceiling stopped it with work left, and rejected with no_control when the body pressed nothing at all. Every loop step carries the rounds it ran.
When the outcome already holds
A step that names a control the page no longer carries comes back rejected. For example, once an item is in the cart, a product page can replace "Add to cart" with "Go to cart", so click add to cart finds nothing.
Two different things cause that, and the caller acts differently on each: the page cannot do the step at all, or the page already shows the step's outcome. The server does not guess from the URL or the title, because that guess marked steps finished that had never run. Jev reads the page instead, and an outcome already in place comes back as rejected with reason already_done.
The series carries on past an already_done step, because the ground the next step stands on is there, whoever put it there. The group still reports rejected, so a run reads honestly, and verified says where it landed.
Call run_action
Steps in tasks run in order on one tab.
Errands that do not depend on each other belong in groups, and they run at the same time, one tab each. Two sites, two accounts, or two separate searches are one call with two groups, never two calls. Use groups first, and fall back to a single series only when every step needs the page the step before it left. Series that share a targetId run one after another, because they share a tab.
{
"groups": [
{
"id": "cart",
"targetId": "TAB_A",
"expect": "subtotal",
"tasks": [
"open the cart page",
"clear cart",
"search colgate toothpaste",
"open the best matching colgate toothpaste from the results",
"click add to cart"
]
},
{
"id": "other",
"targetId": "TAB_B",
"tasks": ["open example.com"]
}
]
}Read the result
The result has is_finished, targetIds, groups, status, summary, and a handoff when something stopped. Each group has its own tasks, steps, and counts. snapshot appears only when you set returnSnapshot, and usage, elapsedMs, and per-step ms only when you set debug.
Every step has its own status.
Status | What it means |
| The step did what it said |
| It did some of the work and stopped with more to do |
| The page answered no. Read |
| Sign-in, password, captcha, or a proxy interstitial. Read |
| The final step completed, and |
| The Jev provider or the network failed. Read |
| The budget or the clock ran out |
| An earlier step in the series stopped this one |
A group's status is worst-case across its steps, so a group where eight of ten steps completed still reads as rejected. Read counts for what actually happened.
{ "status": "rejected", "counts": { "completed": 12, "rejected": 1 } }Later steps depend on earlier ones, so a step that does not complete ends its series, and the rest come back as skipped naming the step that stopped them. Set noFail on a series whose steps are independent, and it runs them all.
One reason is exempt. A step that came back rejected with already_done did not stop anything, because what the next step needs is already on the page.
Why a step was turned down
| Meaning |
| A |
| An |
| The control is gone because the page already shows the outcome |
| Out of stock, sold out, or not delivered here |
| The page offers a different route, such as other sellers |
| The page is not about the wanted thing |
| The page had not finished loading |
| Returned as |
A standing none choice that wins uses the same reasons: no_control for click or press, no_match for pick.
Prove the run worked
A completed status is Jev's claim. Set expect to the text that proves it, and the server reads the final page for that text. Matching ignores case and spacing.
{ "tasks": ["open the cart page"], "expect": "subtotal (3 items)" }Pick text that only the finished state produces. Subtotal (3 items) works. A product name does not. Shops repeat product names in recommendation rails, so that text matches even on an empty cart. The result quotes the words either side of the match in proof, so you can see which it matched.
{ "verified": true, "proof": "…All Carts Subtotal (3 items): ₹509.00 Proceed to Buy…" }The check runs when the final step of the series completed, including under noFail after an earlier rejection. A series whose final step never ran reports proof: not checked, rather than claiming the text was missing from a page it never reached.
Pick up a stopped run
When a series stops early, the result includes a handoff.
{
"resumable": true,
"targetId": "TAB_A",
"url": "https://www.amazon.in/ap/signin",
"stoppedAt": "click add to cart",
"status": "blocked",
"reason": "credentials",
"remaining": ["open the cart page"],
"recent": [{ "step": 4, "operation": "CLICK", "detail": "clicked e42" }]
}The tab is still open and you already share it, so you can finish the step through your own Playwright tools, ask the user, or call run_action again with that targetId and the remaining steps.
Pages that block a step
These reasons come back as blocked with a handoff, because you own the same tab and can still act.
| The page is |
| A sign-in wall |
| Asking for a password or a one-time code |
| Asking the user to prove they are a human |
Finding the words and believing them are separate. The harness spots the words a sign-in wall or a challenge uses, and Jev reads the page and confirms the page is really demanding one. A news story titled "Solving a corn puzzle with CP-SAT" carries the same word a challenge does, and stopping a series on that costs more than the check saves.
This server never fills a password, a one-time code, or a captcha itself. Type the value in the shared tab, hand it to the user, or stop. Pass ordinary strings such as an email or a postcode in values.
A missing control is not blocked. It is rejected with no_control or no_match, as in the reason table above.
When the provider fails
A fault in the Jev API or the network is not a page answer. The step returns error with a reason that names the cause. An HTTP proxy that returns HTML with status 200 is the exception: that comes back as blocked with reason proxy_interstitial, because only you can clear the proxy.
| Typical cause |
| HTTP 429 |
| HTTP 401 or 403 |
| HTTP 402, or a credits or quota message |
| HTTP 404 for a model or endpoint |
| HTTP 5xx |
| Timeout, DNS failure, or connection refused |
| HTTP 200 with a |
| Non-JSON body or empty answers |
The summary includes the HTTP status and a short provider message. The series stops even when noFail is set. handoff.resumable stays true when a tab exists, so you resume with that targetId and the remaining steps instead of redoing work that already finished.
Keep a call short
A call returns after its own timeout, which defaults to 90 seconds. On expiry you get a normal result holding the steps that finished and a handoff naming the rest, so you resume by calling again with that targetId and the remaining steps.
A call your MCP client drops is the case worth avoiding. The browser work still happens, you never see the result, and running the same steps again does them twice. Keep timeoutMs under your client's own transport timeout and resume instead of asking for one long call.
Token usage and timing
usage and elapsedMs appear only when you set debug, alongside tracePath. Every group and task then carries its own ms.
usage sums what each API response reported. Nothing in it is estimated.
Field | Source |
|
|
|
|
| The two above, added |
| Successful Jev calls. A failed decide does not count |
| Calls to the text model that fills a field Jev cannot write |
| Present only when a provider returns a price |
TypeSafe returns tokens and no price, so costUsd is absent on a direct TypeSafe key and present through OpenRouter.
decisions: 0 is normal and not a failure. A search, a destination, and a control labelled exactly what the step said all resolve without a judgment call, so a run made only of those asks Jev nothing. A recent two-site run of 26 steps cost 19 decisions and about $0.001.
Ask Jev without a browser
use_jev_raw sends state and typed questions straight to Jev. Use it whenever a decision has more than one defensible answer and you are about to pick on instinct: which fix to do first, which name to ship when each has a real trade-off, whether a draft meets a bar you can write down, whether a step is risky enough to stop and ask the user.
You get a probability for every option and a confidence, so a close call reads as close. Ask every question you have in one call, because they share the state and answer in parallel. Three questions over a page of state run about a tenth of a cent.
{
"state": { "bug": "Checkout throws on an empty cart.", "fixes": { "guard": "...", "schema": "..." } },
"questions": {
"pick_fix": {
"type": "choice",
"instructions": "Which fix in `fixes` removes the cause of `bug` rather than hiding it?",
"criteria": { "guard": "Adds a runtime check.", "schema": "Makes the broken state unrepresentable." }
}
}
}The answer has the chosen option, a probability for every option, a confidence, and the tokens the API counted. The server publishes a jev://raw-decisions resource with the question types, the rules for writing criteria, and the size limits. Read it before the first call.
Nothing says the state has to be about code. If you are the sort of person who stands in the cereal aisle for ten minutes, your sidekick will happily take that one too. Give it the four job offers, the three flat listings, the cat names, or what to cook tonight, write down what you actually care about in criteria, and it tells you which one it likes and how sure it is. A tenth of a cent for a friend who never answers "I don't know, what do you want to do" is a fair trade. It already picked this package's name, and it was right.
What the server already handles
Leave these out of the plan. The server waits for loads, follows a link that opens its own tab, recovers element references that went stale between the snapshot and the click, and skips invisible controls that carry real labels.
Finding a control and choosing it are separate. When a step names a control, the server collects every control on the page carrying those words, including ones drawn as plain text with no accessibility role. Jev picks one of them, or the standing none option when none fits. A pick step does the same over entry candidates that name the subject. When nothing fits, the step returns rejected instead of pressing something else.
Jev reads text only, so this server works from the page's own text and takes no screenshots. A control drawn without text is found by its DOM text instead.
A step that opens an entry is done only once the page shows that entry. A click that navigates, opens its own tab, or draws an overlay showing the entry all count. A click that leaves the page exactly as it was does not count, so the step tries something else rather than reporting success.
Traps on real sites
Each of these cost real runs in testing. Shopping sites found them, but nothing about them is particular to shopping.
One tab, two storefronts. A site can run more than one storefront in the same tab, and its search box keeps you in whichever one the tab is already in. Results in the wrong storefront open overlays rather than their own pages, so a click step finds nothing. Pass a startUrl that names the storefront you want. For example, Amazon Fresh sits beside the main Amazon store in one tab.
One site, two collections. A site can keep more than one cart, list, or queue, and show a count that spans all of them, so a collection you just emptied still reads as full. The page that lists everything usually shows the other collection read-only, with a link in place of the per-row controls, so a clear step there presses nothing and returns rejected with no_control. Open the collection that owns the items first. For example:
["open the cart page", "click Go to Fresh Cart", "clear cart"]A page that offers a different route. When the usual action is unavailable, a page often puts another control in its place. The step returns rejected with no_control or other_route. That is the page's answer, not a failure to retry. For example, a listing with no default offer shows "See All Buying Options" where "Add to cart" would be.
A near miss in place of a match. A search can answer with something close rather than the thing you named, most often when the real item sits behind a sign-in or in another storefront. Expect no_match, or check the substitution with a read step.
Debug a run
Set debug on a call to write every browser call and Jev decision to a JSONL file. The result names it in tracePath.
To trace every run whatever the caller passes, add --debug to the server command.
{
"mcpServers": {
"jev": {
"command": "npx",
"args": ["--yes", "jev-browser-sidekick-mcp", "--debug"]
}
}
}Traces land in ~/.jev/traces, one file per run.
Run one goal from the CLI
jev-bro run "Open example.com and click More information" --snapshotAdd --expect to check the final page, --task to write a series, and --groups-json for parallel groups.
Call runAction from code
npm install jev-browser-sidekick-mcpimport { runAction } from "jev-browser-sidekick-mcp";
const result = await runAction({
tasks: ["open the cart page", "clear cart"],
expect: "your cart is empty",
});
console.log(result.status, result.verified, result.usage.totalTokens);Available Tools
2 toolsrun_actionA
Drive a browser through steps you write. Jev chooses on each page. Use for search, open, click, and read on live pages, and for several independent errands in one call through groups, one tab each, so two sites never need two calls. Returns a status per step, a handoff when one stops, and the tokens the API reported. Step shapes and status tables are in this server's instructions.
| Name | Required | Description | Default |
|---|---|---|---|
| goal | No | Single step. Use tasks for a series. | |
| debug | No | Record every browser call and Jev answer to the JSONL file named in tracePath. | |
| tasks | No | Steps in order on one tab, one action each. An entry may be a loop. | |
| expect | No | Text that proves the run worked. Checked on the final page. | |
| groups | No | Independent errands, run at the same time, one tab each. Use whenever the work splits across two sites, accounts, or searches, instead of two calls. | |
| noFail | No | Run the rest of the series after a step does not complete. Endpoint faults stop it anyway. | |
| values | No | Text the steps may type, such as an email. Never a password or a one-time code. | |
| groupId | No | Tab group to open a new tab in. | |
| maxSteps | No | Page actions the whole call may spend. Default 40. | |
| startUrl | No | Address to open before the first step. | |
| targetId | No | Tab to work in. Omit to open one, returned in targetIds. | |
| timeoutMs | No | Ceiling for the whole call. Default 90000. On expiry the call returns with the steps that finished plus a handoff naming the rest. Keep it under your client transport timeout: a dropped call still leaves the browser work done. | |
| contextPaths | No | Files the steps may read. | |
| returnSnapshot | No | Include the final page accessibility tree. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden but only partially discharges it: it discloses Jev-driven per-page decisions, per-step status returns, a handoff on stoppage, and API token reporting. It omits safety/profile context such as authentication needs or side-effect/destructive footprint, which matters for a live-browser automation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose is front-loaded in the first sentence and the four sentences each carry information (operations, parallel grouping, return shape, pointer to instructions). The trailing phrase 'so two sites never need two calls' is slightly redundant with the preceding clause.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 14-parameter nested tool with no output schema, the description supplies the missing return structure (per-step status, handoff, tokens) and defers step/status tables to server instructions. That covers the main gaps, though it could say more about failure and auth behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 14 parameters (including groups, tasks, timeoutMs, noFail) are already documented in the schema; that sets the baseline at 3. The description adds the conceptual framing of groups as parallel one-tab-per-errand series, but no syntax or constraints beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening line gives a concrete verb+resource ('Drive a browser through steps you write') and enumerates the operations it covers (search, open, click, read). It is clearly distinguishable in role, though it never names the sibling tool use_jev_raw to sharpen the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states clear use contexts and even an alternative-to-two-calls rule ('several independent errands in one call through groups... so two sites never need two calls'). It stops short of explicit when-NOT-to-use guidance or a direct comparison with use_jev_raw.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
use_jev_rawA
Ask Jev one typed question, or several. No browser.
Use whenever a decision has more than one defensible answer and you are about to pick on instinct. Jev returns a probability for every option plus a confidence, so a close call reads close and a clear one reads clear. Anything you would settle by coin flip and call judgment belongs here.
Worth handing over: which fix to do first, which name or design to ship when each has a real trade-off, whether this text meets a bar you can write down, which of two error messages a stranger reads faster, whether a step is risky enough to stop and ask the user, how to rank candidates, which reading of an ambiguous request the user meant.
Ask every question in one call. They share the state, they answer in parallel, and each extra question costs its own tokens and almost no extra time. Three questions over a page of state run about a tenth of a cent.
The answer gives the chosen option, the probability of every option, a confidence, and the tokens the API counted. Question shapes and criteria rules: jev://raw-decisions
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Pin a Jev version. Default is the configured model. | |
| state | Yes | What every question is about. Plain text, or JSON whose fields the instructions name in backticks. | |
| questions | Yes | Questions keyed by an id you pick. Answers come back under the same ids. Ask them all in one call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries the full burden, and it delivers real behavioral context: answers return a chosen option, per-option probabilities, a confidence, and API-counted tokens; questions share state, run in parallel, and cost roughly a tenth of a cent for three questions. Auth/prerequisites and failure modes are not covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and the core output contract are front-loaded, and the cost/parallelism note is useful. The eight-item "worth handing over" list is longer than needed and pushes length up, but each block still carries usable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with nested question objects, no output schema, and no annotations, the description covers what the response contains and points to jev://raw-decisions for question shapes and criteria rules. It leaves undefined what Jev itself is and any rate-limit or error behavior, but is otherwise sufficient to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so the baseline is 3, but the description adds semantics the schema cannot: every question should be asked in one call, they share state, and batching adds tokens but almost no latency. That meaningfully shapes how the parameters are populated beyond their field descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb and resource ("Ask Jev one typed question, or several") and rules out a browser variant, and the second paragraph makes concrete what Jev actually does: returns probabilities per option plus a confidence. It is clear what the tool is for, though the only listed sibling (run_action) is never distinguished.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger ("whenever a decision has more than one defensible answer and you are about to pick on instinct") plus a concrete list of qualifying scenarios and an implicit exclusion (anything settled mechanically by coin flip is only worth it because it's judgment). No alternative tool is named, so it stops short of full when-not/alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.2.1- Changed
run_action19 fields changed- changed
Input schema / properties / expect / descriptionPrevious value: -"Text that proves the run worked. Read off the final page."New value: +"Text that proves the run worked. Checked on the final page." - changed
Input schema / properties / goal / descriptionPrevious value: -"A single step. Use tasks for a series."New value: +"Single step. Use tasks for a series." - changed
Input schema / properties / groups / descriptionPrevious value: -"Independent errands, run at the same time, one tab each. Use this whenever the work splits across two sites, accounts, or searches instead of calling the tool twice."New value: +"Independent errands, run at the same time, one tab each. Use whenever the work splits across two sites, accounts, or searches, instead of two calls." - changed
Input schema / properties / groups / items / properties / expect / descriptionPrevious value: -"Text that proves this series worked. Read off the final page."New value: +"Text that proves this series worked. Checked on the final page." - changed
Input schema / properties / groups / items / properties / goal / descriptionPrevious value: -"A single step, when tasks is omitted."New value: +"Single step. Use when tasks is omitted." - changed
Input schema / properties / groups / items / properties / id / descriptionPrevious value: -"Label for this series, echoed back in the result."New value: +"Label for this series. Echoed back in the result." - changed
Input schema / properties / groups / items / properties / noFail / descriptionPrevious value: -"Run the rest of this series even after a step that did not complete."New value: +"Run the rest of this series after a step does not complete. Endpoint faults stop it anyway." - changed
Input schema / properties / groups / items / properties / tasks / descriptionPrevious value: -"The steps, in order."New value: +"Steps in order." - added
Input schema / properties / groups / items / properties / tasks / items / $refAdded value: +"#/properties/tasks/items" - removed
Input schema / properties / groups / items / properties / tasks / items / minLengthRemoved value: -1 - removed
Input schema / properties / groups / items / properties / tasks / items / typeRemoved value: -"string" - changed
Input schema / properties / noFail / descriptionPrevious value: -"Run the rest of the series even after a step that did not complete."New value: +"Run the rest of the series after a step does not complete. Endpoint faults stop it anyway." - changed
Input schema / properties / returnSnapshot / descriptionPrevious value: -"Include the final page's accessibility tree."New value: +"Include the final page accessibility tree." - changed
Input schema / properties / targetId / descriptionPrevious value: -"Tab to work in. Omit to open one, which is returned in targetIds."New value: +"Tab to work in. Omit to open one, returned in targetIds." - changed
Input schema / properties / tasks / descriptionPrevious value: -"Steps in order on one tab, one primitive action each."New value: +"Steps in order on one tab, one action each. An entry may be a loop." - added
Input schema / properties / tasks / items / anyOfAdded value: +[ + { + "minLength": 1, + "type": "string" + }, + { + "additionalProperties": false, + "description": "Repeat steps until the page shows `until`. One round per pass.", + "properties": { + "loop": { + "additionalProperties": false, + "properties": { + "maxRounds": { + "description": "Hard ceiling on rounds. Default 12.", + "maximum": 50, + "minimum": 1, + "type": "integer" + }, + "tasks": { + "description": "Steps to run, in order, once per round.", + "items": { + "minLength": 1, + "type": "string" + }, + "maxItems": 8, + "minItems": 1, + "type": "array" + }, + "until": { + "description": "What the finished page shows. Jev reads the live page after every round and answers whether it holds, so name page evidence: \"the cart is empty\", not \"done\".", + "minLength": 1, + "type": "string" + } + }, + "required": [ + "tasks", + "until" + ], + "type": "object" + } + }, + "required": [ + "loop" + ], + "type": "object" + } +] - removed
Input schema / properties / tasks / items / minLengthRemoved value: -1 - removed
Input schema / properties / tasks / items / typeRemoved value: -"string" - changed
Input schema / properties / timeoutMs / descriptionPrevious value: -"Ceiling for the whole call. Default 90000. On expiry the call returns normally with the steps that finished and a handoff naming the rest. Keep it under your own client's transport timeout, because a dropped call still leaves the browser work done."New value: +"Ceiling for the whole call. Default 90000. On expiry the call returns with the steps that finished plus a handoff naming the rest. Keep it under your client transport timeout: a dropped call still leaves the browser work done."
- Changed
use_jev_raw2 fields changed- changed
Input schema / properties / questions / additionalProperties / properties / criteria / descriptionPrevious value: -"Options for choice (max 255), keyed by your own names. Ordered levels for score (max 10), as an array, lowest first. For noul, the two keys true and false, each describing what that answer means."New value: +"Options for choice (max 255), keyed by your own names. Ordered levels for score (max 10), as an array, lowest first. For noul, the keys true and false, each describing what that answer means." - changed
Input schema / properties / questions / additionalProperties / properties / instructions / descriptionPrevious value: -"The whole question, written out. The question id is not sent to the model."New value: +"The whole question, written out. The id is not sent to the model."
2 tool updates
v0.1.2- First observed
run_action - First observed
use_jev_raw
TDQS
Scored across 2 tools
run_action drives a browser while use_jev_raw answers typed decision questions with no browser, so the core split is clear. There is mild overlap since run_action internally has 'Jev choose' on each page, which brushes against use_jev_raw's decision role, but the browser/no-browser boundary keeps them mostly distinct.
Both names are snake_case, but run_action is a clean verb_noun while use_jev_raw embeds a product name plus an adjective, so the pattern is not predictable. Readable, but the two names follow different construction conventions.
Two tools is thin for a browser sidekick, and the surface leans entirely on run_action to multiplex search, click, read, tabs, and groups. It is a defensible facade design, but the low count is borderline for the apparent scope.
run_action covers search, open, click, read, groups, multi-tab errands, and use_jev_raw covers decisions, which spans the main intent. However, there is no explicit surface for page extraction, screenshots, navigation history, or tab cleanup, so some lifecycle operations are unaddressed.
Maintenance
Related MCP Connectors
Run multi-step tasks in a real Chrome browser: persistent environments, live view, human takeover.
Hosted real Google Chrome MCP with per-user persistent state. Navigate, click, type, screenshot.
Automate cloud Chrome—navigate, click, type, screenshot, run code, record screen video
Headless browser primitives for AI agents when sites need real JS rendering.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEnables plain-English browser automation via an MCP server, allowing agents to run objectives or test suites in a real browser without selectors or scripts.103 npm3Apache 2.0
- FlicenseNot gradedqualityCmaintenanceEnables natural language-driven browser automation by launching a real browser, performing interactions, and generating test artifacts like feature/step/page files.-
- AlicenseAqualityBmaintenanceEnables AI agents to delegate complex web browsing goals to a real Chrome instance driven by Jev, completing tasks end-to-end in ~300ms per decision and returning only the final result.11830 npm5MIT
- FlicenseNot gradedqualityBmaintenanceEnables Claude Code and Claude Desktop to control your own Chrome browser, with Jev deciding each click, keystroke, and scroll.5-