Skip to main content
Glama

multi-modal-jev-browser

tests python license

A browser agent that asks a small, cheap decision model what to do next and calls a vision LLM only when that model is unsure. It runs as a CLI, a Python library and an MCP server, so Claude Code, opencode, Cursor and Codex can drive a real browser with one line of config.

The same form filled by three decision models, in real time sped up 2.5x

One form, three decision models, the same browser layer and the same loop. Left: Jev reading the element list, with an LLM fallback: 8 s and $0.0001. Middle: Clef reading the numbered screenshot: 13 s and $0.0053. Right: DeepSeek v4.1 Flash deciding every step: 35 s and $0.0019. All three submit the right values. Real time, played 2.5x faster. Reproduce it with examples/compare_deciders.py.

Why it is cheap

Most browser agents give every click to a large LLM: read the page, think, call a tool, repeat. mmjb splits the job.

  • A decision model answers typed multiple-choice and yes/no questions with probabilities, one call per step. Jev costs $0.000015 a call and answers in about a second. Clef, which also reads the screenshot, costs under a tenth of a cent a step.

  • An ordinary vision LLM takes one step only when the decision model is unsure, claims nothing can move, or the page stops changing.

On the benchmark below, a Jev run needed the LLM fallback in none of 15 tasks.

mmjb with Jev

mmjb with Clef

the same loop with an LLM deciding every step

cost of one form (8 steps)

$0.0001

$0.0060

$0.0015 (DeepSeek v4.1 Flash), $0.0043 (GPT-6 Luna)

cost of 1,000 forms

$0.12

$6.02

$1.48 (DeepSeek v4.1 Flash), $4.30 (GPT-6 Luna)

time for one form

8 s

15 s

53 s (DeepSeek v4.1 Flash), 47 s (GPT-6 Luna)

cost of a 60 step workflow

$0.0009

$0.05

$0.01 (DeepSeek v4.1 Flash), $0.03 (GPT-6 Luna)

sees the screenshot

no

yes, numbered

yes, numbered

The cost and time rows are the measured cost and time per step multiplied by the number of steps. The full tables, the method and the reproduction commands are under Results.

Related MCP server: Cloudflare Playwright MCP

At a glance

Runs as

a CLI (mmjb), a Python library (Agent) and an MCP server (mmjb-mcp, 11 tools)

Decision models

Clef on Cloudflare Workers AI, Jev on OpenRouter, any LiteLLM vision model, or your own class in 20 lines

Fallback

any LiteLLM model; default DeepSeek v4.1 Flash

Each step sends

a screenshot with a numbered box on every control, the element list with roles and values, the page text, your facts and what the last actions did

Handles

new tabs and popups, dropdowns with 40 options, checkboxes, radios, forms that scroll, buttons that stay disabled until the form is valid, confirmations that vanish on reload

Does not handle

CAPTCHAs, canvas apps, drag and drop, content inside iframes, custom dropdowns that need the keyboard

Safety

writes are off by default, destructive verbs are refused unless the goal names them, page text is treated as data

Install

uvx --from git+https://github.com/foklepoint/multi-modal-jev-browser mmjb-mcp

Check the setup

mmjb doctor prints a pass or fail per check with the exact fix

Pick a decision model

you want

use

why

the lowest cost per step

Jev (MMJB_DECIDER=jev)

about $0.000015 a call, about one second a step, enough for forms with clear labels

to see the page as a person does

Clef 27B (MMJB_DECIDER=clef)

reads the numbered screenshot, so unlabeled icons and visual layout count

no Clef or Jev account

any LiteLLM vision model (MMJB_DECIDER=llm)

one key, slower and dearer per step

something else

your own class (MMJB_DECIDER=package.module:Class)

see docs/extending.md

How a step works

  page            observe            one call            act             say what happened
  ──────  ────────────────────  ────────────────────  ─────────────  ────────────────────────
  tab  ─▶  numbered screenshot  ─▶  decision model    ─▶  click /     ─▶  "same url; the page
          + element list            operation            type /          text changed by +27
          + page text               click target         select /        characters"
          + facts                   goal_done            scroll          + the next step record
          + last 8 effects          no_move              back
                                     fact_for
                                          │
                              margin < 0.1, or no move,
                              or 3 dead clicks in a row
                                          ▼
                                 vision LLM fallback

Quick start

git clone https://github.com/foklepoint/multi-modal-jev-browser
cd multi-modal-jev-browser
uv venv --python 3.12
uv pip install -e ".[dev]"
uv run playwright install chromium
export CLOUDFLARE_API_TOKEN=...      # or OPENROUTER_API_KEY, see the table below
export CLOUDFLARE_ACCOUNT_ID=...
uv run mmjb doctor
uv run mmjb run --url https://example.com/contact \
  --goal "Fill in the contact form with the facts and send it." \
  --fact name="Ada Lovelace" --fact message="I would like to hear more." \
  --allow-writes

mmjb doctor checks Python, the browser, every decision model and the fallback, and prints the exact fix for anything that fails. It exits 0 only when the browser works and at least one decision model answers.

Without a clone, for an MCP server:

uvx --from git+https://github.com/foklepoint/multi-modal-jev-browser mmjb-mcp

MCP setup

The server keeps one browser open for the life of the session, so look, click, type and the rest all work on the same page.

claude mcp add mmjb --env CLOUDFLARE_API_TOKEN=your_token --env CLOUDFLARE_ACCOUNT_ID=your_account_id -- uvx --from git+https://github.com/foklepoint/multi-modal-jev-browser mmjb-mcp

Copy-paste config for opencode (opencode.json), for Cursor and Codex, and a generic mcpServers file are in examples/. Every one of them uses the same uvx --from git+https://github.com/foklepoint/multi-modal-jev-browser mmjb-mcp command.

Tools:

Tool

What it does

browse(goal, start_url, facts, allow_writes, max_steps)

runs the whole loop and returns the result plus the last numbered screenshot

look()

the numbered screenshot as an image and the numbered element list as text

click(index, allow_writes)

clicks an element from the last look()

type(index, text, allow_writes)

types text into an element

select(index, option, allow_writes)

picks a dropdown option

scroll(direction)

down, up, top or bottom

goto(url)

opens a URL in the agent's tab

back()

goes back in the tab's history

press(key)

Enter, Tab, Escape and anything else

status()

the current URL, the models in use and how much work has been done

doctor()

the same report as mmjb doctor, as a tool

Choosing and swapping the models

The decision model is the cheap one that makes the decisions. The fallback model is the expensive vision LLM that takes one step when the decision model is unsure. Swap either with one environment variable.

Decision model

Needs

Reads

Notes

Clef via Cloudflare Workers AI

CLOUDFLARE_API_TOKEN, CLOUDFLARE_ACCOUNT_ID

the numbered screenshot and the page

MMJB_CLEF_MODEL=clef is 27B and reads page text. clef-flash is 9B and is too weak to read page text, so it is not the default.

Jev via OpenRouter

OPENROUTER_API_KEY

the element list and the page text, no image

MMJB_JEV_MODEL=typesafe/jev-1.13. The cheapest option, and enough for most forms.

Any LiteLLM vision model

the provider's own key

the screenshot and the page

MMJB_DECIDER=llm and MMJB_DECIDER_MODEL=.... Works with no Clef or Jev account.

With nothing set, the choice is Clef when the Cloudflare variables exist, then Jev when OPENROUTER_API_KEY exists, then a LiteLLM model.

The fallback model is MMJB_FALLBACK_MODEL, default openrouter/deepseek/deepseek-v4.1-flash. Any LiteLLM model string works.

The cost of a step is one decision-model call. In practice most simple forms finish in six to ten steps with no fallback call at all.

Command line

mmjb run --goal "..." --url "..." --fact key=value --fact other=value \
         [--allow-writes] [--max-steps 40] [--budget 300] [--json] [--out shot.jpg]

mmjb look --url "..." [--out shot.jpg]     # one numbered screenshot and the element list
mmjb doctor                                 # check the setup, exit 0 only when it works

mmjb run --json prints one JSON object on stdout and everything else on stderr, so it can be piped straight into another tool. Exit codes: 0 verified, 2 blocked, 3 unverified, 4 step limit, 1 error.

Using it as a library

from mmjb import Agent

agent = Agent(
    goal="Fill in the contact form and send it",
    start_url="https://example.com/contact",
    facts={"name": "Ada Lovelace", "message": "I would like to hear more."},
    allow_writes=True,
)
result = await agent.run()
print(result.outcome, len(result.steps), f"${result.cost:.4f}", result.seconds)

from mmjb import Agent, Result, CloudflareClef, TypesafeJev, LLMDecider also works, and every tool and option has a one-sentence docstring that reads well as an MCP tool description.

Writing your own decision model

A class with one async method. Twenty lines is a working model.

from mmjb.deciders import Answer


class MyDecider:
    name = "mine"

    async def decide(self, state: dict, questions: dict, images: list[str]) -> dict[str, Answer]:
        ...
        return {"operation": Answer(key="operation", kind="choice", choice="CLICK",
                                    probabilities={"CLICK": 0.8, "WAIT": 0.2}, confidence=0.8)}
MMJB_DECIDER=mypackage.decider:MyDecider mmjb run --goal "..." --url "..."

The same works for a fallback (MMJB_FALLBACK=mypackage.fallback:MyFallback), a safety policy (MMJB_POLICY=mypackage.policies:allow_something) and a step hook (Agent(on_step=...)). Installed packages can publish deciders and fallbacks under the mmjb.deciders and mmjb.fallbacks entry point groups, so pip install some-plugin adds one with no change to this repository. Worked examples for all four are in docs/extending.md and in examples/.

Limits

It cannot do these things, and there is no point pretending otherwise:

  • custom dropdowns that need keyboard handling to open, because it only sees the DOM controls a button click would reach

  • canvas applications, games and anything drawn in pixels rather than in HTML

  • CAPTCHAs and browser verification challenges

  • drag and drop

  • anything behind a login it has no credentials for

  • reading a file input the page never shows as a control

  • file uploads are implemented (a file: fact and set_input_files) but only covered by the safety tests, not by an end-to-end run

It also works in the viewport. Content below the fold is listed as an offscreen control and reached by scrolling to it, one step at a time.

Safety

  • allow_writes is false by default. Nothing is typed, selected or uploaded, and no submit, send, save, post or publish button is clicked, until it is on.

  • Even with allow_writes on, delete, remove, log out, pay, buy, upgrade, subscribe and checkout are refused unless the goal itself contains that word.

  • A custom policy can replace these rules entirely.

  • Page text is untrusted data, never instructions, and the rules text handed to every model says so.

  • Keys are read only from environment variables. Nothing is ever written to a file.

Tests

uv run playwright install chromium
uv run pytest -q

The suite is offline: a local test site is served by http.server and a scripted decision model drives the real loop, so the tests make no network call and need no API key. Tests marked live run the same site against a real decision model and skip themselves unless the keys are set:

CLOUDFLARE_API_TOKEN=... CLOUDFLARE_ACCOUNT_ID=... uv run pytest -m live -q -s

Results

Five form tasks of rising difficulty, served from examples/bench_site: a contact form, a demo request that opens in a new tab and keeps its submit button disabled until the form is valid, a three step wizard, a long form that needs scrolling, and a sign up with a 40 country dropdown. Every page posts what was submitted to a local server, and a run counts as correct only when the exact expected values arrived, whatever the agent claims.

Accuracy, speed and cost on the benchmark

decision model

correct submissions

agent said verified

median steps

median seconds

median cost per task

runs that needed the LLM fallback

Jev + LLM fallback

15 of 15 (100%)

15 of 15

8

8.0

$0.0001

0 of 15

Clef 27B (screenshot + elements)

10 of 10 (100%)

10 of 10

8

14.0

$0.0051

1 of 10

DeepSeek v4.1 Flash only

3 of 3 (100%)

3 of 3

5

30.3

$0.0009

0 of 3

GPT-6 Luna only

4 of 4 (100%)

4 of 4

5

31.9

$0.0026

0 of 4

Correct submissions by task

task

Jev + LLM fallback

Clef 27B (screenshot + elements)

DeepSeek v4.1 Flash only

GPT-6 Luna only

contact form

3 of 3

3 of 3

3 of 3

3 of 3

demo request in a new tab

3 of 3

3 of 3

n/a

1 of 1

three step wizard

3 of 3

3 of 3

n/a

n/a

long form with scrolling

3 of 3

1 of 1

n/a

n/a

sign up with a 40 country dropdown

3 of 3

n/a

n/a

n/a

Median steps and seconds by task

task

Jev + LLM fallback

Clef 27B (screenshot + elements)

DeepSeek v4.1 Flash only

GPT-6 Luna only

contact form

5 steps, 5 s, $0.000075

5 steps, 9 s, $0.0031

5 steps, 30 s, $0.0009

5 steps, 31 s, $0.0025

demo request in a new tab

8 steps, 8 s, $0.0001

8 steps, 15 s, $0.0062

n/a

8 steps, 45 s, $0.0048

three step wizard

8 steps, 8 s, $0.0001

8 steps, 14 s, $0.0051

n/a

n/a

long form with scrolling

12 steps, 11 s, $0.0002

13 steps, 30 s, $0.01

n/a

n/a

sign up with a 40 country dropdown

8 steps, 8 s, $0.0001

n/a

n/a

n/a

What one step costs

decision model

cost per step

seconds per step

steps per dollar

Jev + LLM fallback

$0.000015

1.0

66,667

Clef 27B (screenshot + elements)

$0.0008

1.9

1,328

DeepSeek v4.1 Flash only

$0.0002

6.6

5,392

GPT-6 Luna only

$0.0005

5.9

1,861

The same models on longer tasks (measured cost per step, multiplied)

A cost per step measured above, times the number of steps. Real long tasks also hit more unsure steps, so treat the larger columns as a floor for the cheaper models.

decision model

one form (8 steps)

signup flow (25 steps)

complex workflow (60 steps)

1,000 forms (8 steps each)

10,000 workflow steps

Jev + LLM fallback

$0.0001

$0.0004

$0.0009

$0.12

$0.15

Clef 27B (screenshot + elements)

$0.0060

$0.02

$0.05

$6.02

$7.53

DeepSeek v4.1 Flash only

$0.0015

$0.0046

$0.01

$1.48

$1.85

GPT-6 Luna only

$0.0043

$0.01

$0.03

$4.30

$5.37

Time on longer tasks

decision model

one form (8 steps)

signup flow (25 steps)

complex workflow (60 steps)

Jev + LLM fallback

8 s

0.4 min

1.0 min

Clef 27B (screenshot + elements)

15 s

0.8 min

1.9 min

DeepSeek v4.1 Flash only

53 s

2.7 min

6.6 min

GPT-6 Luna only

47 s

2.5 min

5.9 min

Cost compared with GPT-6 Luna only

decision model

cost per step compared with the baseline

Jev + LLM fallback

35.8x cheaper

Clef 27B (screenshot + elements)

1.4x more expensive

DeepSeek v4.1 Flash only

2.9x cheaper

GPT-6 Luna only

1.0x more expensive

How to read these numbers

  • The pages are local and were written for this benchmark. Real sites are messier, so expect more unsure steps and more fallback calls there than here.

  • The LLM-only rows were run on the shortest tasks only, because they are slow. Their cost and time per step come from those runs.

  • Costs: Jev is $0.000015 a call at list price, Clef is the input tokens in the response times $0.24 per million, and the LLM rows and any fallback calls are the cost LiteLLM reports.

  • The tables for longer tasks multiply the measured cost and time per step by the number of steps. They assume the same share of unsure steps as the benchmark.

  • The first pass of this benchmark found that the done check did not recognise "Your account was created" as a confirmation. The word list was extended and the sign up task was run again; the tables use the shipped code.

Reproduce:

uv run python examples/bench.py run --label jev --decider jev --repeats 3
uv run python examples/bench.py run --label clef --decider clef --repeats 3
uv run python examples/bench.py run --label deepseek --decider llm --model openrouter/deepseek/deepseek-v4.1-flash --repeats 3
uv run python examples/bench.py report tmp/bench/*.jsonl
uv run python examples/compare_deciders.py --runs 3        # the three runs behind the GIF

Credits

  • Jev by TypeSafe, for the decision-model idea and the shape of the questions

  • Clef by Cloudflare, on Workers AI

  • inspired by browser-use and jev-ultrafast

  • LiteLLM, for the model-agnostic fallback

  • Playwright, for the browser

This project is unofficial. It is not affiliated with TypeSafe or Cloudflare.

Licence

MIT. See LICENSE.

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables AI assistants to control a browser through a set of tools, allowing them to perform web automation tasks like navigation, typing, clicking, and taking screenshots.
    -
  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables AI assistants to control a browser through a set of tools, allowing them to perform web automation tasks like navigation, typing, clicking, and taking screenshots.
    -