lisa-mcp
OfficialClick on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@lisa-mcpRun QA on the staging app and fix any bugs it finds"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
██╗ ██╗ ███████╗ █████╗
██║ ██║ ██╔════╝ ██╔══██╗
██║ ██║ ███████╗ ███████║
██║ ██║ ╚════██║ ██╔══██║
███████╗ ██║ ███████║ ██║ ██║
╚══════╝ ╚═╝ ╚══════╝ ╚═╝ ╚═╝An autonomous "QA person": Claude drives a real browser through your staging app, finds bugs, and reports them. One engine, three surfaces:
Surface | Entry point | Use it for |
Terminal |
| Watch it work, iterate on missions, one-off checks |
Your agent harness | MCP server ( | "Run QA and fix what it finds" — QA → fix → re-verify loop |
Scheduled / CI | GitHub Actions cron (included) or Docker | Unattended runs, new bugs → Slack |
Inside a harness, lisa defaults to native mode: your agent drives the browser itself using the model access you already pay for, so there's no second API key. More below.
src/core.ts engine: browser primitives, agent loop, dedupe, Slack (no stdout)
src/linear.ts Linear filing: new bug → issue, re-sighting → comment
src/cli.ts terminal app (commander + live action stream)
src/mcp-server.ts MCP stdio server (native primitives or oneshot run_qa)
src/session.ts live browser sessions for native mode (mutex, idle sweep, shutdown)
src/paths.ts config / state / artifact resolution
src/config.ts config loading + validation
src/env.ts .env loading (beside the config, never clobbers process.env)
src/templates.ts template lookup + rendering
src/commands/init.ts `lisa init`
src/commands/install.ts `lisa install`
src/commands/doctor.ts `lisa doctor`
src/commands/update.ts `lisa update`
src/harness/ harness adapters: plan() what would change, then apply it
src/browser.ts lazy Chromium install (on first `lisa run`, not on `npm install`)
src/banner.ts wordmark
templates/ config, starter missions, and the agent briefs
lisa.config.yaml your projects + missionsInstall
npm run setup # installs deps, builds, and npm links — one command, one timeOr the three steps by hand, if you'd rather skip npm link (it puts lisa on your PATH globally):
npm install # deps only — Chromium is not downloaded here
npm run build # → dist/, makes `lisa` and `lisa-mcp` bins
npm link # optional: puts `lisa` on your PATHChromium downloads lazily on the first lisa run (~150MB, one time), not on npm install —
a global install shouldn't pull that unprompted for someone who only needs the terminal app
pointed at an already-wired MCP harness. Set PLAYWRIGHT_SKIP_BROWSER_DOWNLOAD to opt out
entirely (e.g. a machine with its own Chromium already on the expected path); lisa run then
fails with instructions instead of downloading. Run lisa doctor any time to check whether
it's installed without triggering a download.
npm run setup finishes by printing the one command left to run.
Wire your agent — once per machine
lisa install claude-code # or codex | cursor | windsurfThis is not per project. It registers lisa's MCP server in your harness's own
home-directory config (~/.claude.json, ~/.codex/config.toml, …) and installs the
workflow brief as a personal skill, so every repo on the machine is wired from here on.
That works because a user-scope registration carries no --config: the harness launches
the server with your repo as its working directory, and lisa resolves whichever
lisa.config.yaml belongs to it.
Run it again any time to change native or oneshot mode; re-running with nothing to change is a no-op that says so.
lisa install claude-code --projectWrites .mcp.json and .claude/skills/lisa/SKILL.md into the repo so they can be
committed and a teammate gets them on clone. They still need lisa on their PATH, and
they get a per-repo approval prompt in the harness — which is the tradeoff you're making.
Cursor is a partial case in the other direction: its MCP config has a home-directory form
but its rules don't, so even a user-scope install leaves .cursor/rules/lisa.mdc in the
repo. lisa install cursor tells you so rather than reporting a clean "wired".
Then, in each app repo
lisa initOn a machine that's already wired this is the only per-repo step — init detects the
existing install and skips the harness question entirely. On a fresh machine it asks
first how you're going to run lisa: inside an agent harness, or standalone (terminal,
CI, a server). Pick a harness and it asks one follow-up — native or oneshot — then chains
straight into lisa install once the config is written. Pick standalone and nothing
changes from here: same prompts, same .env / ANTHROPIC_API_KEY instructions as always.
Then it asks for a project name, a staging URL, whether the app needs a login, and
which starter mission to begin from — then writes lisa.config.yaml, adds the
credential env vars to .env.example, and makes sure .gitignore covers .env and
.lisa/. Run it again later to add another project.
Every prompt is also a flag, so it scripts:
lisa init --yes --url https://staging.acme.com --name acme-dashboard \
--login --mission auth--yes refuses a URL that doesn't look like a staging host unless you also pass
--non-production. lisa clicks buttons in a real browser; that gate is deliberate.
A non-interactive (--yes) run — the shape a server or a GitHub Actions job would
use — never sees the harness question at all: there's no coding agent on the other end
to wire into, so it defaults to standalone unless you pass --harness <id> explicitly
(or --harness none to say so on a TTY without being asked).
Other flags: --global (write to ~/.config/lisa/config.yaml), --force (replace an
existing config instead of adding to it), --name, --url, --no-login,
--username-env, --password-env, --mission smoke|auth|minimal,
--harness claude-code|codex|cursor|windsurf|none, --mode native|oneshot.
Keeping it up to date
lisa update # pull + rebuild lisa, then refresh the wiring it already installed
lisa update --check # report drift, change nothing (exits 1 when anything is stale)
lisa update --no-build # only refresh the wiringA new lisa version usually ships a new brief, and a brief describing tools that have moved
on is worse than no brief. update re-applies each harness's files at the scope and mode
it's actually wired in — it never wires up a harness you didn't choose, and it names the
ones it skipped. The rebuild half is for a git checkout; installed from npm, upgrade with
npm i -g lisa-cli@latest and lisa update --no-build.
Finally fill in .env:
cp .env.example .env # then editlisa reads .env from the directory its config lives in. Anything already set in the
environment wins, so CI secrets are never overwritten by a checked-out file.
Related MCP server: UI Debugger MCP
1. Terminal
lisa install claude-code # wire your agent — once per machine
lisa init # scaffold a config (see above) — once per repo
lisa update # rebuild lisa + refresh that wiring
lisa list # configured projects (flags unset credentials)
lisa run acme-dashboard # headless run, streams every action live
lisa run acme-dashboard --headed # opens a Chromium window so you can watch
lisa run --all --no-slack --no-linear # everything, print only
lisa run acme-dashboard --json # + the full report as JSON, for piping
lisa run acme-dashboard --mission "…" # one focused run instead of the configured mission
lisa report acme-dashboard # re-print the last report
lisa reset acme-dashboard # forget seen bugs; next run reports all
lisa where # which config + directories are in use
lisa doctor # API key, Chromium, config, harness wiring — all in oneWhile it runs you'll see the agent's one-line reasoning in grey, each browser action (▶ navigate …, ▶ click …), and read_page results flagged red when console errors or failed requests were captured. lisa run exits with code 2 if any new critical bug was found, so CI can gate on it.
--mission is the CLI half of the MCP server's mission_override — it's what makes
re-verifying one fix possible without editing the config, and it's what a CI job uses to
check one flow instead of the whole suite.
lisa run always uses lisa's own agent loop, so it always needs ANTHROPIC_API_KEY.
That's true whichever mode your harness is wired in — the modes below are about what the
harness does, not about the CLI.
2. Inside an agent harness
If you already answered "yes, a harness" during lisa init, this is done — skip ahead.
Otherwise, from your app's repo:
lisa install # pick from a list
lisa install claude-code # or name oneFour adapters. Each writes an MCP registration plus a brief telling the agent how to run QA, triage, fix, and re-verify:
Harness | MCP registration | Agent brief |
|
|
|
|
|
|
|
|
|
|
|
|
Claude Code and Cursor are wired entirely inside the repo, so the wiring travels through
git. Codex and Windsurf keep MCP servers in one user-global file, so only the brief can
travel — a teammate who clones the repo runs lisa install codex once for the server.
Every registration carries an absolute --config path: the harness starts the server
with its own working directory, so config discovery can't be relied on.
Then:
you: run QA on acme-dashboard and fix anything critical agent: (opens a session, drives the login and the flows, screenshots what looks wrong, files a report — then locates the code, patches it, runs tests, and re-verifies with a focused mission)
This is the interesting part: lisa finds it, the coding agent (which has your source) fixes it, then re-verifies against staging.
Native vs oneshot
lisa install asks which of two tool surfaces to register, and they are a genuine
tradeoff rather than a default with a fallback:
|
| |
Second API key | not needed | required — |
Who reasons | your harness's model | lisa's own agent loop |
Cost to your session's context | ~1.5–2k tokens per | ~2k once (the report) |
Tools |
|
|
Both surfaces also get list_qa_projects, get_last_qa_report, reset_qa_state, and
file_linear_issues (below).
Native moves the cost of reasoning off your API bill and onto your context window: a
40-turn mission can eat 80k+ tokens of the session you're working in. The brief tells the
agent to read pages sparingly, and qa_read_page truncates at 6000 characters and 120
elements, but the tradeoff is real — pick oneshot if your context is the scarce resource.
lisa install claude-code --mode oneshot # or --mode native (the default)Never both surfaces at once: a model that can see run_qa will sometimes reach for it,
silently spending the credits a native install was chosen to avoid. (lisa-mcp --tools both
exists for debugging and is not what install writes.)
Two different defaults, on purpose. lisa install writes native — what someone wiring
into a coding agent actually wants — while lisa-mcp's own default stays oneshot, so
every registration written before this existed, none of which carry --tools, keeps
working untouched.
Switching modes changes both the server args and the brief, so lisa install <harness> --status
reports out of date until you re-install. lisa doctor names the mode each harness is
actually wired in.
In native mode the credentials never reach your agent: it passes a role name
(username, password) to qa_fill and lisa types the secret itself. lisa also scrubs its
own credential values out of every tool result, so an app that reflects a password into a
URL or a form field can't leak it into your transcript either. Page text comes back wrapped
in an explicit untrusted-content marker, applied server-side on every read — in native mode
that text is flowing into an agent holding your repo and a shell, so it is a structural
guarantee rather than a line of advice the page gets to argue against.
Everything install can do is a view of the same computed plan, so nothing drifts:
lisa install --list # supported harnesses, and which are on this machine
lisa install --status # detected + wired / out of date / not wired, per harness
lisa install claude-code --dry-run # what would change
lisa install claude-code --print # the file contents, to place by hand
lisa install claude-code --yes # write, no promptsOther flags: --mode native|oneshot (skip the question), --dir <path> (write harness
files somewhere other than the config's directory) and --command "<cmd>" (override how
the harness starts the MCP server — by default lisa-mcp if it's on PATH, otherwise the
dist/mcp-server.js in this checkout).
Re-running install is safe. Files lisa generates whole (SKILL.md, lisa.mdc) are
rewritten; everything else is a surgical edit of the part lisa owns:
JSON (
.mcp.json,.cursor/mcp.json,mcp_config.json) — one key undermcpServers, preserving the file's existing indent width and every other server.TOML (
~/.codex/config.toml) — the[mcp_servers.lisa]table is swapped in place. Your comments, model settings, other servers, and even[mcp_servers.lisa.env]survive. It's text surgery, not a parse-and-re-dump, so the file still looks like the one you wrote.Markdown (
AGENTS.md) — a<!-- lisa:start -->…<!-- lisa:end -->block. The rest of the file is untouchable text.
A file lisa can't confidently read is a hard stop, never an overwrite: unparseable JSON,
mcpServers that isn't an object, [[mcp_servers]] as an array of tables, a duplicate
[mcp_servers.lisa], a half-deleted marker pair. Each exits 1 with one sentence saying what
to fix. The point of merging is to protect the config; guessing at a file we failed to parse
would defeat it.
Codex and Windsurf share one
AGENTS.mdblock rather than each claiming a private one — two lisa sections briefing the same agent differently would be worse than one. So installingcodex --mode oneshotover a nativewindsurfrewrites the block, andlisa install windsurf --statusthen reports out of date. That's status being derived from the plan rather than self-reported: the conflict is visible instead of silent.
3. Scheduled / CI
.github/workflows/lisa-cron.yml runs lisa run <project> per project on weekday mornings, caches dedupe state between runs so Slack only gets new bugs, and uploads screenshots + report-<project>.json as artifacts. Set SLACK_WEBHOOK_URL and your creds as repo secrets. The Dockerfile does the same for Cloud Run / ECS / k8s CronJob.
Config
lisa.config.yaml — written by lisa init, then hand-edited. One entry per project: base_url, allowed_host (optional; defaults to the base_url host, navigation outside it is blocked), credentials_env (env var names, never values), and a plain-English mission. The mission is the whole brief — the more specific it is, the better the report.
An optional top-level linear: block turns on issue filing (see below). Any project can
override it with a linear: of its own — a monorepo files each app's bugs to its own team.
lisa finds it by walking up from the current directory to the repo root, then falling back to ~/.config/lisa/config.yaml. State and artifacts anchor to wherever the config was found — never to your current directory. .env is read from that same directory. Run lisa where to see what resolved.
Env var | Default | Purpose |
| (search) | explicit config path |
|
| model for lisa's own agent loop — no effect in native mode |
|
| hard per-run budget — no effect in native mode |
|
| native mode: close an untouched browser session after this long |
|
| seen-bug fingerprints |
|
| reports + screenshots (namespaced per project) |
| — | suppress the wordmark |
| — | no update nudge, no "what changed" summary, no background check |
| — | required for |
| — | Linear personal API key. Only read when the config has a |
| — | don't lazily download Chromium; fail instead if it's missing |
LISA_MODEL and LISA_MAX_TURNS doing nothing in native mode surprises people, so it's
worth saying plainly: in native mode there is no lisa agent loop to configure. Your
harness's model is the model, and its own turn limits are the budget.
Filing to Linear
A report is ephemeral. A bug nobody fixes in the same session evaporates — no ticket, no
owner, no state. Add a linear: block and every new bug becomes an issue instead:
linear:
api_key_env: LINEAR_API_KEY # a NAME, never the key — same rule as credentials_env
team: ENG # team key or UUID
project: QA Bugs # optional
labels: [qa-agent] # optional; must already exist on the teamThen put the key in .env and lisa doctor will tell you whether it resolves.
The interesting half is what happens on the second sighting. lisa runs on a cron, so a
bug that takes a week to fix is found five times — which without care is five identical
issues. lisa records which fingerprint became which issue in
.lisa/state/<project>.linear.json, so a re-sighting comments "still present as of …" on
the issue it already has. The CI job caches that directory alongside the seen-bug state, so
cron runs comment rather than duplicate. lisa reset <project> clears both — that's what
makes it mean "re-report everything".
Filing is opt-in twice over: no linear: block and no key both mean nothing is ever filed,
and lisa run --no-linear skips it for one run. One run files at most 10 issues; a run that
finds forty bugs is a broken deploy, not forty tickets, and the rest file on the next run.
Inside a harness, the default is the other way around — file_to_linear is false, and the
agent calls file_linear_issues after triage with the bugs it isn't fixing itself. The
ones it just fixed don't need a ticket.
Staying current
lisa tells you when your install is behind and what to run about it — lisa update for a git
checkout, npm i -g lisa-cli@latest for a package install. The check itself runs detached
after a command finishes and is cached for a day, so nothing ever waits on it; the nudge is a
single line on stderr and is suppressed off-TTY and in CI.
Once a newer version actually runs, lisa prints the CHANGELOG sections
between the version you were on and the one you're now on, with Actions required in full.
That's the contract for releases: anything a user has to do goes under that heading, or they
won't be told. Set LISA_NO_UPDATE_CHECK=1 to turn the whole mechanism off.
Run lisa doctor to check all of the above (API key, Chromium, config, harness wiring and
mode) in one shot. With a native harness wired, a missing ANTHROPIC_API_KEY is a ⚠ rather
than a ✗ — nothing is broken, but lisa run and CI would still need one.
Safety rails
Every rail below lives in the engine, not in a prompt, so it applies the same whether lisa's own loop or your harness's model is doing the driving:
Navigation outside
allowed_hostis refused by the tool itself.Page text comes back wrapped in an untrusted-content marker on every read.
Credentials are typed by lisa from a role name, never handed to the model — and lisa's own credential values are scrubbed out of every tool result, so an app that reflects one back can't leak it either.
Unset credential env vars are reported as such, so a missing secret produces "blocked (missing credentials)" rather than a bogus "login is broken" bug.
A failing Slack webhook warns and is recorded as
slack_erroron the report. It never fails the run, changes the exit code, or hides the findings — the report is written first, and the seen-bug state only advances once it is on disk.The same holds for Linear: a failure is recorded as
linear_error, never fails the run, and never loses the issues that did get filed. A bug that couldn't be filed is picked up by the next run — the issue map, not the new/known split, decides what already has a ticket.
Destructive actions are forbidden by instruction rather than by code — the oneshot system prompt and the native session briefing carry the same rules. Reinforce them per-mission where it matters (e.g. "checkout up to but NOT including payment"). Staging + dummy accounts only.
This server cannot be deployed
Maintenance
Related MCP Connectors
Browser-based QA for AI-built software. Test pages with real browsers via agents.
Browser-backed QA with evidence and fix-ready reports for coding agents.
Agentic testing: HyperExecute jobs, test failure triage, SmartUI visual diffs, a11y audits
- openhelmOAuthai.openhelm
Autonomous cloud agent tasks: real browser + your tools, structured evidence-backed results.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceBrowser-based QA testing for AI-built software. Agents open real browsers (via Selenium), navigate pages, fill forms, click buttons, and report findings. Two modes: targeted tests (30-90s) and full-site discovery scans (3-15min).-
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to autonomously debug UIs by delegating high-level stories to a small agent that drives browsers or desktop apps and reports structured pass/fail findings with evidence.83 npm2MIT
- AlicenseNot gradedqualityBmaintenanceEnables autonomous web QA by exposing Playwright browser control as MCP tools for navigation, accessibility snapshotting, interaction, and bug detection.2 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables AI coding agents to autonomously interact with and test web applications in a real browser, providing DOM/Accessibility tree extraction, runtime telemetry, screenshot capture, and Markdown test reports.22 npm1MIT