Skip to main content
Glama
PopoverAI

Browser Automation MCP

by PopoverAI
README.md
# agentic-demo

Narrated demo videos of your web app, recorded by your coding agent.

Ask Claude Code, Codex, Cursor, or any agent that can run shell commands for a walkthrough of a flow. It explores the app with [agent-browser](https://www.npmjs.com/package/agent-browser), writes a short script (the browser commands for each step and a sentence to say over it), and agentic-demo records the script as an mp4 with a voiceover. No screen recorder, no microphone, no editing.

The script stays behind as a file. Nothing improvises during the take, so the same script gives the same video, and when the app changes you re-record instead of redoing the demo.

## Quick Start

Paste this into Claude Code, or any coding agent:

```
Use `agentic-demo` to create a demo of this feature. Start from `npx agentic-demo --help`.
```

The agent does the rest. It will tell you if it needs anything from you, such as an OpenAI API key for the voiceover. You get a video file, plus the script it was made from, so you can ask for a fresh recording whenever the feature changes.

## Before You Record

- **The voiceover costs money.** It is generated by a voice provider (OpenAI by default) and billed to your account with them.
- **Your pages reach your agent's model.** Your coding agent reads the pages it explores, and like anything else it reads, that content goes to the model behind it. agentic-demo itself sends only the narration text, to the voice provider (or the voice endpoint you point it at); the recording and the video stay local. A cloud browser's provider sees the pages too.
- **The video shows whatever is on screen.** Record against demo data, not real customers.
- **Passwords stay out of the script.** The guide tells your agent to sign in through agent-browser's encrypted credential store, so no password is written into the file.

## AGENTS.md / CLAUDE.md

Add this to your project's instructions so your agent reaches for agentic-demo whenever you ask for a demo:

```markdown
## Demo Videos

Use `agentic-demo` to record narrated demo videos of web flows. Run `npx agentic-demo guide` for the workflow before writing a steps file.
```

## Reference

### Installation

```bash
npm install -g agentic-demo          # Global
npm install -D agentic-demo          # Pinned in a project
npx agentic-demo --help              # Without installing
npm install -g agentic-demo@latest   # Update
```

Requires Node.js 22+. Installing downloads an ffmpeg binary (about 77 MB); pnpm 10 blocks that download until you run `pnpm approve-builds`, or point `--ffmpeg` / `FFMPEG_BIN` at another ffmpeg. agent-browser runs as `npx agent-browser`, which uses your installed copy if there is one. The same CLI is also published as [`@popoverai/browser-automation`](https://www.npmjs.com/package/@popoverai/browser-automation).

Versions are 0.x, so a minor version can break things. The [changelog](https://github.com/PopoverAI/browser-automation/blob/main/CHANGELOG.md) lists every change.

### Commands

```bash
agentic-demo <steps.json>            # Record to an mp4 (same as `record <steps.json>`)
agentic-demo validate <steps.json>   # Check a steps file without opening a browser
agentic-demo guide                   # Print the workflow guide for agents
agentic-demo schema                  # Print the steps file's JSON Schema
agentic-demo example                 # Print a starter steps file
```

For the step-by-step workflow, run `agentic-demo guide`.

### Options

| Option                     | Description                                                      |
| -------------------------- | ---------------------------------------------------------------- |
| `-o, --out <dir>`          | Output directory (default: a new temporary directory)            |
| `--tts <provider[:model]>` | Narration provider and model (default: `openai:gpt-4o-mini-tts`) |
| `--tts <url>`              | Get narration from a [voice endpoint](#voice-endpoint) instead   |
| `--voice <voice>`          | Voice id for the narration provider                              |
| `--silent`                 | Silent audio track sized to the narration; no key needed         |
| `--keep`                   | Keep each segment's audio, video and frames next to `final.mp4`  |
| `--session <name>`         | agent-browser session to drive                                   |
| `--agent-browser "<cmd>"`  | How to run agent-browser (default: `npx agent-browser`)          |
| `--trailing-delay <ms>`    | Default `trailingDelay` for every step (default: 1000)           |
| `--max-fps <n>`            | Cap the capture frame rate (default: uncapped)                   |
| `--ffmpeg <path>`          | ffmpeg binary (default: ffmpeg-static's, or `FFMPEG_BIN`)        |
| `--json`                   | Print the summary in `final.json` on stdout as well              |

Without `--json`, stdout is the path to `final.mp4` and progress goes to stderr.

### Output

A recording writes `final.mp4` and, beside it, `final.json`: the video's length, and for each step where it begins in the video and how long it took. Every run writes it, narrated or `--silent`; `--json` prints the same summary.

```jsonc
{
  "videoPath": "/home/me/demo/final.mp4",
  "outputDir": "/home/me/demo",
  "durationSeconds": 16.16, // Length of final.mp4
  "segments": [
    // One per step, in the steps file's order
    {
      "index": 1, // The step's position in the steps file, from 0
      "instruction": "click #red", // The step's commands
      "narrative": "Turn the page red.", // The step's narration
      "startSeconds": 2.323, // Where the step begins in final.mp4
      "captureSeconds": 1.01, // How long its commands took to run
      "narrationSeconds": 1.6, // Length of its narration
      "renderedSeconds": 1.88, // Its length in final.mp4
      "frameCount": 2, // Frames captured; 1 means the step changed nothing on screen
    },
  ],
}
```

`startSeconds` is the time of the step's first frame in `final.mp4`, rounded up to the millisecond, so a player that seeks there shows that step rather than the end of the one before. Use it to open the video at a step. Don't add up `renderedSeconds` instead: joining the steps starts each one a few hundredths of a second later than that sum.

### Steps File

```jsonc
{
  "url": "https://app.example.com/login", // Optional: opened before recording; omit to record the current page
  "speech": { "provider": "openai", "voice": "alloy" }, // Optional: narration settings
  "steps": [
    {
      "narrate": "Sign in with the demo account.", // Said over this step
      "commands": [
        ["auth", "login", "demo", "--no-navigate"], // "demo": credentials saved with `agent-browser auth save`
        ["wait", "--load", "networkidle"],
      ], // agent-browser commands, one argv array each
      "trailingDelay": 1000, // Optional: ms to keep capturing after the last command
    },
  ],
}
```

Also optional: `openArgs` (extra arguments for `open`), `speech.endpoint` (a [voice endpoint](#voice-endpoint) URL, in place of `speech.provider`), `speech.model`, `speech.instructions`, `speech.speed`, `speech.language`, `speech.providerOptions`, and a per-step `speech` block. `agentic-demo schema` prints the full schema.

A step can change the browser's size, with `["set", "viewport", "375", "667"]`, to show a flow on a phone after a desktop. The video keeps one size, the widest and tallest the recording reached, and fits the smaller view inside it, centred on black.

If agent-browser stops sending the browser's picture partway through, the recording stops with an error that names the step, instead of finishing a video that freezes there.

### Narration

Five voice providers ship with the CLI. Each reads its key from its own environment variable:

| Provider          | `--tts`                                                     | Key                                         | Models                                                   |
| ----------------- | ----------------------------------------------------------- | ------------------------------------------- | -------------------------------------------------------- |
| OpenAI            | `openai[:model]` (default `gpt-4o-mini-tts`, voice `alloy`) | `OPENAI_API_KEY`                            | `gpt-4o-mini-tts`, `tts-1-hd`                            |
| ElevenLabs        | `elevenlabs[:model]` (default `eleven_multilingual_v2`)     | `ELEVENLABS_API_KEY`                        | `eleven_v3`, `eleven_flash_v2_5`                         |
| Hume              | `hume`                                                      | `HUME_API_KEY`                              | (single model)                                           |
| Deepgram          | `deepgram[:model]` (default `aura-2`)                       | `DEEPGRAM_API_KEY`                          | `aura`, `aura-2`                                         |
| Vercel AI Gateway | `gateway[:model]` (default `openai/tts-1-hd`)               | `AI_GATEWAY_API_KEY` or `VERCEL_OIDC_TOKEN` | `openai/tts-1`, `fish-audio/s2-pro`, `spacexai/grok-tts` |

Choose the provider and voice in the steps file's `speech` block, or for one run with `--tts` and `--voice`.

The AI Gateway reaches several providers' voices with one key. Its model ids name the provider too, as in `--tts gateway:openai/tts-1`; the [AI Gateway model list](https://vercel.com/ai-gateway/models) has the current set. Without `--voice`, the model's own default voice speaks.

### Voice endpoint

Instead of calling a voice provider itself, agentic-demo can ask an HTTP endpoint for each narration line. Use this when the machine that records should not hold a provider key: your app keeps the key, decides whether to narrate, and returns the audio. Name the endpoint with `--tts https://…` or `"speech": { "endpoint": "https://…" }`, and put the credential your endpoint expects in `AGENTIC_DEMO_TTS_TOKEN`.

Before it opens the browser, agentic-demo narrates the first step, so a refused token or a used-up quota stops the run before any recording. The rest of the lines are requested while the video is rendered, in parallel, one request per step. The first line's audio is reused, so a run makes one request per step in all.

**Request**

```http
POST <endpoint URL>
Authorization: Bearer <AGENTIC_DEMO_TTS_TOKEN>
Content-Type: application/json
Accept: audio/*

{
  "text": "Sign in with the demo account.",
  "model": "gpt-4o-mini-tts",
  "voice": "alloy",
  "instructions": "Friendly product walkthrough; unhurried.",
  "speed": 1.0,
  "language": "en",
  "outputFormat": "mp3",
  "providerOptions": { "openai": { } }
}
```

- `Authorization` is sent only when `AGENTIC_DEMO_TTS_TOKEN` is set.
- `text` is always present. `outputFormat` is `"mp3"` unless the steps file sets another. Every other field is present only when the steps file (or `--voice`) sets it, and is passed through without checking, so your endpoint decides which to honour and which models and voices to allow.
- `model` and `voice` come from the steps file only when its `speech` block names this same endpoint; a voice chosen for a provider is not sent. `--voice` is always sent. A step's own `speech` block overrides `voice`, `instructions`, `speed` and `language` for that line.

**Success:** any `2xx` status with the audio bytes as the body and an `audio/*` `Content-Type` (`audio/mpeg`, `audio/wav`, …). The format is read from the bytes, so any format ffmpeg can decode works. A `2xx` answer without an `audio/*` `Content-Type` is treated as an error.

**Refusal:** any non-`2xx` status. Put the message for the person in the body, as JSON or as plain text:

```http
HTTP/1.1 402 Payment Required
Content-Type: application/json

{ "error": "Your team has used this month's narration quota. Upgrade at https://example.com/billing." }
```

agentic-demo stops, exits with status 1, and prints the `error` string exactly as sent, on its own line after the status. When the body is not JSON with a string `error`, it prints the whole body. It does not retry. A `401` or `403` when `AGENTIC_DEMO_TTS_TOKEN` is unset also says that no credential was sent.

```text
[agentic-demo] the voice endpoint https://app.example.com/api/voice refused to narrate (HTTP 402):
Your team has used this month's narration quota. Upgrade at https://example.com/billing.
```

### License

Apache-2.0. Originally forked from [@browserbasehq/mcp-server-browserbase](https://github.com/browserbase/mcp-server-browserbase).

TDQS

A3.6/5.0

Scored across 14 tools

Disambiguation3/5

Several tools overlap in purpose, particularly the multi-step execution tools (stagehand_agent, stagehand_scenario, stagehand_run_script, stagehand_demo_video) and the action/observation pair (stagehand_act, stagehand_observe). The descriptions do provide distinguishing details, but the boundaries are not immediately obvious.

Naming Consistency2/5

Naming conventions are mixed: some tools use verb_noun (stagehand_session_create, stagehand_run_script), some use bare verbs (stagehand_navigate, stagehand_act), and some use nouns (stagehand_agent, stagehand_scenario). Additionally, there are two different prefixes (stagehand_ vs agent_browser_) without a consistent pattern across the whole set.

Tool Count4/5

14 tools is on the higher end of the typical range. While each tool has a distinct function, the set includes several tools for running multi-step workflows (agent, scenario, run_script, demo_video) that could have been consolidated, making the count slightly over-scoped but still reasonable.

Completeness5/5

The tool set covers the full browser automation lifecycle: session management, navigation, interaction, extraction, screenshots, script execution, and low-level control. There are no obvious missing operations for the domain.

Maintenance

ActivityMaintained
ResponsivenessNo issues