Skip to main content
Glama

Flipbook

CI npm License: MIT

Claude can't watch video. Flipbook records your web app while Claude drives it and hands back a flipbook it can read: a labelled contact sheet, full-resolution key frames, and a timeline correlated with the actions that caused each change.

A Claude Code plugin and a standalone MCP server. Recording needs macOS 15+; analysing a recording — yours, or a Playwright or Cypress test video — works on macOS, Linux and Windows. Works alongside Claude in Chrome without interfering with it.


The problem

Claude in Chrome validates web apps with discrete screenshots. Everything between two screenshots is invisible — spinners, layout shift, flicker, double-submits, and toasts that appear and vanish. A before/after pair can't tell "it worked" from "it worked eventually, badly".

This isn't hypothetical. Driving the bundled test fixture, Claude in Chrome's own screenshot shows:

Order checkout
Idle.
┌─ Order confirmed ─────────────┐
│ Item              Widget Pro  │
│ Quantity                   3  │
│ Total                $147.00  │
└───────────────────────────────┘

Looks like a pass. But the run also showed a two-second loading spinner and a "Saved successfully" toast — and both are missing, because the toast had already disappeared by the time the screenshot was taken.

Flipbook records the same run and returns this instead — one image, the whole timeline:

Contact sheet of the checkout flow: the resting state, three frames of the loading
spinner, then the confirmation panel with a success toast that has already vanished by the
end of the run

Every cell is labelled with its frame number and timestamp, red outlines mark the largest visual changes, and the caption says why each frame was chosen. The spinner (02–04) and the toast (05, top right) are exactly what the screenshot above missed.

That image is real output, not a mockup — regenerate it with node scripts/make-demo.mjs.

gif_creator doesn't close this gap: it stitches together screenshots Claude already took and exports them for a human, returning nothing to the model. And the vision API reads only a GIF's first frame, so an animated GIF isn't something Claude can watch either. (Hand one to analyze_recording and Flipbook will decode every frame of it.)

Related MCP server: ThinkRun

Platform support

Record (start_recording)

Analyse (analyze_recording, get_frames)

macOS 15+

yes — ScreenCaptureKit window capture

yes

macOS 14 and earlier

display capture only (target: "display")

yes

Linux

no

yes

Windows

no

yes

doctor tells you which of the two a machine can do, and why. The analysis half needs only Node 22+ and ffmpeg (bundled). CI runs the full protocol suite on all three platforms.

Install

Flipbook works in any MCP client. The Claude Code plugin is the full experience: it adds the validate-with-recording skill, the /flipbook:record command, and a hook that timestamps every Claude-in-Chrome action into the recording. Everywhere else you get the same eight tools.

Client

How

Claude Code (plugin, recommended)

below

Claude Code (MCP server only)

claude mcp add flipbook -- npx -y @shubhenduvaid/flipbook

Claude Desktop

download flipbook-<version>-darwin-arm64.mcpb from Releases and open it

Cursor

Install in Cursor

VS Code

Install in VS Code

Anything else

the JSON below

{
  "mcpServers": {
    "flipbook": { "command": "npx", "args": ["-y", "@shubhenduvaid/flipbook"] }
  }
}

It is also listed in the MCP Registry as io.github.ShubhenduVaid/flipbook. Whichever you choose, ask for doctor first.

Claude Code plugin

Before you start: macOS 15+ (window capture uses ScreenCaptureKit), Node 22+, Xcode Command Line Tools (xcode-select --install), and Google Chrome. On Linux or Windows the plugin installs and analyses fine; only recording needs the Mac.

1. Add the marketplace and install.

claude plugin marketplace add ShubhenduVaid/flipbook
claude plugin install flipbook@shubhenduvaid

2. Grant Screen Recording to your terminal — System Settings → Privacy & Security → Screen & System Audio Recording — then restart the terminal. This is the one step nothing can do for you, and the one most worth getting right: without it macOS records a blank screen rather than erroring, so a recording looks like it worked and contains nothing.

3. Start Claude and run doctor.

claude --chrome

Then ask Claude to run doctor. It finishes setup on first run — compiling the ScreenCaptureKit recorder and restoring the bundled ffmpeg binary, which Claude Code's installer skips because it installs plugin dependencies with --ignore-scripts — and it checks every prerequisite above, including the permission from step 2.

You're ready when doctor says so. Warnings are fine; failures are not.

4. Record something — see Usage below.

claude plugin list                 # name, version, scope, enabled
claude plugin details flipbook     # what it contributes, and its token cost
git clone https://github.com/ShubhenduVaid/flipbook && cd flipbook
npm install
npm run build:native   # compiles the ScreenCaptureKit recorder
npm run doctor         # preflight every prerequisite

claude --chrome --plugin-dir "$PWD"

--plugin-dir loads the plugin for one session only, and takes precedence over an installed copy — handy for trying a change without disturbing your install.

Updating

Two steps, because they do different things: the first refreshes the catalogue, the second installs from it. Doing only the second is the usual reason an update appears to do nothing.

claude plugin marketplace update shubhenduvaid   # fetch the latest catalogue
claude plugin update flipbook                    # install it

Restart Claude Code afterwards. A running session keeps the version it started with.

Then run doctor once in the new session. It recompiles the native recorder if the plugin shipped a new one, so a version that changes the recorder is fixed by the same command that diagnoses everything else.

claude plugin list        # confirm the new version
git pull
npm install            # in case dependencies changed
npm run build:native   # in case the recorder changed
npm run doctor
  • Version unchanged in claude plugin list — the marketplace catalogue is stale. Run claude plugin marketplace update with no name to refresh every marketplace you have.

  • New tools or parameters missing — the session predates the update. Restart Claude Code; tool schemas are read once at startup.

  • Recording fails or returns nothing after an update — run doctor. A recorder compiled against an older plugin is its most likely cause, and doctor rebuilds it.

  • Still wrong — reinstall cleanly:

    claude plugin uninstall flipbook
    claude plugin install flipbook@shubhenduvaid

    Your recordings live in ~/.flipbook and are untouched by this.

Versions follow semver. Released versions are tagged flipbook--v<version>, so the tag list is the changelog — each annotated tag says what changed and why.

Usage

/flipbook:record the checkout flow shows a spinner, then a confirmation, and no errors

Or just ask — the bundled skill tells Claude when to reach for this. Under the hood:

  1. start_recording — begins capturing the browser window

  2. you or Claude drive the app; every Claude-in-Chrome action is timestamped automatically

  3. stop_recording with a rubric — what "working correctly" means, one criterion per line

  4. Claude judges the evidence and cites frames

A rubric should be observable in pixels and time:

A loading indicator appears within 500ms of clicking Submit.
The loading indicator disappears once results render.
No error toast appears at any point.
The layout does not shift after the results render.

And the verdict cites evidence you can check:

No error toast appears — FAIL. Frame 06 at t=5.25s shows a toast reading "Saved successfully"; it's gone by t=6.75s, which is why the final screenshot looks clean.

Flipbook returns evidence, never a verdict. A tool that answers "PASS" hides its reasoning and can't be argued with.

Three ways to record nothing

Each produces a plausible-looking recording that contains no evidence. All were found the hard way; doctor warns about the first two.

  1. Recording a window whose active tab isn't the one under test. A window paints only its active tab, and Claude in Chrome will happily drive a background tab. Switch to it first.

  2. Recording a window that another window covers. macOS marks it occluded and the browser stops painting it, so you capture frozen browser chrome over a blank page. Two browser windows at nearly the same position are the usual culprit. Behind a full-screen terminal on a different Space is fine; underneath another window on the same Space is not.

  3. Taking a screenshot with size arguments while recording. That overrides the browser's device metrics, so the page repaints into a smaller viewport — or stops painting — while capture keeps rolling at the old size. This one is worse than the other two, because the result isn't blank: it's convincing evidence of a bug that doesn't exist. One reported run showed an 11-second blank load that a control run rendered in ~1.8s. The analysis detects the geometry change and says so.

When actions were recorded but the pixels didn't move, the analysis says so rather than letting you report a false pass. Same for a span between two marks where nothing changed, and for a subject too small in frame to read.

Tools

Tool

Purpose

doctor

Preflight: what this machine can do (record, analyse), ffmpeg + filters, label font, native helper, permission, target window, disk, footprint

start_recording

Capture a window (label, target, title_contains, window_id, fps, max_duration_s)

mark

Annotate the timeline mid-run, and split the per-segment change stats

stop_recording

Stop, analyse, return the evidence against a rubric (roi, clip)

analyze_recording

Same analysis for any .mov/.mp4/.webm/.m4v/.gif or a directory of stills (roi, clip)

get_frames

Full-resolution drill-down at exact timestamps, over a range, or relative to a mark (after_mark, offset, roi)

list_recordings

Browse past sessions, with sizes and total footprint

prune_recordings

Reclaim disk — a dry run unless confirm: true

analyze_recording accepts recordings you made yourself — hand it a QuickTime capture of a bug you can't reproduce on demand.

Test-runner videos, on any platform

Playwright and Cypress already record videos of your end-to-end tests, and Flipbook reads them. This is also how to use it on Linux, Windows and CI, where it cannot record itself:

// playwright.config.js
export default { use: { video: "on" } };   // or "retain-on-failure"
Analyse test-results/checkout-chromium/video.webm against:
  A loading spinner is shown while the order is processing.
  A confirmation panel appears.
  A success toast appears and then disappears.

Test-runner videos show the page without browser chrome; the capture-fault detectors account for that, so a toast in the corner of a dark page is a toast, not a warning.

Framing: roi

An image costs the same whatever it contains, so a small subject in a large window spends most of its resolution on nothing. roi takes a fractional rect or "auto" — which derives the region from the pixels that actually changed — and applies to the three tools above. When the subject is small and no roi was given, the analysis says so and quotes the exact rect to pass.

A cropped frame is announced four ways: a CROPPED VIEW header line, a [CROP …] prefix on every caption, an amber border and badge burnt into every cell, and a roi object in structuredContent that is present whether or not it applied. A crop read as the whole page is exactly the failure mode the geometry detector exists to catch, so it's worth repeating.

Why stills instead of video

There's no video input, animations are explicitly unsupported, and the MCP spec has no video content type (checked in both 2025-06-18 and 2026-07-28). So a recording has to become stills plus text.

The budget was read out of Claude Code's own accounting rather than guessed:

Cost per MCP image

flat 1600 tokens

MCP output budget

25,000 tokens (MAX_MCP_OUTPUT_TOKENS)

Never-truncated zone

50% of budget → ~7 images

Default output is 6 images — one contact sheet plus five detail frames — leaving room for the timeline under the ~12,500-token ceiling. get_frames provides drill-down rather than spending more images up front. A measured run comes in at ~10,100 tokens.

Keyframe selection samples at 128×128 and scores each frame both globally and per-block, because a 64px spinner in a 3000px-wide window moves the whole-frame average by 0.0003 — indistinguishable from noise. Dedupe compares pixels rather than perceptual hashes, which measured zero Hamming distance between an idle page and the same page showing a spinner. When two frames are identical, the earlier one wins, so the moment a state was reached isn't discarded in favour of an identical later frame.

Why ScreenCaptureKit

Display capture records whatever is visually on top — which is your terminal, not the browser. Since the whole point is recording a browser Claude drives in the background, window capture is the only approach that works. It also removes retina scaling and crop arithmetic, and never captures anything but the target window. ffmpeg still does all the analysis, and display capture remains available via target: "display".

What it runs on your machine

Flipbook makes no network calls of its own and never uploads anything; recordings stay in ~/.flipbook. Two things run locally that are worth knowing about, and both happen only when doctor or a recording needs them:

  • npm rebuild ffmpeg-static in the plugin directory, once. Claude Code installs plugin dependencies with --ignore-scripts, which leaves the bundled ffmpeg package without its binary; this restores it (it is ffmpeg-static's own download from its GitHub releases). Skipped whenever a system ffmpeg or FLIPBOOK_FFMPEG is found first.

  • xcrun swiftc compiles native/sckrec.swift, the ScreenCaptureKit recorder, into ~/.flipbook/bin on macOS. It is ~260 lines and you can read it before it runs.

Development

npm run validate       # score the repo against docs/VALIDATION-RUBRIC.md
npm run validate -- --full   # …including the criteria that need macOS + Chrome
npm test               # 169 unit tests, no browser or permission needed
npm run lint           # Biome (via npx — deliberately not a dependency)
npm run lint:manifests # plugin, marketplace, package, registry and .mcpb manifests agree
npm run test:mcp       # 45 MCP protocol checks, any platform (see below)
npm run test:e2e       # fixture: spinner + transient toast must be captured (macOS + Chrome)
npm run test:e2e:occluded  # same, with the browser occluded
npm run pack:mcpb      # build the Claude Desktop extension into dist/

test:mcp drives the real server over stdio. It analyses the macOS fixture recording if one exists, else a synthetic fixture that ffmpeg draws (test/synthetic-fixture.mjs) — blank tab, spinner, confirmation, and a corner toast that vanishes — and asserts the spinner, the panel and the toast are all among the frames selected. That makes the product's central claim a CI check. Point FLIPBOOK_FIXTURE_VIDEO at any recording to run it against that instead.

docs/VALIDATION-RUBRIC.md is the standard this project holds itself to — 38 numbered criteria covering the promises this README makes, naming, security, quality and docs. npm run validate is its executable form and reports each criterion individually; there is no aggregate score, because a percentage would let a security failure be averaged away by passing style checks. docs/ARCHITECTURE.md explains the pipeline and what each module owns.

npm run validate, npm test, npm run test:mcp and the linters run in CI on every push (the protocol suite on Linux, Windows and macOS); test:e2e needs macOS, Chrome and Screen Recording permission, so it's a local check.

Cutting a release

Claude Code users get a new version by the two commands under Updating, which read the marketplace manifest on main. Everyone else gets it from npm, the MCP Registry and GitHub Releases, which the release workflow publishes when the tag is pushed. So a release is: bump, verify, tag.

  1. Bump the version in all six places — package.json, package-lock.json (both the root and packages.""), .claude-plugin/plugin.json, .claude-plugin/marketplace.json, server.json (twice: the server and its npm package) and mcpb/manifest.json. Semver: a new tool or parameter is a minor, not a patch. npm run lint:manifests fails if any of them disagree, and rubric criteria N1 and N3 cover exactly this.

  2. Verify, including the checks CI can't run:

    npm run validate && npm run test:mcp && npm run test:e2e
    claude plugin validate . --strict
  3. Tag it, once the bump is merged to main:

    claude plugin tag --dry-run          # check what it would produce
    claude plugin tag -m "flipbook %s" --push

    claude plugin tag creates the {name}--v{version} tag this project uses, and refuses if plugin.json and the marketplace entry disagree or the tree is dirty — which is why step 1 is worth doing properly. Tag messages are the changelog (and the GitHub release notes), so write them for someone deciding whether to update; add the same entry to CHANGELOG.md.

  4. Watch the release workflow. It refuses a tag that disagrees with package.json, re-runs the tests, then publishes to npm, then to the MCP Registry (which checks the npm package it just published), then attaches the .mcpb to a GitHub release. First-time setup for each destination, and the directories that need a one-off submission, are in docs/PUBLISHING.md.

Biome is invoked through npx rather than added as a devDependency: Claude Code installs plugin dependencies without --omit=dev, so a devDependency would ship into every user's plugin cache. The formatter is off — the linter catches real defects, while reformatting would churn deliberately aligned ffmpeg argument lists.

FLIPBOOK_DEBUG_SELECT=1 traces every keyframe accept/merge decision to stderr, which is what threshold tuning needs.

Variable

Purpose

FLIPBOOK_HOME

Data directory (default ~/.flipbook)

FLIPBOOK_FFMPEG

Use a specific ffmpeg binary

FLIPBOOK_FONT

Font file for contact-sheet labels (default: the first system font found)

FLIPBOOK_DEBUG_SELECT

Trace keyframe selection

.claude-plugin/     plugin + marketplace manifests
server.json         MCP Registry entry
mcpb/               Claude Desktop extension manifest
commands/           /flipbook:record
hooks/              PostToolUse hook correlating Claude-in-Chrome actions
native/sckrec.swift ScreenCaptureKit window recorder
skills/             when and how Claude should use this
src/env/            ffmpeg, native helper, Chrome, doctor, paths
src/capture/        session lifecycle, recorder processes, prune policy
src/analyze/        sampling, delta scoring, selection, segments, capture anomalies,
                    region of interest, sheet, budget, timeline, clips
src/tools/          MCP tool definitions and shared parameter schemas
test/unit/          CI-safe unit tests
test/               protocol suite, synthetic and browser fixtures
scripts/            validate, manifest lint, doctor CLI, .mcpb packer

Recordings are written to ~/.flipbook/sessions/ — outside your project, never auto-uploaded anywhere.

License

MIT © Shubhendu Vaid

Related MCP Connectors

Related MCP Servers