Skip to main content
Glama

screencast

Destructive

Record a browser tab as an MP4 or WebM video to inspect animations or interactions frame by frame; returns the file path and size.

Instructions

Record the tab as a video (.mp4 or .webm) to watch an animation or interaction frame by frame; returns the file path and size, never the frames. action record (default) = start, optional interaction, wait duration_ms, stop. start/stop bracket anything else done meanwhile (max 120 s). Chrome sends a frame only when the page changes; the video keeps the real timing. Needs the tab active and its window visible, and ffmpeg on the server machine (without it: the JPEG folder and the ffmpeg command). Uses chrome.debugger: Chrome shows its debugging bar while recording; in launch mode nobody sees it.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
fpsNoFrame rate of the output video
actionNorecord does start → wait → stop in one callrecord
formatNoVideo containermp4
tab_idNoTarget tab; omitted = last tab navigated in this session, else the active one
qualityNoJPEG quality of the captured frames, 0-100
save_toNoAbsolute path: write the video there and return the path instead of the content
max_widthNoScale frames down to this width in px
duration_msNorecord: how long to record, max 120000
interactionNoAction run right after recording starts

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv1.28.0

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations with operational detail: needs the tab active and the window visible, requires ffmpeg on the server (with a stated fallback of the JPEG folder plus ffmpeg command), the 120 s recording cap, Chrome's change-driven frame emission with real preserved timing, and the chrome.debugger debugging-bar side effect including launch-mode behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Information-dense and front-loaded, with the core purpose in the first clause. Every sentence carries distinct value (return shape, action model, timing semantics, prerequisites, side effects), though the volume of caveats packed into one paragraph is heavy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex nine-parameter tool with a nested interaction object and no output schema, the description covers what matters: the return value (file path and size), prerequisites, fallback behavior, timing semantics, and the debugger side effect. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning on top: it defines the default 'record' action as start + optional interaction + wait + stop, explains that start/stop bracket intervening work, and restates the 120 s duration cap. This clarifies the interaction of parameters rather than merely repeating them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource+artifact: 'Record the tab as a video (.mp4 or .webm).' It states the motivating use case ('watch an animation or interaction frame by frame') and explicitly contrasts its output with sibling frame-returning tools ('never the frames'), so an agent can tell it apart from screenshot without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for when to reach for this tool (frame-by-frame animation/interaction review) and explains the action lifecycle (record = start → interaction → wait → stop; start/stop bracket other work). It does not name alternatives or state when NOT to use it versus siblings like animations/watch, so it falls short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.