Skip to main content
Glama

Marey

Give AI agents eyes for motion.

Marey is an MCP server that captures short screen recordings as a sequence of still frames and turns them into a single, agent-readable contact sheet.

It is named after Étienne-Jules Marey, a pioneer of chronophotography — the study of motion through sequences of images. Marey applies the same idea to AI agents: instead of handing a model a video it cannot reliably inspect frame by frame, it converts motion into a visual sequence the model can reason about.


Why Marey?

AI coding agents are increasingly good at understanding screenshots, but motion is still awkward.

A bug such as:

  • an anchor jumping while it is dragged,

  • a menu flashing and immediately closing,

  • a canvas updating in the wrong order,

  • an animation stuttering between states,

cannot be understood from a single screenshot.

Marey bridges that gap.

You interact            Marey captures               Agent sees

drag / click / type  →  frame 001 · 0.00 s        →  ┌────┬────┬────┐
                         frame 002 · 0.25 s           │ 01 │ 02 │ 03 │
                         frame 003 · 0.50 s           ├────┼────┼────┤
                         frame 004 · 0.75 s           │ 04 │ 05 │ 06 │
                         ...                           └────┴────┴────┘
                                                      contact sheet

Because Marey speaks the Model Context Protocol, an agent can request the recording itself and receive the resulting image directly in context.

No GIF inspection. No manually extracting frames. No dragging a dozen screenshots into chat.


Related MCP server: video-capture-mcp

How it works

  1. An MCP client asks Marey to record the screen.

  2. Marey captures frames at a chosen frame rate.

  3. Each frame is numbered and timestamped.

  4. Marey composes the frames into a contact sheet.

  5. The contact sheet is returned to the agent as MCP image content.

  6. Raw frames remain available on disk for closer inspection or re-stitching.

The idea is deliberately simple:

motion becomes one image containing time.


Tools

Marey exposes three MCP tools:

Tool

What it does

record

Records the screen for a fixed duration and returns a timestamped contact sheet

capture

Captures a single screenshot

list_windows

Lists visible windows that can be targeted for capture

record

Parameter

Default

Description

seconds

5

Recording duration

fps

2

Frames captured per second

region

primary

primary, virtual, or window

title

—

Window-title substring when region is window

delay

0

Delay before capture begins

cols

4

Number of thumbnails per contact-sheet row

thumbWidth

480

Thumbnail width in pixels

The result contains:

  • the contact sheet as MCP image content,

  • a short text summary with capture metadata,

  • raw frames saved locally for later inspection.

capture

Captures one frame immediately using the same region-selection semantics as record.

list_windows

Returns visible window titles, and geometry where available, so an agent can choose a target for window capture.


Example

Once Marey is connected to an MCP client, interaction can be as simple as:

Use Marey to record 6 seconds of the Euclid window at 4 fps while I drag an anchor, then tell me what changes between frames.

The agent receives the complete sequence as a single image and can reason about the transition rather than only the initial state.


Command-line use

Marey also runs standalone, which is handy for verifying your capture backend before wiring up an MCP client:

node src/cli.mjs backend                 # report the detected capture backend
node src/cli.mjs windows                 # list targetable windows
node src/cli.mjs capture --region primary
node src/cli.mjs record --seconds 4 --fps 4 --cols 4 --thumbWidth 320
node src/cli.mjs record --region window --title "Euclid" --seconds 6 --fps 4

A note on frame rate

fps is the target rate. The achievable rate is bounded by how fast the host can grab and encode a frame — on a 2560×1440 primary monitor, a full-screen grab plus PNG save costs a few hundred milliseconds, so the practical ceiling is roughly 2–3 fps at full resolution. Capturing a smaller region (or a single window) is faster. On Windows, an entire recording runs inside one PowerShell process rather than one per frame, so capture is not throttled by process-startup overhead. Frame labels show the actual elapsed time of each frame, so the timeline is always truthful even when the target rate is not met.


Installation

Requirements

  • Node.js 18+

  • A supported screen-capture backend

Zero runtime dependencies. Marey ships with an empty dependencies block — npm install pulls nothing. Everything is built on Node builtins:

  • PNG decode/encode — pure JavaScript over the builtin zlib (src/png.mjs); no sharp, jimp, or pngjs.

  • Thumbnails, compositing, and frame labels — a pure-JS image buffer and a hand-coded 5×7 bitmap font (src/image.mjs, src/font.mjs); no image or font library.

  • MCP protocol — JSON-RPC 2.0 over stdio, hand-rolled (src/jsonrpc.mjs); no MCP SDK.

  • Screen capture — child_process driving the native or command-line backend available on the host (src/capture.mjs).

Clone

git clone https://github.com/anilyesilkaya/marey.git
cd marey
npm install

Connect to an MCP client

Claude Code

claude mcp add marey -- node /absolute/path/to/marey/src/server.mjs

Claude Desktop or another MCP client

Add Marey to the client's MCP configuration:

{
  "mcpServers": {
    "marey": {
      "command": "node",
      "args": ["/absolute/path/to/marey/src/server.mjs"]
    }
  }
}

Capture backends

Marey auto-detects an available screen-capture backend.

Platform

Full / monitor capture

Window capture

Windows 10/11

Built-in PowerShell + System.Drawing

Built-in

Linux · X11

scrot, ImageMagick import, or ffmpeg

ffmpeg + xdotool

Linux · Wayland

grim

Compositor-dependent

macOS

Built-in screencapture

Not yet supported

On Linux, window listing uses wmctrl or xdotool where available. If ffmpeg is on PATH, Marey can use it for X11 capture where supported.


Output

Recordings are stored under captures/:

captures/
└── 20260101-120000/
    ├── frame_001_00000ms.png
    ├── frame_002_00500ms.png
    ├── frame_003_01000ms.png
    └── contactsheet.png

captures/latest-contactsheet.png

The raw frames make it possible to generate a different contact-sheet layout without recording the interaction again.


Design principles

Agent-first

Marey is an MCP server rather than just a screen-recording CLI. The agent can request the visual evidence it needs.

Still images over video

The output is intentionally model-friendly: a numbered, timestamped sequence of frames in one image.

Small surface area

A handful of tools cover the core workflow: record, capture, inspect.

Cross-platform core

Frame composition stays platform-independent while screen capture is delegated to the best backend available on the host.

Zero dependencies

The entire pipeline — PNG codec, image compositing, frame labelling, and the MCP protocol itself — is built on Node builtins. Nothing is pulled from npm, so there is no supply chain to audit, no install step beyond cloning, and no version drift in third-party packages.

Short, deterministic recordings

Fixed duration and frame rate make captures reproducible and easy for agents to request. An optional delay gives the user time to focus the target window before recording starts.


Why the name?

Étienne-Jules Marey used chronophotography to make motion visible by decomposing it into successive images.

Marey does the same thing for AI agents.


License

MIT — see LICENSE.

Available Tools

3 tools
captureA

Capture a single screenshot immediately and return it as an image. Uses the same region-selection semantics as record.

ParametersJSON Schema
NameRequiredDescriptionDefault
delayNoSeconds to wait before capturing (default 0).
titleNoWindow-title substring to match when region is "window".
regionNoCapture region (default primary).

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses the immediate, single-image, image-return behavior, but omits the behavioral traits that matter for a screen-capture tool: OS screen-recording permission requirements, what happens when region='window' matches no title, and whether the call blocks for the delay. Those gaps keep it at a middling score.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero filler, with the core action and the return type front-loaded and the cross-reference to record placed second. Nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter, no-annotation, no-output-schema tool, the description covers the action and the return type, which are the essentials. However it leaves the permission prerequisites and failure behavior for the window region unexplained, so an agent could still invoke it in a context where it silently fails.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and each of the three parameters (delay, title, region) is documented with defaults and the enum values, so the baseline is 3. The description only adds the pointer that region semantics match record, which is useful but does not extend individual parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Capture a single screenshot'), names the immediate single-shot nature, and declares the return form ('return it as an image'). It also implicitly separates itself from the sibling record by tying region semantics to it while implying a one-shot rather than continuous operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'immediately' and by the contrast with record, but there is no explicit when-to-use / when-not statement or guidance on choosing between capture, record, and list_windows (e.g., use list_windows first to find a title). The agent must infer the selection rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_windowsA

List visible windows (title, process, and geometry where available) so a target can be chosen for window capture.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses the scope ('visible windows') and that geometry may be absent ('where available'), but omits permissions, ordering, and an explicit read-only/no-side-effect statement. This is minimally adequate for a simple 0-parameter list tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero waste. The parenthetical return-field list and purpose clause both earn their place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 0-parameter listing tool with no output schema and no annotations, the description supplies the return fields and the use context. It leaves only minor details like ordering or pagination unspecified, which is acceptable at this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and an empty schema, so the baseline is 4. The description does not need to explain parameter meaning, and it appropriately adds none.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('visible windows'), plus the fields returned (title, process, geometry where available). It also positions itself as a discovery step before 'window capture', implicitly distinguishing it from the sibling capture tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'so a target can be chosen for window capture' clearly implies when to use it: before attempting capture. It does not name an alternative tool or give explicit exclusions, but the context is clear enough for an agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

recordA

Record the screen for a fixed duration and return a numbered, timestamped contact sheet (a grid of still frames) as an image. Use this to inspect motion: dragging, animations, flashes, reordering, anything that cannot be understood from a single screenshot. Raw frames are also saved to disk.

ParametersJSON Schema
NameRequiredDescriptionDefault
fpsNoFrames captured per second (default 2).
colsNoThumbnails per contact-sheet row (default 4).
delayNoSeconds to wait before recording starts (default 0).
titleNoWindow-title substring to match when region is "window".
regionNoprimary monitor, the full virtual desktop, or a window (default primary).
secondsNoRecording duration in seconds (default 5).
thumbWidthNoThumbnail width in pixels (default 480).

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden and does well: it discloses the fixed-duration behavior, the contact-sheet output format, and the side effect that raw frames are written to disk. It omits any mention of screen-recording permissions or cleanup of saved frames.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core action and output, followed by the discriminating use case and a side-effect note. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description does the heavy lifting by describing the returned image and the on-disk frames, and the schema fully covers the 7 parameters. Remaining gaps (permissions, frame retention) are minor for a read-only capture helper.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with all 7 parameters documented in the schema, so the baseline is 3. The description only loosely echoes duration and the grid layout (cols) and adds no syntax or constraint detail beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (record) and resource (the screen) plus the exact deliverable: a numbered, timestamped contact sheet image. It also implicitly separates this from the sibling 'capture' by framing it around 'anything that cannot be understood from a single screenshot.'

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly gives the usage condition — inspect motion such as dragging, animations, flashes, reordering — which is clear context for choosing this over a static screenshot. It does not name the alternative tool (capture) directly, so routing relies on inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.1.0
    • First observedcapture
    • First observedlist_windows
    • First observedrecord

TDQS

A4.1/5.0

Scored across 3 tools

Disambiguation5/5

The three tools have clearly distinct purposes: capture for immediate stills, record for motion over time, and list_windows for target discovery. No overlap or confusion in their descriptions.

Naming Consistency4/5

All names are lowercase snake_case, but 'record' and 'capture' are bare verbs while 'list_windows' follows a verb_noun pattern. This is a minor inconsistency that remains readable and predictable.

Tool Count5/5

Three tools are well-scoped for a screen capture and recording server, and each earns its place. The count is minimal but sufficient, falling comfortably within the typical 3–15 range.

Completeness4/5

The core lifecycle of capturing stills, recording motion, and selecting windows is covered. A minor gap exists in directly capturing a named window without first listing and manually choosing a region, but agents can work around it.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    F
    maintenance
    A cross-platform MCP server that allows AI agents to capture screenshots of specific windows, displays, or regions for native application testing. It provides tools to list active windows and monitors, enabling precise visual verification and interaction during automated workflows.
    32 npm
    1
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    MCP server for capturing screenshots of desktop windows on Windows. Allows AI assistants to see what's on screen for UI development, debugging, and iterating on designs.
    MIT