Skip to main content
Glama

agent-recorder-mcp

Local MCP server that records the computer an agent is using and returns a replay link when the task is done.

The recording stays on the agent machine. The model receives a link, not the video. FastFileLink is the delivery transport: the local file remains until the recipient finishes downloading.

This version targets the Linux X11 display used by GrokBot. It does not record the user's own desktop from another machine.

Install

Give the assistant one URL. Install the skill from tag v0.1.0 first. The skill then installs the MCP connector for whichever assistant is running, and it carries the rule that browser and computer-use work is recorded.

mkdir -p ~/.grok/skills/agent-session-recording
curl -fsSL \
  -o ~/.grok/skills/agent-session-recording/SKILL.md \
  https://raw.githubusercontent.com/nuwainfo/agent-recorder-mcp/v0.1.0/.grok/skills/agent-session-recording/SKILL.md

Claude uses ~/.claude/skills/agent-session-recording/SKILL.md. The same file is the source for every assistant. uv selects Python 3.11 or newer when the skill installs the connector.

The agent computer also needs ffmpeg and an X11 size tool (xdpyinfo or xwininfo, usually in x11-utils).

The first connector approval, each FastFileLink share, and each Google Drive upload ask for approval. The skill tells the assistant to expect those prompts.

Useful environment variables:

Variable

Role

AGENT_RECORDER_DIR

Recording root. Default: ~/.agent-recorder

AGENT_RECORDER_FFMPEG

Explicit ffmpeg binary

Recordings are written to <root>/recordings/rec_YYYYMMDD_HHMMSS_<id>.mp4.

Related MCP server: markupR MCP Server

Tools

Tool

Purpose

startRecording

Select this agent's display and start FFmpeg. Defaults: 5 fps, 7200 seconds.

recordingStatus

Duration, size, display, and state. idle when nothing is active.

finishRecording

Stop, finalize, and hand the file to a transfer strategy.

confirmUpload

Store the Google Drive link and apply after_upload cleanup.

abortRecording

Stop without sharing. Deletes the partial file unless deletePartial is false.

cleanupRecording

Delete one local recording by recordingId and stop its share.

finishRecording defaults to delivery=ffl and cleanup=after_download. delivery=local with cleanup=manual returns a file: URI and leaves the file in place. delivery=google_drive with cleanup=after_upload returns localPath for the assistant's Google Drive connector. confirmUpload stores the Drive URL and then deletes the local file.

A second startRecording while one recording is active returns that recording and sets alreadyRecording to true.

Display selection

startRecording looks at X11 sockets and the process tree that spawned this server (the agent, an intermediate launcher such as uvx, and browsers those processes started). It prefers a single browser display in that tree. Otherwise it uses the nearest DISPLAY on that chain.

It does not fall back to another socket. If the display cannot be identified, or a frame cannot be grabbed from it, the tool fails with:

No active agent display could be identified.

A blank frame on the identified display is still recorded. The browser often opens after recording starts.

Cleanup

Creating an FFL link does not delete the file. With after_download, the server waits until FFL reports /download/complete or /webrtc/transfer/complete, stops the share, then deletes the local file. A full HTTP download, including curl, is one of those reports. If the MCP process exits before that event, the file stays and a later finishRecording can share it again.

maxDurationSeconds stops FFmpeg and keeps a playable file so finishRecording can still share it. A crashed recorder does not leave FFmpeg running past that limit or after its owner process is gone.

Skill

.grok/skills/agent-session-recording/SKILL.md is the install guide and the recording rule. It tells the agent to install the pinned connector when startRecording is missing, to call startRecording before browser or computer-use work, and to call finishRecording before the final answer.

Development

On this machine, tests use the recoding-mcp conda environment.

conda activate recoding-mcp
pip install -e .
python -m unittest discover -s tests -p "*Test.py" -v

tests/Smoke.py is a manual check for a Linux agent computer. It records for a few seconds. It does not open a browser; watch the resulting file and confirm it shows the agent display.

python tests/Smoke.py --seconds 5 --delivery local --cleanup manual

Network share tests are not part of the default suite.

Out of scope

Dashboard, accounts, tamper-proof storage, action traces, OCR, a custom browser, and Windows/macOS/Wayland capture. This server records, shares, and deletes. Google Drive upload is performed by the assistant's own Drive connector.

Available Tools

6 tools
abortRecordingAbortrecordingA

Stop a recording without sharing it.

deletePartial defaults to true and removes the incomplete file. Use this when the task is cancelled or the capture is the wrong display.

ParametersJSON Schema
NameRequiredDescriptionDefault
recordingIdNo
deletePartialNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses the destructive default ('deletePartial defaults to true and removes the incomplete file') and that nothing is shared, but is silent on permissions, whether the aborted recording remains listed/recoverable, and what happens when deletePartial is false.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core action and then the destructive-default detail. Minor waste in the awkward mid-sentence line break, but nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and the destructive behavior is disclosed. However, for a mutation tool with zero annotations the description leaves the recordingId parameter and post-abort state undefined, which is a real invocation risk.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It explains deletePartial beyond the schema (default true, removes the incomplete file), but recordingId is completely unaddressed despite being nullable with a null default — an agent cannot tell whether to pass it or what null means.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Stop a recording') and adds the distinguishing qualifier 'without sharing it,' which separates it from finishRecording/confirmUpload. It never names a sibling explicitly, so an agent still has to infer why it would pick this over cleanupRecording, which likely also discards files.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit triggering condition: 'Use this when the task is cancelled or the capture is the wrong display.' No when-not guidance or named alternative (e.g., finishRecording for a good capture) is provided, so the routing is clear but not fully closed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cleanupRecordingCleanuprecordingB

Delete a local recording and stop any share still serving it.

Pass the recordingId. This removes the file even if a recipient has not downloaded it yet.

ParametersJSON Schema
NameRequiredDescriptionDefault
recordingIdYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it does disclose non-obvious side effects: termination of an active share and file removal even for undownloaded recipients. However, it omits irreversibility, required permissions, and any failure behavior for an unambiguously destructive delete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the action and its main side effect; nothing is redundant. The line breaks read slightly awkwardly but no sentence is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers the action, the required id, and the key sharing side effect. For an unannotated destructive tool it would ideally also flag irreversibility and permission requirements, but the essentials are present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single required parameter, so the description has to compensate. 'Pass the recordingId' tells the agent which identifier is required but adds no format, source, or provenance guidance, which is thin for the only input.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Delete a local recording') plus an accompanying side effect ('stop any share still serving it'), which an agent can understand immediately. It does not explicitly distinguish itself from the close sibling abortRecording, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description never says when to use this versus abortRecording or finishRecording, which are the plausible alternatives for ending a recording session. Usage is only implied by the word 'Delete' and the post-hoc 'cleanup' framing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

confirmUploadConfirmuploadA

Store the link from the connected Google Drive tool.

Call this after finishRecording(delivery="google_drive") once the Drive tool has uploaded localPath. url is that Drive file link. cleanup after_upload then deletes the local recording.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
recordingIdYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses a side effect (cleanup 'after_upload' deletes the local recording) and the external dependency on the Drive tool, but says nothing about idempotency, what happens if url is wrong, or permission/auth requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, front-loaded with the core action and immediately followed by the trigger condition. The trailing cleanup sentence is relevant but sits slightly apart from the call instructions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation. For a simple two-parameter state-recording tool, the sequencing, dependency, and url meaning are the essentials, and they are covered; only recordingId semantics and failure behavior are unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and neither parameter is described in the schema, so the description must compensate. It explains url ('that Drive file link') but leaves recordingId entirely undefined beyond its name, covering only half the inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb and object ('Store the link from the connected Google Drive tool') and anchors it to a specific recording workflow, so an agent can distinguish it from startRecording/abortRecording/recordingStatus. It never explicitly says it finalizes/confirms an upload for a given recordingId, leaving the exact resource slightly implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit precondition and ordering: 'Call this after finishRecording(delivery="google_drive") once the Drive tool has uploaded localPath.' The alternative (calling before the Drive upload exists) is ruled out by that phrasing, so no inference is required.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

finishRecordingFinishrecordingA

Stop the recording and hand it to a transfer strategy.

delivery defaults to ffl. cleanup defaults to after_download. The local file stays until FastFileLink reports that the download finished, then the share stops and the file is deleted. delivery local with cleanup manual returns a file URI and leaves the file in place.

delivery google_drive with cleanup after_upload does not call Google. It returns localPath and uploadStatus pending. Upload that file with the connected Google Drive tool, then call confirmUpload with the Drive link.

Put the returned url in the final answer as the operation replay. If this call fails, tell the user that no replay was produced and include the reason.

ParametersJSON Schema
NameRequiredDescriptionDefault
cleanupNoafter_download
deliveryNoffl
recordingIdNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so richly: it discloses file lifecycle ('stays until FastFileLink reports the download finished, then the share stops and the file is deleted'), the google_drive branch that deliberately does not call Google and returns uploadStatus pending, and what to do on failure. This is exactly the behavioral context an agent needs to invoke it safely.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and defaults, then organized paragraph-by-paragraph around delivery/cleanup combinations. Slightly long, but nearly every sentence carries behavioral information; only minor tightening is possible.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be enumerated, yet the description still tells the agent what comes back (url, localPath, uploadStatus pending) and how to use it in the final answer. Combined with the delivery/cleanup matrix and the failure instruction, nothing needed to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does for delivery (ffl/local/google_drive) and cleanup (after_download/manual/after_upload) with concrete outcome semantics for each combination. The third parameter, recordingId, is left undescribed, which keeps it from a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Stop the recording and hand it to a transfer strategy'), which clearly distinguishes it from startRecording, abortRecording, and cleanupRecording. It could have been explicit about how it differs from abortRecording (does it preserve the file?), but the core action is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context: defaults (delivery ffl, cleanup after_download), the local/cleanup-manual path, and the google_drive/after_upload path with an explicit follow-up instruction to call confirmUpload. It does not state when to prefer abortRecording or cleanupRecording instead, so the alternative-selection guidance is incomplete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

recordingStatusRecordingstatusA

Return the recording state, duration, display, and file size.

Omit recordingId to ask about the active recording. The state is idle when nothing is recording.

ParametersJSON Schema
NameRequiredDescriptionDefault
recordingIdNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, so the description carries the full burden. It usefully discloses the default-target behavior and that state is 'idle' when nothing is recording, and 'Return' implies a read-only operation, but it omits error behavior for an unknown/expired recordingId and any permission or rate-limit context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with what is returned, then the parameter rule, then the idle-state edge case. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be enumerated, yet the description names them anyway. For a single-parameter read tool this is nearly complete; only failure modes for an invalid recordingId are missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and recordingId is documented only as a nullable string, so the description's rule 'omit recordingId to ask about the active recording' supplies the parameter's actual semantics. It adds real meaning the schema cannot convey, though it does not describe accepted ID formats.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (Return) plus the exact resource and its fields (recording state, duration, display, file size). It is unmistakably distinct from the sibling mutators startRecording/finishRecording/abortRecording/cleanupRecording/confirmUpload.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells the agent when to omit recordingId (to query the active recording), which is the main usage decision for this tool. It does not state exclusions or errors, but the sibling names make the alternative operations obvious.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

startRecordingStartrecordingA

Start recording the agent computer display before browser or computer-use actions.

Call this before the first browser or computer action. The display is selected from this agent process tree. A display belonging to another session is not used.

fps defaults to 5. maxDurationSeconds defaults to 7200. When the limit is reached the capture stops and the file is kept so finishRecording can still share it.

If a recording is already active, or is waiting to be shared, that recording is returned and a second recorder is not started.

ParametersJSON Schema
NameRequiredDescriptionDefault
fpsNo
maxDurationSecondsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses display selection from the agent's own process tree, that another session's display is never used, the exact defaults, the behavior when maxDurationSeconds is reached (capture stops, file retained for finishRecording), and the no-duplicate-recorder semantics. This is rich behavioral context beyond what any structured field provides.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short paragraphs, front-loaded with the action and the timing rule, then escalating to edge-case behavior. Every sentence adds information an agent needs; nothing is padded or repeated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. Given two optional parameters, no annotations, and a set of closely related sibling tools, the description covers timing, defaults, cross-session safety, duration-limit handling, and idempotency — everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the schema only lists bare types with defaults, so the description must compensate. It supplies the default values and, more usefully, the consequence of maxDurationSeconds being reached, though it never explains what fps controls or states any valid range for either parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('start recording the agent computer display') and scopes it to browser/computer-use actions, which cleanly separates it from finishRecording, abortRecording, recordingStatus, and cleanupRecording. An agent can identify the tool without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit timing guidance: 'Call this before the first browser or computer action,' plus the when-not case that an already-active or pending recording is returned instead of starting a second one. It does not name sibling alternatives for the status/finish/abort flows, so routing is implied rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 6 tool updatesv0.1.0
    • First observedabortRecording
    • First observedcleanupRecording
    • First observedconfirmUpload
    • First observedfinishRecording
    • First observedrecordingStatus
    • First observedstartRecording

TDQS

A3.9/5.0

Scored across 6 tools

Disambiguation4/5

Each tool maps to a distinct lifecycle stage (start, status, finish, abort, cleanup, confirmUpload), so boundaries are mostly clear. Minor overlap exists between cleanupRecording, abortRecording(deletePartial), and finishRecording(cleanup) since all three can delete the local file, which could cause some hesitation.

Naming Consistency4/5

Five of six tools use a consistent verb+Recording/Upload camelCase pattern (startRecording, finishRecording, abortRecording, cleanupRecording, confirmUpload). recordingStatus breaks the verb-first convention by being noun-first, a minor deviation.

Tool Count5/5

Six tools is well-scoped for a capture-and-share recorder, with each tool earning its place across the start/stop/abort/share/cleanup lifecycle. No redundant or bloated entries.

Completeness4/5

The surface covers the full recording lifecycle from start through sharing and cleanup, including the Google Drive handoff path. It lacks discoverability operations like listing past recordings or retrieving an existing file, but core workflows are complete.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers