Skip to main content
Glama

seshat

An MCP server that records a screen and narrates the take.

seshat captures an output (or a region of one) into a silent artifact, ingests the timeline streams that other tool servers publish while the take runs, and then muxes a synthesized narration track — aligned to those real events, with styled captions burned in — over the recorded video. The prose is the caller's; the timing, the speech, the captions and the container are seshat's.

It is named after the Egyptian goddess of writing, measurement and record-keeping.

What it is, and what it is not

  • It is a recorder and a narrator. It owns capture, the take timeline, speech synthesis, caption layout and muxing.

  • It is not a desktop controller. It never clicks, types, or focuses anything. The actions a take documents are performed by other tool servers; they reach seshat as published event streams. That split is deliberate: a recoder that also drives the desktop cannot record a desktop driven by anything else.

  • It embeds no model. Narration prose is written by the calling agent and passed in. seshat is deterministic: same inputs, same timeline, same schedule.

Related MCP server: demo-vid-mcp

Requirements

  • Python 3.10+ (standard library only — dependencies = [] is a hard contract).

  • wf-recorder for capture, ffmpeg + ffprobe for finalization, muxing and probing. swaymsg (or a SWAYSOCK exposing the same JSON) to resolve output geometry. A missing binary is reported as a clear tool error, never an import failure.

  • Optional: edge-tts (keyless, network) or piper (offline, needs a voice model) for narration; tesseract for OCR of scene keys.

make check runs the unit tests, bytecode compilation, and the file-length check. make help lists every target; make ci-local runs the whole gate plus a distribution build, and make runtime shows the streams, recordings and active take when something looks wrong.

Install

uv venv
uv pip install -e .
seshat --doctor     # what this host can record and narrate with
seshat --self-test  # non-mutating checks

Register with an MCP harness:

hermes mcp add seshat --command /path/to/venv/bin/seshat
codex mcp add seshat -- /path/to/venv/bin/seshat

Recording a take

recording_start(format="mp4", max_duration_seconds=120)
   ... the demonstration happens ...
recording_stop()
recording_status()        # poll until phase == completed | failed
recording_timeline()      # the ingested events, with recording-relative t_ms

One take may be active at a time. Artifacts are written to $XDG_RUNTIME_DIR/seshat/recordings (0700, files 0600) and do not survive a logout; copy anything you want to keep.

A completed take reports two durations, and they are deliberately different: capture_elapsed_seconds is how long capture ran (including the recorder's shutdown, which can take seconds), and media_duration_seconds is the playable picture. latest_event_ms and events_beyond_media say how far the ingested timeline reaches and how much of it has no picture to point at. A take stopped by its own deadline can lose its tail, so the two are not expected to match.

make check                         # unit tests, compilation, file-length ceiling
SESHAT_INTEGRATION=1 make integration   # records the live screen; opt in explicitly

The timeline stream contract

seshat cannot timestamp actions it does not perform. Any tool server that wants its actions to be narratable publishes them — one JSON object per line — to $XDG_RUNTIME_DIR/seshat/streams/<source>.jsonl:

{"at_monotonic": 12345.678, "tool": "click", "ok": true,
 "payload": {"x": 640, "y": 360}, "source": "computer-use-sway"}
  • at_monotonic is required: CLOCK_MONOTONIC seconds, the value of time.monotonic() in the emitting process. That clock is host-wide, which is what makes an unrelated process's timestamps directly comparable with the recording epoch — no handshake, no session id.

  • tool is required. ok (default true), payload (default {}) and source (default: the file stem) are optional.

  • Each emitter owns its file and should truncate it when a new session starts.

  • Events are filtered to the take's own capture window. Anything published before recording_start or after capture stopped describes something the video does not contain, and is not an anchor.

  • Malformed, oversized or unreadable lines are counted in the timeline's sources report and skipped. A broken stream can never destroy a take.

Pass timeline_sources to recording_start to ingest specific files instead of everything in the streams directory; explicit paths are validated when the take starts, not when it is finalized.

Narrating a take

Read the timeline first, then write prose against the event ids that are really there:

recording_timeline()
recording_voiceover(segments=[
  {"anchor": {"event_id": 1}, "text": "..."},
  {"anchor": {"at_ms": 8200},  "text": "..."}
])
recording_status()        # poll until phase == completed | failed

recording_voiceover synthesizes each segment, builds one audio track placed at the resolved anchors (with offset_ms, optional tempo compression via fit="compress", and lead-silence trimming), muxes it over the existing video, and — unless subtitles=false — burns styled ASS captions. The video is only re-encoded when captions are burned; otherwise it is stream-copied. The result is re-validated: exactly one video stream plus exactly one audio stream.

Anchors are checked against the picture, not the timeline. Both event_id and at_ms are validated against the playable video extent — the video stream's own duration, falling back to the container's — so speech cannot be placed after the last frame. An event id that exists but resolves past the end of the video is refused, with both numbers in the message, rather than anchored into silence.

If recording_timeline reports no events, the demonstration was driven by a tool server that does not publish a stream. Anchor segments with at_ms, or use recording_scenes for approximate cuts. Scene cuts are secondary evidence; they are never the sync source.

What the events do and do not prove

An event records that a tool call was dispatched, not that it visibly worked. The driving server is the only party that knows whether its action had the intended effect; a driver that returns sent: true without verifying the result will publish an event for an action that did nothing. Two open issues in the sibling project are exactly this shape: a modified key that arrived as an unmodified one, and a keystroke delivered to a window that had stolen focus.

So narrate what the take shows. Use event ids to place a line in time, and ground its content in the video, in the driver's own verified payload, or in something you observed yourself — never in the mere existence of an event.

Security notes

  • edge-tts sends the narration prose to Microsoft. It is keyless but network dependent. Install piper and a voice model (voice, or SESHAT_PIPER_MODEL) to keep narration on the host.

  • Narration text is the caller's; seshat never reads it from the screen, so nothing visible on screen becomes narration unless the agent says so.

  • Recordings and streams live under $XDG_RUNTIME_DIR/seshat with 0700 / 0600 permissions.

  • computer-use-sway — drives a Sway session and publishes the timeline stream that this server ingests. Recording and narration were extracted from that project so that both could stand on their own.

Available Tools

7 tools
recording_scenesA

Optional, approximate fallback anchors for a completed take: ffmpeg scene-cut timestamps, with optional keyframe OCR via tesseract. Use when no timeline stream was published for the take; scene cuts are secondary evidence, never the sync source.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoTake id from recording_start; defaults to the latest take.
ocrNo
thresholdNo
max_scenesNo

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully characterizes the results as optional, approximate, and secondary evidence, and names the underlying tools (ffmpeg, tesseract), which hints at cost/quality. However, it says nothing about runtime cost of OCR, pagination/limits, or what a result looks like, and there is no output schema to cover that gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the core purpose before the conditional usage note and the evidence-quality caveat. Every clause adds decision-relevant information with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a four-parameter tool with no annotations and no output schema, the description covers purpose, trigger, and result-quality but omits the meaning of threshold/max_scenes and any notion of return shape or cost. Adequate but with clear gaps given the structured fields do little of the work.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 25% and two of four parameters (threshold, max_scenes) have no semantic description anywhere. The description does add real meaning for 'ocr' by saying it is optional keyframe OCR via tesseract, partially compensating, but leaves the tuning parameters unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (scene-cut timestamps / fallback anchors for a completed take) and the mechanism (ffmpeg scene cuts, optional tesseract OCR). It also implicitly distinguishes itself from the recording_timeline sibling by framing itself as the fallback when no timeline stream exists.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger ('Use when no timeline stream was published for the take') and an explicit exclusion/limitation ('scene cuts are secondary evidence, never the sync source'). The agent knows exactly when to reach for this versus the timeline tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

recording_startA

Start recording an active output (or a region within it) into a silent artifact: H.264 MP4 by default, or AV1 WebM, or a constrained GIF fallback. Returns immediately; call recording_stop to finish, then poll recording_status until the phase is completed or failed. The pointer cursor is always included. Narration anchors come from the timeline streams ingested for this take, so the tool server that drives the demonstration has to publish one.

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNomp4: silent H.264 MP4 at 30 fps (default, widest compatibility); webm: silent AV1 WebM at 30 fps; gif: constrained fallback at 12 fps, 960 px maximum width.mp4
outputNoExact Sway output name; inferred when exactly one output is active.
regionNoRegion fully contained in the selected output.
timeline_sourcesNoStream files to ingest for this take. Defaults to every *.jsonl file in streams/ under the seshat runtime directory. Paths are validated when the take starts.
max_duration_secondsNoAutomatic stop deadline. Defaults to 60 (mp4/webm) or 15 (gif); GIF recordings are capped at 15 seconds.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the burden and discloses key traits: it is asynchronous (returns immediately), records a silent artifact, always includes the pointer cursor, and depends on timeline streams being published for narration anchors. It does not cover permission requirements, error behavior, or where artifacts are stored.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four tightly written sentences, front-loading the action and format options, then the workflow, then two key behavioral notes. Every sentence contributes information without repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter recording tool with no annotations and no output schema, the description supplies the necessary workflow, format context, async behavior, and the timeline dependency. It leaves some operational details like permissions or artifact destination unaddressed, but is otherwise complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds operational meaning beyond the schema by tying timeline_sources to narration anchors and requiring the driving tool server to publish a stream, and by clarifying region selection within the active output.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: start recording an active output or region into a silent artifact, with formats enumerated. It distinguishes the action from siblings by naming recording_stop and recording_status as the workflow continuation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly explains that the call returns immediately and that the agent must call recording_stop, then poll recording_status until completion or failure. This gives explicit next steps and alternatives, but does not state when not to use the tool or how it differs from other recording_* siblings such as recording_scenes.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

recording_statusB

Report the take lifecycle phase (idle, recording, stopping, processing, narrating, completed, failed) plus live progress or final artifact metadata.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoTake id from recording_start; defaults to the latest take.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose meaningful behavior: the phase state machine (idle through failed) and that output varies between live progress and final artifact metadata. However, it never states that this is a non-mutating read, nor any polling/rate considerations, leaving the safety profile implicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the parenthetical state list is dense but earns its place by enumerating the exact return domain.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only status tool with one optional parameter and no output schema, the description covers the return content well by naming phases and distinguishing live vs. final metadata. The main missing element is workflow context (when/how to poll) rather than anything an agent needs to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single 'id' parameter is fully documented in the schema (source tool and default-to-latest behavior). The description adds no additional semantics beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a concrete verb and resource ('Report the take lifecycle phase') and enumerates the exact phase values, so an agent knows precisely what it returns. It does not explicitly differentiate itself from siblings like recording_timeline, which may also surface progress, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance: nothing tells the agent to poll this after recording_start, how often to poll, or whether it competes with recording_timeline. Usage is only weakly implied by the 'lifecycle phase' framing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

recording_stopA

Gracefully stop the active recording and begin finalizing the requested artifact. Returns the stopping/processing state; poll recording_status for the final artifact path and metadata.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does well: it discloses that the operation is asynchronous, that the immediate return is a stopping/processing state rather than the artifact, and that polling recording_status is required. It omits permission requirements and error/duplicate-call behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the action and followed by the follow-up procedure; every clause earns its place with no redundant restatement of the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly explains that the return is an intermediate state and where the real output comes from (recording_status), which is the essential missing piece. Minor gaps remain around failure modes, but for a zero-parameter action this is well covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema leaves nothing to explain and the baseline is 4. The description correctly does not invent parameter details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (stop) and resource (the active recording) plus the consequential next step (finalizing the artifact), which cleanly separates it from the sibling recording_start. An agent can tell immediately what this does without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context: use it to stop the active recording, and then poll recording_status for the final artifact path and metadata. It does not state the when-not case (e.g., behavior when no recording is active, or whether calling it twice is safe), so it stops short of full alternative/exclusion guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

recording_timelineA

Return the take's monotonic event timeline: every event ingested from the take's timeline streams, filtered to the capture window, with a recording-relative t_ms, the emitting source, and a compact payload. Use event ids as narration anchors. A sidecar JSON copy is written next to the artifact on completion.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoTake id from recording_start; defaults to the latest take.

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose real behavior: monotonic ordering, capture-window filtering, recording-relative timestamps, and the side effect that a sidecar JSON copy is written next to the artifact on completion. It stops short of stating state prerequisites, permissions, or whether the call is safe during an active recording, leaving meaningful gaps for a no-annotation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The verb and resource are front-loaded in the first clause, and the three sentences each carry content (what is returned, a usage tip, a side effect). The narration-anchor sentence is slightly tangential but still earns its place as practical guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must convey the return shape, and it does: recording-relative t_ms, emitting source, compact payload, plus a sidecar artifact. The only omissions are operational prerequisites and volume/pagination expectations for what could be a long event stream.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'id' parameter is fully documented there, including its default to the latest take. The description alludes to 'the take's' timeline but adds no syntax, format, or defaulting detail beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Return) and resource (the take's monotonic event timeline), and scopes it precisely: events from timeline streams, filtered to the capture window. The resource is clearly distinct from siblings like recording_scenes or recording_status, so an agent can route without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when the tool is useful (after a take exists, to obtain a chronological event stream) and hints at a downstream use ('Use event ids as narration anchors'). However, it names no alternatives and gives no when-not guidance, and never states whether the take must be stopped or can still be recording.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

recording_voiceoverA

Attach a scripted narration track to a completed take: the caller supplies the prose, the server synthesizes speech, aligns segments to timeline anchors (or to an explicit at_ms), and muxes the audio over the existing video stream, optionally burning styled captions. Starts an async 'narrating' phase; poll recording_status. GIF cannot carry audio.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoTake id from recording_start; defaults to the latest take.
fitNonatural
voiceNo
engineNoauto
tail_msNo
segmentsYes
offset_msNo
subtitlesNoBurn styled captions from the narration into the video (re-encodes the video). Set false to copy the video without captions.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the async phase, the required polling via recording_status, and a media-type limitation. It stops short of stating permission requirements, what happens to any existing audio track, or failure/partial-alignment behavior, so it is not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense, front-loaded sentences with no filler; the core action leads and the workflow/limitation follow. Jargon ('muxes', 'narrating') compresses well but borders on under-explaining for a non-expert agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter mutation-style tool with no annotations and no output schema, the description covers the workflow shape and the async contract but leaves several parameters and the return/result behavior unaddressed. Adequate to attempt a correct call, thin on the details needed to call it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25%, so the description has to compensate and only partly does: it clarifies that segments carry prose plus timeline anchors or an explicit at_ms, and that captions can be burned. The enum/general parameters (fit, engine, voice, tail_ms, offset_ms) are left to the schema's bare names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Attach a scripted narration track to a completed take') and then spells out the actual mechanism: synthesize speech, align to anchors, mux over the video stream. This is distinct enough from recording_scenes/timeline that an agent can route to it without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear operating context ('completed take', async 'narrating' phase, 'poll recording_status') and one hard exclusion ('GIF cannot carry audio'). It does not name alternatives for cases where narration is not the right call, but the preconditions and follow-up workflow are explicit enough to act on.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

seshat_infoB

Inspect the recorder: active take, stream and recording directories, the streams that will be ingested, the capture outputs, and the binaries narration depends on.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. 'Inspect' implies a non-mutating diagnostic read, and the enumeration of what it surfaces partially substitutes for the absent output schema. However, it never explicitly states that it is read-only/side-effect free or describes any caveats about when the reported state may be stale.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that names the action first and then lists what is inspected. It is efficient, though the trailing enumeration of five items is dense and mixes distinct output categories rather than prioritizing them.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only diagnostic tool with no output schema, the description usefully enumerates the categories of information returned, so an agent knows what to expect. It falls short only on when-to-use and on explicit confirmation of side-effect-free behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline of 4 applies. There is nothing for the description to disambiguate, and the schema correctly declares an empty object with additionalProperties false.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ('Inspect') and enumerates the resources returned: active take, stream/recording directories, ingest streams, capture outputs, and narration binaries. It is clear what the tool does, though the name 'seshat_info' is opaque and the description does not distinguish this inspection tool from the sibling recording_status, whose scope likely overlaps.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to call this versus alternatives such as recording_status or recording_timeline. Usage is only implied by the word 'Inspect', and no prerequisites, conditions, or exclusions are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 7 tool updatesv0.1.1
    • First observedrecording_scenes
    • First observedrecording_start
    • First observedrecording_status
    • First observedrecording_stop
    • First observedrecording_timeline
    • First observedrecording_voiceover
    • First observedseshat_info

TDQS

A3.9/5.0

Scored across 7 tools

Disambiguation5/5

Each tool has a distinct role: start/stop/status for lifecycle, timeline for event anchors, scenes for fallback visual anchors, voiceover for narration, and seshat_info for environment inspection. No two tools appear interchangeable or likely to be confused.

Naming Consistency4/5

Six of seven tools use the recording_ prefix, which is consistent and predictable. The single seshat_info tool breaks the pattern, but the deviation is minor and still readable.

Tool Count5/5

Seven tools is well-scoped for a recording and narration lifecycle. Each tool earns its place with no redundant or bloated operations.

Completeness4/5

The core lifecycle is covered: start, stop, status, timeline, fallback scenes, voiceover, and environment info. Minor gaps like listing or deleting completed artifacts exist, but they are not essential for the stated recording/narration purpose.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    Turns your AI host into a product videographer — scripted screen recordings of your own web app with a gliding cursor, camera zooms, highlight callouts, captions, and branded transitions, plus marketing-grade screenshots. Automatic dark-frame cleanup and MP4/GIF export. Free, MIT, 100% local — no account, no API keys.
    14
    75 npm
    3
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables AI agents to create polished product demo videos by controlling a real Chromium browser, recording actions, and rendering 1080p MP4s with narration, captions, and styled overlays.
    1
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables agents to record narrated, zoom-animated product demo videos of websites entirely on-device, including outlining pages, scripting tours, recording MP4s, reviewing frames, and extracting embedded recipes to fork or rerender demos.
    7 npm
    4
    MIT