quickcast-mcp
Uses existing Cloudflare R2 watch URLs for screen recording sessions, enabling timestamped video links and metadata to be included in generated reports and documentation.
Formats and generates structured GitHub bug reports with timestamped video reproduction links, environment tables, and failure descriptions for engineering issues.
Formats and generates Jira issue content with timestamped video reproduction links, environment details, and failure descriptions for engineering tickets.
QuickCast MCP (Screen-to-Action Protocol)
Autonomous AI Protocol for Screen Recordings
Turn screen recordings directly into structured GitHub bug reports with video timestamps, executable Playwright E2E tests, and Standard Operating Procedure (SOP) documentation.
Part of the QuickCast Screen-Recorder & Watermark & Resize Studio ecosystem.
⚡ Features
🎥
get_latest_session: Instantly fetches recording metadata, duration, Cloudflare R2 watch URL, and user interaction markers from the latest session.📋
list_recent_sessions: Inspects recent screen recordings with durations, timestamps, and shareable preview links.🐛
format_github_bug_report: Formats an engineering-ready GitHub or Jira issue containing timestamped reproduction links (e.g.[00:15](watchUrl#t=15)), environment tables (URL, OS, resolution, browser), and failure descriptions.🎭
scaffold_playwright_test: Automatically outputs an executable Playwright reproduction test (in TypeScript or JavaScript) mirroring the exact interaction flow.📖
generate_sop_guide: Synthesizes a clean Markdown Standard Operating Procedure (SOP) for team training and workflow documentation.📥
ingest_session: Allows AI agents or scripts to register and store new recording telemetry into the local profile database (~/.quickcast/sessions.json).
Related MCP server: Krometrail
📦 Installation & Setup
1. Claude Desktop Configuration
Add the following to your claude_desktop_config.json:
Windows:
%APPDATA%\Claude\claude_desktop_config.jsonmacOS:
~/Library/Application Support/Claude/claude_desktop_config.json
{
"mcpServers": {
"quickcast": {
"command": "node",
"args": [
"C:\\Users\\Milmann\\.gemini\\antigravity\\scratch\\quickcast-mcp\\src\\index.js"
]
}
}
}2. Cursor Configuration
Add to your project's .cursor/mcp.json or global Cursor MCP Settings:
{
"mcpServers": {
"quickcast": {
"command": "node",
"args": [
"C:\\Users\\Milmann\\.gemini\\antigravity\\scratch\\quickcast-mcp\\src\\index.js"
]
}
}
}🛠️ MCP Tools Reference
Tool | Description | Inputs |
| Get latest recording details & R2 watch URL |
|
| List last 10 recordings with durations & links |
|
| Generate markdown bug report with video timestamps |
|
| Generate executable Playwright E2E script |
|
| Generate SOP documentation |
|
| Store session telemetry to local DB |
|
🔒 Safety & Privacy
100% Client-Side & Local: Runs as a local
stdioserver on your machine.Zero Cloud Costs: Uses your existing Cloudflare R2 links or local recording files. No external AI API keys or third-party cloud brokers required.
Decoupled Architecture: Strictly isolated from extension recording engines and web apps.
📄 License
Available Tools
6 toolsformat_github_bug_reportB
Format an automated, high-fidelity GitHub/Jira bug reproduction issue markdown document with timestamped video links, environment tables, and failure points from a QuickCast session.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | Additional developer notes or context. | |
| watch_url | No | Optional QuickCast video watch URL. | |
| session_id | No | Optional session ID. If omitted, uses latest recording. | |
| issue_title | No | Custom issue title (e.g. "[Bug]: Checkout button unresponsive on Safari"). | |
| session_data | No | Optional raw session JSON string. | |
| actual_behavior | No | What actually occurred in the recording. | |
| expected_behavior | No | What the user expected to happen. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full behavioral burden. It discloses the generated content and implies a formatting/output operation rather than a state mutation, but it never states whether it writes to disk, posts an issue, requires an existing session, or what happens when no session is available.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with zero filler, front-loaded with the verb and the deliverable, followed by the specific components produced. It is long but every clause earns its place; a short lead sentence about the session prerequisite would improve structure further.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter, all-optional tool with no annotations and no output schema, the description covers what the document contains but not what the agent receives back, where the document goes, or how the sibling session-lookup tools relate to it. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% across all 7 parameters, so the schema already documents each field, making 3 the baseline. The description gestures at a few of them (video links map to watch_url, failure points to actual/expected_behavior) but adds no format or constraint details beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb (format) and a precise resource (GitHub/Jira bug reproduction markdown document) and enumerates the artifacts it assembles: timestamped video links, environment tables, failure points from a QuickCast session. That is far more informative than a restatement of the name, though it never explicitly distinguishes itself from generate_sop_guide or scaffold_playwright_test.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance, no prerequisites, and no routing to alternatives. The phrase 'from a QuickCast session' hints at the input context but leaves the agent to infer whether it should first call get_latest_session or list_recent_sessions from the sibling set.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_sop_guideB
Generate a clean, structured Standard Operating Procedure (SOP) or training walkthrough document with timestamped video checkpoints from a QuickCast session.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | Optional custom document title. | |
| session_id | No | Optional session ID. If omitted, uses latest recording. | |
| session_data | No | Optional raw session JSON string. | |
| workflow_name | No | Name of the workflow (e.g. "Customer Refund Process in Admin Panel"). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions output content ('clean, structured... with timestamped video checkpoints') but omits permissions, side effects, whether it writes files, required session context, or any safety profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. It efficiently conveys the core action and output.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description should compensate more. It does not explain what happens when optional parameters are omitted, what the generated document contains beyond a high-level phrase, or what the return value looks like.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters. The description adds no additional meaning or syntax for any parameter, which matches the baseline of 3 when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Generate') and resource ('Standard Operating Procedure (SOP) or training walkthrough document') with additional qualifiers ('timestamped video checkpoints from a QuickCast session'). This clearly distinguishes it from sibling tools like format_github_bug_report, scaffold_playwright_test, and ingest_session.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description only states what the tool does; it gives no explicit when-to-use guidance, prerequisites, or alternatives. There is no indication of when to choose this tool over list_recent_sessions, get_latest_session, or other siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_latest_sessionB
Retrieve metadata, duration, markers, and Cloudflare R2 watch URL of the latest or specified QuickCast recording session.
| Name | Required | Description | Default |
|---|---|---|---|
| watch_url | No | Optional Cloudflare R2 watch URL to search for. | |
| session_id | No | Optional specific QuickCast session ID (e.g., "qc_rec_1740001234"). If omitted, retrieves latest recording. | |
| session_data | No | Optional raw JSON string of a QuickCast session payload to parse directly. | |
| session_file | No | Optional path to custom quickcast-sessions.json file. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. 'Retrieve' makes the read-only nature clear and it discloses exactly which fields are returned, but it says nothing about failure modes (e.g., no sessions existing), R2/auth requirements, or which input source takes precedence. Adequate but with real gaps for a zero-annotation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, well-formed sentence that front-loads the verb and lists the returned content without padding. Slightly dense field enumeration but nothing wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully names the returned fields, which compensates well. But it never explains how the four optional selectors (watch_url, session_id, session_data, session_file) relate or which takes precedence when several are supplied, leaving an operational gap for a 4-param tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four optional parameters are already documented in the schema, and the description adds only the notion of 'latest or specified' rather than field syntax or precedence. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('Retrieve') and resource ('QuickCast recording session') and even enumerates the returned fields (metadata, duration, markers, R2 watch URL). It is clearly a single-session retrieval, distinguishing it from list_recent_sessions and ingest_session. It stops short of naming a sibling explicitly, so no differentiation credit beyond inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'latest or specified' implies when the tool is used (returns newest by default, or a targeted one when identified), which is useful implied context. However, no alternative tool is named and no condition tells the agent when to prefer list_recent_sessions or ingest_session instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ingest_sessionB
Ingest and store a QuickCast recording session into the local session storage (~/.quickcast/sessions.json) so AI assistants can reference it.
| Name | Required | Description | Default |
|---|---|---|---|
| markers | No | Array of interaction markers with timeSec, label, selector, etc. | |
| filename | No | Original recording filename. | |
| watchUrl | Yes | Cloudflare R2 watch URL for the recording. | |
| sessionId | No | Unique session identifier (e.g. "qc_rec_1740001234"). | |
| environment | No | Environment metadata (url, browser, resolution, os). | |
| durationSeconds | No | Total length in seconds. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the behavioral burden, and it does disclose the concrete storage target (~/.quickcast/sessions.json). However, it is silent on write semantics such as overwrite vs append behavior, duplicate sessionId handling, or permission requirements for that file.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single well-formed sentence with the action front-loaded and the storage side effect attached at the end. Nothing is wasted, though a brief usage clause would have justified more length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a write tool with nested marker/environment objects and no annotations or output schema, the description covers the destination but leaves idempotency, overwrite behavior, and return expectations unaddressed. The rich schema carries most of the load, so it is adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for all six parameters, so the schema already documents watchUrl, markers, environment, etc. The description adds no syntax or format detail beyond what the schema provides, which is the baseline 3 case.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (ingest/store) and resource (QuickCast recording session) plus the destination store, which clearly separates it from read-oriented siblings like list_recent_sessions and get_latest_session. It stops short of explicitly naming those siblings, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The clause 'so AI assistants can reference it' implies the downstream purpose and thus when to use it, but there is no explicit when-to-use instruction, no prerequisites, and no exclusion relative to the sibling retrieval tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_recent_sessionsB
List recent QuickCast screen recording sessions with IDs, timestamps, durations, and shareable watch URLs.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of sessions to return (default: 10). | |
| session_file | No | Optional path to custom quickcast-sessions.json file. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral burden. It implies a read-only list and discloses the returned content shape (IDs, timestamps, durations, watch URLs), but says nothing about ordering, pagination, or that it reads from a local sessions file.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no waste; since there is no output schema, naming the returned fields is the most valuable content it could carry in that space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates the return fields and covers the resource well. It falls short only on default ordering and any routing versus get_latest_session, which are minor gaps for a simple list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (limit, session_file) are already documented in the schema. The description adds no syntax or behavior beyond what the schema provides, matching the baseline for fully-covered schemas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List recent QuickCast screen recording sessions') and enumerates returned fields, so the agent knows what it gets. However, it does not distinguish itself from the close sibling get_latest_session, leaving the scope difference (recent list vs. single latest) to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no mention of the obvious alternative get_latest_session, and no note about when the session_file override matters. The agent is left to guess selection criteria among siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scaffold_playwright_testA
Generate an executable Playwright (TypeScript or JavaScript) E2E reproduction test based on the interactions, clicks, and URLs captured during the QuickCast recording.
| Name | Required | Description | Default |
|---|---|---|---|
| language | No | Language to output ("typescript" or "javascript"). Default is "typescript". | typescript |
| test_name | No | Name for the Playwright test block. | |
| session_id | No | Optional session ID. If omitted, uses latest recording. | |
| target_url | No | Target starting URL (overrides recorded page URL if provided). | |
| session_data | No | Optional raw session JSON string. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It reveals that the output is an executable Playwright test and the input source, but says nothing about permissions, side effects (e.g., writing a file), error conditions, or what the tool returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that is front-loaded with the core action and output type, with no redundant or unnecessary phrasing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description adequately states what the tool generates and from what data, which is enough for an agent to select it. However, with no output schema and no annotations, it should ideally explain the return format or whether a file is written, leaving a notable gap for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters are already documented in the input schema. The description adds no parameter-level detail beyond mentioning TypeScript or JavaScript and the recorded interactions, making the baseline score of 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (Generate), resource (executable Playwright E2E reproduction test), and input source (interactions, clicks, URLs from QuickCast recording). This clearly distinguishes it from sibling tools like generate_sop_guide and format_github_bug_report.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the phrase 'based on the interactions, clicks, and URLs captured during the QuickCast recording,' which suggests using it after a recording exists. However, it does not state when to choose this tool over alternatives or any exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v1.0.0- First observed
format_github_bug_report - First observed
generate_sop_guide - First observed
get_latest_session - First observed
ingest_session - First observed
list_recent_sessions - First observed
scaffold_playwright_test
TDQS
Scored across 6 tools
Each tool has a distinct purpose: listing vs. retrieving session metadata, ingesting, and three different output-artifact generators (SOP, bug report, Playwright test). The only mild overlap is between get_latest_session and list_recent_sessions, but their descriptions clarify the distinction well.
All six tools follow a consistent snake_case verb_noun pattern (generate_sop_guide, list_recent_sessions, get_latest_session, format_github_bug_report, scaffold_playwright_test, ingest_session). The verbs vary appropriately to match each action without breaking the predictable convention.
Six tools is a well-scoped set for a session-recording processing server, covering retrieval, ingestion, and three distinct artifact generators without redundancy or padding.
The surface covers ingestion, listing, retrieval, and several generation workflows, which fits the domain well. Minor gaps remain, such as no delete/update or search-by-content operation, but agents can work around these.
Maintenance
Related MCP Connectors
Approved test intent, reviewed Playwright automation and run evidence, inside your editor.
Screen recording & video platform: search, share, transcribe, translate videos & AI meeting notes
Loom for agents: AI agents record narrated product demos, plus transcripts, summaries, and search.
Screen recording, meeting notes, and voice dictation - all with AI
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables AI coding agents to capture screen and voice recordings, extract timestamped frames, and receive structured Markdown reports with context for bug fixing and UI feedback.22 npm18MIT
- AlicenseNot gradedqualityAmaintenanceGives AI coding agents eyes into running applications by recording browser activity and providing session investigation tools for debugging.8 npm3MIT
- FlicenseAqualityBmaintenanceEnables automated testing and demo video creation using Playwright, with login, session orchestration, narration, and video compilation.8-
- AlicenseNot gradedqualityBmaintenanceRecords browser sessions and provides structured context (clicks, console, network, screenshots) for AI coding agents to understand and act upon.2MIT