Recordly MCP Server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Recordly MCP Servercopy demo_粗剪 and cut the 20s-26s range"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Recordly MCP Server
基于MCP编辑 Recordly工程副本:剪掉已经审核的口误/长停顿/重说区间,修改或删除已有缩放效果,并把工程保存到软件可以直接列出的Projects目录。
这是非官方兼容工具,不是Recordly官方MCP。基于 ClipAgent 的MCP代码改造,原软件来自 Recordly;来源、修改及许可证见 NOTICE.md 和 LICENSE.md。
安装
需要Node.js 20+。媒体分析需要系统FFmpeg和ffprobe;工程读写本身不需要它们。
git clone https://github.com/AzenYes/recordly-mcp-server.git
cd recordly-mcp-server
npm ci
npm run build
npm test在Codex的MCP配置中添加:
[mcp_servers.recordly]
command = "node"
args = ["C:/tools/recordly-mcp-server/dist/index.js"]将路径替换为实际克隆位置,重连MCP。其他支持stdio的客户端可使用相同command和args。
默认解析Recordly的用户数据目录,支持 recordings-settings.json 中的自定义recordingsDir。可选环境变量:RECORDLY_USER_DATA_DIR、FFMPEG_PATH、FFPROBE_PATH。
Related MCP server: video-agent-mcp
工具
工具 | 能力 |
get_paths | 返回实际录制目录和Projects目录 |
list_projects / open_project | 列出工程、读取完整状态与revision |
copy_project | 在Projects创建不覆盖原文件的新副本 |
get_media_info / detect_silence / scan_frames | 查看素材、检测停顿候选、抽查源画面 |
cut_ranges | 将已审核的源时间区间从现有clipRegions中剪除,返回成片时间映射 |
list_zooms | 读取已有缩放的ID、时间、深度、焦点 |
update_zoom | 修改单个缩放的倍率档位、焦点、时间、自动/手动模式 |
delete_zooms | 按ID或区间删除已有缩放,不删除录屏内容 |
剪切及缩放写入支持 dryRun、expectedRevision 和工程旁 .mcp-backups 备份。所有时间单位均为源视频毫秒,不是剪辑后成片时间。先保存并关闭编辑器中的对应副本,再让MCP写入,完成后重开;原软件不遵守MCP内部的写入锁。
{"project":"demo_粗剪","timelineModel":"recordly-1.4-source-time","ranges":[{"startMs":20000,"endMs":26000}],"dryRun":true}{"project":"demo_粗剪","zoomIds":["实际缩放ID"],"dryRun":true}实测与兼容边界
工程编辑接口已在Windows Recordly 1.4.0 实测;测试媒体不随仓库公开。
cut_ranges只支持已验证的version=2、单源、1倍速、按源时间排序的clipRegions。不同软件版本也可能使用version=2;调用方仍必须确认实际时间轴语义。检测到sourceStartMs、变速或与保留区间冲突的trimRegions时拒绝写入;接受原生保存时生成的冗余trim-gap镜像。时间轴中被剪掉的位置仍显示为空隙,原生导出会接起保留片段。不要拖动片段填空,以免改变源区间。
没有内置“识别口误”或“检测晃动”AI;调用方需要审核剪切和缩放修改区间。
本MCP不提供成片渲染接口;修改完成后可使用Recordly原生导出。
验证
npm testNode测试使用真实stdio MCP客户端和合成工程,验证备份、并发写入、版本保护、未知字段保留、Projects落盘、剪切映射与音轨设置。
Available Tools
11 toolscopy_projectB
Create a new editable project copy directly in the desktop Projects directory. Preserves media references, assigns a new projectId, and never overwrites.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| project | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does meaningfully disclose behavior beyond the name: it preserves media references, assigns a fresh projectId, and never overwrites. It omits permissions/authorization needs, what happens on a name collision, and any return/error behavior, keeping it below 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; every clause (location, preservation, new id, non-overwrite) adds actionable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and no output schema, the description covers the mutation behavior reasonably but leaves the two parameters unexplained and says nothing about failure modes or the resulting copy's identity. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the two required params ('project' and 'name'), so the description must compensate and does not. It explains the generated projectId but never clarifies that 'project' is the source to copy or what 'name' designates for the copy.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a new editable project copy') and scopes it to the desktop Projects directory, which clearly separates it from siblings like open_project, list_projects, and delete_zooms. It does not explicitly name or contrast against alternatives, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance or mention of alternative tools (e.g. open_project for viewing vs copy_project for duplicating). The description only says what it does when invoked, leaving the agent to infer that this is the duplication path.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cut_rangesA
Subtract reviewed SOURCE-time ranges from existing 1x clipRegions. Only the tested Recordly 1.4.0 single-source timeline model. Never moves zoom times or automatically decides which speech to cut. Backups and source-to-output mapping included.
| Name | Required | Description | Default |
|---|---|---|---|
| dryRun | No | ||
| ranges | Yes | ||
| project | Yes | ||
| timelineModel | Yes | ||
| expectedRevision | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose meaningful traits: backups are created, source-to-output mapping is included, zoom times are preserved, and the tool will not auto-select speech. It omits permission/auth requirements, reversibility beyond backups, and failure behavior, so it falls short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the core action and each sentence carrying a distinct piece of information (operation, applicability, scope limits, side effects). Dense but no filler; the terseness trades some readability for economy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description covers side effects well but leaves gaps: dryRun (preview semantics), expectedRevision (concurrency/conflict behavior), and what the caller gets back are unexplained. It is adequate but not complete for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It adds real meaning for two params — 'SOURCE-time ranges' clarifies the time base for ranges, and 'Recordly 1.4.0' maps to the timelineModel const. But project, dryRun (important for a cut operation), and expectedRevision are entirely undocumented in both description and schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Subtract') and resource ('reviewed SOURCE-time ranges from existing 1x clipRegions'), so the core operation is identifiable. However, the jargon-heavy phrasing ('clipRegions', '1x') and lack of explicit differentiation from siblings like delete_zooms or detect_silence keep it from a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an applicability constraint ('Only the tested Recordly 1.4.0 single-source timeline model') and a scope exclusion ('Never moves zoom times or automatically decides which speech to cut'), which implicitly tells the agent it must supply reviewed ranges. But it names no alternatives and does not say when to prefer this over sibling tools, so guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_zoomsA
Remove existing zoom effects, leaving the video intact. Select EITHER explicit zoomIds OR a half-open source-time startMs/endMs range. Range matching defaults to overlap and deletes whole matching zoom regions, including portions outside the requested range; use dryRun to inspect, or contained to select only fully contained effects. No selector is rejected. Automatic backup on change.
| Name | Required | Description | Default |
|---|---|---|---|
| endMs | No | Source/project time in milliseconds, exclusive. Supply together with startMs. | |
| match | No | overlap | |
| dryRun | No | Preview the exact change without writing or creating a backup. | |
| project | Yes | Project name or absolute .recordly path. Save/close this project in Recordly before writing; reopen afterwards. Prefer a working copy. | |
| startMs | No | Source/project time in milliseconds, inclusive. Supply together with endMs. | |
| zoomIds | No | ||
| expectedRevision | No | Optional SHA-256 revision from list_zooms; reject if the project changed. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations the description must carry the full load and it does well: it discloses that whole matching regions are deleted including portions outside the requested range, that no-selector calls are rejected, that dryRun writes nothing and creates no backup, and that backups are automatic on change. Auth/permission behavior and revision-conflict handling are left to the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense sentences, each earning its place, with the destructive scope warning front-loaded ahead of the selector details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a seven-parameter destructive tool with no output schema, the description covers selection, matching semantics, preview, and backup; only permission requirements and conflict behavior are unstated, and the latter is documented on expectedRevision in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Adds real meaning beyond the 71% schema coverage: mutual exclusivity of zoomIds vs the range, half-open startMs/endMs semantics stated in prose, the overlap-vs-contained distinction, and dryRun's side-effect-free nature.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource ('remove existing zoom effects') with an explicit scope clarification ('leaving the video intact') that separates it from siblings like update_zoom and list_zooms.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear selection rules: EITHER zoomIds OR a startMs/endMs range, with dryRun recommended for inspection. It does not name update_zoom or explain when a caller should modify rather than delete, so it stops short of full alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
detect_silenceB
Report candidate quiet intervals in the supplied audio, not automatic cuts. Separate mic times require adding its startDelayMs before editing.
| Name | Required | Description | Default |
|---|---|---|---|
| noiseDb | No | ||
| filePath | Yes | ||
| minDurationSec | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses that results are candidates, not cuts (non-destructive advisory behavior), and surfaces a real domain caveat about startDelayMs for separately-timed mics. It still omits detection thresholds, permissions, and output shape, so it is only partially transparent for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the advisory scope front-loaded and no filler. The second sentence is somewhat context-dependent (it assumes knowledge of what startDelayMs belongs to), costing a point.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, no annotations, and 0% schema coverage mean the agent gets no picture of the returned interval list, units, or how noiseDb/minDurationSec shape results. The description leaves significant gaps for a three-parameter detection tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never explains noiseDb (default -38 dB) or minDurationSec (default 1s), so the two parameters that control detection are undocumented anywhere. The only named field, startDelayMs, is not a parameter of this tool, so it adds context but no parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Report') and resource ('candidate quiet intervals in the supplied audio') and immediately scopes it as advisory rather than an editing action, which distinguishes it from the cut-oriented siblings in the namespace. It does not name a specific sibling like cut_ranges, so it falls short of full differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'not automatic cuts' implies this tool is for inspection/advisory reporting while a cutting tool would be used for edits, which is usable implied guidance. However, it never explicitly states when to reach for this versus cut_ranges, nor any prerequisites such as requiring an open project or loaded media.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_media_infoC
Inspect source video or separate microphone audio.
| Name | Required | Description | Default |
|---|---|---|---|
| filePath | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It does not say whether the operation is read-only, what kind of metadata is reported, what happens with unsupported formats, or whether the file must be part of an open project.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no filler, front-loaded with the verb. It is efficient, though the brevity comes at the cost of the information the other dimensions lack.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and an undocumented parameter, the description needed to explain return content and input expectations. For a 1-parameter inspection tool it says almost nothing an agent could act on beyond the name.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single parameter filePath has no description in the schema. The description mentions 'source video' and 'microphone audio' but never clarifies what filePath should point to for either case, leaving the only parameter effectively undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The verb 'Inspect' with the resource 'source video or separate microphone audio' gives a rough sense of the tool, but it never says what information is returned or why a caller would want it. It also does not distinguish itself from siblings such as scan_frames or detect_silence, which also examine media files.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool, when not to, or which sibling to prefer. The phrase 'source video or separate microphone audio' hints at two input cases but does not explain how the caller signals which one applies.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_pathsB
Resolve the actual recordings and Projects directories, including custom recordings-settings.json.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. 'Resolve' implies a read-only lookup, but it never states that nothing is mutated, whether a missing recordings-settings.json causes an error or a fallback, or whether the resolved path can differ from a default. Only the existence of the custom-settings override is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no padding. It is efficient, though the trailing clause about recordings-settings.json is dense jargon that could be split for readability without losing brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must describe the return value, and it only gestures at it ('the actual recordings and Projects directories'). It does not say whether one or two paths come back, in what format, or what happens when the custom settings file is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. There is nothing for the description to disambiguate, and it correctly avoids inventing arguments.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Resolve') and resource ('recordings and Projects directories'), plus the twist that a custom recordings-settings.json is honored. An agent can tell this is a path/config lookup rather than a media operation, so it is distinguishable from the Zoom-oriented siblings, though the description never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No statement of when to call this tool, when not to, or which sibling it substitutes for. The agent must infer that it exists to discover directory locations before invoking file-based siblings, and nothing confirms or denies that inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_projectsB
List projects in the desktop Projects directory.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden, yet it only names the source directory. It says nothing about whether the directory is created if missing, whether results are sorted, recursion depth, or error behavior when the directory does not exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler or redundancy. It is appropriately sized, though its brevity reflects thin content rather than efficient density.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a low-complexity read-only listing with no output schema, so the description need not explain return values in depth. Still, for an unannotated tool the agent would benefit from knowing the shape of results (names, paths) and what happens when the directory is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there are no parameter semantics to document and the baseline of 4 applies. Nothing in the description misleads about inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('projects') plus the location scope ('desktop Projects directory'). However, it does not differentiate itself from siblings such as list_zooms or open_project, leaving the boundary between listing and opening projects implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives like open_project or get_paths, nor any stated prerequisites or exclusions. The agent must infer the trigger condition entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_zoomsA
List existing zoom IDs, times, depth and focus before editing. Optional half-open source-time range selects overlapping regions. depth is a discrete level, not an arbitrary scale. Native Recordly export preserves its animation; MCP ffmpeg export differs.
| Name | Required | Description | Default |
|---|---|---|---|
| endMs | No | Source/project time in milliseconds, exclusive. Supply together with startMs. | |
| project | Yes | Project name or absolute .recordly path. Save/close this project in Recordly before writing; reopen afterwards. Prefer a working copy. | |
| startMs | No | Source/project time in milliseconds, inclusive. Supply together with endMs. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full disclosure burden. 'List' implies a read-only operation and the field list plus half-open range semantics add real behavioral detail, but permissions, ordering, pagination, and project-locking behavior are not addressed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first two sentences are front-loaded and dense, but the trailing sentence about Recordly export versus MCP ffmpeg export concerns output tooling, not listing zooms, so it reads as context bleed rather than information an agent needs to call this tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly enumerates the returned fields and explains range-selection semantics. It omits return ordering and pagination, but is otherwise sufficient for a 3-parameter listing tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, and the description still adds meaning beyond it: the range is 'half-open' and 'selects overlapping regions' rather than requiring containment, and depth is clarified as a discrete level. This is genuine value on top of the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource (list zooms) plus the exact fields returned (IDs, times, depth, focus). It hints at the editing workflow that siblings update_zoom and delete_zooms serve, but never names them explicitly, so sibling differentiation is only implied.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Before editing' implies a read-before-write workflow, which is useful context for choosing this over update_zoom/delete_zooms. However there is no explicit when-to-use/when-not guidance and no stated prerequisites for invoking it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
open_projectC
Read project state and revision before source-time edits.
| Name | Required | Description | Default |
|---|---|---|---|
| project | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does not disclose whether opening a project mutates or locks state, whether it is idempotent, what happens if the project is already open, or what errors can occur.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single nine-word sentence with no filler, and the read/state framing is front-loaded. It is efficient, though the brevity contributes to the gaps scored elsewhere rather than to wastefulness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and an undocumented required parameter, the description should explain the identifier format and what 'state and revision' are returned as. As written, an agent cannot reliably construct the call or predict the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there is one required parameter, 'project'. The description never says whether this is a name, ID, or file path, nor how to obtain a valid value, so it fails to compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a verb ('Read') and a resource ('project state and revision'), but the resource does not match the tool name 'open_project' — it never says the project is opened/loaded or what 'opening' implies. It gives no differentiation from the ten sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'before source-time edits' supplies a concrete trigger condition, which is more than most descriptions offer. However, it names no alternative tool and gives no exclusions or prerequisites, leaving the when-not case entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scan_framesB
Inspect raw source frames around candidate speech cuts or distracting zooms. These frames do not show the project zoom overlay; compare native previews too.
| Name | Required | Description | Default |
|---|---|---|---|
| endMs | No | ||
| startMs | No | ||
| filePath | Yes | ||
| sampleCount | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose one genuinely useful caveat: these frames omit the project zoom overlay, so native previews should be compared. It says nothing about whether the call is read-only, how frames are returned (format/count), or any cost or rate concerns, leaving major behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight, front-loaded sentences with no filler; the core inspection purpose leads and the overlay caveat follows. The trailing 'compare native previews too' is slightly loose but still earns its place as a usage cue.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations, no output schema, and zero parameter documentation, the description should do more. It covers what is inspected and one important visual caveat, but omits the return shape (frame images vs. metadata) and any parameter guidance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for four parameters, so the description must compensate and does not. The mention of 'candidate speech cuts' loosely hints at startMs/endMs ranges, but filePath, sampleCount (default 6, max 20), and the units/bounds of the time parameters are entirely unaddressed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('inspect') and resource ('raw source frames') with a scoping qualifier ('around candidate speech cuts or distracting zooms'). It is clearly distinguishable from a generic media-info tool, though it never names a sibling to contrast against.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'around candidate speech cuts or distracting zooms' implies a diagnostic context, gesturing at workflows tied to detect_silence and list_zooms. However, it never states when to use this instead of alternatives such as get_media_info, nor any exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_zoomA
Edit one EXISTING zoom by ID: adjust depth, focus, start/end, or auto/manual mode. Unspecified and unknown fields are preserved. Use mode=manual with a fixed focus to avoid automatic cursor-follow for this region. Saves an automatic backup; does not render or change video cuts.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | ||
| depth | No | Discrete zoom strength. Lower is less magnified; see list_zooms depthScales. | |
| endMs | No | ||
| dryRun | No | Preview the exact change without writing or creating a backup. | |
| focusX | No | ||
| focusY | No | ||
| zoomId | Yes | ||
| project | Yes | Project name or absolute .recordly path. Save/close this project in Recordly before writing; reopen afterwards. Prefer a working copy. | |
| startMs | No | ||
| expectedRevision | No | Optional SHA-256 revision from list_zooms; reject if the project changed. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden and does it well: it discloses merge semantics (unspecified/unknown fields preserved), an automatic backup on save, and a clear non-effect (does not render or change video cuts). Missing are auth/permission requirements and what a revision conflict does, but the destructive-scope picture is solid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three front-loaded sentences plus a targeted usage hint; every clause carries information with no filler. Slightly dense run-on in the first sentence, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description covers the key concerns an agent needs: what can be changed, that unspecified data survives, that a backup exists, and that video cuts are untouched. Gaps remain around revision conflict behavior and project open/close sequencing (partially in the schema), keeping it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 40%, so the description should compensate: it names the fields being edited (depth, focus, start/end, mode) and hints at mode semantics. It adds nothing for expectedRevision, dryRun, project path handling, or duration units (ms), leaving several parameters documented only sparsely or not at all.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (edit) and resource (one EXISTING zoom by ID) and enumerates the editable facets (depth, focus, start/end, mode). It contrasts cleanly with siblings like list_zooms and delete_zooms, so an agent can route correctly without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives one concrete usage hint (use mode=manual with fixed focus to avoid auto cursor-follow), which is genuinely actionable. However it never states when to reach for this tool versus list_zooms or delete_zooms, nor any preconditions beyond what the schema already says about project handling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
11 tool updates
v0.2.0- First observed
copy_project - First observed
cut_ranges - First observed
delete_zooms - First observed
detect_silence - First observed
get_media_info - First observed
get_paths - First observed
list_projects - First observed
list_zooms - First observed
open_project - First observed
scan_frames - First observed
update_zoom
TDQS
Scored across 11 tools
Most tools target clearly distinct resource+action pairs (list_zooms vs update_zoom vs delete_zooms, list_projects vs open_project vs copy_project). Minor potential confusion between get_paths and list_projects, and between get_media_info and scan_frames, but the descriptions clarify intent well.
Consistent verb_noun snake_case throughout (get_paths, list_zooms, cut_ranges, detect_silence). Slight deviation in singular/plural for the same resource (list_zooms/delete_zooms vs update_zoom), but the overall pattern is predictable and readable.
11 tools is well-scoped for a video project/zoom editing assistant, covering path resolution, project management, zoom editing, and media analysis without bloat. Each tool appears to earn its place.
The zoom lifecycle has list/update/delete but no tool to create a new zoom, a notable gap for a zoom-editing tool. Project handling lacks a from-scratch create (only copy) and no delete, and there is no render/export step, though agents can partly work around these.
Maintenance
Related MCP Connectors
Edit video in your open VidTL browser editor: cuts, subtitles, audio cleanup, effects, export.
- VidmoatOAuthcom.vidmoat
AI video editor: create projects, edit timelines, add captions and effects, and render videos.
Transcript-based audio editing: transcribe audio, edit by word ID, export edited audio.
Validate video cut plans and parse timed subtitles for Laqta’s local browser video editor.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceEnables AI agents and MCP clients to programmatically edit video projects on a local desktop editor, with 119 tools for multitrack editing, effects, captions, audio, and batch auto-editing, producing reviewable and reversible real timeline edits.AGPL 3.0
- AlicenseNot gradedqualityBmaintenanceEnables MCP clients to run project-scoped video editing workflows: propose and approve editing strategies, apply validated plans, review immutable versions, and export final renders via FFmpeg, with durable persistence and approval gates.1MIT
- AlicenseNot gradedqualityCmaintenanceEnables editing inside an open Premiere Pro project by reading the real timeline, applying cuts and transcript-based cleanup in place, and verifying results with rendered frames from the Program Monitor.MIT
- AlicenseNot gradedqualityCmaintenanceEnables LLM agents to read and edit local CapCut desktop projects directly, with an ffmpeg-based preview loop to verify edits before committing.1MIT