fcpxml-mcp-server
FCPXML MCP
The bridge between Final Cut Pro and AI. 53 tools that turn timeline XML into structured data Claude can read, edit, and generate.
Why This Exists
After directing 350+ music videos (Chief Keef, Migos, Masicka), I noticed the same editing bottlenecks on every project: counting cuts manually, extracting chapter markers one by one, hunting flash frames by scrubbing, building rough cuts clip by clip.
These are batch operations that don't need visual feedback. Export the XML, let Claude handle the tedium, import the result. That's the entire philosophy.
Related MCP server: fcp-mcp
See It In Action
You: "Run a health check on my wedding edit"
Claude: ✓ Analyzed WeddingFinal.fcpxml
├─ 247 clips · 42:18 total · 24fps · 1920×1080
├─ 3 flash frames detected (clips 44, 112, 198)
├─ 2 unintentional gaps at 12:04 and 31:47
├─ 14 duplicate source clips
└─ Health score: 72/100
You: "Fix the flash frames and gaps, then add chapter markers from
this transcript"
Claude: ✓ Extended adjacent clips to cover 3 flash frames
✓ Filled 2 gaps by extending previous clips
✓ Added 18 chapter markers from transcript
→ Saved: WeddingFinal_modified.fcpxmlImport the modified XML back into Final Cut Pro. Every change is non-destructive — your original file is never touched.
What Claude Actually Sees
This is the magic trick. When you export XML from Final Cut Pro, your timeline becomes structured data that Claude can reason about:
<!-- What FCP exports -->
<asset-clip ref="r2" offset="342/24s" name="Interview_A"
start="120s" duration="720/24s" format="r1">
<marker start="48/24s" duration="1/24s" value="Key quote"/>
<keyword start="0s" duration="720/24s" value="Interview"/>
</asset-clip># What Claude works with (after parsing)
Clip(
name="Interview_A",
offset=TimeValue(342, 24), # timeline position: 14.25s
start=TimeValue(120, 1), # source in-point: 2:00
duration=TimeValue(720, 24), # 30 seconds
markers=[Marker(value="Key quote", start=TimeValue(48, 24))],
keywords=["Interview"]
)Every time value stays as a rational fraction — 720/24s, not 30.0 — so trim, split, and speed operations have zero rounding error across any frame rate. Comparisons use cross-multiplication (a/b < c/d → a*d < c*b) to stay in integer-land end to end. Denominators are always normalized to positive values at construction, so sign lives on the numerator and cross-multiplication is always correct. Addition and subtraction share a single _binop() code path that handles same-denominator fast paths and LCM alignment in one place.
How It Works
┌──────────┐ ┌──────────────────────────────┐ ┌──────────┐
│ Final Cut│ │ parser.py → Python objects │ │ Final Cut│
│ Pro │─XML─>│ writer.py → Modify & save │─XML─>│ Pro │
│ │ │ rough_cut.py→ Generate new │ │ │
└──────────┘ │ diff.py → Compare │ └──────────┘
│ export.py → Resolve / FCP7 │
└──────────────────────────────┘
▲
Claude Desktop / MCP clientExport from FCP —
File → Export XML...Ask Claude — analyze, edit, generate, QC, export
Import back —
File → Import → XML
What This Is NOT
Not a plugin — it doesn't run inside Final Cut Pro
Not real-time — you work with the XML between exports
Not for creative calls — color, framing, motion still need your eyes
Quick Start
1. Clone & Install
git clone https://github.com/DareDev256/fcp-mcp-server.git
cd fcp-mcp-server
pip install -e .2. Configure Claude Desktop
Add to ~/Library/Application Support/Claude/claude_desktop_config.json:
Using uv (recommended):
{
"mcpServers": {
"fcpxml": {
"command": "uv",
"args": ["--directory", "/path/to/fcp-mcp-server", "run", "server.py"],
"env": { "FCP_PROJECTS_DIR": "/Users/you/Movies" }
}
}
}Using pip:
{
"mcpServers": {
"fcpxml": {
"command": "python",
"args": ["/path/to/fcp-mcp-server/server.py"],
"env": { "FCP_PROJECTS_DIR": "/Users/you/Movies" }
}
}
}3. Use It
Export XML from Final Cut Pro, open Claude Desktop, and ask it to work with your timeline.
When To Use This
Good For | Not Ideal For |
Batch marker insertion (100 chapters from a transcript) | Creative editing decisions (no visual feedback) |
QC before delivery (flash frames, gaps, duplicates) | Real-time adjustments (export/import cycle) |
Data extraction (EDL, CSV, chapter markers) | Fine-tuning cuts (faster directly in FCP) |
Template generation (rough cuts from tagged clips) | Anything visual (color, framing, motion) |
Automated assembly (montages from keywords + pacing) | |
Timeline health checks (validation, stats, scoring) |
Prompt Cookbook
Copy-paste these into Claude Desktop. Each one maps to a real tool chain under the hood.
Analysis
"Give me a full breakdown of ProjectX.fcpxml — clips, duration, frame rate, markers, everything"
"Show me pacing analysis for my timeline — where are the slow sections?"
"Export an EDL and CSV of all clips with timecodes"QC & Fixes
"Run a health check on my timeline and fix anything under 2 frames"
"Find all gaps and flash frames, then auto-fix them"
"Are there any duplicate source clips I can consolidate?"Markers & Chapters
"Add chapter markers from this transcript: [paste transcript]"
"Import markers from my-subtitles.srt onto the timeline"
"List all markers and export them as YouTube chapter timestamps"Generation
"Build a 60-second rough cut from clips tagged 'Interview' — medium pacing"
"Generate a montage from all B-roll clips with accelerating pacing"
"Create an A/B roll: Interview_A as primary, B-roll cuts every 8 seconds"Cross-NLE & Reformat
"Export this timeline for DaVinci Resolve"
"Convert to FCP7 XML so I can open it in Premiere"
"Reformat my 16:9 timeline to 9:16 for Instagram Reels"Under the Hood
When you say "Run a health check on my wedding edit", Claude chains these tools:
analyze_timeline → stats, frame rate, resolution
detect_flash_frames → clips under threshold duration
detect_gaps → unintentional silence/black
detect_duplicates → repeated source media
validate_timeline → structural health score (0-100)Each tool returns structured text that Claude synthesizes into the summary you see. No magic — just batch XML queries that would take 20 minutes by hand.
Pre-Built Prompts
Select these from Claude's prompt menu (⌘/) — they chain multiple tools automatically.
Prompt | What It Does |
qc-check | Full quality control — flash frames, gaps, duplicates, health score |
youtube-chapters | Extract chapter markers formatted for YouTube descriptions |
rough-cut | Guided rough cut — shows clips, suggests structure, generates |
timeline-summary | Quick overview — stats, pacing, keywords, markers, assessment |
cleanup | Find and auto-fix flash frames and gaps |
All 53 Tools
Category | Tools | What It Does |
Analysis | 11 | Stats, clips, markers, keywords, EDL/CSV, pacing |
Multi-Track | 3 | Connected clips, compound clips, secondary lanes |
Roles | 4 | List, assign, filter, export stems |
QC & Validation | 4 | Flash frames, duplicates, gaps, health score |
Editing | 9 | Markers, trim, reorder, transitions, speed, split |
Batch Fixes | 3 | Auto-fix flash frames, rapid trim, fill gaps |
Comparison | 1 | Diff two timelines — added/removed/moved/trimmed |
Reformat | 1 | Aspect ratio conversion (9:16, 1:1, 4:5, custom) |
Silence | 2 | Detect and remove silence candidates |
NLE Export | 2 | DaVinci Resolve v1.9, FCP7 XMEML v5 |
Generation | 3 | Rough cuts, montages, A/B roll |
Beat Sync | 2 | Import beat markers, snap cuts to beats |
Import | 2 | SRT/VTT subtitles, YouTube chapters → markers |
Audio | 1 | Add audio clips, music beds at any lane |
Compound | 2 | Create/flatten compound clips |
Templates | 2 | Pre-built timeline structures (intro/outro, lower thirds, music video) |
Effects | 1 | List FCP transition effects with UUIDs |
53 |
Analysis — 11 tools
list_projects · analyze_timeline · list_clips · list_library_clips · list_markers · find_short_cuts · find_long_clips · list_keywords · export_edl · export_csv · analyze_pacing
Multi-Track — 3 tools
list_connected_clips · add_connected_clip · list_compound_clips
Roles — 4 tools
list_roles · assign_role · filter_by_role · export_role_stems
QC & Validation — 4 tools
detect_flash_frames · detect_duplicates · detect_gaps · validate_timeline
Editing — 9 tools
add_marker · batch_add_markers · insert_clip · trim_clip · reorder_clips · add_transition · change_speed · delete_clips · split_clip
Batch Fixes — 3 tools
fix_flash_frames · rapid_trim · fill_gaps
Comparison · Reformat · Silence
diff_timelines · reformat_timeline · detect_silence_candidates · remove_silence_candidates
NLE Export — 2 tools
export_resolve_xml (DaVinci Resolve FCPXML v1.9) · export_fcp7_xml (Premiere Pro / Resolve / Avid XMEML v5)
Generation — 3 tools
auto_rough_cut · generate_montage · generate_ab_roll
Beat Sync — 2 tools
import_beat_markers · snap_to_beats
Import — 2 tools
import_srt_markers · import_transcript_markers (supports SMPTE HH:MM:SS:FF with frame-accurate placement)
v0.6.0 — Audio, Compound, Templates, Effects — 6 tools
list_effects · add_audio · create_compound_clip · flatten_compound_clip · list_templates · apply_template
Environment Variables
Variable | Required | Default | Description |
| No |
| Root directory for FCPXML file discovery via |
| No | — | Route LLM calls through any OpenAI-compatible proxy (LiteLLM, OpenRouter, Ollama, vLLM) |
Compatibility
Component | Supported Versions |
FCPXML format | v1.8 – v1.11 |
Final Cut Pro | 10.4+ |
Python | 3.10, 3.11, 3.12 |
MCP protocol | 1.0 |
Export targets | |
→ DaVinci Resolve | FCPXML v1.9 |
→ Premiere Pro / Avid | FCP7 XMEML v5 |
Architecture
fcp-mcp-server/ ~8.9k lines Python
├── server.py MCP entry point — 53 tools, 5 prompts, resource discovery
│ _resolve_io_paths() / _setup_modifier() / _setup_generator()
│ _format_clip_table() / _markdown_table() / _format_batch_result()
│ _raw_markers_to_batch()
│ _detect_flash_frames() / _detect_gaps() / _detect_duplicate_groups()
│ consolidate path validation, QC detection, rendering, handler boilerplate
├── fcpxml/
│ ├── README.md Developer guide — TimeValue, clip hierarchy, modifier patterns
│ ├── models.py TimeValue, Timecode, Clip, ConnectedClip, MarkerType, Timeline
│ ├── parser.py FCPXML → Python (spine, connected clips, roles, markers)
│ ├── writer.py Modify & write (markers, trim, gaps, transitions, silence)
│ │ FCPXMLModifier: index-based editing (clips/resources/formats dicts)
│ │ FCPXMLWriter: generate new FCPXML from Python objects
│ │ Helpers: _resolve_asset, _absorb_into_neighbor, _ripple_from_index
│ ├── rough_cut.py Generate timelines (rough cuts, montages, A/B roll)
│ ├── diff.py Timeline comparison engine (identity matching, threshold docs)
│ ├── export.py DaVinci Resolve v1.9 + FCP7 XMEML v5 export
│ ├── safe_xml.py Centralized defusedxml wrappers (XXE/entity-bomb protection) + serialize_xml()
│ └── templates.py Template system (intro/outro, lower thirds, music video)
├── tests/ 912 tests across 18 suites
│ ├── test_models.py TimeValue math, Timecode formatting, MarkerType contracts
│ ├── test_parser.py FCPXML parsing, connected clips, edge cases
│ ├── test_writer.py Clip editing, marker writing, speed changes
│ ├── test_fcpxml_writer.py FCPXMLWriter generation from Python objects
│ ├── test_server.py MCP tool handlers, dispatch, path validation
│ ├── test_rough_cut.py Rough cut generation, montage, A/B roll
│ ├── test_diff.py Moved clips, transitions, markers, clip identity
│ ├── test_export.py Attribute stripping, compound flattening, audio tracks
│ ├── test_features_v05.py Multi-track, roles, diff, reformat, export
│ ├── test_features_v06.py Audio, compound clips, templates, effects, validation
│ ├── test_marker_pipeline.py Marker builder, batch modes, output format
│ ├── test_speed_cutting.py Speed cutting, montage config, pacing curves
│ ├── test_security.py Input validation, XML sanitization, XXE protection
│ ├── test_edge_cases.py Boundary arithmetic, clip collisions, split/diff edges
│ ├── test_diversity.py Boundary conditions across diff, models, validation
│ ├── test_refactored_helpers.py _index_elements, _iter_spine_clips, serialize_xml edges
│ └── test_targeted_gaps.py Targeted branch coverage for diff, export, models
├── docs/
│ └── WORKFLOWS.md 8 production workflow recipes
└── examples/
└── sample.fcpxml 9 clips, 24fps — test fixtureSecurity
Every tool handler is hardened against adversarial input — critical for MCP servers where prompts may be LLM-generated, not human-typed.
Layer | Protection |
File I/O | Path traversal blocked, null bytes rejected, symlinks resolved, 100 MB size limit |
Output sandbox | All generation, write, export, beat sync, subtitle, and reformat handlers enforce |
Subprocess bounds |
|
Speed validation |
|
Directory listing | Confined to |
XML parsing |
|
JSON depth limit | Iterative BFS depth checker rejects payloads nested beyond 50 levels — immune to RecursionError even at ~1000 nesting |
Batch limits | Marker batch operations capped at 10,000 entries — prevents memory exhaustion from adversarial payloads with millions of markers |
Inline text limits | Inline transcript arguments capped at ~1 MB — file-based inputs go through |
Symlink filtering |
|
Marker strings | Sanitized via |
Role values | Stripped of control characters before XML attribute assignment |
URI parsing | MCP resource URIs parsed via |
Output suffixes | Path separators and special characters stripped — no traversal via suffix injection |
Marker types |
|
132 security-specific tests across test_security.py covering XXE, path traversal, sandbox boundaries, output path anchoring, input validation, subprocess bounds, minidom hardening, JSON depth limits, role sanitization, ffmpeg parameter bounds, symlink filtering, file count caps, and write-handler sandbox enforcement. Ruff S (bandit) rules enforced in CI — S314/S320 block unsafe XML parsing, S105 catches hardcoded passwords, S108 flags insecure temp paths. Security events (null bytes, sandbox escapes, unhandled exceptions) are logged via Python logging for audit trails.
Timestamp Parsing — How Import Tools Place Markers
All subtitle and transcript import tools (import_srt_markers, import_transcript_markers) funnel through a single internal function: _parse_timestamp_parts() in server.py. Understanding it matters when timestamps don't land where you expect.
Supported Formats
Format | Example | Parts | Result |
Minutes:Seconds |
| 2 | 90.0s |
H:MM:SS |
| 3 | 3930.0s |
HH:MM:SS.ms |
| 3 | 135.5s |
SMPTE (HH:MM:SS:FF) |
| 4 | 3610.5s @ 24fps |
The SMPTE 4-part format converts the frame component to fractional seconds: frames / frame_rate. The default rate is 24fps — pass frame_rate= to override for 25fps (PAL) or 30fps (NTSC) projects.
The Import Pipeline
SRT / VTT / YouTube chapters / plain transcript
│
▼
parse_srt() / parse_vtt() / parse_transcript_timestamps()
│ │ │
└────────────────┴──────────────────────┘
│
split on ':'
│
▼
_parse_timestamp_parts(parts, frame_rate=24.0)
│
▼
total seconds (float)
│
▼
marker placed on timelineEdge Cases
Unrecognized part counts (1 part, 5+ parts) return
None— the marker is silently skipped, not placed incorrectlyZero frame rate — falls back to base seconds (frames ignored) rather than dividing by zero
Milliseconds — only carried in 3-part format via
float()on the seconds component ("15.500"→15.5)Frame rounding — SMPTE frames are divided exactly (
12/24 = 0.5), not rounded to the nearest frame boundary. The resulting float is converted to FCPXML's rationalTimeValuedownstream, preserving precision
Why This Matters
Before v0.6.20, the 4-part SMPTE parser silently dropped frames — 01:00:10:12 became 3610.0s instead of 3610.5s. At 24fps, that's up to ~0.96 seconds of drift per marker. If you imported a subtitle file with SMPTE timecodes, every marker was slightly off. This was subtle enough to pass QC but visible when scrubbing.
Design Principles
Principle | Implementation |
Rational time, never floats | All durations are fractions ( |
Non-destructive by default | Modified files get |
Single source of truth |
|
Security-first | 10-layer defense-in-depth across all 53 handlers — see Security for the full matrix |
Dispatch, not conditionals |
|
Documentation
Guide | What's Inside |
8 production recipes — QC pipelines, beat-synced assembly, cross-NLE handoffs, documentary A/B roll | |
How this server composes with GitNexus, filesystem, and memory MCP servers | |
Full version history from v0.1.0 to present |
Testing
uv run --extra dev pytest tests/ -v # or: python3 -m pytest tests/ -v
ruff check . --exclude docs/ # lint — must pass before committing795 tests across 17 suites covering models, parser, writer, FCPXMLWriter generation, server handlers, rough cut generation, speed cutting & pacing curves, marker pipeline, refactored helper functions (_index_elements, _iter_spine_clips, _find_spine_clip_at_seconds, _require_clip, _require_spine_clip, _resolve_asset, serialize_xml), recent fix regressions (rapid_trim directions, min_duration, offset recalculation, interval timing accuracy, gap skipping, duplicate-name clip operations across trim/speed/split/delete/markers), security hardening (XXE, entity expansion, path traversal, sandbox boundaries, minidom defense-in-depth, JSON depth limits, input validation, ffmpeg bounds, write-handler sandboxing), connected clips, roles, diff, export, compound clip flattening, audio track generation, templates, effects, boundary conditions, and backward compatibility.
Requirements
Python 3.10+ · Final Cut Pro 10.4+ (FCPXML 1.8+) · Claude Desktop or any MCP client
Dependencies (auto-installed):
mcp,defusedxmlSee Compatibility for full version matrix
Roadmap
Core FCPXML parsing (v1.8–1.11)
Timeline analysis, markers, EDL/CSV export
Clip editing (trim, reorder, split, speed, transitions)
QC tools (flash frames, gaps, duplicates, health scoring)
Generation (rough cuts, montages, A/B roll, beat sync)
MCP Prompts + Resources (auto-discovery)
Subtitle & transcript import as markers
Multi-track (connected clips, compound clips, roles)
Timeline diff + social media reformat
Silence detection & cleanup
Cross-NLE export (DaVinci Resolve, Premiere Pro, Avid)
Audio sync detection
Premiere Pro native XML support
Known Issues
Issue | Impact | Workaround |
Still images crash FCP | PNG/JPEG assets referenced directly in FCPXML crash Final Cut Pro on import ( | Convert stills to short MOVs before referencing: |
Non-standard timebases | FCP rejects time values with denominators outside its standard set (e.g. | Fixed in v0.5.29 — TimeValue arithmetic now uses LCM, and speed changes snap to frame boundaries in 2400-tick timebase. |
Malformed frameDuration crash | A | Fixed in v0.6.23 — writer now validates both numerator and denominator, falling back to 30.0 fps. |
Duplicate clip names corrupt edits | When multiple spine clips share the same name (e.g. | Fixed in v0.6.37–0.6.39 — all methods now resolve clips via |
Contributing
PRs welcome. If you're a video editor who codes (or a coder who edits), let's build this together.
Credits
Built by @DareDev256 — former music video director (350+ videos), now building AI tools for creators.
License
MIT — see LICENSE.
Available Tools
13 toolsdeliverA
Get the edit out: export to other NLEs, CSV, EDL and stems, reformat, relink media, or push straight into a running Final Cut Pro. Exports and push refuse a cut with no rendered preview of its current state (pass confirm_unreviewed=true to ship anyway). Actions: export_csv, export_edl, export_fcp7_xml, export_resolve_xml, export_role_stems, reformat_timeline, relink_media, push_to_fcp, list_fcp_libraries.
| Name | Required | Description | Default |
|---|---|---|---|
| args | No | Arguments for the chosen action, e.g. {"filepath": "/path/to/project.fcpxml"}. | |
| action | Yes | Which operation to run. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It usefully discloses that exports and push refuse to run on a cut without a rendered preview and explains the confirm_unreviewed=true escape hatch. However, it does not mention side effects, mutating behavior of reformat_timeline or relink_media, permissions, or output behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded with the core purpose, followed by a critical behavioral warning. The action list overlaps with the enum but still usefully summarizes the tool's breadth without excessive detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a multi-action dispatcher with a nested args object and no output schema. The description does not explain per-action required arguments, return values, or how args should be structured for each action. The confirm_unreviewed detail helps but is not enough to make the tool safely callable for all nine actions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds value beyond the schema by explaining the confirm_unreviewed=true flag and connecting action names to concrete output formats and workflows, which helps disambiguate the generic args object.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb phrase ('Get the edit out') and resource (a cut/edit), and clearly differentiates it from siblings like inspect, diagnose, preview, and watch by focusing on export, reformat, relink, and push operations. It also names concrete outputs and actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear context for when to use the tool—when you need to export or move an edit elsewhere. However, it does not explicitly state when not to use it or name alternatives for related tasks, instead leaving that mostly to sibling names and inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
diagnoseA
Find problems in a timeline: gaps, flash frames, duplicates, dead air, and beat structure. Read-only. Run before editing. Actions: validate_timeline, detect_gaps, detect_duplicates, detect_flash_frames, detect_silence_candidates, detect_media_silence, detect_beats, find_short_cuts, find_long_clips, diff_timelines.
| Name | Required | Description | Default |
|---|---|---|---|
| args | No | Arguments for the chosen action, e.g. {"filepath": "/path/to/project.fcpxml"}. | |
| action | Yes | Which operation to run. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral burden; it explicitly flags the tool as read-only and limits its behavior to the listed detection/validation actions. It does not disclose output format or per-action side-effect nuances, but for a no-annotation tool the core non-mutating behavior is clearly stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with purpose, then usage, then a compact action list, with no filler. The action list is long but all entries are load-bearing for action selection, and the whole text remains scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The definition is adequate for a read-only diagnostic tool: purpose, read-only behavior, and action vocabulary are present. However, the 10 actions are heterogeneous and have no per-action argument breakdown or output description, so an agent may need supplemental information to invoke some actions with confidence.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already covers both parameters fully, and the description's action list mirrors the enum. It adds no per-action argument semantics beyond the schema's example, so the description earns the baseline score of 3 rather than extra credit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Find problems in a timeline') and enumerates concrete problem types and actions, so the tool's diagnostic role is clear. It does not explicitly differentiate itself from sibling tools like inspect or find, but the action vocabulary makes the intended scope fairly unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Read-only. Run before editing.' gives a clear temporal usage rule and implies it is a pre-edit analysis step rather than a mutation. It does not name alternatives or state when not to use it, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
editA
Change clips on the timeline: insert, delete, trim, split, reorder, retime, remove silence, and attach audio, B-roll or transitions. Writes a new file. Actions: insert_clip, delete_clips, trim_clip, split_clip, reorder_clips, change_speed, rapid_trim, add_transition, add_audio, add_connected_clip, assign_role, fill_gaps, fix_flash_frames, remove_silence_candidates, remove_media_silence.
| Name | Required | Description | Default |
|---|---|---|---|
| args | No | Arguments for the chosen action, e.g. {"filepath": "/path/to/project.fcpxml"}. | |
| action | Yes | Which operation to run. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It does disclose a key trait: 'Writes a new file', which suggests a non-destructive output model. However, it does not describe failure behavior, prerequisites, permissions, or whether the original file is left untouched, so the disclosure is only partial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with purpose, immediately states the key side effect, and then lists all supported actions. The action list is long but useful; no sentences are wasted, though the list partially duplicates the enum in the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a complex tool with 15 actions, a generic nested args object, no output schema, and no annotations. The description covers the operation set and the new-file side effect, but it leaves per-action argument requirements, output location, and return behavior unexplained, which is a significant gap for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so a baseline of 3 applies. The description adds an example of args ('filepath') and names the action enum, but it does not explain what arguments each specific action needs beyond the generic 'args' object.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ('Change clips on the timeline') and enumerates concrete operations, so an agent understands exactly what this tool does. The action list also separates it from read-only or output-oriented siblings like inspect, preview, and deliver.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it: whenever timeline clip edits are needed. It does not explicitly contrast it with sibling tools or state when not to use it, leaving the routing decision mostly to the action names and general intent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
findA
Natural-language shot search. find_shots (filepath, query, limit=10, visual=false, clip_name?) ranks moments by what was said (transcript), what was logged (names, keywords, notes, markers, audio events) and, when a local vision model is installed, what the frames look like — the first line of every result names which tiers answered and why one could not. find_index (captions=true, backend) warms scenes and captions and reports which clips have no transcript (it never transcribes or goes online). find_to_timeline (min_source_separation=1, output_path?) assembles the hits into a _found selects reel and reports its diversity score. Actions: find_index, find_shots, find_to_timeline.
| Name | Required | Description | Default |
|---|---|---|---|
| args | No | Arguments for the chosen action, e.g. {"filepath": "/path/to/project.fcpxml"}. | |
| action | Yes | Which operation to run. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral disclosure burden. It reveals ranking criteria, the optional local vision model dependency, the constraint that find_index never transcribes or goes online, and what result elements are reported. It falls short of clarifying side effects for find_to_timeline, which appears to create a reel, and does not mention auth or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is information-dense and front-loaded with the core purpose, then systematically covers each action. The final 'Actions:' enumeration is slightly redundant after the prose has already introduced all three actions, but it is harmless and aids quick parsing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers all three actions, their parameters, key constraints, and output hints, which is especially important given there is no output schema. Some details remain implicit, such as exact query syntax, clip_name format, and valid backend values, but the agent has enough context to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although the schema only defines a generic args object, the description names the action-specific arguments and their meanings, such as filepath, query, limit, visual, clip_name, captions, backend, min_source_separation, and output_path. This goes well beyond the schema's bare example and effectively compensates for the lack of per-action schema detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear purpose, 'Natural-language shot search,' then enumerates three concrete actions with distinct resources and behaviors. Each action is described with a specific verb and expected outcome, making the tool's scope immediately understandable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives strong routing guidance by explaining what each of the three actions does and when it applies, including boundaries like 'it never transcribes or goes online.' It does not explicitly compare against sibling tools such as index or transcript, but the action-level guidance is sufficient for most invocation decisions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generateB
Build new timeline structure from source clips: rough cuts, montages, A/B roll, templates and compound clips. Actions: auto_rough_cut, generate_ab_roll, generate_montage, apply_template, create_compound_clip, flatten_compound_clip, import_edl_json.
| Name | Required | Description | Default |
|---|---|---|---|
| args | No | Arguments for the chosen action, e.g. {"filepath": "/path/to/project.fcpxml"}. | |
| action | Yes | Which operation to run. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure, but it only lists action names like flatten_compound_clip and import_edl_json. It does not explain side effects, whether operations mutate existing projects, required project state, or return behavior, leaving the tool mostly a black box.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded: one clear purpose sentence followed by a compact action list. There is no filler, and the structure makes the tool's scope immediately scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool exposes seven distinct operations with a nested args object, yet the description provides no per-action argument guidance, output expectations, or usage caveats. No annotations or output schema compensate, so the description is not complete enough for reliable invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the action enum and args object are already documented. The description repeats the action list but adds no additional meaning about what arguments each action needs or how args should be structured beyond the schema's single example.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and resource: 'Build new timeline structure from source clips', and enumerates seven concrete operations. It is clear what the tool does, though it does not explicitly differentiate itself from sibling tools like edit or organize.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The purpose statement and action list imply when to use the tool—when constructing new timeline structures—but no explicit conditions, exclusions, or alternatives are mentioned. The agent must infer the boundary against sibling tools rather than being told.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
indexA
The analysis cache under ~/.fcp-mcp/index.db. index_status shows what is cached and how old it is; index_build warms it for every source in a timeline (args: filepath, with_transcript); index_clear drops it. Every other tool works with the index off (FCP_MCP_INDEX=off) — this only makes the second question fast. Actions: index_status, index_build, index_clear.
| Name | Required | Description | Default |
|---|---|---|---|
| args | No | Arguments for the chosen action, e.g. {"filepath": "/path/to/project.fcpxml"}. | |
| action | Yes | Which operation to run. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it discloses the key behaviors: status reads cache metadata, build warms the cache, and clear is destructive ('drops it'). It also signals low risk by stating every other tool works with the index off. Return values and build-overwrite behavior are left implicit, but the core operation semantics are honestly described.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded with the resource and optionality, then lists the three actions. The final action list is slightly redundant with the earlier sentence, but overall every sentence contributes useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers all three actions and the cache's optional nature, but there is no output schema and it does not explain what each action returns. It also leaves 'the second question' vague and only states args for index_build, so an agent must infer the arg behavior for status and clear.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already covers both parameters with 100% description coverage, so the baseline is 3. The description adds that index_build takes filepath and with_transcript, which helps because the schema's args object is generic, but it does not explain what with_transcript means or whether status/clear accept args. This is only a modest improvement over the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies the resource (~/.fcp-mcp/index.db) and enumerates three concrete operations (index_status, index_build, index_clear), so an agent can tell this is the index-management tool. It is specific and not a tautology, though it does not explicitly contrast itself with a sibling alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The sentence 'Every other tool works with the index off... this only makes the second question fast' gives an explicit selection signal: the index is optional and purely for performance. It does not spell out when to prefer one sub-action over another in detailed scenarios, but the contextual guidance is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inspectA
Read a timeline or project without changing it. Use this first to understand what you are working with. Actions: list_projects, analyze_timeline, analyze_pacing, list_clips, list_markers, list_roles, list_keywords, list_effects, list_templates, list_library_clips, list_compound_clips, list_connected_clips, filter_by_role.
| Name | Required | Description | Default |
|---|---|---|---|
| args | No | Arguments for the chosen action, e.g. {"filepath": "/path/to/project.fcpxml"}. | |
| action | Yes | Which operation to run. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure; the phrase 'without changing it' adequately communicates that this is a non-mutating, read-only operation. However, it does not describe return shapes, failure modes, or differences between the 13 actions, which leaves meaningful gaps for a dispatcher-style tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with purpose and usage guidance, then compactly lists the supported actions. The action list repeats the enum but is useful for quick comprehension in a dispatcher tool, and there is no wasteful prose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
As a 13-action multiplexer with no annotations and no output schema, the description gives a solid high-level orientation but lacks per-action argument requirements and expected results. An agent would likely call obvious actions like list_projects correctly, but less obvious actions like filter_by_role are under-specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: both 'action' and 'args' have descriptions, and 'args' includes an example. The description itself adds no per-action parameter detail, so it remains at the baseline 3 without compensating for the generic args object.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description begins with a clear read-only verb and resource ('Read a timeline or project without changing it'), which immediately distinguishes it from mutating sibling tools like edit, mark, and generate. It further lists all 13 concrete actions it supports, leaving no confusion about what the tool is for.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says 'Use this first to understand what you are working with,' giving a strong contextual trigger for selecting this tool. It does not explicitly name excluded situations or direct the agent to alternative siblings for modifications, but the read-only phrasing vs. mutation-oriented siblings makes the boundary clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
markB
Add or import markers and chapters, including from SRT/VTT subtitles, transcripts, and beat analysis. Actions: add_marker, batch_add_markers, import_srt_markers, import_transcript_markers, import_beat_markers, snap_to_beats.
| Name | Required | Description | Default |
|---|---|---|---|
| args | No | Arguments for the chosen action, e.g. {"filepath": "/path/to/project.fcpxml"}. | |
| action | Yes | Which operation to run. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral disclosure burden. It does state that the tool adds/imports markers and chapters and lists the action names, which conveys basic mutating behavior. But it does not disclose side effects, whether existing markers are modified, prerequisites, or what happens after an import or snap operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded with the core purpose, followed by a compact action list. It is easy to skim, though the action list largely duplicates the schema's enum and adds little new information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a dispatch-style tool with six distinct actions and a nested args object, yet the description only names the actions and does not explain per-action arguments, expected inputs, or return behavior. An agent could not confidently call the correct action with correct arguments based solely on this definition.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is high at 100%, so the baseline is 3 even without parameter explanations in the description. The description adds source-type context for actions like SRT/VTT and transcript imports, but it does not explain what arguments each action requires beyond the generic filepath example in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb + resource ('Add or import markers and chapters') and enumerates the supported source types and action names. It is clearly a marker/chapter tool, so it is distinguishable from siblings like 'transcript' and 'edit', though it does not explicitly contrast itself with any sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage contexts via 'including from SRT/VTT subtitles, transcripts, and beat analysis', and the action list hints at the available operations. However, it gives no explicit guidance on when to choose this tool over an alternative, nor when one action should be used instead of another.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
organizeA
Bulk library logging and the ledger. organize_auto proposes keywords per clip from cached captions and the transcript (apply=true writes them). Select clips with clip_name (glob), keyword and/or role, then organize_keywords (keywords, mode add|remove|replace), organize_rate (rating favorite|rejected|clear) or organize_roles (audio_role, video_role); each writes a _organized copy. history lists every recorded operation for the file's folder (limit); undo (n) moves the outputs of the last n writes into the journal's undone/ folder — it never deletes and refuses when a file changed since it was written. Actions: organize_auto, organize_keywords, organize_rate, organize_roles, history, undo.
| Name | Required | Description | Default |
|---|---|---|---|
| args | No | Arguments for the chosen action, e.g. {"filepath": "/path/to/project.fcpxml"}. | |
| action | Yes | Which operation to run. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so well. It discloses that apply=true writes keywords, that each operation writes a _organized copy, and that undo moves outputs to a journal folder, never deletes, and refuses when files changed since writing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and information-rich, but the single block of text contains several run-on clauses and the final 'Actions:' list partially duplicates what was already described inline and in the schema enum. It is not poorly sized, but structure and conciseness could be improved with bullets or clearer separation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex multi-action tool with no output schema and no annotations, the description covers the main actions, selection criteria, side effects, and undo safety. Minor gaps remain, such as exact return values for history and precise required arguments per action, but the description gives enough context for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema only has action and args, but the description enriches both by explaining each action's relevant parameters such as clip_name, keyword, role, keywords, mode, rating, audio_role, video_role, and limit. It adds real meaning beyond the schema, though it does not provide a complete per-action argument spec.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as a bulk library logging and ledger utility, and enumerates specific operations like organize_auto, organize_keywords, organize_rate, organize_roles, history, and undo. This distinguishes it from the sibling tools by domain and action, though it does not name a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a concrete workflow: select clips, then apply keywords, ratings, or roles; it also explains when history and undo are relevant. However, it does not explicitly compare this tool to alternatives or state when not to use it, so the usage guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
previewA
See the edit without opening Final Cut Pro. Render a proxy video of the timeline, a contact sheet of every cut, a single frame, or a filmstrip-plus-waveform check read from the SOURCE MEDIA rather than from the XML. Run preview_check to confirm a fix actually landed. Actions: preview_render, preview_sheet, preview_frame, preview_check, preview_timeline.
| Name | Required | Description | Default |
|---|---|---|---|
| args | No | Arguments for the chosen action, e.g. {"filepath": "/path/to/project.fcpxml"}. | |
| action | Yes | Which operation to run. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of explaining behavior. It usefully reveals that previews read from SOURCE MEDIA rather than from the XML and that Final Cut Pro need not be opened. However, it does not disclose whether rendered proxies or check outputs persist as files, what side effects occur, or what output format the agent should expect.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with the main value proposition, and every sentence adds information. It avoids restating the tool name or schema boilerplate and ends with a clean action list.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has five distinct actions, no output schema, and no annotations, yet the description never explains what each action returns or what arguments each action needs apart from a generic filepath example. An agent selecting between preview_render and preview_timeline would still lack enough detail to invoke the right action correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%: both action and args have descriptions, and action is an enum. The description reinforces the action meanings by naming the five action values and clarifying that the check reads source media, but it does not add detailed semantics for the args object beyond the schema's single example.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear, concrete purpose: 'See the edit without opening Final Cut Pro.' It then enumerates exactly what the tool produces (proxy video, contact sheet, single frame, filmstrip-plus-waveform check) and lists all five actions. This is a specific verb-plus-resource statement that distinguishes preview from sibling tools like inspect or diagnose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a practical context for use: preview output can be generated from source media without launching Final Cut Pro, and preview_check should be run 'to confirm a fix actually landed.' It does not explicitly exclude cases or name alternative sibling tools, but it clearly indicates when preview is the right choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scenesA
Shot boundaries. detect_scenes reads every source in a timeline and reports cuts in timeline time (args: filepath, clip_name?, backend auto|content|adaptive|ffmpeg, threshold?, min_scene_len=0.5); scenes_to_markers writes a marker at each cut; scenes_split cuts the clips there. PySceneDetect via the scenes extra, ffmpeg without it — the result names which one answered. Actions: detect_scenes, scenes_to_markers, scenes_split.
| Name | Required | Description | Default |
|---|---|---|---|
| args | No | Arguments for the chosen action, e.g. {"filepath": "/path/to/project.fcpxml"}. | |
| action | Yes | Which operation to run. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does a good job: it discloses that detect_scenes reads every source, reports cuts in timeline time, writes markers, and performs splits. It also reveals the backend dependency on PySceneDetect vs ffmpeg and that the result names which backend answered, though it does not describe the exact return shape or whether splitting is destructive/reversible.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loads 'Shot boundaries,' then packs the three actions plus backend behavior into a compact format. The inline argument list inside the first sentence makes parsing slightly harder, but every clause contributes useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a multi-action tool with no output schema and no annotations, the description leaves gaps: it does not describe the return value structure for detect_scenes, and it does not spell out the arguments required for scenes_to_markers and scenes_split. The action enum and filepath are covered, but an agent could still be unsure how to invoke two of the three actions correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema only describes args as a generic object, so the description adds substantial meaning by enumerating filepath, clip_name, backend, threshold, and min_scene_len with optionality and a default. However, it does not specify which arguments apply to scenes_to_markers and scenes_split, so it is not fully complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool's domain with 'Shot boundaries' and names three distinct actions with specific outcomes: detect_scenes reports cuts, scenes_to_markers writes markers, and scenes_split cuts clips. This differentiates it from sibling tools like 'mark' and 'edit' by giving each action a specific verb and resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context by laying out the three operations and their sequential relationship: first detect, then optionally write markers or split. It does not explicitly name sibling alternatives or state when not to use this tool, so it falls just short of giving full exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transcriptB
Transcribe source media locally and edit the timeline by what was SAID rather than by timecode. Also removes filler words. Actions: transcribe_media, edit_by_transcript, remove_filler_words, transcript_pack.
| Name | Required | Description | Default |
|---|---|---|---|
| args | No | Arguments for the chosen action, e.g. {"filepath": "/path/to/project.fcpxml"}. | |
| action | Yes | Which operation to run. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden; it does reveal local processing and the 'by what was SAID' approach. But it never states whether edit_by_transcript or remove_filler_words mutates the timeline/project, whether changes are reversible, or what transcript_pack returns. Important behavioral gaps remain for a tool that implies editing side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences plus an action list; the description front-loads the core value proposition and is easy to scan. The filler-word sentence slightly duplicates the remove_filler_words action name but still provides immediate comprehension.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a multi-action tool with no annotations and no output schema, and the description defines each action only by name, not by input/output or side-effect behavior. An agent would struggle to know which action to select in a given situation or what result to expect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers both parameters fully, including enum values and an args example, so the baseline is 3. The description repeats the action names but adds little meaning about expected args, return effects, or how each action customizes the args object.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete capability—transcribe media locally and edit the timeline by spoken content rather than timecode—and lists the accepted actions, making the tool's purpose and scope clear. It distances itself from generic edit and generate siblings by emphasizing transcript-based editing and filler-word removal. However, it doesn't explicitly contrast itself with the sibling 'edit' tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Use cases are implied: editing by transcript or removing filler words. But there is no explicit when-to-use/when-not-to-use guidance, no named alternatives, and no exclusions. This is minimum viable but leaves routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
watchA
Close the round-trip. Watch the Final Cut Pro export folder, detect the XML the moment it lands, and diff it against the last one seen. Pair with deliver.push_to_fcp for a full loop: push in, edit, Cmd-E, watch_pull. Actions: watch_start, watch_status, watch_stop, watch_pull.
| Name | Required | Description | Default |
|---|---|---|---|
| args | No | Arguments for the chosen action, e.g. {"filepath": "/path/to/project.fcpxml"}. | |
| action | Yes | Which operation to run. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must carry the behavioral burden. It does disclose stateful behavior: watching the folder, detecting the XML, and diffing against the last seen one. However, it does not explain the side effects or lifecycle semantics of the four actions, such as what watch_start begins, what watch_status reports, or what watch_pull returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: it opens with the workflow value, then states the mechanism, then the companion tool, then the action list. The phrase 'Close the round-trip' is slightly jargon-heavy, but each sentence earns its place and there is no redundant filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a multi-action dispatcher with no output schema and no annotations, so the description needs to explain each action's individual contract. It does not: watch_start, watch_status, and watch_stop are only named, not described, and there is no indication of return values, state persistence, or prerequisites beyond the general workflow pairing with deliver.push_to_fcp.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description lists the action enum values but adds no semantic detail beyond the schema's own 'Which operation to run' and the args example. It does not clarify which arguments each action requires.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Watch the Final Cut Pro export folder, detect the XML the moment it lands, and diff it against the last one seen.' This makes the tool's core function unmistakable and distinguishes it from siblings like deliver.push_to_fcp by framing it as the feedback step in a round-trip workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit workflow guidance: 'Pair with deliver.push_to_fcp for a full loop: push in, edit, Cmd-E, watch_pull.' This tells the agent when in the process the tool belongs. It does not explicitly state when not to use it or name alternatives, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
75 tool updates
v0.22.1- Removed
add_audio - Removed
add_connected_clip - Removed
add_marker - Removed
add_transition - Removed
analyze_pacing - Removed
analyze_timeline - Removed
apply_template - Removed
assign_role - Removed
auto_rough_cut - Removed
batch_add_markers - Removed
change_speed - Removed
create_compound_clip - Removed
delete_clips - Added
deliver - Removed
detect_beats - Removed
detect_duplicates - Removed
detect_flash_frames - Removed
detect_gaps - Removed
detect_media_silence - Removed
detect_silence_candidates - Added
diagnose - Removed
diff_timelines - Added
edit - Removed
edit_by_transcript - Removed
export_csv - Removed
export_edl - Removed
export_fcp7_xml - Removed
export_resolve_xml - Removed
export_role_stems - Removed
fill_gaps - Removed
filter_by_role - Added
find - Removed
find_long_clips - Removed
find_short_cuts - Removed
fix_flash_frames - Removed
flatten_compound_clip - Added
generate - Removed
generate_ab_roll - Removed
generate_montage - Removed
import_beat_markers - Removed
import_srt_markers - Removed
import_transcript_markers - Added
index - Removed
insert_clip - Added
inspect - Removed
list_clips - Removed
list_compound_clips - Removed
list_connected_clips - Removed
list_effects - Removed
list_fcp_libraries - Removed
list_keywords - Removed
list_library_clips - Removed
list_markers - Removed
list_projects - Removed
list_roles - Removed
list_templates - Added
mark - Added
organize - Added
preview - Removed
push_to_fcp - Removed
rapid_trim - Removed
reformat_timeline - Removed
relink_media - Removed
remove_filler_words - Removed
remove_media_silence - Removed
remove_silence_candidates - Removed
reorder_clips - Added
scenes - Removed
snap_to_beats - Removed
split_clip - Removed
transcribe_media - Added
transcript - Removed
trim_clip - Removed
validate_timeline - Added
watch
7 tool updates
v0.13.1- Added
add_marker - Added
assign_role - Added
batch_add_markers - Added
detect_gaps - Added
export_edl - Added
list_compound_clips - Added
list_projects
10 tool updates
v0.13.0- Removed
add_marker - Removed
assign_role - Removed
batch_add_markers - Removed
detect_gaps - Added
edit_by_transcript - Removed
export_edl - Removed
list_compound_clips - Removed
list_projects - Added
remove_filler_words - Added
transcribe_media
3 tool updates
v0.12.0- Added
detect_beats - Added
detect_media_silence - Added
remove_media_silence
3 tool updates
v0.9.0- Added
list_fcp_libraries - Added
push_to_fcp - Added
relink_media
53 tool updates
v0.6.63- First observed
add_audio - First observed
add_connected_clip - First observed
add_marker - First observed
add_transition - First observed
analyze_pacing - First observed
analyze_timeline - First observed
apply_template - First observed
assign_role - First observed
auto_rough_cut - First observed
batch_add_markers - First observed
change_speed - First observed
create_compound_clip - First observed
delete_clips - First observed
detect_duplicates - First observed
detect_flash_frames - First observed
detect_gaps - First observed
detect_silence_candidates - First observed
diff_timelines - First observed
export_csv - First observed
export_edl - First observed
export_fcp7_xml - First observed
export_resolve_xml - First observed
export_role_stems - First observed
fill_gaps - First observed
filter_by_role - First observed
find_long_clips - First observed
find_short_cuts - First observed
fix_flash_frames - First observed
flatten_compound_clip - First observed
generate_ab_roll - First observed
generate_montage - First observed
import_beat_markers - First observed
import_srt_markers - First observed
import_transcript_markers - First observed
insert_clip - First observed
list_clips - First observed
list_compound_clips - First observed
list_connected_clips - First observed
list_effects - First observed
list_keywords - First observed
list_library_clips - First observed
list_markers - First observed
list_projects - First observed
list_roles - First observed
list_templates - First observed
rapid_trim - First observed
reformat_timeline - First observed
remove_silence_candidates - First observed
reorder_clips - First observed
snap_to_beats - First observed
split_clip - First observed
trim_clip - First observed
validate_timeline
TDQS
Scored across 13 tools
The thirteen tools map onto distinct workflow stages—inspect, diagnose, edit, deliver, preview, watch, etc.—so an agent can usually pick the right one by intent. Minor overlap exists between inspect and diagnose (both read-only analysis) and between scenes_split and edit.split_clip, but the descriptions separate general understanding from problem-finding and scene-boundary work.
All tool names follow a consistent single-lowercase-word style, which is predictable after seeing a few. The main deviation is grammatical: most are verbs (inspect, edit, mark, watch, find) while transcript, index, and scenes are nouns, and the actual operations live inside as snake_case actions.
With 13 tools, the server sits comfortably in the well-scoped range, and each tool represents a distinct phase of FCP XML editing: analysis, diagnostics, editing, marking, generation, transcription, delivery, preview, round-trip, indexing, scene detection, organization, and search. None feels redundant; the breadth is justified by the size of the domain.
The surface is unusually complete for the domain, covering inspection, diagnosis, editing, generation, metadata logging, transcription, export, preview, watch-loop integration, caching, scene detection, organization, and natural-language search. Read-only tools have write counterparts, output paths are verifiable, and there are no obvious dead-end workflows.
Maintenance
Related MCP Connectors
MCP server for Clipkit — gives AI agents a video toolbox via the Clipkit schema.
Hosted MCP tools for FFmpeg-style video and audio processing through FFMPEG API.
MCP server for the FFmpeg Micro video transcoding API — create, monitor, download transcodes.
MCP server for generating rough-draft project plans from natural-language prompts.
Related MCP Servers
- FlicenseCqualityDmaintenanceEnables comprehensive remote control and automation of Final Cut Pro through 99 tools covering timeline editing, project management, and AI-powered features. It facilitates complex workflows including media organization, color grading, and FCPXML generation using AppleScript and JXA automation.10011-
- AlicenseBqualityAmaintenanceThe most capable MCP server for Final Cut Pro — 88 tools covering FCPXML editing, live FCP control, parametric puppets, and media analysis.8844 PyPI9MIT
- AlicenseNot gradedqualityFmaintenanceComprehensive MCP server for DaVinci Resolve with 295+ tools to control projects, timelines, editing, color grading, rendering, and more via natural language.106 PyPI6MIT
- AlicenseAqualityBmaintenanceGive any MCP client a real video editor — 32 typed tools over ffmpeg, Whisper and MediaPipe, plus an optional local UI with a drag-and-drop timeline.38MIT