Skip to main content
Glama

🎬 CapCut MCP Server (Upgraded by Antigravity)

이 ν”„λ‘œμ νŠΈλŠ” CapCut Pro μ˜μƒ νŽΈμ§‘μ„ AI μ—μ΄μ „νŠΈκ°€ μ™„λ²½ν•˜κ²Œ μžλ™ν™”ν•  수 μžˆλ„λ‘ μ§€μ›ν•˜λŠ” Model Context Protocol (MCP) μ„œλ²„μž…λ‹ˆλ‹€. κΈ°μ‘΄ μ˜€ν”ˆμ†ŒμŠ€ 기반 μœ„μ— **우리 νŒ€(Antigravity & Luca Agent)**의 ν•„μš”μ— 맞좰 λ Œλ”λ§ νŒŒμ΄ν”„λΌμΈ, μ•„μ‹œμ•„κΆŒ 폰트 지원 및 νŠΈλžœμ§€μ…˜ 처리 λ‘œμ§μ„ λŒ€ν­ μ—…κ·Έλ ˆμ΄λ“œν–ˆμŠ΅λ‹ˆλ‹€! πŸš€

✨ 무엇이 μ—…κ·Έλ ˆμ΄λ“œ λ˜μ—ˆλ‚˜μš”? (Antigravity's Touch)

  1. VectCutAPI μ•ˆμ •μ„± κ·ΉλŒ€ν™” πŸ› οΈ

    • Python 기반의 둜컬 λ Œλ”λ§ μ„œλ²„μΈ VectCutAPIλ₯Ό λ‚΄μž₯ν•˜κ³ , save_draft_impl.py λ“±μ˜ μ €μž₯ λ‘œμ§μ„ κ°œμ„ ν•˜μ—¬ μΊ‘μ»· μ΄ˆμ•ˆ(Draft)이 μ•ˆμ „ν•˜κ²Œ μ‚¬μš©μž 폴더에 κ½‚νžˆλ„λ‘ μˆ˜μ •ν–ˆμŠ΅λ‹ˆλ‹€.

  2. λ‹€κ΅­μ–΄(CJK) 폰트 μ™„λ²½ 지원 🌏

    • κΈ°λ³Έ System 폰트 μ—λŸ¬λ₯Ό ν•΄κ²°ν•˜κ³ , SourceHanSansCN_Regular λ“± λ‹€κ΅­μ–΄/ν•œκΈ€ 지원 폰트λ₯Ό λ§€ν•‘ν•˜μ—¬ μžλ§‰ 깨짐 없이 κΉ”λ”ν•˜κ²Œ 좜λ ₯λ˜λ„λ‘ 폰트 엔진을 μ—…κ·Έλ ˆμ΄λ“œν–ˆμŠ΅λ‹ˆλ‹€.

  3. κ³ κΈ‰ νŠΈλžœμ§€μ…˜ 및 ν‚€ν”„λ ˆμž„ 보강 πŸŒ€

    • Dissolve, Mix λ“± μΊ‘μ»· 고유의 λŒ€μ†Œλ¬Έμž ꡬ뢄 νŠΈλžœμ§€μ…˜ μ—λŸ¬λ₯Ό ν•΄κ²°ν•˜κ³ , 쀌인/μ€Œμ•„μ›ƒ λ“± add_video_keyframe을 ν™œμš©ν•œ λ―Έμ„Έ μ• λ‹ˆλ©”μ΄μ…˜ μ œμ–΄λ₯Ό κ°•ν™”ν–ˆμŠ΅λ‹ˆλ‹€.

  4. 유튜브 μ‡ΌμΈ  μžλ™ν™” μ΅œμ ν™” πŸ“±

    • 9:16 λΉ„μœ¨(μ„Έλ‘œν˜•) λ Œλ”λ§ 및 λ””μ‘ΈλΈŒ νŠΈλžœμ§€μ…˜μ„ μžλ™ μ μš©ν•˜μ—¬ μŒμ›μ— 맞좰 이미지λ₯Ό λ‘€λ§ν•˜λŠ” 유튜브 νŒŒμ΄ν”„λΌμΈμ— μ΅œμ ν™”λ˜μ—ˆμŠ΅λ‹ˆλ‹€.


Related MCP server: capcut-mcp

πŸš€ μ‹œμž‘ν•˜κΈ° (Installation & Setup)

CapCut MCPλŠ” 2단계 ꡬ쑰둜 μ‹€ν–‰λ©λ‹ˆλ‹€. Python μ„œλ²„(VectCutAPI)κ°€ μΊ‘μ»· μ΄ˆμ•ˆμ„ λ§Œλ“€κ³ , Node.js μ„œλ²„(MCP)κ°€ μ—μ΄μ „νŠΈμ™€ μ†Œν†΅ν•©λ‹ˆλ‹€.

1. Python λ°±μ—”λ“œ μ‹€ν–‰ (VectCutAPI)

cd VectCutAPI
pip install -r requirements.txt
python capcut_server.py

기본적으둜 http://localhost:9000 (λ˜λŠ” 9001)μ—μ„œ λŒ€κΈ°ν•©λ‹ˆλ‹€.

2. Node.js MCP μ„œλ²„ μ—°κ²°

μ•ˆν‹°κ·Έλž˜λΉ„ν‹° mcp_config.json에 λ‹€μŒ 섀정을 μΆ”κ°€ν•˜λ©΄ μ—μ΄μ „νŠΈκ°€ μžλ™μœΌλ‘œ μΈμ‹ν•©λ‹ˆλ‹€.

{
  "mcpServers": {
    "capcut-mcp": {
      "command": "node",
      "args": ["C:/Users/User/Documents/capcut-mcp/dist/index.js"],
      "env": {
        "CAPCUT_API_URL": "http://localhost:9000"
      }
    }
  }
}

πŸ› οΈ μ œκ³΅λ˜λŠ” 도ꡬ (Available Tools)

AI μ—μ΄μ „νŠΈλŠ” λ‹€μŒ 11κ°€μ§€ 도ꡬλ₯Ό 톡해 λΉ„λ””μ˜€λ₯Ό ν”„λ‘œκ·Έλž˜λ° λ°©μ‹μœΌλ‘œ νŽΈμ§‘ν•©λ‹ˆλ‹€.

  1. capcut_create_draft: μƒˆλ‘œμš΄ ν”„λ‘œμ νŠΈ μ΄ˆμ•ˆ 생성 (HD, 4K, μ„Έλ‘œν˜• λ“±)

  2. capcut_add_video: μ˜μƒ μΆ”κ°€ 및 μ»·νŽΈμ§‘, 배속, λ³Όλ₯¨ 쑰절

  3. capcut_add_audio: λ°°κ²½μŒμ•…, BGM 및 νŽ˜μ΄λ“œμΈ/아웃 적용

  4. capcut_add_text: 타이틀, ν…μŠ€νŠΈ μΆ”κ°€ (색상, 그림자, μ• λ‹ˆλ©”μ΄μ…˜)

  5. capcut_add_image: 이미지 μΆ”κ°€ 및 배치 (크기, νšŒμ „)

  6. capcut_add_subtitle: SRT ν˜•μ‹μ˜ μžλ§‰ μžλ™ 생성 및 싱크 λ§€ν•‘

  7. capcut_add_keyframe: λΆ€λ“œλŸ¬μš΄ 쀌인/μ€Œμ•„μ›ƒ λ“± ν‚€ν”„λ ˆμž„ μ• λ‹ˆλ©”μ΄μ…˜ 생성

  8. capcut_add_effect: λΈ”λŸ¬, λΉ„λ„€νŒ… λ“± ν™”λ©΄ μ΄νŽ™νŠΈ μΆ”κ°€

  9. capcut_add_sticker: μŠ€ν‹°μ»€ 및 이λͺ¨μ§€ μ‚½μž…

  10. capcut_save_draft: μ™„λ£Œλœ ν”„λ‘œμ νŠΈλ₯Ό μΊ‘μ»· λ°μŠ€ν¬νƒ‘ μ•±μœΌλ‘œ μΆ”μΆœ

  11. capcut_get_duration: μ˜μƒ/μŒμ›μ˜ μ •ν™•ν•œ 길이(메타데이터) 확인


πŸ“– μžλ™ν™” μ˜ˆμ‹œ (Workflow Example)

# μ—μ΄μ „νŠΈ νŒŒμ΄ν”„λΌμΈ λ™μž‘ μ˜ˆμ‹œ
draft = req("/create_draft", {"width": 1080, "height": 1920}) # μ„Έλ‘œν˜• μ‡ΌμΈ 
req("/add_audio", {"draft_id": draft["draft_id"], "audio_url": "music.wav", "start": 0, "end": 60})
req("/add_image", {"draft_id": draft["draft_id"], "image_url": "scene.png", "transition": "Dissolve"})
req("/add_subtitle", {"draft_id": draft["draft_id"], "srt": "lyrics.srt", "font": "SourceHanSansCN_Regular"})
req("/save_draft", {"draft_id": draft["draft_id"]})

πŸ™ Acknowledgments

  • Based on the original VectCutAPI framework.

  • Upgraded and Maintained by the Antigravity & Luca Agent Team for hyper-automated video production.

Available Tools

11 tools
capcut_add_audioAdd Audio to DraftA

Add audio track to draft with volume and fade effects.

This tool adds background music or sound effects to the video timeline.

Args:

  • draft_id (string): The draft ID

  • audio_url (string): URL to audio file (mp3, wav, aac, m4a, flac, ogg)

  • start (number): Start time in seconds

  • end (number): End time in seconds

  • volume (number): Audio volume 0.0-1.0 (default: 1.0)

  • fade_in (number): Fade in duration in seconds (default: 0)

  • fade_out (number): Fade out duration in seconds (default: 0)

  • response_format ('markdown' | 'json'): Output format

Examples:

  • Add background music: audio_url="https://...", volume=0.5

  • Add with fade: fade_in=2, fade_out=2

ParametersJSON Schema
NameRequiredDescriptionDefault
draft_idYesThe ID of the draft to add audio to
audio_urlYesURL to the media file
startYesStart time in seconds
endYesEnd time in seconds
volumeNoAudio volume (0.0 to 1.0)
fade_inNoFade in duration in seconds
fade_outNoFade out duration in seconds
response_formatNoOutput format: 'markdown' for human-readable or 'json' for machine-readablemarkdown

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate mutating (readOnlyHint=false) and non-destructive (destructiveHint=false). The description adds the accepted audio formats (mp3, wav, etc.) but does not disclose whether audio is appended or replaces existing tracks, or any limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is structured with a summary sentence and an Args section. While the Args replicates schema info, it is clear and the examples are helpful. Could be slightly more concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 8 parameters and no output schema. The description covers all parameters with examples, explaining start/end times and volume/fade. It does not explain return values, but for a simple addition tool this is acceptable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, baseline 3. The description enhances parameters by listing audio formats (mp3, wav, aac, m4a, flac, ogg) for audio_url and provides examples showing volume and fade usage, adding value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool adds an audio track to a draft, with volume and fade effects. It distinguishes from sibling tools like capcut_add_video (adds video) and capcut_add_effect (adds effects).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool vs alternatives. The description implies it is for adding audio (background music or sound effects), but does not compare to other add tools or specify prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

capcut_add_effectAdd Visual EffectA

Apply visual effects to video segments.

This tool adds effects like blur, sharpen, brightness adjustments, and more.

Args:

  • draft_id (string): The draft ID

  • effect_name (string): Effect to apply (blur, sharpen, brightness, contrast, saturation, vignette, grain, glitch)

  • start (number): Start time in seconds

  • end (number): End time in seconds

  • intensity (number): Effect intensity 0.0-1.0 (default: 0.5)

  • response_format ('markdown' | 'json'): Output format

Examples:

  • Add blur: effect_name="blur", intensity=0.7

  • Increase brightness: effect_name="brightness", intensity=0.8

  • Add vignette: effect_name="vignette", intensity=0.4

ParametersJSON Schema
NameRequiredDescriptionDefault
draft_idYesThe ID of the draft to add effect to
effect_nameYesName of the effect to apply
startYesStart time in seconds
endYesEnd time in seconds
intensityNoEffect intensity (0.0 to 1.0)
response_formatNoOutput format: 'markdown' for human-readable or 'json' for machine-readablemarkdown

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate this is a write operation (readOnlyHint false) but no destructive or idempotent info. The description adds no behavioral details beyond what the schema provides, such as whether effects stack or replace. It adequately describes the action but does not enhance transparency beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, well-structured with Args and Examples sections. Every sentence adds value, and the main purpose is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (6 params) and no output schema, the description provides examples and covers all parameters. It lacks information on return values or side effects, but for a straightforward apply-effect tool, it is nearly complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so parameters are well-documented. The description repeats parameter info and provides examples, adding marginal value. Baseline 3 is appropriate as it does not significantly expand on the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The title and description clearly state the tool applies visual effects to video segments, listing specific effects like blur, sharpen, brightness. This distinguishes it from sibling tools that add audio, images, etc.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool (for visual effects on video segments) and excludes alternatives like adding audio or images. It does not explicitly state when not to use, but context from sibling tools makes it clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

capcut_add_imageAdd Image to DraftA

Add image overlay to video with positioning, scaling, rotation, and animation.

This tool adds static or animated images to the video timeline.

Args:

  • draft_id (string): The draft ID

  • image_url (string): URL to image file (jpg, jpeg, png, gif, webp, bmp)

  • start (number): Start time in seconds

  • end (number): End time in seconds

  • position_x (number): Horizontal position 0.0-1.0 (default: 0.5)

  • position_y (number): Vertical position 0.0-1.0 (default: 0.5)

  • scale (number): Scale multiplier 0.1-5.0 (default: 1.0)

  • rotation (number): Rotation angle 0-360 degrees (default: 0)

  • animation (string): Animation effect (optional)

  • response_format ('markdown' | 'json'): Output format

Examples:

  • Add logo: image_url="https://...", position_x=0.9, position_y=0.1, scale=0.3

  • Add rotating image: rotation=45, animation="zoom_in"

ParametersJSON Schema
NameRequiredDescriptionDefault
draft_idYesThe ID of the draft to add image to
image_urlYesURL to the media file
startYesStart time in seconds
endYesEnd time in seconds
position_xNoHorizontal position (0.0 to 1.0)
position_yNoVertical position (0.0 to 1.0)
scaleNoScale multiplier
rotationNoRotation angle in degrees
animationNoAnimation effect to apply
response_formatNoOutput format: 'markdown' for human-readable or 'json' for machine-readablemarkdown

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate non-readonly, non-destructive, non-idempotent, open-world. The description adds context about adding overlay to the timeline, which is helpful. No contradictions with annotations. It does not detail return values or error handling, but annotations already cover behavioral hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured with an Args section listing parameters clearly. Examples are included without unnecessary verbosity. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 10 parameters (4 required) and no output schema, the description adequately explains how to use the tool with examples. It does not describe the output format or return values, but the response_format parameter covers that. Missing some situational context (e.g., prerequisites like draft existence).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds value by providing default values, ranges (e.g., 0.1-5.0 for scale), and examples for parameters like position_x and rotation. This goes beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool adds an image overlay to video with positioning, scaling, rotation, and animation. It distinguishes itself from sibling tools (e.g., add_audio, add_video) by specifying the media type (image) and its capabilities.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides examples and explains the tool's functionality, but it does not explicitly state when to use this tool over alternatives. Sibling names imply different media types, so usage context is somewhat clear, but lacking explicit 'when not to use' guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

capcut_add_keyframeAdd Keyframe AnimationA

Add keyframe-based property animation to tracks.

This tool creates smooth animations by interpolating between keyframe values.

Args:

  • draft_id (string): The draft ID

  • track_name (string): Name of track to animate

  • property_types (string[]): Properties to animate (scale_x, scale_y, alpha, rotation, position_x, position_y)

  • times (number[]): Keyframe times in seconds (at least 2)

  • values (string[]): Values for each keyframe (same length as times)

  • response_format ('markdown' | 'json'): Output format

Examples:

  • Fade in: property_types=["alpha"], times=[0, 2], values=["0.0", "1.0"]

  • Zoom in: property_types=["scale_x", "scale_y"], times=[0, 2], values=["0.5", "1.5"]

  • Rotate: property_types=["rotation"], times=[0, 3], values=["0", "360"]

ParametersJSON Schema
NameRequiredDescriptionDefault
draft_idYesThe ID of the draft to add keyframes to
track_nameYesName of the track to animate
property_typesYesProperties to animate (e.g., scale_x, scale_y, alpha, rotation)
timesYesKeyframe times in seconds
valuesYesValues for each keyframe
response_formatNoOutput format: 'markdown' for human-readable or 'json' for machine-readablemarkdown

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are minimal (no readOnly or destructive hints), and the description adds that it creates animations but does not disclose behaviors like whether existing keyframes are overwritten or if operations are reversible.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with a summary, explanation, argument list, and examples. However, the argument list partially repeats schema details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Missing output schema and no description of return values. Also lacks information on prerequisites (e.g., draft must exist) and error handling, which would be helpful for such a tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the description adds value with examples that demonstrate how to combine properties, times, and values for effects like fade-in, zoom, and rotation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Add keyframe-based property animation to tracks' with an explanation of interpolation. This distinguishes it from sibling tools that add audio, effects, etc.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not explicitly state when or when not to use this tool, but the examples and context make it clear it's for animation. No exclusions or alternatives are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

capcut_add_stickerAdd Sticker to DraftA

Add sticker/emoji overlay with positioning and transformation.

This tool adds decorative stickers or emojis to the video.

Args:

  • draft_id (string): The draft ID

  • sticker_url (string): URL to sticker image

  • start (number): Start time in seconds

  • end (number): End time in seconds

  • position_x (number): Horizontal position 0.0-1.0 (default: 0.5)

  • position_y (number): Vertical position 0.0-1.0 (default: 0.5)

  • scale (number): Scale multiplier 0.1-5.0 (default: 1.0)

  • rotation (number): Rotation angle 0-360 degrees (default: 0)

  • response_format ('markdown' | 'json'): Output format

Examples:

  • Add corner sticker: position_x=0.9, position_y=0.1, scale=0.2

  • Add rotating emoji: rotation=15, scale=0.5

ParametersJSON Schema
NameRequiredDescriptionDefault
draft_idYesThe ID of the draft to add sticker to
sticker_urlYesURL to the media file
startYesStart time in seconds
endYesEnd time in seconds
position_xNoHorizontal position (0.0 to 1.0)
position_yNoVertical position (0.0 to 1.0)
scaleNoScale multiplier
rotationNoRotation angle in degrees
response_formatNoOutput format: 'markdown' for human-readable or 'json' for machine-readablemarkdown

TDQS

A3.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=false, destructiveHint=false, idempotentHint=false, openWorldHint=true. The description does not add behavioral context beyond stating it 'adds decorative stickers or emojis'. No mention of side effects, authorization needs, or state changes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with a concise summary, followed by a clear Args list and relevant examples. Every sentence adds value without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers parameter usage, defaults, and examples comprehensively. However, it does not describe the return value or response structure, which would be useful given no output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with descriptions for all 9 parameters. The description adds value through practical examples (e.g., corner sticker positioning) and clarifies response_format defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool adds a sticker/emoji overlay with positioning and transformation. It distinguishes from siblings like capcut_add_text and capcut_add_image through the specific resource and behavior.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage via examples and defaults, but lacks explicit guidance on when to use this tool versus alternatives like capcut_add_text or capcut_add_image. No when-not-to-use or prerequisite information is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

capcut_add_subtitleAdd Subtitles to DraftB

Add subtitles from SRT file content with styling options.

This tool imports subtitles in SRT format and applies styling.

Args:

  • draft_id (string): The draft ID

  • srt_content (string): SRT formatted subtitle content

  • font (string): Font family name (optional)

  • font_size (number): Font size 12-100 (default: 36)

  • font_color (string): Hex color (default: #FFFFFF)

  • background_enabled (boolean): Enable background (default: true)

  • background_color (string): Background hex color (default: #000000)

  • response_format ('markdown' | 'json'): Output format

Example SRT format: 1 00:00:01,000 --> 00:00:03,000 Welcome to my video

2 00:00:03,500 --> 00:00:05,000 Subscribe for more content

ParametersJSON Schema
NameRequiredDescriptionDefault
draft_idYesThe ID of the draft to add subtitles to
srt_contentYesSRT formatted subtitle content
fontNoFont family name
font_sizeNoFont size in points
font_colorNoFont color in hex format#FFFFFF
background_enabledNoEnable background behind text
background_colorNoBackground color in hex format#000000
response_formatNoOutput format: 'markdown' for human-readable or 'json' for machine-readablemarkdown

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations already indicating the tool is not read-only or destructive, the description adds minimal behavioral context. It does not mention potential side effects, permission requirements, or what happens to existing subtitles.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description includes a detailed Args list and example, making it somewhat long. While structured, it could be more concise by omitting redundant parameter details already in schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, but the description does not explain what the tool returns (e.g., updated draft, confirmation). It only mentions response_format as a parameter, not the actual output. This leaves the agent uncertain about the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, providing baseline 3. The description adds value with an example SRT format and listing default values in the Args section, which helps understanding beyond schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool adds subtitles from SRT content with styling options. It distinguishes itself from sibling tools like capcut_add_text, which likely handles manual text entry.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies use when you have SRT content but does not explicitly state when to use this versus other text/subtitle tools. No alternatives or exclusion criteria provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

capcut_add_textAdd Text to DraftA

Add styled text overlay to video with positioning, colors, shadows, and animations.

This tool creates text elements with full styling control including fonts, colors, backgrounds, shadows, and animations.

Args:

  • draft_id (string): The draft ID

  • text (string): Text content to display (1-500 characters)

  • start (number): Start time in seconds

  • end (number): End time in seconds

  • font (string): Font family name (optional)

  • font_size (number): Font size 12-200 (default: 48)

  • font_color (string): Hex color e.g., #FFFFFF (default: #FFFFFF)

  • background_color (string): Background hex color (optional)

  • background_alpha (number): Background opacity 0.0-1.0 (default: 0.8)

  • shadow_enabled (boolean): Enable shadow (default: false)

  • shadow_color (string): Shadow hex color (default: #000000)

  • position_x (number): Horizontal position 0.0-1.0 (default: 0.5 center)

  • position_y (number): Vertical position 0.0-1.0 (default: 0.5 center)

  • animation (string): Animation effect (fade_in, slide_up, slide_down, slide_left, slide_right, zoom_in, bounce)

  • response_format ('markdown' | 'json'): Output format

Examples:

  • Add title: text="Welcome", font_size=72, position_y=0.2, animation="fade_in"

  • Add subtitle: text="Subscribe!", font_size=36, background_color="#000000"

ParametersJSON Schema
NameRequiredDescriptionDefault
draft_idYesThe ID of the draft to add text to
textYesThe text content to display
startYesStart time in seconds
endYesEnd time in seconds
fontNoFont family name
font_sizeNoFont size in points
font_colorNoFont color in hex format#FFFFFF
background_colorNoBackground color in hex format
background_alphaNoBackground opacity (0.0 to 1.0)
shadow_enabledNoEnable text shadow
shadow_colorNoShadow color in hex format#000000
position_xNoHorizontal position (0.0 to 1.0, where 0.5 is center)
position_yNoVertical position (0.0 to 1.0, where 0.5 is center)
animationNoAnimation effect to apply
response_formatNoOutput format: 'markdown' for human-readable or 'json' for machine-readablemarkdown

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate this is not read-only or destructive. The description adds that the tool 'creates text elements with full styling control', accurately reflecting its mutating behavior. No contradiction or missing behavioral traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured: a concise summary sentence, a second sentence elaborating styling control, a detailed args list, and practical examples. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (15 parameters, no output schema), the description covers input thoroughly with schema and examples. However, it does not mention the return value (e.g., whether it returns a text element ID or success status), which is a gap for a tool without output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, but the description adds value by providing defaults, ranges, and examples for parameters (e.g., position_x/y default 0.5 = center, animation enum values). This exceeds what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Add styled text overlay to video' with explicit mention of positioning, colors, shadows, and animations. It distinguishes from sibling tools (e.g., capcut_add_audio, capcut_add_video) by focusing on text overlay.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not explicitly state when to use this tool vs alternatives, but the examples (e.g., adding title or subtitle) and the resource focus imply its appropriate context. It lacks explicit 'when not to use' or comparison to siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

capcut_add_videoAdd Video to DraftA

Add a video clip to an existing draft with timing, volume, and effects.

This tool adds video content to the timeline with support for transitions, speed adjustments, and volume control.

Args:

  • draft_id (string): The draft ID from create_draft

  • video_url (string): URL to video file (mp4, mov, avi, mkv, webm, flv)

  • start (number): Start time in seconds (>= 0)

  • end (number): End time in seconds (> 0)

  • volume (number): Audio volume 0.0-1.0 (default: 1.0)

  • transition (string): Optional transition effect (fade_in, fade_out, dissolve, wipe, slide, zoom)

  • speed (number): Playback speed 0.1-10x (default: 1.0)

  • response_format ('markdown' | 'json'): Output format

Examples:

  • Add background video: draft_id="abc123", video_url="https://...", start=0, end=10

  • Add with slow motion: speed=0.5

  • Add with fade in: transition="fade_in"

ParametersJSON Schema
NameRequiredDescriptionDefault
draft_idYesThe ID of the draft to add video to
video_urlYesURL to the media file
startYesStart time in seconds
endYesEnd time in seconds
volumeNoAudio volume (0.0 to 1.0)
transitionNoTransition effect to apply
speedNoPlayback speed multiplier
response_formatNoOutput format: 'markdown' for human-readable or 'json' for machine-readablemarkdown

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate non-read-only and non-destructive behavior. The description adds transparency by detailing supported video formats, parameter defaults, and typical modifications (adding to timeline). No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with a summary paragraph, Args list, and Examples. It is front-loaded with the main purpose and avoids redundancy. Could be slightly shortened but is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 8 parameters and no output schema, the description covers core functionality, required/optional params, and examples. However, it lacks details on error handling or prerequisites beyond draft_id. Still adequate for a clear use case.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema has 100% coverage, but the description adds value by listing supported video formats (mp4, mov, etc.) not in schema, and provides practical examples that clarify parameter combinations. Defaults are also reiterated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool adds a video clip to an existing draft with specific capabilities (timing, volume, effects). It distinguishes itself from siblings like capcut_add_audio by explicitly mentioning 'video clip' and provides a specific verb-resource pair.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage after creating a draft and provides examples, but lacks explicit guidance on when to use this tool vs alternatives like capcut_add_audio or capcut_add_image. No when-not-to-use conditions are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

capcut_create_draftCreate CapCut DraftA

Create a new video editing draft with specified dimensions and frame rate.

This tool initializes a new draft project that can be edited by adding videos, audio, text, images, and effects.

Args:

  • width (number): Video width in pixels (360-4096, default: 1920)

  • height (number): Video height in pixels (360-4096, default: 1080)

  • fps (number): Frames per second (24-120, default: 30)

  • response_format ('markdown' | 'json'): Output format (default: 'markdown')

Returns: { "draft_id": string, // Unique draft identifier for subsequent operations "width": number, // Video width "height": number, // Video height "fps": number, // Frame rate "duration": number, // Current duration (starts at 0) "created_at": string // ISO timestamp }

Examples:

  • Create HD draft: params with width=1920, height=1080

  • Create vertical video: params with width=1080, height=1920

  • Create 4K draft: params with width=3840, height=2160

ParametersJSON Schema
NameRequiredDescriptionDefault
widthNoVideo width in pixels
heightNoVideo height in pixels
fpsNoFrames per second
response_formatNoOutput format: 'markdown' for human-readable or 'json' for machine-readablemarkdown

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate this is not readOnly or destructive, consistent with creating a new draft. The description adds that it initializes a project, but does not disclose that multiple calls create multiple independent drafts (non-idempotent).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is structured with Args, Returns, and Examples, but it repeats schema information. Could be more concise by omitting redundant defaults.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description provides a detailed return example. Sibling tools are clearly different, and the description covers the essential behavior of creating a draft.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with full parameter descriptions. The description mostly repeats schema info, but the Examples section adds practical context for choosing width/height combinations.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Create') and clearly identifies the resource ('video editing draft'). It distinguishes itself from sibling tools (which add elements) by focusing on creating the draft itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this is for starting a new project, but does not explicitly state when not to use it (e.g., if a draft already exists). The return of a draft_id provides implicit guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

capcut_get_durationGet Media DurationA
Read-onlyIdempotent

Get duration and metadata of video or audio file.

This tool analyzes media files to retrieve duration, format, and resolution information.

Args:

  • url (string): URL to media file

  • response_format ('markdown' | 'json'): Output format

Returns: { "duration": number, // Duration in seconds "format": string, // File format "width": number, // Video width (if video) "height": number // Video height (if video) }

Examples:

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesURL to the media file to analyze
response_formatNoOutput format: 'markdown' for human-readable or 'json' for machine-readablemarkdown

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and other hints. The description adds behavioral detail by specifying the return format (duration, format, width, height) and response format options, complementing annotations without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with Args, Returns, and Examples sections, and is front-loaded with the purpose. It is clear but slightly verbose; could be trimmed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With full schema coverage, annotations, and explicit return format, the description provides sufficient context for correct tool selection and invocation, though no output schema is provided.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%. The description adds meaning by explaining the url parameter as 'URL to media file' and response_format with examples, going beyond the schema's basic definitions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The title and description clearly state the tool retrieves duration and metadata of video/audio files, listing specific return fields. It is distinct from sibling tools which are all additive operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides usage context like 'Check video length before adding' and 'Verify audio duration', but does not explicitly exclude alternative tools or give when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

capcut_save_draftSave DraftA
Idempotent

Save the draft to a file that can be imported into CapCut.

This tool finalizes the draft and generates a folder that can be copied to the CapCut drafts directory.

Args:

  • draft_id (string): The draft ID to save

  • response_format ('markdown' | 'json'): Output format

Returns: { "draft_url": string, // Path to the saved draft folder "status": "saved" }

The draft folder starts with "dfd_" and should be copied to:

  • Windows: C:\Users<username>\AppData\Local\CapCut\User Data\Projects\Draft Content

  • macOS: ~/Library/Containers/com.lemon.lvpro/Data/Documents/JianyingPro/User Data/Projects/Draft Content

ParametersJSON Schema
NameRequiredDescriptionDefault
draft_idYesThe ID of the draft to save
response_formatNoOutput format: 'markdown' for human-readable or 'json' for machine-readablemarkdown

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate idempotentHint=true and destructiveHint=false. The description adds that the tool 'finalizes the draft' and provides specific path details for the saved folder, which clarifies the output location and naming convention ('dfd_'). No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with a brief overview, then parameters, return format, and path details. It could be slightly more concise by removing the OS-specific paths, but it remains clear and informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (2 parameters, no output schema), the description fully covers what the tool does, what it returns, and how to use the output. It also provides context about the folder name and installation paths, making it self-contained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the description repeats the parameters without adding new semantic constraints or details beyond the schema. It lists them but does not explain allowed values or relationships (e.g., 'draft_id must be an existing draft').

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'save' and the resource 'draft' and explains that the output is a folder for CapCut import. It distinguishes from siblings like capcut_add_audio or capcut_create_draft, which have different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains what the tool does but does not explicitly state when to use it over alternatives or when not to use it. The context implies it is used after creating a draft, but no exclusion criteria or alternative tools are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

A3.9/5.0
Disambiguation4/5

Most tools have distinct purposes (draft creation, media addition, effects, saving), but add_image and add_sticker are very similar (both overlays with positioning/scaling), and add_text vs add_subtitle serve different text use cases but could still cause confusion for an agent.

Naming Consistency5/5

All tools follow a consistent 'capcut_verb_noun' pattern: create_draft, add_video, add_audio, add_image, add_sticker, add_text, add_subtitle, add_effect, add_keyframe, get_duration, save_draft. The verbs are uniform and descriptive.

Tool Count5/5

With 11 tools, the set is well-scoped for a video editing assistant. Each tool covers a core operation (create, add media types, effects, keyframes, metadata, save) without being bloated.

Completeness3/5

The tool set covers creating drafts and adding various elements, but lacks delete or update operations for draft elements, no export or rendering tool, and no draft listing. Agents have no way to remove mistakes or finalize a video, limiting workflow autonomy.

Maintenance

ActivityInactive
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides MCP servers for AI-driven video editing. Enables offline editing of CapCut drafts and remote control of Adobe Premiere Pro via UXP plugin, with shared media analysis for beat detection and transcription.
    MIT
  • A
    license
    B
    quality
    C
    maintenance
    An MCP server that lets Claude read and edit CapCut desktop draft projects, enabling manipulation of video clips, text, audio, and images with atomic saves and validation.
    17
    2
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    An MCP server that reads and builds CapCut projects locally, enabling natural language queries about project contents, missing media, and creation of new edits including beat-synced cuts.
    6
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/eery1677-lab/capcut-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server