Skip to main content
Glama

Character Swap: create

create_character_swap
Destructive

Swap a character image onto a driving video, optionally restoring selected source-video ranges in the final result. ASYNC — returns a projectId; poll get_job. SPENDS CREDITS (min 25; ×2 for 1080p; +5 voice change; +2 for engine "pro"). Needs a character-image r2Key (from swap_import_character for an image already on VidGuy, or swap_upload for raw bytes) and a driving-video r2Key from swap_upload. A driving-video r2Key is reusable: upload the source once, then call this repeatedly with different characters to batch one cut across many avatars. get_video on a finished swap returns its exact settings (ranges, keys) to replay.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
nameNo
engineNoVisual engine. "standard" (default): any length up to 2 minutes. "pro": higher movement accuracy with lipsync, +2 credits; every swapped section (video minus keep-original ranges) must be 14 s or shorter or the request is rejected — add keepOriginalVisualRanges to split longer footage.
promptNo
saveAudioNoKeep the driving video's original audio in the result (default true). Set false to drop it. Ignored when voiceChangeVoiceId is set — voice change always keeps audio.
resolutionNo
sceneMatchNo
characterR2KeyYescharacter-image r2Key from swap_upload
characterFileNameYes
drivingVideoR2KeyYesdriving-video r2Key from swap_upload
voiceChangeVoiceIdNoIf set, enables voice change with this ElevenLabs voice id (+5 credits)
drivingVideoFileNameYes
drivingDurationSecondsYesDriving video length (drives cost)
keepOriginalVisualRangesNoUp to 20 ranges to keep exactly as filmed, e.g. screen recordings. Times are milliseconds and each range must be at least 250 ms. The swap applies everywhere else; audio processing is unchanged.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changed
    • addedInput schema / properties / engine
      Added value: +{
      +  "description": "Visual engine. \"standard\" (default): any length up to 2 minutes. \"pro\": higher movement accuracy with lipsync, +2 credits; every swapped section (video minus keep-original ranges) must be 14 s or shorter or the request is rejected — add keepOriginalVisualRanges to split longer footage.",
      +  "enum": [
      +    "standard",
      +    "pro"
      +  ],
      +  "type": "string"
      +}
  2. First observed

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide only readOnlyHint=false and destructiveHint=true, which the description does not contradict. Beyond that, the description discloses critical runtime behavior: it is ASYNC (returns a projectId, poll get_job), it spends credits with a detailed cost table (min 25, ×2 for 1080p, +5 voice, +2 pro), and it mentions that get_video returns exact settings for replay. This far exceeds the annotation coverage and is essential for correct invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: purpose, async nature, cost, key sourcing, reusability, and replay capability. It front-loads the core action and then layers constraints logically. There is no filler or repetition of schema fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 13 parameters, 5 required, no output schema, and a complex workflow (async, costing, multiple input sources), the description is remarkably complete. It explains how to obtain inputs, what the async flow looks like, cost scaling, and how to retrieve results (get_video). The only minor omission is a detailed return schema, but the description points to get_job and get_video, which is sufficient for an agent to proceed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 54%, so the description must compensate. It explains the provenance of characterR2Key and drivingVideoR2Key (linking to specific sibling tools), the cost implications of resolution and voiceChangeVoiceId, and the engine-specific constraints (pro requires sections ≤14s or keepOriginalVisualRanges). Not every parameter is detailed, but the most critical ones for correct usage are clarified, adding real value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-resource pair ('Swap a character image onto a driving video') and adds an optional refinement (restoring ranges). It clearly distinguishes from sibling tools like swap_upload (which uploads) and swap_import_character (which imports) by stating it consumes their outputs. The async and credit-callout further separate it from simple read tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit prerequisites: it names the exact source tools (swap_import_character or swap_upload for the character key, swap_upload for the driving video) and explains when to use each. It also notes the r2Key is reusable for batching. It does not explicitly state 'do not use if X', but the guidance is strong and actionable, so a 4 is warranted.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources