video_character_swap
视频换脸
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| image_url | No | 目标人脸参考图 URL | |
| video_url | No | ||
| callback_id | No | 自定义追踪 ID,webhook URL 后台配置 | |
| workspace_id | No |
视频换脸
| Name | Required | Description | Default |
|---|---|---|---|
| image_url | No | 目标人脸参考图 URL | |
| video_url | No | ||
| callback_id | No | 自定义追踪 ID,webhook URL 后台配置 | |
| workspace_id | No |
Changes observed during successful MCP inspections.
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare a non-read-only, open-world, non-idempotent task, but the description adds nothing: no mention of asynchronous job submission, callback/webhook delivery implied by callback_id, processing time, or failure modes. For a generation tool this is a complete disclosure gap, with no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four characters is not conciseness but under-specification; there is no front-loaded context or structure to evaluate, and the schema's duplicate "视频换脸任务。" adds nothing either.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter, open-world, non-idempotent video generation task with no output schema, the agent needs to know input formats, async/callback behavior, and constraints. None of this is present, so the definition is inadequate to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50%: image_url and callback_id are documented in the schema, while video_url and workspace_id carry no description anywhere. The description supplies no parameter meaning at all, so it fails to compensate for the uncovered half.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"视频换脸" is a direct restatement of the tool name video_character_swap rather than an independent explanation of what the tool does. It conveys the resource (face swap on video) but adds no scope, verb nuance, or sibling differentiation against the many other video_* tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use, when-not-to-use, or alternative guidance. Nothing tells an agent how this differs from video_edit, video_magic_eraser, or video_generate, so selection must be inferred from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Add one secure layer between your agents and this server.