CraftStory MCP Server
OfficialSummary: CraftStory MCP server generates talking-avatar videos from a single photo plus audio (or description), via text-to-speech, job polling, and optional upscaling.
Discover options:
list_models(model status, limits, prices),list_voices(library + optional cloned voices),list_avatars(custom avatars and their scenes).Make speech:
create_audio_clipfrom text (≤2000 chars + voice) or by uploading a WAV/MP3/M4A recording.Estimate cost:
preview_costfor a CraftStory 2.0 job before creating it.Generate a CraftStory 2.0 video: any-length talking video from a photo (
image_url/image_path) or avatar scene, with resolution, gestures, lipsync mode, faceswap and motion hints.Generate a MiniMax H3 clip: up to 15 s from one photo, either description-driven (
basic) or audio-driven with lip-sync (reference), with optional reference files.Track jobs:
get_job_status(progress, failure reason, refund flag),wait_for_job(bounded polling up to 55 s, repeat as needed),get_job_result(full record + signed video URL valid 7 days).Upscale output:
upscale_videoto 1080p (CraftStory 2.0 720p source) or 2x (H3) as a new job.Guided flow: the
talking_video_from_photoprompt walks through script + photo end-to-end.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@CraftStory MCP Servermake a talking-avatar video of ~/photos/portrait.jpg saying hello"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
CraftStory MCP server
Generate talking-avatar videos from Claude, Cursor, Claude Code or any other MCP client, using the CraftStory API:
CraftStory 2.0 - a talking video of any length from one photo plus an audio clip (script + voice, your own recording, or a custom avatar). 8-15 minutes per video.
MiniMax H3 - a clip of up to 15 s from one photo: description-driven with generated sound, or audio-driven with lip-sync. 1-3 minutes.
You need Node.js 20 or newer, a CraftStory account on a plan with API access, and an API key (app: Account -> API Access, keys look like sk-cs-...). Generations are billed in credits exactly like in the app; failed jobs are refunded.
Config
Standard MCP configuration, identical for every client (Claude Desktop, Cursor, Windsurf, VS Code, Zed...):
{
"mcpServers": {
"craftstory": {
"command": "npx",
"args": ["-y", "@craftstory/mcp"],
"env": {
"CRAFTSTORY_API_KEY": "sk-cs-..."
}
}
}
}Related MCP server: EchoFM Studio
Install
Claude Code
claude mcp add --scope user craftstory -e CRAFTSTORY_API_KEY=sk-cs-... -- npx -y @craftstory/mcp(--scope user makes it available in every project; drop it to install for the current project only.)
Claude Desktop
Add to claude_desktop_config.json (Settings -> Developer -> Edit Config), then restart Claude Desktop; the server shows up under the tools icon:
{
"mcpServers": {
"craftstory": {
"command": "npx",
"args": ["-y", "@craftstory/mcp"],
"env": { "CRAFTSTORY_API_KEY": "sk-cs-..." }
}
}
}Cursor / other clients
Same shape in ~/.cursor/mcp.json (user-level, so the key never lands in a repository): command npx, args ["-y", "@craftstory/mcp"], env CRAFTSTORY_API_KEY. In a project-level .cursor/mcp.json prefer "CRAFTSTORY_API_KEY": "${env:CRAFTSTORY_API_KEY}" and keep the real key in your shell environment.
Environment variables: CRAFTSTORY_API_KEY (required), CRAFTSTORY_API_BASE (optional, default https://api.craftstory.com/api/v1; must be https unless CRAFTSTORY_ALLOW_HTTP=1).
Tools
list_models: Models, status (a paused model answers 503), limits and priceslist_voices: Library voices;include_clonedadds your cloned voiceslist_avatars: Your custom avatars, or the scenes of one avatarcreate_audio_clip: Speech from text + voice, or upload a local recordingpreview_cost: Credit estimate for a CraftStory 2.0 videocreate_craftstory2_video: Start a CraftStory 2.0 job (photo or avatar scene + audio clips)create_minimax_h3_video: Start a MiniMax H3 job (basic or reference mode)get_job_status: Status, percentage, failure reason, refund flagget_job_result: Full record with the signed video URL (valid 7 days)wait_for_job: Bounded polling (default 45 s, max 55 s); call again whilestateisrunningupscale_video: New job with the upscaled result (CraftStory 2.0 720p -> 1080p, H3 2x)
Plus the prompt talking_video_from_photo (script + photo) that walks the model through the whole flow.
Example
Make a portrait video of the person in
/Users/me/photos/portrait.jpgsaying "Welcome to our spring collection", calm gestures.
The assistant will: list_voices -> create_audio_clip -> wait_for_job(audio-clip) -> preview_cost -> create_craftstory2_video (resolution 720_1280, gestures calm) -> wait_for_job(craftstory-2) a few times -> get_job_result -> the video URL.
Long jobs: wait_for_job returns within timeout_s (default 45 s, max 55 s) plus a few seconds for the final result fetch; every request inside it is capped by the time left, so it stays under the 60 s tool-call limit of most clients. A CraftStory 2.0 video needs several calls; that is by design so agent runtimes do not time out. A CraftStory 2.0 video needs several calls; that is by design so agent runtimes do not time out.
Local files vs URLs
Photos accept image_url (JPG/PNG) or image_path (absolute path or ~/...; JPG/PNG/HEIC, up to 20 MB, uploaded as multipart). Recordings and extra references are local paths too. Anything you pass as a path is uploaded to the CraftStory API, so keep your client's tool-approval prompts on (see SECURITY.md).
Development
npm install
npm run build
CRAFTSTORY_API_KEY=sk-cs-... npm run smoke # live check over stdio: lists tools, makes an audio clip, previews cost
# add "-- --h3" and SMOKE_IMAGE=/path/to/photo.jpg to also render a 5 s MiniMax H3 clip (costs 17 credits)
npm testFull API reference: https://api.craftstory.com/api/v1/docs/public/ and the curl walkthrough at https://api.craftstory.com/api/v1/docs/samples/curl/.
Privacy Policy
This server runs on your machine and keeps no data of its own. It sends your API key and the inputs you pass to tools
(text, photos, audio, reference files) only to the CraftStory API at CRAFTSTORY_API_BASE (default
https://api.craftstory.com), where they are processed under the CraftStory privacy policy:
https://craftstory.com/privacy/. Generated videos and uploaded files are stored in your CraftStory account and can be
deleted there or via the API; nothing is retained locally by the server. No analytics or third-party services are
called by the server itself. Questions: support via https://craftstory.com/contacts/.
License
MIT
Available Tools
11 toolscreate_audio_clipCreate an audio clip (speech from text, or upload a recording)A
The soundtrack every video model takes as input. Either text (up to 2000 characters) plus exactly one voice (voice_id from list_voices, or voice_user_id for a cloned voice), or file_path to upload a local WAV/MP3/M4A recording. Returns the clip id; it is ready when wait_for_job(model='audio-clip') reports done (usually seconds). Longer scripts: create several clips and pass all ids to create_craftstory2_video in order.
| Name | Required | Description | Default |
|---|---|---|---|
| text | No | Script to synthesize (<= 2000 chars) | |
| voice_id | No | Library voice id (from list_voices) | |
| file_path | No | Local path of a recording to upload instead of text | |
| voice_user_id | No | Cloned voice id (from list_voices with include_cloned) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (mutating, open-world, non-idempotent), so the description's job is to add operational context, which it does: it discloses the async nature, that the returned id is only usable once wait_for_job(model='audio-clip') reports done, and the typical latency. It does not mention auth requirements or failure modes, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences with no filler; the input constraint is stated before the async/returns detail. Slightly information-heavy for one paragraph but each clause carries a distinct constraint or pointer, so little is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description supplies the missing return contract (clip id) and the completion protocol (poll wait_for_job with the audio-clip model). Together with the covered parameters and annotations, an agent has everything needed to call and consume this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, establishing a baseline of 3, but the description adds real semantic value beyond the schema: it states that exactly one voice must accompany text, that the voice sources are mutually exclusive alternatives to file_path, and where each id comes from. This mutual-exclusivity constraint is not expressed in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource and explicitly frames the two modes: synthesize speech from text with a voice, or upload a local recording via file_path. The opening line ('The soundtrack every video model takes as input') immediately situates it relative to the video-creation siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete selection rules between the two modes (text + exactly one voice, or file_path), names the prerequisite tools (voice_id from list_voices, voice_user_id for cloned voices), and routes longer scripts to multiple clips composed into create_craftstory2_video. Explicit when-to-use with named alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_craftstory2_videoCreate a CraftStory 2.0 talking video (photo + audio)A
Start a craftstory-2 generation: a photo of a person (image_url or image_path, or a custom avatar scene via scene_id) speaks the given audio clips with lip-sync, gestures and natural motion; any length. resolution is WIDTH_HEIGHT (480_832 / 720_1280 portrait, 832_480 / 1280_720 landscape); 1080p is available afterwards via upscale_video. Credits are charged on create (see preview_cost) and refunded if the job fails. Returns the job id and initial status; generation takes 8-15 minutes, so call wait_for_job(model='craftstory-2') repeatedly until it reports done, then get_job_result for the video URL.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Label, used as the download file name | |
| faceswap | No | Identity pass on the result (default true; off for custom avatars) | |
| gestures | No | How much the avatar moves (default normal) | |
| scene_id | No | Custom avatar scene id (from list_avatars) used instead of a photo | |
| avatar_id | No | Custom avatar id (from list_avatars); its trained model drives identity | |
| image_url | No | Public URL of the photo (JPG/PNG) | |
| image_path | No | Absolute local path of the photo to upload (JPG/PNG/HEIC, <= 20 MB); ~/ is expanded | |
| resolution | Yes | ||
| user_prompt | No | Optional motion / scene hint | |
| lipsync_mode | No | craftstory (default) / sync_so (alternative engine) / empty (no lip-sync) | |
| audio_clip_ids | Yes | Audio clip ids (from create_audio_clip), played in order |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, and the description adds important non-obvious behavior: credits are charged on create and refunded on failure, generation takes 8-15 minutes, and the agent must poll wait_for_job(model='craftstory-2') then fetch get_job_result. That workflow guidance is valuable and goes beyond the annotations, though it does not cover rate limits or concurrency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph that front-loads the core purpose, then enumerates key options and the post-create workflow. It is efficient with little filler, though the length and semicolon-heavy structure make it slightly dense rather than crisply front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a non-idempotent, credit-charging, multi-minute generation tool with no output schema, the description covers the essential lifecycle: input choices, resolution semantics, cost behavior, and the polling/result workflow. It is nearly complete, missing only explicit success/failure payload details, which are partly handled by the named get_job_result sibling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is already high at 91%, so the schema documents most parameters (image_url, image_path, scene_id, avatar_id, gestures, etc.). The description adds meaning by explaining that image_url or image_path provide the photo, scene_id is for a custom avatar scene, resolution is WIDTH_HEIGHT with named portrait/landscape values, and audio_clip_ids are played in order - useful framing that supplements the schema rather than repeating it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource ('Start a craftstory-2 generation') and immediately specifies the mechanism: a photo of a person speaks audio clips with lip-sync, gestures, and motion. This clearly distinguishes it from the sibling create_minimax_h3_video and create_audio_clip by naming the craftstory-2 model and the photo+audio inputs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context for when to use this tool (photo + audio clips + lip-sync) and points to alternatives for related steps (scene_id/custom avatars via list_avatars, upscale_video for 1080p, preview_cost for cost preview). However, it does not explicitly state when NOT to use it or directly compare against the sibling video generator create_minimax_h3_video.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_minimax_h3_videoCreate a MiniMax H3 clip (up to 15 s)A
Start a minimax-h3 generation from one photo. mode='basic': user_prompt (scene description) + requested_duration_s (5-15); the model animates the photo and generates the soundtrack itself. mode='reference': one audio_clip_id drives the clip with lip-sync (first 15 s billed); user_prompt is optional; up to 8 extra image / 3 video / 2 audio reference_files with reference_captions keep a product or background consistent. Output is 768 px on the short side, orientation follows the photo. Cost 3.3 credits per billed second, charged on create. Returns the job id; call wait_for_job(model='minimax-h3') until done (1-3 min).
| Name | Required | Description | Default |
|---|---|---|---|
| mode | Yes | ||
| name | No | ||
| image_url | No | ||
| image_path | No | ||
| user_prompt | No | Scene / motion description (required in basic mode) | |
| audio_clip_id | No | Reference mode: the clip that drives the video | |
| reference_files | No | Reference mode: local paths of extra reference images/videos/audio | |
| reference_captions | No | One caption per reference file, same order | |
| requested_duration_s | No | Clip length in basic mode (default 8) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (not read-only, open-world, non-idempotent, non-destructive). The description adds valuable context beyond that: cost (3.3 credits/billed second, charged on create), output resolution (768 px short side), orientation behavior, and the async flow with wait_for_job. It could disclose more about failure/retry but is unusually transparent for a create tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then mode-specific guidance, then cost/output/async details in a compact flow. It's dense but every clause earns its place; only slightly overloaded with parenthetical constraints.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 9 params, 56% schema coverage, no output schema, and an async job flow, the description supplies the missing operational details (cost, output size, job id + wait_for_job poll). It is essentially complete for correct invocation, though the image input fields are lightly documented.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 56%, so the description must carry weight, and it does: it explains that user_prompt is the scene description and required in basic mode, that audio_clip_id drives the clip in reference mode, the limits on reference_files (up to 8 image / 3 video / 2 audio), and reference_captions ordering. Minor gaps remain for name/image_url/image_path/requested_duration defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource ('Start a minimax-h3 generation from one photo') plus a clear distinction between the two modes, exactly what differentiates it from sibling video tools like create_craftstory2_video. The mode semantics (basic vs reference) make the tool's behavior unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly describes when to use each mode ('basic': user_prompt + duration; 'reference': audio_clip_id for lip-sync), and notes that user_prompt is optional in reference mode. It doesn't route to sibling tools like create_craftstory2_video or upscale_video, but the mode conditions are explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_job_resultGet a finished job (video URL and details)ARead-only
Full record of a job. For craftstory-2 the video is in video, for minimax-h3 in video_url, for audio clips in file; all are signed URLs valid for 7 days (call again for a fresh link). Also returns the parameters used and the credits charged.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| model | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint/openWorldHint annotations, the description discloses real operational behavior: signed URLs expire after 7 days and must be re-fetched, and the response includes input parameters and credits charged. It stops short of saying what happens when the job is not yet finished.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense but well-structured sentence: identity first, then the model-to-field mapping, then URL expiry behavior, then extra returned data. Every clause earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so reasonably by naming the model-specific URL fields and the parameters/credits returned. It omits other record fields and the failure mode for an unfinished job, so it is solid but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates by mapping each model enum value to where its output lands (video, video_url, file), which is meaningful semantics the schema lacks. The id parameter is left to the schema's uuid format with no added meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb+resource ('Full record of a job') and the title adds the key disambiguator 'finished job'. However, the body text alone does not distinguish it from close siblings like get_job_status or wait_for_job, leaving that routing to the title.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance or mention of alternatives such as get_job_status or wait_for_job. The only implied condition is that the job must be finished, carried by the title rather than the description text.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_job_statusGet a job's statusARead-only
Status of a video job or audio clip: status, status_percentage, status_failed, credits_refunded. Terminal states: done; failed*, rejected_* and not_pass_moderation (audio) are failures. Prefer wait_for_job, which polls for you.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| model | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover safety (readOnly/openWorld), and the description earns credit by disclosing terminal-state semantics: done is success, while failed*, rejected_* and not_pass_moderation are failures, plus the credits_refunded signal. It doesn't say how long results are retained or what a non-terminal percentage means.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Compact and front-loaded: what it returns first, then failure semantics, then the recommended alternative. The telegraphic style ('failed*', 'rejected_*') is dense but efficient; the asterisk notation is mildly ambiguous.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, listing the return fields and failure states is genuinely necessary and present. However, for a 2-parameter tool with 0% schema coverage, the complete silence on how to supply model/id leaves a real gap for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the full burden for model and id, yet neither is mentioned. The phrase 'a video job or audio clip' only loosely hints at the model enum values without explaining that model must match the original creation call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb+resource and enumerates the returned status fields (status, status_percentage, status_failed, credits_refunded), which clearly separates it from a result-fetching tool. It stops short of explicitly distinguishing itself from the sibling get_job_result, so an agent still has to infer that division of labor.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names an alternative and the condition that selects it: prefer wait_for_job because it polls for you. It gives no when-not guidance (e.g., use this instead for a single non-blocking check) and doesn't mention get_job_result.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_avatarsList custom avatars (and their scenes)ARead-only
Custom avatars trained in the CraftStory app that craftstory-2 can generate with (pass an id as avatar_id). Each avatar may carry a default voice {id, voice_kind}: voice_kind 'user' means send it as voice_user_id, 'library' as voice_id in create_audio_clip. Pass avatar_id to list that avatar's scenes; a scene id can replace the photo (scene_id) in create_craftstory2_video.
| Name | Required | Description | Default |
|---|---|---|---|
| avatar_id | No | Return the scenes of this avatar instead of the avatar list |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/openWorldHint, so safety is covered. The description adds genuine behavioral context the annotations cannot: the voice_kind-to-parameter mapping and the fact that supplying avatar_id switches the tool into scene-listing mode. It omits return shape/pagination, which for an unpaginated list tool is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with what the resource is, then the actionable id hand-offs. Dense but every clause carries routing information; slightly heavy on nested parameter names, though none are removable without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must carry the return semantics, and it does: avatars, their optional default voice object, and scenes when avatar_id is given. Missing only pagination/format detail, which is minor for a simple list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already documents avatar_id, so the baseline is 3. The description goes beyond it by explaining the mode-switch consequence and how the returned id feeds create_audio_clip and create_craftstory2_video, adding downstream meaning the schema lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (custom avatars trained in CraftStory) and the sibling model it serves (craftstory-2), plus the dual listing mode (avatars vs. an avatar's scenes). An agent can distinguish it from list_models and list_voices without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains how to consume the output — pass the id as avatar_id, route voice_kind 'user' to voice_user_id and 'library' to voice_id in create_audio_clip, and use scene_id in create_craftstory2_video. It does not explicitly say when to prefer this over list_voices/list_models, so there is no exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_modelsList CraftStory video modelsARead-only
Catalogue of the video models behind this server with their status, modes, limits and credit prices. Two models today: craftstory-2 (a talking video of any length from one photo plus an audio clip; 8-15 min) and minimax-h3 (a clip of up to 15 s from one photo, either description-driven with generated sound or audio-driven with lip-sync; 1-3 min). Call this first when unsure which model fits, or to check that a model is not paused.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint and openWorldHint; the description adds substantive behavior by disclosing the return contents (status, modes, limits, credit prices) and the pause/pricing state that a follow-up call must respect. It does not cover auth requirements or whether the model roster can change, but for a zero-parameter read-only listing this is substantial added value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the catalogue purpose, then the per-model breakdown, then the usage cue. Dense but each clause carries information an agent uses for model selection; the parenthetical model details are somewhat heavy but not redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must convey what a caller learns. It does so by naming the models and the dimensions returned (status, modes, limits, prices) and by telling the agent when the result matters (paused checks, model fit). Nothing needed to invoke a zero-param listing correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so the baseline is 4. The description adds no parameter guidance, and none is needed; the schema is trivially complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (the catalogue of video models on this server) and enumerates exactly what it exposes: status, modes, limits and credit prices. It even names the two models with their capabilities and durations, so an agent can distinguish this discovery tool from execution siblings like create_craftstory2_video without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to call this first when unsure which model fits, or to verify a model is not paused, which is clear actionable context. It stops short of naming the alternative execution tools to route to after resolving the choice, so the when/when-not framing is present but not fully closed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_voicesList voices for text-to-speechARead-only
Library voices (id, name, language, gender) usable as voice_id in create_audio_clip. With include_cloned=true also returns the account's own cloned voices, usable as voice_user_id. Voices are cloned in the CraftStory app, not via the API.
| Name | Required | Description | Default |
|---|---|---|---|
| include_cloned | No | Also return the account's cloned voices (default false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so safety is covered. The description adds genuinely useful non-obvious context: cloned voices resolve to voice_user_id rather than voice_id, and cloning happens in the CraftStory app, not via the API. It omits pagination/ordering/size limits, keeping it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with what the tool returns, then the parameter branch, then the constraint about where cloning occurs. No filler; every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description carries the return-value burden, and it does so by naming the fields and the id mapping. With one optional param fully documented and annotations covering the safety profile, nothing material is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and there is only one optional boolean, so the baseline is 3. The description goes beyond the schema's 'also return the account's cloned voices' by explaining that those voices are used as voice_user_id, adding real semantic value about the parameter's effect.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (list voices) and enumerates the returned fields (id, name, language, gender), making it clearly distinct from siblings like list_models and list_avatars. It also pins down the exact downstream parameter each voice type feeds, which is unusually precise.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the condition that selects the cloned-voice branch (include_cloned=true) and cross-references create_audio_clip as the consumer. It lacks an explicit 'when not to use this' or a named alternative among siblings, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preview_costEstimate the credit cost of a CraftStory 2.0 videoARead-only
Credits a craftstory-2 job would cost for the given audio clips and settings, without creating anything. Rate per second of audio: 480p 2.2 (2 with lipsync_mode=empty), 720p 3.3 (3 with empty); rounded up per job. MiniMax H3 is a flat 3.3 credits per billed second.
| Name | Required | Description | Default |
|---|---|---|---|
| resolution | Yes | ||
| lipsync_mode | No | ||
| audio_clip_ids | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint/openWorldHint; the description goes beyond them by confirming no job is created and disclosing the actual billing behavior: per-second rates by resolution, the lipsync_mode=empty discount, rounding up per job, and the flat MiniMax H3 rate. That is real behavioral disclosure, though it doesn't say whether costs are estimated tolerance or exact.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then the rate table in two dense sentences. Telegraphic phrasing ('2.2 (2 with lipsync_mode=empty)') is compact but slightly cryptic, and the MiniMax sentence sits oddly in a CraftStory-scoped definition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does identify the return unit (credits, rounded up per job), which is the key completeness need for a cost-estimate tool. However it omits how audio_clip_ids drive duration, whether errors occur on invalid clips, and leaves the enum values unresolved, so an agent can call it but not fully predict inputs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry parameter meaning. It partially compensates by mapping rate tiers to 480p/720p and describing the effect of lipsync_mode=empty, but audio_clip_ids is never explained (duration source? billing per clip?), and the enum values 480_832/832_480 and sync_so/craftstory are not mapped to their human-readable meanings.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Credits a craftstory-2 job would cost') and explicitly scopes it as a dry run ('without creating anything'), which cleanly separates it from siblings like create_craftstory2_video and create_minimax_h3_video.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by 'without creating anything' — an agent can infer this is a pre-flight estimate before calling a create tool. There is no explicit when-to-use/when-not, no stated alternatives, and it oddly mixes MiniMax H3 pricing into a tool whose title scopes it to CraftStory 2.0.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upscale_videoUpscale a finished video (new job)A
Creates a NEW job with the upscaled result; the original stays. craftstory-2: only 720p sources, resolution 1080_1920 (from 720_1280) or 1920_1080 (from 1280_720). minimax-h3: always 2x, no resolution needed. Costs 0.2 credits per second. Poll the returned id with wait_for_job.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| model | Yes | ||
| resolution | No | craftstory-2 only |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare non-readonly, non-idempotent, non-destructive, open-world. The description adds real value beyond them: the original is preserved, cost is 0.2 credits/second, and the job must be polled via wait_for_job. It doesn't cover failure/auth behavior, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense, front-loaded, telegram style with no filler; the key scoping fact leads and the polling instruction closes. Slightly clipped phrasing costs a point on readability but every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description covers the essential return behavior ('poll the returned id with wait_for_job') plus cost and model constraints. Adequate for the tool's complexity, with only the `id` parameter's origin left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%, so the description must compensate, and it does: it explains that minimax-h3 needs no resolution, that craftstory-2 only accepts 720p sources, and maps the resolution enum values to source resolutions. The meaning of the `id` parameter is only implied ('the original stays').
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: creating a NEW upscale job while the original stays, which cleanly separates it from the sibling create_*_video tools. It names no sibling explicitly, but the 'new job / original stays' scoping makes the intent unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives per-model selection guidance (craftstory-2 source/resolution constraints, minimax-h3 always 2x) and tells the agent to poll with wait_for_job. It stops short of explicit when-not-to-use routing versus the create_*_video siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_for_jobWait for a job (bounded polling)ARead-only
Polls a job's status every few seconds for up to timeout_s (default 45, max 55 - most MCP clients cut a tool call at 60 s) and returns as soon as it is terminal. If it returns state='running', call it again - craftstory-2 jobs take 8-15 minutes, minimax-h3 1-3 minutes, audio clips seconds. Reports progress notifications when the client supports them.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| model | Yes | ||
| timeout_s | No | How long this call may wait (default 45, max 55) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint and openWorldHint; the description adds real behavioral detail beyond them, including the polling cadence, the early-return-on-terminal semantics, the default/max timeout, the reason for the 55 s cap (clients cut calls at 60 s), progress-notification support, and typical run durations. This is exactly the kind of context an agent needs to avoid misusing a long-running poll.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, front-loaded with the polling behavior and return condition, then the re-invocation rule and the duration table. Every clause carries operational information; nothing is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and only partial parameter documentation, the description carries the return-contract burden and does it well by describing the 'running' vs terminal outcome and progress notifications. It stops short of naming the terminal states or explaining what the agent should do with a terminal result, which would fully close the loop.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%, so the description must compensate, and it does: it explains timeout_s's default and max and the client-cutoff rationale behind them, and it attaches expected durations to each model enum value, giving the enum practical meaning. The required id parameter still has no description anywhere (only format=uuid), which keeps this below a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: polls a job's status and returns as soon as it reaches a terminal state, with the bounded-timeout behavior named up front. It implicitly separates itself from the one-shot get_job_status by describing the re-invocation loop, but it never names that sibling explicitly, so an agent must infer the distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete when-to-call guidance: re-invoke whenever the result is state='running', and it supplies per-model duration expectations (craftstory-2 8-15 min, minimax-h3 1-3 min, audio clips seconds) so the agent can budget calls. It does not explicitly contrast this with get_job_status or state when polling is unnecessary, so the alternative selection is left partly to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
11 tool updates
v0.1.2- First observed
create_audio_clip - First observed
create_craftstory2_video - First observed
create_minimax_h3_video - First observed
get_job_result - First observed
get_job_status - First observed
list_avatars - First observed
list_models - First observed
list_voices - First observed
preview_cost - First observed
upscale_video - First observed
wait_for_job
TDQS
Scored across 11 tools
Most tools have clearly distinct roles, and the two video-creation tools are unambiguously split by model (craftstory-2 vs minimax-h3). The trio get_job_status / wait_for_job / get_job_result overlaps somewhat, but descriptions explicitly differentiate single-check, polling, and full-record use cases.
Consistent snake_case verb_noun pattern throughout (list_models, create_audio_clip, get_job_status, wait_for_job, upscale_video). Model-specific names like create_craftstory2_video follow the same convention and clearly encode the target model.
11 tools is well-scoped for a media-generation service, with each tool earning its place across discovery, creation, polling, and post-processing. No redundant or filler tools.
Covers the full lifecycle: model/voice/avatar discovery, audio creation, cost preview, generation for both models, status/result polling, and upscaling. Minor gaps like job cancellation, job listing, or asset deletion are absent but not blocking for core workflows.
Maintenance
Related MCP Connectors
Plan, compare, price, generate, and recover AI video from compatible MCP clients.
Create and manage AI image and video generations through Quriov's fixed public MCP tools.
Multi-model AI image and video generator. 14 models behind one OAuth-secured MCP endpoint.
Generate images, video, audio and short films with 140+ AI models from any MCP client.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenance🎬 Enterprise-grade MCP Server for Creatify AI - 12 tools for AI video generation: avatar videos, URL-to-video, AI shorts, custom avatars, script generation, advanced lip-sync with emotion control. Complete API coverage with semantic versioning.30 npm23MIT
- FlicenseNot gradedqualityCmaintenanceEnables AI assistants to manage audio-story series: create, edit, narrate, publish, and generate marketing content like UGC videos and reels through a secure MCP endpoint.-
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to generate professional storyboards and videos from scripts or creative descriptions via MCP-compatible clients.21 npmMIT
- FlicenseNot gradedqualityBmaintenanceEnables AI clients like Claude and ChatGPT to generate images and videos, animate images, create lip-synced videos, list TTS voices, and manage media via remote MCP tools.-