dora-mcp
Server Details
Multi-model AI image and video generator. 14 models behind one OAuth-secured MCP endpoint.
- Status
- Healthy
- Uptime
- 55.1% over 43 days
- OAuth
- Works in Glama
- Last Tested
- Transport
- Streamable HTTP · MCP 2024-11-05
- URL
- Repository
- Saga-Labs/dora-mcp
- GitHub Stars
- 0
- Server Listing
- Dora Video MCP
TDQS
Scored across 4 tools
The four tools have clearly distinct purposes: two generation tools separated by media type (image vs video), a polling tool for job status, and a model listing tool. There is no overlap or confusing boundary between any of them.
All tool names follow a consistent snake_case verb_noun pattern: check_job, generate_image, generate_video, list_models. This makes the set predictable and easy to navigate.
Four tools are well-scoped for a generation service. Each tool earns its place, and complex options are handled through parameters rather than additional tools, avoiding bloat.
The surface covers the core lifecycle for both image and video generation: generate, poll for status, and list available models. Minor gaps exist, such as job cancellation or history retrieval, but these are not essential for typical workflows.
Available Tools
4 toolscheck_jobAInspect
Check the status of a generation job. Returns status (pending|done|failed) and result_url when done. Wait at least 30 seconds between polls — generations take 1-5 minutes typically.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden, and it does disclose the poll interval (30s) and typical job duration (1-5 minutes), which are genuinely useful behavioral constraints not derivable from the schema. It omits failure semantics (what 'failed' means, retry guidance), but the core timing behavior is well-communicated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the purpose, followed by the return shape and the key timing constraint. Zero wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-param polling tool with no annotations or output schema, the description covers purpose, return values (status enum and result_url), and timing behavior. Missing only the failure/retry path and where job_id originates, minor for a tool of this simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter (job_id) with 0% schema description coverage. The description doesn't explain what a job_id is or where it comes from (presumably returned by generate_image/generate_video). Schema coverage is low, but with a single obvious parameter the baseline of 3 is reasonable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('check the status of a generation job') and distinguishes itself from siblings like generate_image/generate_video by being a polling/status tool rather than a generation tool. An agent can immediately tell this is not a generator.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives implicit operational context ('Wait at least 30 seconds between polls') that guides invocation, and implies this is used after starting a generation. It does not explicitly name the sibling tool that starts a job, but the polling guidance is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_imageAInspect
Generate an image. Default model: nano_banana_2 (good quality, ~10 coins, supports reference photo). Use z_image (1 coin) only if the user wants the cheapest possible result. Use nano_banana_pro (18-24) only when the user explicitly asks for face preservation or maximum quality. Default resolution 1K — never use 2K/4K unless the user explicitly asks. If aspect ratio matters and the user did not say, ASK whether they want 9:16 (vertical), 16:9 (horizontal), or 1:1 (square). For multi-character or character-plus-outfit shots, pass extra anchor images in reference_image_urls — Nano Banana 2 / Pro accept up to 7 inputs. Returns job_id; poll check_job until status=done.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | ||
| prompt | Yes | ||
| resolution | No | 1K cheapest; 4K up to 50% more. Default 1K. | |
| aspect_ratio | No | ||
| reference_image_url | No | Primary reference image. Optional. For multi-ref pass reference_image_urls instead or in addition. | |
| reference_image_urls | No | Extra reference images (Nano Banana 2 / Pro accept up to 7 total including reference_image_url). Use for multi-character, face + outfit, style references. Order matters. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses per-model cost (~10 coins, 1 coin, 18-24), the default resolution, the 7-input reference limit, and that the tool is asynchronous (returns job_id; poll check_job until done). These are the behavioral traits an agent cannot infer from the schema alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Despite its length, every sentence is actionable guidance with no filler, and the most important constraint (default model and default resolution) is front-loaded. Sentence-level imperative style keeps it scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description fills that hole by stating the tool returns a job_id and must be polled via check_job. Combined with cost, defaults, limits, and refusal-to-guess guidance on aspect ratio, an agent has everything needed to invoke this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%, so the description usefully compensates by explaining model trade-offs, the resolution default, the ask-first policy for aspect_ratio, and the semantics of reference_image_urls (multi-character, up to 7 total, order matters). It leaves four enum models (google_nano_banana, google_nano_banana_edit, gpt4o_image, gpt_image_2) completely unmentioned, which is a real gap given the 7-value enum.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific verb+resource ('Generate an image') and the rest makes the operation unmistakable. It also names the related sibling 'check_job' as the polling counterpart, so an agent can distinguish this from generate_video and check_job without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing rules for every meaningful choice: default model nano_banana_2, z_image only for cheapest, nano_banana_pro only for face preservation/max quality, and never 2K/4K unless asked. It also states the condition under which the agent should stop and ask about aspect ratio, which is exactly the when/when-not guidance this dimension rewards.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_videoAInspect
Generate a video. Default model: seedance_1_5_pro at 480p, 8s, no audio (cheapest reasonable quality). Before calling, if the user has not specified them, briefly ASK for: (1) aspect ratio (9:16 vertical, 16:9 horizontal, 1:1 square), (2) duration (4/8/12s), (3) whether they want audio, (4) resolution (480p/720p/1080p) — higher costs more. Skip the questions only when the user already gave you those details or asked for "the cheapest". For character consistency across shots (e.g. a recurring person), pass extra face/body anchor photos in reference_image_urls — Seedance 2 / 2-fast use up to 7 references. MOTION CONTROL: model kling_2_6_motion transfers the motion of a driving video onto the person in your image — pass the person photo in image_url AND the driving clip URL in motion_video_url (both required); set duration to the driving clip length (billed ~11 coins/s, min 3s). Returns job_id; poll check_job until status=done.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | ||
| prompt | Yes | ||
| duration | No | ||
| image_url | No | First-frame / primary input image for image-to-video. Optional. For multi-ref, also pass reference_image_urls. | |
| resolution | No | 480p cheapest; 1080p ~5x cost. Default 480p. | |
| with_audio | No | Generate sound. SeeDance/Kling support audio — typically doubles cost. Default false. | |
| aspect_ratio | No | ||
| motion_video_url | No | Driving/reference VIDEO URL for Motion Control (model kling_2_6_motion only). The motion of this clip is applied to the person in image_url. Required for kling_2_6_motion; ignored by other models. Must be a publicly reachable mp4/mov. | |
| reference_image_urls | No | Extra reference images (Seedance 2 / 2-fast accept up to 7 total including image_url). Use for character face/body anchors, multiple people, style references. Order matters — most important reference first. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it does most of it: it states default model/resolution/duration/audio, cost multipliers for audio and resolution, a concrete billing rate for motion control (~11 coins/s, min 3s), the required parameter pairing for kling_2_6_motion, reference-count limits per model, and the async contract (returns job_id, poll check_job until status=done). It omits auth requirements, failure/retry behavior, and quota limits, so it falls short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The default configuration is front-loaded, which is the right priority, and most sentences carry actionable content (costs, required pairs, polling). It is dense and long, with a few parenthetical asides that could be trimmed, but little is pure filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter async generation tool with no annotations and no output schema, the description covers the essentials an agent needs: defaults for unspecified inputs, the clarifying-question protocol, model-specific parameter requirements, cost signals, and the return/polling contract. It stops short of error handling, authorization, and output shape beyond job_id, so a 5 is not warranted.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 56%, so the description must compensate, and it does add real meaning: it names the default model, explains that duration options are 4/8/12s, that aspect ratio maps to vertical/horizontal/square, that audio roughly doubles cost, and that resolution tiers scale cost to 1080p at ~5x. It does not explain the remaining enum members (grok_*, veo_3, kling_2_6) or the full duration enum, leaving some gaps against the 56% coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource ("Generate a video") and immediately scopes it with defaults and model family. It clearly separates this from generate_image and check_job, the latter named explicitly as the follow-up poll tool. An agent knows exactly what this tool produces.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-do-what rules: ask for the four unspecified options before calling, skip the questions if the user already supplied them or asked for "the cheapest", and use kling_2_6_motion only under a specific stated condition. Alternatives and their trigger conditions are spelled out rather than implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_modelsAInspect
List Dora generation models, cheapest-first, with per-call coin cost. Call this when the user asks what models are available, when cost matters, or when you want to verify a model name. Defaults: nano_banana_2 for images, seedance_1_5_pro for videos — only step up if the user explicitly asks for higher quality, face preservation, or audio.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it discloses return ordering (cheapest-first), the presence of per-call coin cost, and the default-model policy for each media type — meaningful behavioral context. It still doesn't state whether the list is paginated, cached, or requires auth, so it falls just short of fully self-sufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, all front-loaded: what it returns, when to call it, and the defaults/escalation rule. Every clause carries decision-relevant information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no parameters, the description must convey the shape of the result, which it does (model list, cost ordering, coin cost). The main remaining gap is that it never states the response format or whether the catalog is exhaustive/stable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is no parameter surface for the description to explain; the baseline of 4 applies. The description correctly wastes no space on input semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List Dora generation models') and adds distinguishing traits the siblings don't have: cheapest-first ordering and per-call coin cost. An agent can immediately tell this discovery tool apart from generate_image and generate_video.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly enumerates the three triggering situations (user asks what's available, cost matters, verifying a model name) and then gives actionable routing advice: default to nano_banana_2 for images and seedance_1_5_pro for videos, stepping up only on explicit user requests for higher quality, face preservation, or audio.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
- First observed
check_job - First observed
generate_image - First observed
generate_video - First observed
list_models
Related MCP Connectors
Image, video, music and text generation across 100+ models through one endpoint.
Generate images with any major model — one API key, one prepaid balance, one MCP.
Create images & video from any MCP agent — 17 models, spend limits, one URL.
Best Image and video generation: 20+ models (Kling, Seedance, Veo, NB, FLUX.2), OAuth, pay-per-use.
Related MCP Servers
AlicenseNot gradedqualityDmaintenanceEnables AI agents and developers to generate images, videos, audio, and text using 100+ models via MCP or REST with a single API key.4MIT- AlicenseAqualityCmaintenanceHosted multi-model AI media + chat MCP server. Generates images, video, audio, face-swaps and talking-avatars, and chats across 300+ models (Claude, GPT, Gemini, DeepSeek…) - all from one balance and one API key.16MIT
- AlicenseAqualityCmaintenanceMulti-provider AI video, speech, music, and transcription MCP server enabling video generation, image-to-video, TTS, music creation, and speech-to-text via a unified interface.3MIT
- AlicenseAqualityDmaintenanceMCP server for multi-provider AI image generation (AWS Bedrock, OpenAI, Google Gemini) enabling image generation, transformation, and editing through a unified interface.41MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.