@vidofy/mcp
OfficialAn MCP server that lets you browse and run Vidofy media generation — images, video, audio, speech, lipsync, effects and more — billed to your own Vidofy account.
List what Vidofy can generate (
list_modes): text-to-image, image-to-video, lipsync, text-to-speech and other modes.Browse models (
list_models,get_model): see available models per mode, their credit costs, durations, and full input contracts / file limits.Estimate cost before spending (
estimate_cost): check what a planned generation will charge, since prices can vary greatly with settings.Generate media (
generate): the only tool that spends your balance — submit a generation, optionally make output public, and pass file paths or reuse earlier generation IDs.Track progress (
get_status): poll a running generation and get how long to wait before checking again.Retrieve results (
get_result): get finished media as an image/preview plus the link (private links expire; public links are permanent).Check account (
get_balance,get_usage): see coin balance, expiring credits, and recent generation costs/spending.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@@vidofy/mcpGenerate a 5-second video of a cat surfing using Kling 3.0"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
@vidofy/mcp
MCP server for Vidofy — generate images, video, audio and speech from Claude, ChatGPT, VS Code, Gemini CLI, Cursor, Hermes, or any MCP client, billed to your own Vidofy account, at the same prices the website charges.
Over 570 models, including Veo 3.1, Kling 3.0, Flux 2, Seedance 2.5, Wan 2.7, Hailuo 2.3, Runway, Luma Ray 2, Qwen Image 3.0, Vidu Q3 and LTX 2 — text-to-video, image-to-video, text-to-image, image editing, video and photo effects, lipsync, text-to-speech and voice cloning. The agent browses the catalogue, prices a generation before running it, and follows one to its result.
Status: first public release.
This server is for personal Vidofy accounts. It takes one credential,
VIDOFY_TOKEN, and spends your own coins — the same balance the website spends, at the same prices. There is no other billing mode:VIDOFY_API_KEYis refused at startup, and that is a decision, not a feature waiting on a release.
Tools
Tool | What it does | Spends |
| What Vidofy can generate: text-to-image, image-to-video, lipsync, speech… | no |
| The models in one mode, with each one's credit cost and rough duration | no |
| One model's full input contract: a JSON Schema, its file slots and their limits | no |
| What a generation will cost, before running it | no |
| Runs it. The only tool that spends the balance. | yes |
| Whether a generation has finished | no |
| The finished media | no |
| Coins left, and how many expire with the subscription | no |
| Recent generations and what they cost | no |
The usual order is list_modes → list_models → get_model → estimate_cost → generate
→ get_status → get_result.
generate is the only tool without readOnlyHint, which is what tells a client to ask the user
before running it. It charges at submit, not on success, and returns immediately with an id —
a generation takes from ~30 seconds to several minutes, so the agent polls get_status rather
than holding the call open. Output is private by default; pass public: true only when the
user asked for a permanent public link.
File inputs take a path on the machine running the server. The package reads the user's own file and streams it with the submit — it never makes a temporary copy — and checks the extension and size against that model's own limits first, so a file the server would reject never leaves the disk.
Not exposed, deliberately: checkout, auto top-up, purchases, referrals, the daily reward. Nothing in this package can buy coins or change a plan, however it is prompted.
Related MCP server: media-gen-mcp
Setup
There are two ways in. Take the first one unless your client cannot do it.
1. Remote connector — nothing to install
Give your client this URL:
https://vidofy.ai/mcp-appYou sign in in your browser and approve once. No token to copy, nothing to keep in a config file, and nothing to update when this package changes.
Client | How |
Claude.ai · Claude Desktop | Settings → Connectors → Add custom connector → paste the URL → Connect, then approve the sign-in. They share one list: add it in either and it appears in both. Available on every plan, including Free — where you get one connector. |
ChatGPT | Settings → Connectors → add a custom connector (no such option? turn on Developer Mode in Settings first) → paste the URL → Connect, then approve. On a Business or Enterprise workspace an administrator adds it for everyone. |
VS Code | Add an MCP server of type |
Claude Code · Codex · Cursor | Each accepts a remote MCP server URL. Follow that client's own MCP documentation and give it the URL above. |
Then ask it: "list Vidofy modes" to confirm the connection, and "make me a 5-second clip of a red bicycle" — it prices the generation before running it.
Some clients cannot take this route, and it is worth knowing why before you try. Signing in
here needs a client that identifies itself with a published metadata document — an https URL
the authorization server fetches. Clients that instead expect to register themselves at a
registration_endpoint, or to be handed a client_id and client_secret you created by hand,
have nothing to work with: this server issues neither. Gemini CLI and Hermes are both in
that group today. Take route 2 — it is not a lesser path, just a different way of proving who
you are.
2. Local stdio server — works with every client here
Create a personal MCP token at vidofy.ai → Studio → Account → MCP Access. It is shown once.
// claude_desktop_config.json (Cursor: .cursor/mcp.json — same shape)
{
"mcpServers": {
"vidofy": {
"command": "npx",
"args": ["-y", "@vidofy/mcp"],
"env": { "VIDOFY_TOKEN": "vmt_..." }
}
}
}npx fetches it on first run. Prefer a pinned copy? npm i -g @vidofy/mcp, then:
{ "command": "vidofy-mcp", "env": { "VIDOFY_TOKEN": "vmt_..." } }The entry is the same everywhere — the file it goes in, and what the outer key is called, are not. Check yours against your client's own documentation before you paste:
Client | File | Outer key |
Claude Desktop |
|
|
Cursor |
|
|
Gemini CLI |
|
|
VS Code |
|
|
Hermes |
|
|
So VS Code wants:
// .vscode/mcp.json
{
"servers": {
"vidofy": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@vidofy/mcp"],
"env": { "VIDOFY_TOKEN": "vmt_..." }
}
}
}and Hermes wants the same thing in YAML:
mcp_servers:
vidofy:
command: npx
args: ["-y", "@vidofy/mcp"]
env:
VIDOFY_TOKEN: "vmt_..."Both paths reach the same account, the same models and the same balance. The difference is only where the process runs and how you prove who you are.
One client shows more than the others. A generation normally comes back as text the model
reads out. A host that supports MCP Apps gets a live card instead — the picture or clip
itself, its progress while it runs, and a download button — and the server offers it only to a
host that says it can render one. Claude and VS Code both do (in VS Code, turn on
chat.mcp.apps.enabled). Everywhere else the same result arrives as text and, for an image,
an inline picture. Nothing is missing; it is just quieter.
Environment
Variable | Required | What it does |
| yes | Personal MCP token ( |
| no | Override the origin the server talks to — an origin only, no path. Defaults to |
VIDOFY_API_KEY is recognised only in order to be refused: a vky_… key bills a
different balance, which this server does not serve. Setting it stops startup with a message
naming the token to use instead — and setting both is refused too, since the two bill
different balances and no precedence rule is worth having to remember.
Development
npm install
npm run build
npm run inspect # MCP Inspector — spends nothingVIDOFY_API_BASE points it at a different origin, if you are running one.
Nothing here writes to stdout. With stdio transport, stdout is the protocol channel — a
single stray console.log() puts a non-JSON line in the stream and the client drops the
connection with an error that explains nothing. Diagnostics go to stderr via the log() helper
in src/index.ts.
Layout
src/config.ts credential + mode + base URL, validated at startup
src/backend.ts the only place that talks HTTP: auth, retries, multipart, errors
src/schema.ts one model's m_options → a JSON Schema the agent can fill in
src/map/b2c.ts both response shapes → one; strips the provider cost
src/tools/info.ts list_modes, list_models, get_model
src/tools/generation.ts estimate_cost, generate, get_status, get_result
src/tools/account.ts get_balance, get_usage
src/index.ts the server: stdio transport, tool registration, annotations
server.json MCP registry manifest (name must match package.json "mcpName")Licence
MIT
Available Tools
9 toolsestimate_costARead-onlyInspect
What a generation will cost, before running it. Pass the SAME input you will pass to generate — on some models the price varies 20x with the settings. Show this to the user before spending their balance.
| Name | Required | Description | Default |
|---|---|---|---|
| input | No | The same input object you would pass to generate. Price varies enormously with it — on some models by 20× between the cheapest and dearest settings — so pass the real values, not an empty object. | |
| model | Yes | Model slug, as given to get_model. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint and openWorldHint, so the description does not need to restate safety. It adds meaningful behavioral context beyond annotations: cost can vary 20x by settings, the real input must be supplied, and this is a pre-execution step that should be surfaced to the user before any spend.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with no filler. The core purpose is front-loaded, followed by the most critical usage instruction, and then the user-facing action. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only estimator with two documented parameters, the description is complete: it states the purpose, the required input relationship, the variability risk, and the expected user interaction. The lack of an output schema is acceptable because the description makes the returned concept clear.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds important parameter guidance by emphasizing that the input must be identical to the generate call and should not be an empty object. This reinforces and extends the schema's own parameter descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states what the tool does: it reports what a generation will cost before running it. It also distinguishes itself from the sibling generate tool by explicitly framing itself as the pre-run cost check, so an agent can select it without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit usage context: pass the same input you would pass to generate, and show the result to the user before spending their balance. It does not enumerate alternatives or exclusions, but the timing and input guidance are concrete enough to route the agent correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generateAInspect
Run a generation. THIS SPENDS THE USER'S BALANCE — call estimate_cost first and tell them the price. Charged when the job is submitted, not when it succeeds. Returns immediately with an id; the media is not ready yet. Check get_status — which waits for you — and obey the check_again_in_seconds it returns, then call get_result. Typical waits are 30 seconds for audio and one to three minutes for an image or a video, so tell the user it is running rather than checking over and over. Output is private unless public:true — except on a free account, where it is always published and watermarked. To chain — animate an image you just made, lipsync a video, and so on — pass {"from_generation": ""} as the file input instead of a path. Never download a Vidofy result just to upload it back.
| Name | Required | Description | Default |
|---|---|---|---|
| input | Yes | The model's inputs, using the names from get_model. A file input takes EITHER a path to a file on this machine (pass the user's own file, never a copy you made), OR — to reuse something Vidofy already made — an object {"from_generation": "<id of an earlier generation>"}. Reuse is the right way to chain: it costs no upload and no download, because the file is already in Vidofy's storage. Example: {"image": {"from_generation": "97803476830366548"}}. | |
| model | Yes | Model slug, as given to list_models / get_model. | |
| public | No | Publish the result to the public Vidofy CDN, permanently. Default false. On a subscriber account false keeps the output private, served through a link that expires. On a free account the result is published and watermarked either way, and this flag changes nothing. Only pass true if the user asked for a public, permanent link. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
It discloses balance charging, when the charge occurs, asynchronous result readiness, typical wait times, privacy behavior, free-account watermarking, and chaining semantics. This is substantial behavioral context beyond the annotations, with no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but dense, and the cost warning is front-loaded as the most critical fact. Every sentence contributes a necessary behavioral or workflow constraint, so the length is justified.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
It covers the complete call path — prerequisite estimate_cost, submission behavior, polling via get_status, wait estimates, result retrieval via get_result, privacy, and chaining. With no output schema, nothing essential is left for the agent to infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even though schema coverage is 100%, the description enriches both main inputs: file input can be a user path or a from_generation object, and public is clarified with subscriber vs free-account outcomes. It makes the parameters' real-world consequences explicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It names a concrete action, 'Run a generation', and immediately frames the lifecycle: an id is returned, then get_status waits, then get_result fetches the result. This clearly distinguishes generate from sibling tools like estimate_cost, get_status, and get_result.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit sequence guidance: call estimate_cost first, surface the price, submit, then poll get_status and obey check_again_in_seconds before get_result. It also tells the agent when not to do something — avoid repeated polling and avoid re-uploading a Vidofy result by using from_generation instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_balanceARead-onlyInspect
The account balance, and how much of it expires with the current subscription. Check before a costly generation.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so no contradiction exists. The description adds useful behavioral context by noting that some balance expires with the current subscription, which is not conveyed by the annotations. It does not cover rate limits or authentication, but the simple read-only nature is well disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences that immediately state what the tool returns and why to use it. No filler or repetition; every phrase earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only tool, the description provides the core information an agent needs: what the balance is, the expiring portion, and when to check it. Since there is no output schema, a more explicit statement of the return shape would improve completeness, but the current description already implies the returned values.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema is fully covered by default. The description correctly avoids inventing parameter details. Baseline for a no-parameter tool is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as returning the account balance and the portion that expires with the current subscription. It does not use an imperative verb like 'get', but 'The account balance, and how much of it expires...' unambiguously describes the resource and scope. It is distinguishable from siblings like estimate_cost and get_usage because it focuses specifically on balance and expiring credits.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear usage trigger: 'Check before a costly generation.' This tells the agent when to invoke the tool, though it does not explicitly mention alternatives or when not to use it. The guidance is strong enough for selection in most cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_modelARead-onlyInspect
Everything needed to call generate on one model: a JSON Schema for its inputs, which file slots it takes and their size limits, and notes the schema cannot express. Call this after list_models and before generate. Its credits_from is the model's cheapest possible price, not this request's — quote estimate_cost.
| Name | Required | Description | Default |
|---|---|---|---|
| model | Yes | The model slug from list_models (e.g. "nano-banana-2-t2i"). The slug, not the model_key — they differ in case and separators. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true and openWorldHint=true, the annotations already cover the safety profile. The description adds valuable behavioral context beyond annotations: the meaning of credits_from ('the model's cheapest possible price, not this request's') and the composition of the returned model definition. It does not contradict the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two tight sentences with no filler. It front-loads what the tool returns, then gives workflow order and a cost-related warning in the second sentence. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Since there is no output schema, the description adequately summarizes the return contents (input JSON Schema, file slots/limits, and notes the schema cannot express) and clarifies the credits_from semantics. It also gives the surrounding workflow (list_models → get_model → generate). It does not discuss error cases or invalid model slugs, but the schema's minLength and slug guidance cover the main calling pitfall.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already documents the only parameter with 100% coverage, including the slug-vs-model_key distinction and an example. The tool description itself adds no parameter-level information beyond saying 'one model,' so the schema carries the semantic burden. A baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states what the tool provides: 'a JSON Schema for its inputs, which file slots it takes and their size limits, and notes the schema cannot express,' and scopes it to 'one model.' It distinguishes itself from list_models and estimate_cost through workflow placement and the credits_from clarification. It could be slightly more explicit with a verb like 'retrieves,' but the purpose is unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit workflow guidance: 'Call this after list_models and before generate.' It also directs cost questions away from this tool by saying 'quote estimate_cost' instead of relying on credits_from. This is clear when-to-use and alternative-selection guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_resultARead-onlyInspect
The finished media, returned as an image you can actually see — an image comes back as itself, a video as its poster frame. Also gives the link: whether it lasts depends on the generation, a private one is signed and expires in about 8 hours, a public one is a permanent CDN URL. The url_note field says which — read it before telling the user to save the link, and call this again for a fresh one if it expired. Pass include_preview:false to get the link alone.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The id returned by generate. | |
| include_preview | No | Return the media itself as an image, so it appears in the conversation and you can see it. Default true. An image comes back as itself; a video comes back as its poster frame, with the video at url. Pass false when you only need the link — it saves roughly 1,400 tokens. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint and openWorldHint annotations, the description discloses important behavior: private links are signed and expire in ~8 hours, public links are permanent CDN URLs, url_note indicates which case applies, and re-calling refreshes an expired link. This is rich, practical behavioral detail that annotations alone do not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core behavior and each sentence adds useful information about previews, link expiry, url_note, and the include_preview option. It is slightly dense and conversational, but there is no wasted content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description adequately covers return behavior: image vs poster frame, link, url_note, and expiration. It could more explicitly state when to call this relative to generate/get_status, but the essential operational details are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters. The description reinforces include_preview:false behavior but adds little semantic meaning beyond what the schema's parameter descriptions already state, including the token savings.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the resource (finished media) and the specific operation (returning a visible image/preview plus the link). It distinguishes get_result from siblings like get_status by framing it around finished media and the result of generation, not status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: read url_note before telling the user to save the link, re-call if expired, and pass include_preview:false when only the link is needed. It does not explicitly name alternative tools or exclusion conditions, but the usage context is strong enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_statusARead-onlyInspect
Whether a generation has finished. Returns done:true once it reaches a final state, then call get_result. This call WAITS for up to 10 seconds before answering, so it is never instant and never needs repeating straight away. While a job is running the answer carries check_again_in_seconds — wait at least that long before calling again. Nothing finishes in under 10 seconds and most media takes one to three minutes, so calling in a tight loop only spends the user's context to learn nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The id returned by generate. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, covering safety and open-world assumptions. The description adds valuable behavioral context BEYOND annotations: it WAITS up to 10 seconds before answering, never finishes in under 10 seconds, and tells the agent to wait check_again_in_seconds before calling again. It also explicitly warns that tight loops 'only spends the user's context to learn nothing', which is a behavioral disclosure about latency and polling semantics. Minor gap: no mention of error states or what happens if the 10-second wait elapses without reaching a final state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, all dense with useful information and zero waste. The core purpose is front-loaded in the first sentence. Every subsequent sentence earns its place: done:true behavior, wait behavior, check_again_in_seconds guidance, and the anti-tight-loop warning. No redundant phrasing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter polling tool with readOnlyHint=true and openWorldHint=true annotations, the description is largely complete. It covers polling semantics (wait up to 10s, check_again_in_seconds), completion signal (done:true), and next step (call get_result). Minor gap: no mention of behavior when the id is invalid/expired or when a generation fails, but for a simple status check with no output schema, the main usage loop is well covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema describes the id parameter as 'The id returned by generate', which is 100% coverage for the single parameter. The description references the polling flow but doesn't add new meaning about the id parameter beyond what the schema provides. Baseline 3 is appropriate since schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool checks whether a generation has finished, uses a specific verb (get) and resource (status), and distinguishes itself from get_result by instructing to call get_result once done:true. This differentiates it from siblings without requiring schema inspection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit when-to-use guidance: call it after generate to poll for completion, and explicitly instructs to call get_result once done:true. It also gives clear when-not-to-use guidance by warning against tight loops and stating 'never needs repeating straight away', with a concrete wait time (check_again_in_seconds). This is exemplary usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_usageARead-onlyInspect
Recent generations and what they cost: totals for the window plus the rows behind them. Use it to answer "what have I spent".
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | How many days back to look. Default 30. | |
| limit | No | Rows to return. Default 20. | |
| offset | No | Rows to skip, for paging. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint, so the safety profile is covered. The description adds that results include both aggregate totals and detailed rows, which is useful, but it does not discuss pagination behavior or any limits beyond what the schema already provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the tool's core purpose and followed by a concrete use case. Every phrase earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with three optional parameters and fully described schema properties, the description is adequate. It communicates the return shape (totals and rows) and the intended question, leaving little ambiguity for invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (days, limit, offset) are already documented. The description adds no parameter-specific meaning, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource and action: recent generations and their costs, with totals plus detail rows. This clearly distinguishes it from siblings like estimate_cost (estimation) and get_balance (account balance).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an explicit use case: answering 'what have I spent'. It does not mention exclusions or alternatives, but the context is clear enough for an agent to select it over cost estimation or balance tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_modelsARead-onlyInspect
List the models available in one mode, with the rough duration of each and credits_from — the CHEAPEST that model can cost, for comparing models against each other. It is not the price of a request and is often several times under it; only estimate_cost answers that. Use the mode code from list_modes (e.g. "t2i").
| Name | Required | Description | Default |
|---|---|---|---|
| mode | Yes | Mode code from list_modes — e.g. t2i, t2v, i2v, lipsync. Not the long name. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry readOnlyHint=true and openWorldHint=true, so the read-only safety profile is declared. The description adds valuable behavioral context beyond the annotations: the warning that credits_from is the cheapest possible cost and 'is often several times under' the real price. That warning prevents a realistic misuse. Minor gap: it doesn't clarify the unit of the 'rough duration,' but the key semantics are well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences of clean, front-loaded prose: the primary purpose first, the critically misleading caveat about credits_from second, and the usage instruction last. Every sentence carries weight and the caveat is placed where it matters most.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema, so the description must explain the return semantics itself — and it does, covering the costs items, their comparability, the cost-overstatement caveat, and the expected unit of the mode input. The only gap is that it does not clarify the units of the rough duration, and it does not note whether the returned model list can be passed to get_model. For a one-parameter tool these are minor omissions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% — the input schema already explains that mode is the code from list_modes with examples (t2i, t2v, i2v, lipsync). The description restates this in 'Use the mode code from list_modes (e.g. "t2i")' but adds no new semantic information beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource — 'List the models available in one mode' — and specifies exactly what is returned: rough duration and credits_from. It also distinguishes itself from its sibling estimate_cost by explicitly saying 'It is not the price of a request... only estimate_cost answers that,' so an agent cannot confuse the two tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description tells the agent when to use this tool ('for comparing models against each other'), provides the prerequisite input rule ('Use the mode code from list_modes'), and gives a named alternative with the condition that selects it ('only estimate_cost answers that' for request pricing). This is explicit, actionable guidance covering both when-to and when-not-to.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_modesARead-onlyInspect
List what Vidofy can generate — text-to-image, image-to-video, lipsync, text-to-speech and so on. Start here, then call list_models with the mode code.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, covering the safe, open-ended nature. The description adds the sequencing context (start here) but does not disclose any additional behavioral traits beyond what annotations provide, such as return format or unfiltered listing behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero waste. The main verb and resource appear in the first few words, followed by inline examples, and the crucial next-step instruction at the end. Perfectly front-loaded and every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only list tool with no output schema, the description gives enough context: what it lists, the categories it covers, and how the return value should be used downstream. It stops short of describing the exact payload structure, but asking for more would be overkill for this simple case.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the parameter baseline is 4 per the rubric. The description wisely mentions the 'mode code' in the usage guidance, which helps the follow-up tool, but since there are no params to explain, this is a solid 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states that the tool lists the generation modes Vidofy supports, with concrete examples (text-to-image, image-to-video, etc.), and explicitly distinguishes it from list_models by calling itself the starting point. This leaves no ambiguity about the tool's job.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit, actionable guidance: 'Start here, then call list_models with the mode code.' This tells the agent exactly when to use this tool and which sibling to invoke next, making the intended workflow unavoidable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.1.0- First observed
estimate_cost - First observed
generate - First observed
get_balance - First observed
get_model - First observed
get_result - First observed
get_status - First observed
get_usage - First observed
list_models - First observed
list_modes
TDQS
Scored across 9 tools
Each tool occupies a distinct stage in the workflow: discovery (list_modes, list_models, get_model), cost estimation, generation, status polling, result retrieval, and account queries. get_balance and get_usage are clearly separated as current balance versus spending history.
Tool names are consistently lowercase snake_case and follow a clear verb_noun pattern (list_*, get_*, estimate_cost, generate). The single bare verb 'generate' still fits the imperative command style and does not create confusion.
Nine tools map cleanly onto the full generation lifecycle—discover modes and models, inspect schemas, estimate cost, submit, poll, retrieve, and check account state. There are no redundant or filler tools.
The set covers discovery, cost estimation, paid generation, asynchronous status checks, result retrieval, and balance/usage monitoring end-to-end. Chaining outputs is also supported via from_generation, so there are no obvious dead ends in the core workflow.
Maintenance
Related MCP Connectors
Generate images, video, audio and short films with 140+ AI models from any MCP client.
1Generate AI video from any MCP client. Pick the model, see the per-second price before you spend.
Plan, compare, price, generate, and recover AI video from compatible MCP clients.
Multi-model AI image and video generator. 14 models behind one OAuth-secured MCP endpoint.
Related MCP Servers
- AlicenseAqualityDmaintenanceProvides access to Vidu's video generation models for creating high-quality videos from text, images, and reference content. It enables users to generate creative video content directly within MCP-compatible applications like Claude and Cursor.55MIT
- FlicenseNot gradedqualityBmaintenanceEnables AI clients like Claude and ChatGPT to generate images and videos, animate images, create lip-synced videos, list TTS voices, and manage media via remote MCP tools.-
- AlicenseNot gradedqualityCmaintenanceExposes video generation and editing tools to MCP-capable clients, enabling text-to-video, image-to-video, video editing, reusable voice and character profiles, and task status queries through natural language.4MIT
- FlicenseAqualityBmaintenanceEnables AI image and video generation, editing, and history tracking inside MCP clients using OpenAI, Google Gemini, and Veo models, with flexible API key or cookie-based configuration.121-