Tanvo MCP
OfficialEnables song generation through Suno V6: agents can describe a song, supply their own lyrics, or go instrumental, and each run returns two takes as MP3 files with cover art, a title and the lyrics (the generate_music tool).
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Tanvo MCPTurn my cat photo into a Renaissance oil painting and save it to ~/Pictures"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Tanvo MCP
The MCP server for Tanvo, an AI image, video and music studio. Ask your assistant for "a renaissance portrait of my cat from this photo", "turn this selfie into a boxed action figure", "a 10-second clip of this product spinning" or "a birthday song for Sam", and it runs it on Tanvo and hands you the file.
Works in Claude Desktop, Claude Code, Cursor, VS Code, Windsurf, Cline, Codex, Gemini CLI and any other MCP client. No API key needed to try it: the house image engine runs on the free tier.
What it can do
80 ready-made apps with 350+ preset looks. Pet portraits, figurines, old photo restoration, AI baby, product photos, outfit try-on, trend videos (hotel lobby duet, crying filter, dance), birthday and love songs. Each look is a tuned prompt with a real example output, so you know what you will get.
Images on Nano Banana 2 / Pro, Seedream 5, GPT Image 2.5, Qwen Image, Grok Imagine and the free house engine: text to image, edits and multi-photo compositions.
Video on Veo 3.1, Kling 3.0, Seedance 2.5, Wan, Minimax and more: text to video, image to video, start and end frames, native sound.
Songs on Suno V6: two takes per run, each an MP3 with cover art, a title and the lyrics.
Local photos are uploaded for you, and results can be saved straight into a folder.
Related MCP server: media-gen-mcp
Tools
Tool | What it does |
| Search the apps by what the user wants ("royal pet portrait", "fix an old photo") |
| An app's inputs and looks, with example outputs, prompts and cost |
| Run an app's look on the user's photos |
| Text to image, or edit / combine photos |
| Text or image to video (returns an id; videos take minutes) |
| Write a song: describe it, bring lyrics, or go instrumental |
| Status and outputs of a run; can wait and save files |
| Every model, its options and its credit cost |
| Free tier or account, and the credits left |
| A link that opens the app, look or studio ready to go |
Install
The server runs with npx, so there is nothing to install first (Node 20+).
Claude Code
claude mcp add tanvo -- npx -y tanvo-mcpAdd -e TANVO_API_KEY=sk_... to run on your account.
Claude Desktop, Cursor, Windsurf, Cline, VS Code (the mcpServers block of the client's MCP config)
{
"mcpServers": {
"tanvo": {
"command": "npx",
"args": ["-y", "tanvo-mcp"],
"env": { "TANVO_API_KEY": "sk_..." }
}
}
}Gemini CLI
gemini extensions install https://github.com/tanvoai/tanvo-mcpCodex (~/.codex/config.toml)
[mcp_servers.tanvo]
command = "npx"
args = ["-y", "tanvo-mcp"]
env = { TANVO_API_KEY = "sk_..." }AI agents that install servers themselves (Cline and others) can follow llms-install.md.
Free tier and API keys
Without a key the server uses Tanvo's anonymous free tier: the house image engine (studio-image-v1), watermarked, a few images per machine per day. Everything else (apps on premium models, video, songs) needs a key. When a call needs one, the server replies with a link that opens the same app or prompt in the browser, where new accounts start with free credits.
Create a key under Settings → API keys. A key runs on your own account: the same models, prices and plan as the website, and every run lands in your history. Each run's price is shown by get_app / list_models, and failed runs are refunded automatically.
Variable | Meaning |
| Your API key (optional) |
| Save every result into this folder (optional; each call can also pass |
| Another deployment, e.g. a local build (default |
| Fix the free-tier id instead of the one stored in |
Examples
Make my dog look like a Renaissance oil painting. ~/Pictures/rex.jpg
find_apps → get_app ai-pet-portrait-generator → generate_from_app with look renaissance and the photo.
Write a birthday song for my sister Ana, she loves surfing and terrible puns. Save it to ~/Music.
generate_music in Describe mode with save_to: "~/Music".
Animate this product photo into a slow 360° spin, 9:16.
find_apps "product spin" → generate_from_app product-spin, then get_generation until it is done.
Privacy
Prompts and photos you pass are sent to tanvo.ai to run the job, under the privacy policy. Content is moderated before anything runs. The server stores nothing except the anonymous free-tier id.
Develop
npm install
npm run build
TANVO_BASE_URL=http://localhost:3210 node dist/index.jsDocs: tanvo.ai/developers/mcp · API reference: tanvo.ai/developers/api · OpenAPI: tanvo.ai/openapi.yaml
Also from Tanvo
tanvo-python:
pip install tanvotanvo-js:
npm install @tanvoai/sdk
License
MIT
Available Tools
10 toolsaccountAccount and creditsARead-only
Who the server runs as (an API key account or the anonymous free tier) and how many credits are left.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so the safe-read profile is covered. The description adds that the result reveals the auth mode (API key vs anonymous free tier), which is useful context, but says nothing about freshness/caching of the credit count or whether the call itself costs credits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; both the identity and credit facets are stated immediately and nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must carry the return-value burden, and it does name the two returned facts (account identity and remaining credits). It stops short of describing the exact shape or when the identity would differ, leaving a small gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4; there is nothing for the description to disambiguate beyond what the empty schema already conveys.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (the server's account identity) and the two things it returns: the acting account (API key vs anonymous free tier) and remaining credits. This is specific and a reader can distinguish it from the generation-oriented siblings, though it doesn't explicitly contrast itself with any sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: an agent infers it should call this to check identity or credit balance before generating. There is no explicit when-to-use, when-not-to-use, or named alternative (e.g., 'check this before generate_image to avoid credit failures').
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_appsFind a Tanvo appARead-only
Search Tanvo's ready-made AI apps (photo effects, trends, portraits, product shots, video effects, songs) by what the user wants, e.g. 'royal pet portrait', 'action figure of me', 'old photo restoration', 'dancing video', 'birthday song'. Each app has tuned preset looks with real example outputs. Use get_app for the looks, then generate_from_app to run one.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | Only apps that make this kind of media | |
| limit | No | ||
| query | No | What the user wants, in their words. Leave empty to list trending apps. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds useful behavioral context beyond that: results are curated apps with tuned preset looks and real example outputs, and an empty query switches the call from search to trending listing. No pagination or result-field detail, but the added context is genuine.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with purpose, then query examples, then the follow-up workflow. Every sentence earns its place with zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter, no-output-schema search tool this is nearly complete: the agent knows what it searches, how to phrase queries, what the results conceptually are, and which siblings to invoke next. Minor gaps are the undocumented limit default/cap and the concrete shape of returned app records.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67% (kind and query documented, limit not). The description reinforces query semantics with natural-language examples mapped to the image/video/music categories covered by the kind enum, which helps the agent formulate calls. It adds nothing for limit, keeping it short of a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search) and resource (Tanvo's ready-made AI apps), enumerates the content categories, and gives five concrete example queries that make the scope unmistakable. An agent can distinguish it from generate_image or list_models without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent through the workflow: 'Use get_app for the looks, then generate_from_app to run one,' and the query parameter notes that an empty string lists trending apps. Clear context and named alternatives, though it does not state when *not* to use this tool (e.g. for raw model selection via list_models).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_from_appRun a Tanvo appA
Run one of Tanvo's apps with one of its looks on the user's photos, e.g. the 'renaissance' look of a pet portrait app. Photos go in the order of the app's inputs (see get_app). Uses the look's tuned prompt and model; details adds the user's own touch. Charges the look model's credits (shown by get_app) and refunds failures. Needs TANVO_API_KEY unless the look runs on the free engine.
| Name | Required | Description | Default |
|---|---|---|---|
| app | Yes | App slug from find_apps | |
| look | No | Look id from get_app; defaults to the first look | |
| wait | No | Wait for the result. Defaults to true for images and songs, false for video. | |
| photos | No | Photos as local file paths or public https URLs. Local files are uploaded for you. | |
| details | No | Optional personal detail woven into the prompt, e.g. 'for my sister Ana' or 'wearing a red scarf' | |
| save_to | No | Optional folder to download the results into (defaults to TANVO_OUTPUT_DIR when set) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare openWorldHint, so the description carries the full behavioral burden and does so well: it discloses credit charging based on the look model, refunds on failure, the auth requirement (TANVO_API_KEY), and that the look's tuned prompt and model are used. These are non-obvious operational facts an agent needs before invoking.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense sentences, front-loaded with the core action and example, then constraints (ordering, prompt/model, cost/refund, auth). No filler or restated name/title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description covers cost, refunds, auth, ordering, and defaults well enough to call the tool. It leaves one gap: it doesn't explain the async path when wait=false (presumably retrieving via get_generation), which matters given the video default is not to wait.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3; the description adds real meaning beyond the schema by explaining that photos are ordered according to the app's inputs (see get_app) and that `details` is a personal touch woven into the prompt. That clarifies usage of two parameters in ways the schema does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource ('Run one of Tanvo's apps with one of its looks on the user's photos') with a concrete example that clarifies the distinction from the generic generate_image sibling. An agent can tell this is preset/look-driven generation rather than free-form prompting.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Routes the agent to get_app for input ordering and look ids, and conditionally to find_apps via the schema ('App slug from find_apps'), plus a note that TANVO_API_KEY is needed unless the free engine is used. It never states when NOT to use this tool versus generate_image/generate_video, so it is clear context without explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_imageGenerate or edit an imageA
Text to image, or an edit / combination when images are given. Returns the image (inline when small) and its URL. The free tier runs studio-image-v1 without a key; other models need TANVO_API_KEY. Credits are charged per run and refunded if it fails.
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | Wait for the image (usually under a minute) | |
| model | No | Model id from list_models | studio-image-v1 |
| aspect | No | 1:1, 16:9, 9:16, 4:3, 3:4 … as the model allows (default: the model's) | |
| format | No | ||
| images | No | Photos as local file paths or public https URLs. Local files are uploaded for you. | |
| prompt | Yes | What to make: subject, setting, light and style. For edits, say what to change and what to keep. | |
| save_to | No | Optional folder to download the results into (defaults to TANVO_OUTPUT_DIR when set) | |
| resolution | No | 1K, 2K or 4K where the model supports it (default: the model's) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare openWorldHint, so the description carries the behavioral burden and does it well: it discloses the return form (inline when small, plus URL), the authentication requirement (TANVO_API_KEY for non-free models), and the billing behavior (credits charged per run, refunded on failure). That is exactly the cost/auth/error context an agent needs before invoking.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with zero filler, front-loaded with the core capability and mode switch before dropping into auth and billing details. Every sentence carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly takes on the return-value explanation (inline image plus URL) and covers setup and cost. The one gap is async behavior: `wait` defaults to true, but the description never tells the agent what to do when wait=false (poll get_generation), which is a meaningful omission for a long-running generation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 88%, so the schema already documents prompt, images, model, aspect, resolution, save_to, wait and format. The description adds one genuinely new semantic — that `images` turns the call into an edit/combination — but contributes nothing about `wait`, `save_to` defaults, or aspect/resolution constraints beyond the schema. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening clause names a specific verb pair (generate/edit) and resource (image), and explicitly states the condition that switches modes: an edit or combination 'when `images` are given.' It is immediately distinguishable from siblings like generate_video, generate_music, and generate_from_app.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the two usage modes via the `images` condition and notes the free-tier model, which is useful routing context. But it never says when to prefer this over siblings (e.g. generate_from_app) nor names list_models as the way to discover valid model ids, and it is silent on the non-blocking path (wait=false, then get_generation).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_musicWrite a songA
Makes two takes of a song on Suno V6, each an MP3 with cover art, a title and the lyrics as sung. Modes: Describe (prompt is one sentence about the song; lyrics are written for you), Lyrics (prompt is your lyrics, with [Verse] [Chorus] [Bridge] tags), Instrumental (prompt describes the sound). 60 credits a run; needs TANVO_API_KEY.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Describe | |
| wait | No | Wait for the songs (usually about a minute) | |
| style | No | Genre, instruments, tempo, voice, e.g. 'acoustic pop, hand claps, warm female vocal' | |
| title | No | Song title (Lyrics and Instrumental modes) | |
| vocal | No | Lean towards a male or female voice (Lyrics mode) | |
| prompt | Yes | ||
| save_to | No | Optional folder to download the results into (defaults to TANVO_OUTPUT_DIR when set) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations supply only openWorldHint=true, so the description carries the behavioral load — and it delivers: credit cost ('60 credits a run'), auth requirement ('needs TANVO_API_KEY'), and what a run actually produces (two takes, each an MP3 with cover art, title, and sung lyrics). It doesn't cover async retrieval when wait=false, but cost and auth disclosure is substantive context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with what the tool produces before the mode table and the cost/auth constraints. Every clause carries information — output format, mode semantics, pricing, credential requirement — with no filler or restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly compensates by describing the return (two takes, MP3, cover art, title, lyrics) and adds cost/auth/mode detail. The one gap is that it never mentions the async path — when wait=false the agent presumably needs get_generation to fetch results — which matters for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 71%, and the description adds real meaning on the highest-stakes parameter by explaining how prompt and mode interact, which the enum alone cannot convey — the same string means a one-line brief in Describe but full tagged lyrics in Lyrics. The remaining params (wait, save_to, title, vocal) are left to the schema, which is acceptable at this coverage level.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Makes two takes of a song on Suno V6') plus the concrete output (two MP3s with cover art, title, and lyrics as sung). This clearly distinguishes it from siblings generate_image, generate_video, and generate_from_app without needing to name them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The mode breakdown ('Describe' = one-sentence prompt, lyrics written for you; 'Lyrics' = your lyrics with [Verse]/[Chorus]/[Bridge] tags; 'Instrumental' = prompt describes the sound) tells an agent exactly how to choose the mode and what to put in prompt for each. There are no explicit exclusions or routing rules versus sibling generation tools, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_videoGenerate a videoA
Text to video, or animate a still when image is given (optionally ending on end_image). Video takes one to several minutes, so by default this returns at once with an id for get_generation. Needs TANVO_API_KEY.
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | Wait for the clip (up to 8 minutes) | |
| audio | No | Native sound on models that offer it | |
| image | No | Start frame: a local path or a public https URL | |
| model | No | Model id from list_models | studio-video-v1 |
| aspect | No | 16:9, 9:16, 1:1 … (default: the model's) | |
| prompt | Yes | One scene: subject, action, camera movement, light and style | |
| save_to | No | Optional folder to download the results into (defaults to TANVO_OUTPUT_DIR when set) | |
| duration | No | Seconds; one of the model's options (default: the model's) | |
| end_image | No | End frame on models that support it | |
| resolution | No | 480p, 720p, 1080p or 4K where supported (default: the model's) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only openWorldHint annotated, the description carries the burden and delivers useful traits: generation takes one to several minutes, the call returns immediately with an id by default, and it needs TANVO_API_KEY. It stops short of disclosing cost or the exceptions to the default async return.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly packed sentences with zero filler; the capability, the latency/default behavior, and the auth requirement are all front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description covers the essential return contract (id for get_generation) and auth. It omits how save_to affects returns, but the schema covers that parameter fully.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds only the image/end_image pairing semantics (start frame vs ending frame), which the schema largely already conveys per-parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb+resource ('Text to video') and clarifies the alternate mode ('animate a still when `image` is given'), which cleanly differentiates it from the sibling generate_image tool without requiring a schema read.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains the default behavior and routes the agent onward ('by default this returns at once with an id for get_generation'), which is clear usage context. It does not explicitly contrast when to choose this over generate_image or generate_from_app, so it falls short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_appShow an app and its looksARead-only
One Tanvo app: the photos it needs, its preset looks (each with an example output and the prompt behind it), the model it runs on and the link that opens it in the browser.
| Name | Required | Description | Default |
|---|---|---|---|
| app | Yes | App slug from find_apps |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so safety is covered. The description goes further by disclosing the actual payload contents (preset looks with example outputs and underlying prompts, model, browser link), which is meaningful context beyond the structured fields, though it says nothing about behavior for an invalid slug.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with 'One Tanvo app', with the returned fields listed compactly. It is appropriately sized for a read tool, though the parenthetical about example output and prompt is slightly dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of explaining return values and does so thoroughly, covering all major response components. Only the failure mode for an unknown slug is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'app' parameter is documented as 'App slug from find_apps', so the schema carries the load. The description adds no syntax or format detail beyond restating that it concerns one app; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource ('One Tanvo app') and enumerates exactly what it returns: photos, preset looks with example output and prompt, model, and browser link. An agent can distinguish it from find_apps (search) and get_generation (single generation), though the verb is implicit rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use sentence or named alternative, but the content listing implies a detail-fetch operation for an app the agent already knows about. The workflow link (slug comes from find_apps) lives in the schema parameter description, not the tool description, so guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_generationCheck a generationBRead-only
Status and outputs of a run started earlier (poll this for videos).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| wait | No | Wait up to 2 minutes for it to finish | |
| save_to | No | Optional folder to download the results into (defaults to TANVO_OUTPUT_DIR when set) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark it read-only and open-world. The description adds that this observes an earlier-started run and should be polled for videos. It does not disclose return shape, status transitions, pagination, or whether wait blocks, leaving behavior largely to schema and annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded, with a parenthetical that adds polling guidance. No filler, though the structure is minimal and the parenthetical could be slightly clearer.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only polling tool with no output schema, the description identifies what is returned (status, outputs) and when to poll. However it omits what id refers to, expected status values, and how wait/save_to interact, leaving gaps an agent must infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, and the description mentions none of the three parameters. id is undocumented in both schema and description; wait and save_to have schema text but the description adds no meaning beyond that. It fails to compensate for the missing id semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (a generation run) and action (retrieve status/outputs) in a short sentence, and distinguishes itself from siblings like generate_video by being the polling endpoint. It stops short of explicitly naming alternatives, so it is clear but relies on the parenthetical for differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Includes an implicit when-to-use cue: poll this after starting a run, specifically for videos. It does not identify alternatives or state when not to use it, so guidance is present but incomplete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_modelsList modelsARead-only
Every image, video and music model with the options each accepts (aspect, resolution, duration, format, audio), how many reference images it takes, whether the free tier covers it, and its credit cost at default settings.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | Filter by media kind |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint, so safety is covered. With no output schema, the description usefully carries the return-content burden, disclosing that options (aspect, resolution, duration, format, audio), reference-image capacity, free-tier eligibility and default credit cost are all included — an agent knows exactly what it gets back.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with the scope ('Every image, video and music model') front-loaded before the payload list. The parenthetical field enumeration is long but earns its place given the absence of an output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless, read-only catalog call with no output schema, the description covers the return payload thoroughly. It could note whether the full catalog is returned at once or whether the kind filter is required for reasonable performance, which is the only remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter (kind) exists and schema description coverage is 100%, so the enum values and their meaning are fully documented in the schema. The description adds nothing about filtering behavior, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource precisely (image, video and music models) and enumerates the facets returned, which cleanly separates it from the generate_* siblings. The verb is only implicit (the title supplies 'List'), and the sentence is framed as a return-value inventory rather than an action, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to call this versus alternatives, nor any prerequisite or exclusion. An agent can infer it is a discovery/catalog call to run before generate_image or generate_video, but that routing is left entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
open_in_tanvoOpen in the browserARead-only
A link that opens Tanvo ready to go: an app on a chosen look, or the image / video / music studio with a model and prompt filled in. Use it when the user wants to upload photos and tweak settings themselves, or has no API key.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App slug from find_apps | |
| kind | No | Studio to open when no app is given | |
| look | No | Look id from get_app | |
| model | No | Model id to open the studio on | |
| prompt | No | Prompt to prefill |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true is consistent with returning a link, and the description adds what the link is preconfigured with, which is genuine value beyond the annotation. It does not disclose whether the link is shareable, expiring, or auth-bound, nor whether the tool launches a browser itself versus returning a URL for the agent to present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with what the tool yields before the usage condition. Slightly dense in the middle clause, but every sentence carries content and none is redundant with the title or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must carry the return contract, and it only partially does: 'A link' is stated, but it is ambiguous whether the tool opens the browser or returns a URL the agent should surface. With 5 optional parameters and a prefill contract, a sentence on precedence between app/look and kind/model would close the remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds relational meaning the schema lacks: 'app on a chosen look' and 'studio with a model and prompt filled in' explain how the parameters combine into the two distinct modes. It stops short of explaining behavior when conflicting params (e.g. both app and kind) are supplied.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (a link that opens Tanvo preconfigured) and enumerates the two modes: an app on a chosen look, or a studio with model/prompt prefilled. It is distinguishable from the generate_* siblings because it produces a UI link rather than a generation, though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear positive trigger: 'Use it when the user wants to upload photos and tweak settings themselves, or has no API key.' That condition implicitly routes the agent away from generate_image/video/music, but no alternative is named and no explicit when-not case is stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
10 tool updates
v0.2.0- First observed
account - First observed
find_apps - First observed
generate_from_app - First observed
generate_image - First observed
generate_music - First observed
generate_video - First observed
get_app - First observed
get_generation - First observed
list_models - First observed
open_in_tanvo
TDQS
Scored across 10 tools
The generation tools are cleanly separated by media type (image, video, music) and by source (generate_from_app vs generate_image), and discovery tools (find_apps, get_app, list_models) have distinct roles. The only mild tension is generate_image vs generate_from_app (both yield images) and open_in_tanvo vs the generate_* family (both can set up a creation), but descriptions make the boundaries clear.
Eight of ten tools follow a consistent snake_case verb_noun pattern (generate_image, generate_video, generate_music, find_apps, get_app, get_generation, list_models, generate_from_app). Deviations are minor: 'account' is a bare noun, and 'open_in_tanvo' is verb_preposition_noun, both still readable and unambiguous.
Ten tools is well-scoped for a media-generation service: discovery (2), details (1), generation (4), status (1), account (1), and a browser escape hatch (1). Each tool earns its place with no redundancy.
The surface covers the full create-and-poll lifecycle across image, video, music, and preset apps, plus model catalog and account/credit visibility. Minor gaps remain: no cancel/delete for in-flight generations and no credit top-up or history listing, though agents can work around these.
Maintenance
Related MCP Connectors
Generate images, video, audio and short films with 140+ AI models from any MCP client.
Create images & video from any MCP agent — 17 models, spend limits, one URL.
Create and manage AI image and video generations through Quriov's fixed public MCP tools.
Multi-model AI image and video generator. 14 models behind one OAuth-secured MCP endpoint.
Related MCP Servers
- AlicenseAqualityAmaintenanceEnables AI image and video generation using Midjourney through the AceDataCloud API. It supports comprehensive features including image creation, transformation, blending, editing, and video generation directly within MCP-compatible clients.164,844 PyPI11MIT
- FlicenseNot gradedqualityBmaintenanceEnables AI clients like Claude and ChatGPT to generate images and videos, animate images, create lip-synced videos, list TTS voices, and manage media via remote MCP tools.-

@vidofy/mcpofficial
AlicenseAqualityBmaintenanceEnables generating images, video, audio, and speech from MCP clients using your own Vidofy account, with access to hundreds of models for text-to-video, image-to-video, image editing, lipsync, text-to-speech, and voice cloning.955 npm2MIT- AlicenseNot gradedqualityBmaintenanceEnables any MCP client to search and submit tasks across 350+ multimodal image, video, audio, and 3D models, run ComfyUI workflows and AI apps, upload/download files, and chat with LLMs. It supports stdio MCP integration and optional scope trimming for image, video, or audio only.1MIT