ModelsLab
Server Details
Generate images, video, speech and music, and run LLM chat, across 10,000+ models on ModelsLab.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 23 tools
Most tools target distinct modalities and actions, but fetch-generation overlaps with the specialized fetch-image, fetch-video, and fetch-audio tools. Descriptions help clarify usage, but the redundancy could still cause confusion when an agent chooses which fetch tool to use.
All tools use lowercase kebab-case and are readable, but the naming patterns vary (verb-noun, X-to-Y, noun-noun, bare noun). This is mostly consistent with minor deviations from a single verb_noun convention.
With 23 tools covering text, image, video, audio, music, and model listing, the count is on the high side but reasonable given the platform's broad multi-modal scope. Each tool appears to target a distinct capability, though the fetch-generation redundancy suggests one tool could be removed.
The surface covers core generation, transformation, and retrieval workflows across all major modalities. Minor gaps exist, such as no explicit cancel-generation tool and limited specialized editing operations like upscaling or background removal, but these are not critical dead ends.
Available Tools
23 toolschat-completionChat CompletionBInspect
Chat with AI language models.
Send messages to an LLM and receive AI-generated responses.
Supports various models and configuration options.
A message may carry a PDF as a {"type":"file","file":{"filename":...,"file_data":"data:application/pdf;base64,..."}} content part.| Name | Required | Description | Default |
|---|---|---|---|
| n | No | Number of completions to generate (1-10). | |
| seed | No | Random seed for reproducible results. | |
| stop | No | Up to 4 sequences where the API will stop generating. | |
| top_k | No | Top-k sampling parameter. | |
| top_p | No | Nucleus sampling parameter (0-1). | |
| plugins | No | OpenRouter plugins. Send [{"id":"file-parser","pdf":{"engine":"native"}}] alongside a {"type":"file"} content part to have a PDF attachment parsed. | |
| messages | Yes | Array of message objects with "role" (system/user/assistant) and "content" keys. | |
| model_id | Yes | The LLM model ID to use (e.g., "gpt-4", "claude-3"). | |
| max_tokens | No | Maximum number of tokens to generate. | |
| temperature | No | Sampling temperature (0-2). Higher values make output more random. | |
| response_format | No | Response format configuration (e.g., {"type": "json_object"}). | |
| presence_penalty | No | Penalty for new topics (-2 to 2). | |
| frequency_penalty | No | Penalty for frequent tokens (-2 to 2). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide only openWorldHint=true, so the description carries most of the behavioral burden. It does add non-obvious context: PDF attachment handling via a typed file content part, tying into the plugins/file-parser mechanism. However, it says nothing about cost, streaming, rate limits, or error behavior, and 'Supports various models and configuration options' is filler.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and only four short sentences. The PDF detail earns its place, but 'Supports various models and configuration options' is a low-value sentence that could be cut.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 13 parameters, nested content objects, and no output schema, the description covers the essential action and the notable PDF attachment case but omits the shape of the response (choices, streaming), model selection cost implications, and error handling. Adequate but with clear gaps for a tool this complex.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 13 parameters are already documented, which sets the baseline at 3. The description supplements this by showing the exact file content-part shape for messages, adding real value beyond the schema, but it adds nothing for the sampling/knob parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: chat with AI language models, send messages and receive generated responses. It clearly identifies the text-LLM capability amid media-generation siblings, but never names a sibling or explicitly contrasts itself (e.g. against list-models), so differentiation is implied rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use/when-not guidance and no routing to alternatives. The only practical hint is that a message may carry a PDF, which is a capability note rather than selection guidance. Model choice is left entirely to the caller with just two example IDs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
dubbingDubbingAInspect
Create dubbed audio content. Takes a video and translates/dubs the audio from one language to another. Returns a request ID that can be used with fetch-audio to retrieve results.
| Name | Required | Description | Default |
|---|---|---|---|
| webhook | No | URL to receive webhook notification when dubbing completes. | |
| model_id | Yes | The model ID to use for dubbing. | |
| track_id | No | Custom tracking ID for the request. | |
| init_video | Yes | URL or base64 string of the video to dub. | |
| output_lang | Yes | Target language code for the dubbed output. | |
| source_lang | Yes | Source language code (e.g., "en", "es", "fr"). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare openWorldHint=true, so the description carries the rest. It usefully discloses the asynchronous job pattern: the call returns a request ID that must be redeemed via fetch-audio, which is meaningful behavioral context beyond the annotation. It does not mention cost, latency, or auth requirements, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each earning its place: the action, the transformation, and the return value. The core purpose is front-loaded and nothing is redundant. Minor leading whitespace in the raw string is immaterial.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly explains the return value (a request ID) and how to consume it via fetch-audio, closing the biggest gap for an async tool. It does not cover failure modes, dubbing duration, or webhook behavior, which the webhook parameter implies but the description never elaborates.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six parameters (model_id, init_video, source_lang, output_lang, webhook, track_id) are already documented in the schema. The description restates only that a video and two languages are involved and adds no format, enum, or constraint detail beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('Create dubbed audio content') and clarifies the transformation: a video's audio is translated/dubbed between languages. It also routes the agent to the follow-up tool 'fetch-audio', which helps separate it from pure text-to-speech or speech-to-speech siblings, though it never explicitly contrasts itself with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied — the agent must infer that this is the tool for dubbing a video's audio. The one concrete guideline is that the returned request ID should be used with fetch-audio, which tells the agent how to complete the workflow but not when to choose dubbing over alternatives like lip-sync or speech-to-speech.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fetch-audioFetch Audio ResultARead-onlyInspect
Retrieve the status and results of an audio generation request. Use the request ID returned from text-to-speech, speech-to-text, music-generation, and other audio tools.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The request ID returned from a previous audio generation call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already establishes this is a safe, non-mutating read. The description adds value by disclosing that the call returns both status and results, implying polling semantics, but it says nothing about what happens while a request is still pending or how failures are reported. Adequate but thin beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, and the core operation is front-loaded before the sourcing guidance. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read-only fetch tool with no output schema, the description covers the operation, the input source, and the existence of status/result data. The only gap is sibling routing against fetch-generation, which is minor given the audio scoping.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter documented at 100% schema coverage, so the schema already carries the meaning. The description reinforces it by telling the agent the ID comes from an audio tool, but adds no format or constraint detail beyond that. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description pairs a specific verb ("Retrieve") with a specific resource ("status and results of an audio generation request"), which is far better than a tautology. It scopes itself to audio, which separates it from fetch-image and fetch-video, but it never explicitly distinguishes itself from the sibling fetch-generation tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence gives clear context: call this with the request ID produced by text-to-speech, speech-to-text, music-generation, and other audio tools. That is a real when-to-use signal, though there is no explicit exclusion telling the agent when to prefer fetch-generation instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fetch-generationFetch Generation StatusARead-onlyInspect
Check the status of a queued or processing generation request on the ModelsLab V7 API.
Use this tool when a generation tool returns a `status` of `"processing"`, and call it again
while the status stays `"processing"`. Image jobs usually finish in seconds; video and music
jobs can take several minutes, so tell the user the job is still running if it has not
finished after a few polls.
Pass the `id` from the original generation response and the `type` matching the category
of the original request (e.g. "images" for text-to-image, image-to-image, or inpaint-image).
The response will contain:
- `status`: "success", "processing", or "error"
- `output`: array of URLs when status is "success"| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric id returned by a generation tool when status was "processing". | |
| type | Yes | The generation category matching the original request: "images" for image tools, "videos" for video tools, "audios" for audio tools. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so safety is covered. The description goes beyond by disclosing polling behavior, expected latency per media type, and what the agent should tell the user while waiting — useful operational context. It stops short of noting rate limits or a max-polling cutoff.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose, then usage, then parameter guidance, then return shape — a logical order. Slightly verbose in the polling-timing paragraph, but every part earns its place given there is no output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by enumerating the return fields (`status`, `output`) and their values, and it covers the full call lifecycle including re-polling. An agent has everything needed to invoke and interpret this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description reinforces the type-to-category mapping with concrete examples ('images' for text-to-image/image-to-image/inpaint-image), which is helpful routing context, but it largely restates what the schema property descriptions already say.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Check the status of a queued or processing generation request.' Clear what it does. However, it never distinguishes itself from the fetch-image, fetch-video, and fetch-audio siblings, which an agent could easily confuse with this generic poller — the `type` parameter is the only hint at the difference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit trigger ('use this tool when a generation tool returns a status of "processing"'), explicit loop condition ('call it again while the status stays "processing"'), and explicit user-facing guidance about job duration. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fetch-imageFetch Image ResultARead-onlyInspect
Retrieve the status and results of an image generation request. Use the request ID returned from text-to-image, image-to-image, or inpaint-image tools.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The request ID returned from a previous image generation call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true already tells the agent this is a safe, non-mutating read, so the safety burden is covered. The description adds that it returns 'status and results', hinting at an async polling pattern, but does not explain possible states (pending/failed/complete), whether polling is required, or how partial results are represented.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the primary action front-loaded and the prerequisite second. No filler, no redundancy with the title or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description is the only place to learn what comes back, and 'status and results' is thin for a polling-style fetch tool: it omits the possible status values and what a caller should do when the generation is not yet ready. The input side is adequately covered via the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the sole 'id' parameter is already documented as the request ID from a previous image generation call. The description marginally enriches this by naming the specific tools that produce the ID, but adds no format or validation detail beyond the schema. Baseline 3 applies when the schema carries the semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (retrieve) and resource (status and results of an image generation request), which clearly separates it from generic fetchers like fetch-generation and from the other media fetchers (fetch-video, fetch-audio). It does not, however, explicitly name a sibling it is not, so the differentiation is implied by the 'image generation request' scope rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to use the request ID returned from text-to-image, image-to-image, or inpaint-image, giving the agent the prerequisite and provenance of the input. It stops short of when-not guidance (e.g., which fetch-* to use for non-image generations), so no explicit exclusions are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fetch-videoFetch Video ResultARead-onlyInspect
Retrieve the status and results of a video generation request. Use the request ID returned from text-to-video, image-to-video, video-to-video, lip-sync, or motion-control tools.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The request ID returned from a previous video generation call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint=true, and the description is consistent with that non-mutating profile. It notes that status and results are returned, but for an asynchronous job endpoint it omits important behavior such as in-progress/failed states or the need to poll until completion.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero padding, with the core action front-loaded before the sourcing guidance. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing return values, and 'status and results' is only a thin sketch of what comes back. It is adequate for invocation but incomplete about async state semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'id' parameter is fully documented as the request ID from a previous video generation call. The description largely restates that, adding only the list of generator tools that can source the ID, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Retrieve the status and results of a video generation request'), which clearly separates it from fetch-audio and fetch-image by resource type. It does not, however, distinguish itself from the sibling fetch-generation, leaving some ambiguity about which fetch endpoint applies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent to call it with a request ID produced by text-to-video, image-to-video, video-to-video, lip-sync, or motion-control, which supplies clear triggering context. It stops short of stating when-not to use it or naming fetch-generation as the alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
image-to-imageImage to ImageAInspect
Transform existing images based on text prompts. Takes an input image and modifies it according to the provided prompt. Returns a request ID that can be used with fetch-image to retrieve results.
| Name | Required | Description | Default |
|---|---|---|---|
| width | No | Output image width in pixels (512-1024). | |
| height | No | Output image height in pixels (512-1024). | |
| prompt | Yes | Text description of how to transform the image. | |
| samples | No | Number of images to generate (1-4). | |
| webhook | No | URL to receive webhook notification when generation completes. | |
| model_id | Yes | The model ID to use for image transformation. | |
| strength | No | Transformation strength (0-1). Higher values mean more change from the original. | |
| track_id | No | Custom tracking ID for the request. | |
| init_image | Yes | Input image URL or base64 string to transform. | |
| aspect_ratio | No | Aspect ratio for the output image. | |
| negative_prompt | No | Text describing what to avoid in the output. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare openWorldHint=true, so the description carries most of the behavioral burden. It usefully discloses the async pattern (returns a request ID, results retrieved via fetch-image), but omits cost/latency, auth requirements, and what happens to the original image. This partial disclosure is adequate but thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with the core action front-loaded, followed by input semantics and the async return path. No filler or repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description appropriately explains that the immediate return is a request ID and that fetch-image retrieves the result, closing the main gap for an async tool. Coverage is nearly complete, missing only edge conditions such as failure or timeout behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% across 11 parameters, including ranges and defaults, so the schema does all parameter work. The description adds no extra meaning about strength, samples, or negative_prompt beyond what the schema already says, giving the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Transform existing images based on text prompts') and clarifies that an input image is modified, which implicitly separates it from text-to-image. It stops short of naming the closest siblings (text-to-image, inpaint-image, image-to-video), so an agent must infer which one fits.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The mention that the returned request ID is used with fetch-image gives a clear downstream routing hint, and 'takes an input image' implies usage. However, there is no explicit when-to-use vs when-not guidance relative to text-to-image or inpaint-image, which are the natural alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
image-to-videoImage to VideoAInspect
Animate static images into videos. Takes input image(s) and creates a video based on the prompt. Returns a request ID that can be used with fetch-video to retrieve results.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | No | Text description of the video motion/transformation. | |
| webhook | No | URL to receive webhook notification when generation completes. | |
| duration | No | Video duration in seconds (minimum 4). | |
| model_id | Yes | The model ID to use for video generation. | |
| portrait | No | Generate in portrait orientation. | |
| track_id | No | Custom tracking ID for the request. | |
| init_audio | No | URL of audio to sync with the video. | |
| init_image | Yes | Input image URL or base64 string to animate. | |
| resolution | No | Output resolution preset. | |
| aspect_ratio | No | Aspect ratio for the video. | |
| negative_prompt | No | Things to avoid in the generated video. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare openWorldHint=true, so the description carries most of the burden. It correctly discloses the asynchronous pattern: a request ID is returned and fetch-video must be used to retrieve results. It omits details on the webhook parameter's interplay with polling, cost, or timing, but the key async contract is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with what the tool does, then input, then the return contract. No filler, though the leading whitespace/trailing structure is slightly loose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter async generation tool with no output schema, the description covers the essential return contract (request ID + fetch-video). It leaves out how webhook interacts with polling, model discovery via list-models, and any notion of latency or cost, leaving the agent to infer the full workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 11 parameters are already documented in the schema, establishing a baseline of 3. The description adds no parameter-level detail (e.g., model_id selection or duration constraints) beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: animate static images into videos, and clarifies the input is image(s) with a prompt. This implicitly separates it from text-to-video and video-to-video, though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The retrieval path is spelled out (use fetch-video with the returned request ID), which is genuinely useful workflow guidance. However, there is no guidance on when to choose this over sibling tools like image-to-image, video-to-video, or motion-control.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inpaint-imageInpaint ImageAInspect
Edit specific areas of images using masks. Provide an image, a mask indicating the area to edit, and a prompt describing the desired changes. Returns a request ID that can be used with fetch-image to retrieve results.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Text description of what to paint in the masked area. | |
| webhook | No | URL to receive webhook notification when generation completes. | |
| model_id | Yes | The model ID to use for inpainting. | |
| strength | No | Strength of the transformation (0-1). Higher values mean more change. | |
| track_id | No | Custom tracking ID for the request. | |
| init_image | Yes | URL or base64 string of the input image to edit. | |
| mask_image | Yes | URL or base64 string of the mask image (white areas will be edited, black areas preserved). | |
| negative_prompt | No | Things to avoid in the generated content. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only openWorldHint=true in annotations, the description carries most of the behavioral burden and delivers the key trait: this is asynchronous, returning a request ID rather than the image, retrieved later via fetch-image. It omits cost, rate limits, auth, and whether status can be polled instead of fetched.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short lines, front-loaded with the core action, then the required inputs, then the return-path. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter generation tool with full schema coverage and no output schema, the description covers the essential lifecycle: inputs required, async return, and retrieval tool. Minor gap in not mentioning the webhook alternative to fetching results, but nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 8 parameters are already documented in the schema, establishing a baseline of 3. The description restates the mask/prompt/image triad and confirms the mask semantics (edit vs preserve), adding little beyond what the schema already states verbatim.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Edit specific areas of images using masks') and immediately disambiguates from text-to-image/image-to-image by naming the mask-based workflow. An agent can identify this as the masked-region editing tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives the required operating context (supply image, mask, prompt) and names the follow-up tool fetch-image for retrieving results, which is clear temporal guidance. It stops short of stating when NOT to use it, e.g. versus image-to-image for whole-image edits or song-inpaint for audio.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lip-syncLip SyncAInspect
Sync video with audio for lip movements. Takes a video and audio file and syncs the lip movements to match the audio. Returns a request ID that can be used with fetch-video to retrieve results.
| Name | Required | Description | Default |
|---|---|---|---|
| webhook | No | URL to receive webhook notification when processing completes. | |
| model_id | Yes | The model ID to use for lip sync. | |
| track_id | No | Custom tracking ID for the request. | |
| init_audio | Yes | URL or base64 string of the audio file to sync lips with. | |
| init_video | Yes | URL or base64 string of the input video containing the face to sync. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only openWorldHint=true declared, the description usefully discloses the asynchronous execution pattern: a request ID is returned and must be redeemed via fetch-video, and the webhook parameter signals optional completion notification. It still omits processing duration, failure behavior, and any auth or cost considerations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and the return/retrieval flow. Sentence one and two slightly overlap, but nothing is wasted and the async contract comes before supporting detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an async generation tool with no output schema, the description adequately covers the call-then-fetch lifecycle. It stops short of guidance on choosing model_id, input size or format limits, or what happens on failure, which an agent would need to call this reliably.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters (including model_id, init_video, init_audio) are already documented in the schema. The description adds no format, size, or model-selection detail beyond what the schema provides, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: syncing video lip movements to an audio track. It is clear what the tool produces, but it does not distinguish itself from the closely related 'dubbing' sibling, which an agent could easily confuse it with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: it says the tool takes a video plus audio and returns a request ID for fetch-video, which routes the agent to the retrieval sibling. There is no explicit when-to-use vs when-not guidance or any mention of alternatives such as dubbing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list-modelsList ModelsARead-onlyInspect
List available AI models on the ModelsLab platform. Filter by category (imagen, video, audio, llm, 3d), provider, tags, and more. Returns model IDs that can be used with generation tools.
| Name | Required | Description | Default |
|---|---|---|---|
| nsfw | No | Set to false to exclude NSFW models, true to include them. Defaults to user preference. | |
| sort | No | Sort order: "recommended" (default), "latest", "most-used". | recommended |
| limit | No | Maximum number of models to return (1-100). | |
| search | No | Search models by name, ID, description, or tags. | |
| feature | No | Filter by product feature: "imagen" (images), "videofusion" (videos), "audiogen" (audio/voice), "llmaster" (LLMs), "threedverse" (3D). | |
| category | No | Filter by model category (e.g., "stable_diffusion", "stable_diffusion_xl", "flux", "llm", "video", "voice_cloning"). | |
| provider | No | Filter by model provider (e.g., "modelslab", "civitai"). | |
| subcategory | No | Filter by model subcategory (e.g., "lora", "controlnet", "embeddings", "checkpoint"). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true already establishes this as a safe read, so the description's main added value is that it returns model IDs intended for downstream generation tools. It does not disclose pagination/limit interaction or whether the result set is exhaustive, so the added context is modest.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and resource, then filters, then return value. No filler, though the middle sentence's filter list is loose ('and more').
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With readOnlyHint covering the safety profile, 100% schema coverage for all 8 optional params, and the description explaining what the call yields, an agent has enough to invoke it correctly. The absence of an output schema is partly offset by the description naming the returned model IDs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 8 parameters and the baseline is 3. The description's 'category (imagen, video, audio, llm, 3d)' actually mirrors the feature parameter's domain rather than the category parameter's values (stable_diffusion, flux, etc.), which introduces mild ambiguity instead of adding clarifying meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List available AI models on the ModelsLab platform') and clarifies the value of the output ('model IDs that can be used with generation tools'). It does not explicitly distinguish itself from the sibling list-providers, leaving that differentiation to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by noting the returned IDs feed generation tools, and it gestures at the available filters, but it never states when to prefer this tool over list-providers or how to chain it into a generation call. Usage is implied rather than instructed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list-providersList ProvidersARead-onlyInspect
List all available model providers on the ModelsLab platform. Returns provider names with model counts for each. Use provider names to filter models in the list-models tool.
| Name | Required | Description | Default |
|---|---|---|---|
| feature | No | Filter providers by product feature: "imagen" (images), "videofusion" (videos), "audiogen" (audio/voice), "llmaster" (LLMs), "threedverse" (3D). | |
| category | No | Filter providers by model category (e.g., "stable_diffusion", "flux", "llm", "video"). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover only readOnlyHint=true; the description adds the return shape (provider names plus per-provider model counts), which is beyond what the annotation or schema declares. It says nothing about pagination or rate limits, but for a simple read-only enumeration those are minor gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and followed by the return content and the integration hint. Every sentence carries distinct information with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-required-parameter read-only listing tool with no output schema, the description covers action, return content, and usage chain adequately. The only omission is any acknowledgement of the optional filtering params, which the schema handles.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description does not mention the 'feature' or 'category' filters at all and even says 'List all available', so it adds no semantic guidance beyond the schema's own parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List all available model providers'), names the platform, and describes the payload ('provider names with model counts'). It also implicitly separates itself from the sibling list-models by framing providers as the input to that tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit downstream use case: 'Use provider names to filter models in the list-models tool,' which tells the agent when to reach for this tool (before filtering models). It stops short of naming when not to use it or how it relates to the generation siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
motion-controlMotion ControlBInspect
Control motion in video generation. Uses an image and video to create motion-controlled output. Returns a request ID that can be used with fetch-video to retrieve results.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Processing mode: std (standard) or pro (professional). | |
| prompt | No | Optional text prompt (max 2500 characters). | |
| webhook | No | URL to receive webhook notification when processing completes. | |
| model_id | Yes | The model ID to use for motion control. | |
| track_id | No | Custom tracking ID for the request. | |
| init_image | Yes | URL or base64 string of the input image (character/subject). | |
| init_video | Yes | URL or base64 string of the video for motion reference. | |
| keep_original_sound | No | Keep original sound from the video. | |
| character_orientation | Yes | Use character orientation from image or video. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only openWorldHint declared, the description carries most of the behavioral burden. It usefully discloses the async nature (returns a request ID) and the fetch-video handoff, but says nothing about cost, generation latency, permissions, or what happens if the optional webhook is omitted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with the core purpose front-loaded and the return/consumption detail last. There is a small amount of wasted leading whitespace and no further padding, so it is efficient without being exemplary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter, 4-required generation tool with no output schema, the description covers the essential async return contract via the request ID and fetch-video pointer. It remains thin on the required-input expectations and mode/model_id selection, leaving gaps that annotations (openWorldHint only) cannot fill.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3 — every parameter, including enums for mode, keep_original_sound, and character_orientation, is already documented in the schema. The description only restates the image/video roles, adding no syntax or format detail beyond it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Control motion in video generation') and clarifies the input combination (image + video) that produces motion-controlled output. It does not explicitly distinguish itself from close siblings like video-to-video or image-to-video, so it lands at clear-but-undifferentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides one useful workflow cue — the returned request ID is consumed by fetch-video — which implies the async two-step pattern. However, it gives no guidance on when to choose this over video-to-video, image-to-video, or lip-sync, nor any prerequisites for the required image/video inputs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
music-generationMusic GenerationBInspect
Create music from text prompts. Generates original music based on your description, optional tags, and lyrics. Returns a request ID that can be used with fetch-audio to retrieve results.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | Array of genre/style tags for the music. | |
| lyrics | No | Lyrics to include in the generated song. | |
| prompt | Yes | Text description of the music to generate. | |
| webhook | No | URL to receive webhook notification when generation completes. | |
| model_id | Yes | The model ID to use for music generation. | |
| track_id | No | Custom tracking ID for the request. | |
| music_length_ms | No | Duration of the music in milliseconds. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only include openWorldHint, so the description must carry behavioral context. It discloses that generation is asynchronous by returning a request ID to be used with fetch-audio, which is useful. However, it omits other traits such as authentication needs, rate limits, or handling of webhooks, leaving gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three concise sentences with no waste, and the core purpose is front-loaded. Every sentence adds value: what it does, what it works with, and how to retrieve results.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a music generation tool with 7 parameters and no output schema, the description covers the essential workflow: creation from prompts and retrieval via request ID. It could better highlight the asynchronous nature and webhook option, but the schema already documents those parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all parameters are already documented. The description references prompt, tags, and lyrics but adds no syntax, format, or constraint details beyond what the schema provides. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Create music from text prompts') and clarifies that it generates original music based on description, tags, and lyrics. It distinguishes itself from retrieval tools by mentioning fetch-audio, but does not explicitly differentiate from other generation siblings like sound-generation or song-extender.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no explicit when-to-use or when-not-to-use guidance. It mentions fetch-audio for retrieving results, but does not compare to alternatives or state prerequisites. Only implied usage: generate music when you have a text prompt.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
song-extenderSong ExtenderAInspect
Extend existing music tracks. Takes an existing audio file and extends it from either the beginning or end. Returns a request ID that can be used with fetch-audio to retrieve results.
| Name | Required | Description | Default |
|---|---|---|---|
| side | Yes | Which side to extend: left (beginning) or right (end). | |
| tags | No | Array of genre/style tags. | |
| lyrics | No | Lyrics for the extended portion. | |
| prompt | No | Optional text description for the extended portion. | |
| webhook | No | URL to receive webhook notification when extension completes. | |
| model_id | Yes | The model ID to use for song extension. | |
| track_id | No | Custom tracking ID for the request. | |
| init_audio | Yes | URL or base64 string of the audio file to extend. | |
| crop_duration | No | Duration to crop from the original. | |
| extend_duration | No | Duration to extend in seconds. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare openWorldHint=true, leaving the description to carry most of the behavioral burden. It usefully discloses the asynchronous contract (returns a request ID, results retrieved via fetch-audio), which an agent could not infer from the schema. It does not mention auth, cost, or duration limits, so it stops short of full transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with no filler, and the core action and its constraint are front-loaded before the retrieval note. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly compensates by explaining that the return is a request ID consumed by fetch-audio. Combined with 100% schema coverage this is nearly complete; only error/limit behavior is missing, which is minor for a generation task.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all ten parameters are already documented in the schema, including the side enum meanings. The description adds no syntax, format, or default information beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Extend existing music tracks') plus the scope constraint of extending from either the beginning or end. It is reasonably distinguishable from siblings like song-inpaint or music-generation, though it never explicitly contrasts itself with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a clear input precondition ('takes an existing audio file') and an explicit follow-up route ('returns a request ID that can be used with fetch-audio'), which is genuine workflow guidance. However, it gives no guidance on when to choose this over song-inpaint or music-generation, and no prerequisites for model_id.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
song-inpaintSong InpaintBInspect
Edit specific sections of songs. Takes an audio file and regenerates a specific section defined by start and end times. Returns a request ID that can be used with fetch-audio to retrieve results.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | Array of genre/style tags. | |
| lyrics | No | Lyrics for the regenerated section. | |
| prompt | No | Optional text description for the regenerated section. | |
| webhook | No | URL to receive webhook notification when inpainting completes. | |
| model_id | Yes | The model ID to use for song inpainting. | |
| sections | Yes | Array of 2 numbers: [start_time, end_time] in seconds for the section to regenerate. | |
| track_id | No | Custom tracking ID for the request. | |
| init_audio | Yes | URL or base64 string of the audio file to edit. | |
| instrumental | No | Generate instrumental only (no vocals). | |
| selection_crop | No | Return only the regenerated section. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide only openWorldHint=true, so the description carries most of the burden. It does disclose a genuinely useful behavioral trait — that this is asynchronous, returning a request ID retrieved later via fetch-audio — but says nothing about cost, latency, or whether regeneration is reversible/destructive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the purpose before the mechanism and the retrieval path. No filler; only slight redundancy between 'edit specific sections' and 'regenerates a specific section'.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter async generation tool with no output schema and thin annotations, the description covers the core workflow and the fetch-audio handoff, which is the most important missing piece. It leaves the webhook callback, selection_crop, and instrumental behavior to the schema, which is acceptable but not rich.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 10 parameters including sections as [start_time, end_time]. The description only restates the section semantics, adding no syntax, format, or constraint detail beyond the structured field.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('edit specific sections of songs') and clarifies the mechanism (regenerates a section defined by start/end times from an audio file). An agent can distinguish it from inpaint-image and song-extender by domain noun alone, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a clear use case (regenerating a time-bounded section of existing audio) and usefully routes the agent to fetch-audio for results. But it gives no explicit when-not guidance or comparison against close siblings like song-extender or music-generation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sound-generationSound GenerationBInspect
Generate sound effects from text descriptions. Creates audio sound effects based on your prompt. Returns a request ID that can be used with fetch-audio to retrieve results.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Text description of the sound effect to generate. | |
| webhook | No | URL to receive webhook notification when generation completes. | |
| duration | No | Duration of the sound effect in seconds. | |
| model_id | Yes | The model ID to use for sound generation. | |
| track_id | No | Custom tracking ID for the request. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only provide openWorldHint=true, so the description carries most of the burden. It does add real behavioral context: generation is asynchronous and yields a request ID, and a webhook parameter exists for completion notification. It does not disclose cost, latency, permissions, or whether generation can be cancelled.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first two sentences are near-duplicate restatements of the same fact ("Generate sound effects from text descriptions" / "Creates audio sound effects based on your prompt"), so only the third sentence carries unique information. Front-loading is fine, but roughly half the text is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter tool with full schema coverage and no output schema, the async request-ID return and fetch-audio handoff are the key missing pieces the description supplies, and it supplies them. It is silent, however, on model selection guidance and completion timing, leaving some agent-facing questions open.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and every parameter (prompt, model_id, webhook, duration, track_id) is documented in the schema, so the baseline is 3. The description adds no syntax, format, or constraint detail beyond what the schema already states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ("Generate sound effects from text descriptions") that distinguishes it from music-generation and text-to-speech siblings. However, the second sentence restates the same fact rather than sharpening the differentiation, so the boundary against music-generation is left implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use guidance or exclusions relative to siblings like music-generation. The one useful workflow cue is that the result must be retrieved via fetch-audio, which routes the agent to the follow-up tool but says nothing about when this tool is the right choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
speech-to-speechSpeech to SpeechAInspect
Voice conversion and transformation. Takes an audio file and converts it to a different voice. Returns a request ID that can be used with fetch-audio to retrieve results.
| Name | Required | Description | Default |
|---|---|---|---|
| webhook | No | URL to receive webhook notification when conversion completes. | |
| model_id | Yes | The model ID to use for voice conversion. | |
| track_id | No | Custom tracking ID for the request. | |
| voice_id | Yes | The target voice ID to convert to. | |
| init_audio | Yes | URL or base64 string of the audio file to transform. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only carry openWorldHint=true, so the description bears most of the disclosure burden. It usefully reveals the async/request-ID retrieval pattern and, implicitly via the schema, webhook support, but says nothing about auth requirements, rate limits, latency, or what happens on failure. The async behavior is the main value added.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded sentences with no filler; purpose comes first, then mechanism, then the return/retrieval contract. Minor leading whitespace/indentation is cosmetic noise but nothing is bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly supplies the return contract (a request ID retrievable via fetch-audio), which is exactly the missing piece an agent needs. It is nearly complete for a 5-param async tool, lacking only error/timeout and permission context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters are already documented in the schema, and the baseline is 3. The description adds only the general notion that init_audio is transformed to a target voice, no format or constraint detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('voice conversion', 'takes an audio file and converts it to a different voice'), which distinguishes it from text-to-speech by making clear the input is audio, not text. It does not explicitly name or contrast with the closest siblings (text-to-speech, dubbing), so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the async workflow — 'returns a request ID that can be used with fetch-audio to retrieve results' — which tells the agent it must pair this with fetch-audio. However, it gives no guidance on when to choose this over siblings like dubbing or text-to-speech, and states no exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
speech-to-textSpeech to TextAInspect
Transcribe audio to text. Takes an audio file and converts it to text transcription. Returns a request ID that can be used with fetch-audio to retrieve results.
| Name | Required | Description | Default |
|---|---|---|---|
| webhook | No | URL to receive webhook notification when transcription completes. | |
| model_id | Yes | The model ID to use for speech-to-text. | |
| track_id | No | Custom tracking ID for the request. | |
| init_audio | Yes | URL or base64 string of the audio file to transcribe. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare openWorldHint, so the description carries most of the burden, and it usefully discloses that this is an asynchronous operation returning a request ID rather than a transcript. It does not mention auth needs, rate limits, or format/length constraints on the audio, leaving some gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded sentences with the core purpose first. The second sentence ('Takes an audio file and converts it to text transcription') largely restates the first, so there is minor redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly explains the return value (a request ID) and the webhook/async flow, which is the key thing an agent needs. It is adequate for a straightforward transcription tool, though it omits supported audio formats and model selection context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are already documented in the input schema. The description adds no format, size, or model-choice guidance beyond what the schema states, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Transcribe audio to text'), which is clearly distinct from sibling synthesis tools like text-to-speech or speech-to-speech. However, it never names those siblings explicitly, so differentiation relies on the reader's inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use or when-not-to-use statement relative to alternatives such as speech-to-speech. The final sentence implies a workflow ('use with fetch-audio to retrieve results'), which gives implicit guidance on the follow-up step but not on tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
text-to-imageText to ImageAInspect
Generate images from text prompts using AI models. Returns a request ID that can be used with fetch-image to retrieve results. Supports various AI image generation models.
| Name | Required | Description | Default |
|---|---|---|---|
| width | No | Image width in pixels (512-1024). | |
| height | No | Image height in pixels (512-1024). | |
| prompt | Yes | Text description of the image to generate. | |
| samples | No | Number of images to generate (1-4). | |
| webhook | No | URL to receive webhook notification when generation completes. | |
| model_id | Yes | The model ID to use for image generation (e.g., "flux-dev", "sdxl"). | |
| track_id | No | Custom tracking ID for the request. | |
| aspect_ratio | No | Aspect ratio for the image (e.g., "1:1", "16:9", "9:16"). | |
| negative_prompt | No | Text describing what to avoid in the image. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare openWorldHint, so the description carries most of the burden — and it delivers the key behavioral trait: the call is asynchronous and returns a request ID rather than an image, retrieved later via fetch-image. It does not cover auth, rate limits, or what happens if the request fails, keeping it below a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with the core action front-loaded and the async retrieval flow second. The closing sentence ('Supports various AI image generation models') is mild filler since model_id already implies this, so it is not maximally tight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully compensates by explaining that the return value is a request ID consumed by fetch-image. Combined with 100% schema coverage and the openWorldHint annotation, an agent has enough to call it correctly; the webhook and failure behavior remain unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (width, height, samples, webhook, negative_prompt, aspect_ratio, track_id, model_id) is already documented with examples and ranges. The description adds no format, default, or interaction detail beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Generate images from text prompts') and names the mechanism ('using AI models'). This cleanly separates it from siblings like image-to-image, text-to-video, and inpaint-image, which take different inputs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description routes the agent downstream by naming fetch-image as the retrieval tool and explaining the request-ID handoff, which is genuine usage guidance for an async tool. It stops short of stating when to prefer this over alternatives such as image-to-image or a specific model family, so it lacks explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
text-to-speechText to SpeechAInspect
Convert text to natural speech audio. Takes text and generates realistic speech using the specified voice. Returns a request ID that can be used with fetch-audio to retrieve results.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | The text to convert to speech. | |
| webhook | No | URL to receive webhook notification when generation completes. | |
| model_id | Yes | The model ID to use for text-to-speech. | |
| track_id | No | Custom tracking ID for the request. | |
| voice_id | Yes | The voice ID to use for speech generation. | |
| temperature | No | Temperature for voice variation (0-1). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only supply openWorldHint, so the description carries most of the burden and it does disclose the key behavioral trait: generation is asynchronous and returns a request ID rather than audio. It omits permission/auth requirements, latency expectations, and whether the request is durable, but the async contract is the critical disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action, then the mechanism, then the retrieval path. No filler or restated name/title padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully explains that the return value is a request ID and how to use it, closing the biggest gap an agent would face. It stops short of covering error behavior or the webhook-vs-poll choice, but is sufficient to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six parameters (prompt, voice_id, model_id, webhook, track_id, temperature) are already documented in the schema. The description adds no syntax, format, or constraint details beyond what the schema provides, which is the expected baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb+resource: 'Convert text to natural speech audio,' and specifies the inputs that drive generation ('using the specified voice'). This is unambiguously distinct from the nearby speech-to-speech, speech-to-text, and text-to-video siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear workflow context by naming fetch-audio as the follow-up tool for retrieving results, which tells the agent this is an asynchronous submit step. It does not, however, explain when to prefer this over alternatives such as speech-to-speech or when a webhook should be used instead of polling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
text-to-videoText to VideoAInspect
Generate videos from text descriptions. Creates AI-generated videos based on your text prompt. Returns a request ID that can be used with fetch-video to retrieve results.
| Name | Required | Description | Default |
|---|---|---|---|
| fps | No | Frames per second for the output video. | |
| width | No | Video width in pixels (512-1024). | |
| height | No | Video height in pixels (512-1024). | |
| prompt | Yes | Text description of the video to generate. | |
| webhook | No | URL to receive webhook notification when generation completes. | |
| duration | No | Video duration in seconds (minimum 4). | |
| model_id | Yes | The model ID to use for video generation. | |
| portrait | No | Generate in portrait orientation. | |
| track_id | No | Custom tracking ID for the request. | |
| init_audio | No | URL of audio to sync with the video. | |
| resolution | No | Output resolution preset. | |
| aspect_ratio | No | Aspect ratio for the video (e.g., "16:9", "9:16"). | |
| camera_fixed | No | Keep camera position fixed during generation. | |
| enhance_prompt | No | Use AI to enhance the prompt. | |
| generate_audio | No | Generate audio for the video. | |
| negative_prompt | No | Things to avoid in the generated video. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only supply openWorldHint=true, so the description carries most of the burden. It usefully discloses that generation is asynchronous and yields a request ID rather than a video directly, which is the single most important behavioral fact here. It omits cost, latency, and rate-limit behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first two sentences are near-duplicates ('Generate videos from text descriptions' vs 'Creates AI-generated videos based on your text prompt'), wasting a line. The return-value/next-step sentence is the only one carrying new information, so the text is not front-loaded efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 16 parameters, no output schema, and only a weak openWorld annotation, the description does cover the one thing the schema cannot: that the call returns a request ID to be polled via fetch-video. But it says nothing about model choice, defaults for the 14 optional params (fps, resolution, aspect_ratio), or generation constraints, leaving real gaps for a high-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% across all 16 parameters, so the schema already documents fps, resolution, aspect_ratio, negative_prompt, etc. The description adds no parameter-level meaning beyond that, which is the expected baseline when the schema does the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Generate videos from text descriptions') and confirms it creates AI-generated video from a prompt, which cleanly separates it from siblings like image-to-video and video-to-video. It never explicitly names those alternatives, but the resource is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the async workflow by noting the result is a request ID consumed by fetch-video, which is genuinely useful context. However, it gives no guidance on when to pick this tool over image-to-video or video-to-video, nor any prerequisites (e.g., model selection via list-models).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
video-to-videoVideo to VideoBInspect
Transform existing videos with AI. Takes input video(s) and modifies them based on the prompt. Returns a request ID that can be used with fetch-video to retrieve results.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Random seed for reproducible results (0-4294967295). | |
| prompt | Yes | Text description of how to transform the video. | |
| webhook | No | URL to receive webhook notification when generation completes. | |
| duration | No | Video duration in seconds (minimum 4). | |
| model_id | Yes | The model ID to use for video transformation. | |
| track_id | No | Custom tracking ID for the request. | |
| init_image | No | Optional guidance image URLs applied in order across the input video. | |
| init_video | Yes | Input video URL to transform. | |
| aspect_ratio | No | Aspect ratio for the output video. | |
| negative_prompt | No | Things to avoid in the generated video. | |
| image_timestamps | No | Optional seconds into the video where each init_image applies, in the same order. Images without a timestamp are spread evenly across the clip. | |
| public_figure_threshold | No | Threshold for public figure detection. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only supply openWorldHint, so the description carries most of the burden. It usefully discloses the asynchronous contract (returns a request ID retrieved later via fetch-video), which is real added value. It omits cost, latency, auth, and what happens on failures, so it is adequate but incomplete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and then the async retrieval path. No filler, though the middle sentence largely restates the first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter generation tool with no output schema, the description covers the essential input-modify-then-fetch loop. It does not note that three parameters are required, nor explain model selection or duration/aspect-ratio interactions, leaving gaps an agent must fill from the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every one of the 12 parameters is already documented in the schema. The description adds no syntax, format, or interaction detail (e.g. how init_image pairs with image_timestamps) beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Transform existing videos with AI') and clarifies that it takes an input video and modifies it per a prompt. This implicitly separates it from text-to-video and image-to-video siblings, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by 'existing videos' and 'based on the prompt', and it hints at the async workflow by pointing to fetch-video. However there is no explicit when/when-not guidance or statement of prerequisites such as which model IDs are valid or the minimum duration.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
23 tool updates
- First observed
chat-completion - First observed
dubbing - First observed
fetch-audio - First observed
fetch-generation - First observed
fetch-image - First observed
fetch-video - First observed
image-to-image - First observed
image-to-video - First observed
inpaint-image - First observed
lip-sync - First observed
list-models - First observed
list-providers - First observed
motion-control - First observed
music-generation - First observed
song-extender - First observed
song-inpaint - First observed
sound-generation - First observed
speech-to-speech - First observed
speech-to-text - First observed
text-to-image - First observed
text-to-speech - First observed
text-to-video - First observed
video-to-video
Related MCP Connectors
Image, video, audio, face-swap, talking avatars and chat across 300+ AI models, one balance.
Generate image, video, audio, 3D and vector media with 100+ AI models. Pay per generation.
500+ AI models in one account: find and price models, make images, video, speech and text.
Generate images, video, music, voice and 3D through one API. 30 tools, 200+ models.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceGenerate and refine AI images/audio/video through natural conversation.409Apache 2.0
- AlicenseNot gradedqualityCmaintenanceGenerate 3D models from text or image. Browse 10K+ free 3D models. AI creative platform with API1MIT
- FlicenseNot gradedqualityDmaintenanceAll AI Models in One API 500+ AI Models: https://www.cometapi.com/-
- AlicenseAqualityCmaintenanceHosted multi-model AI media + chat MCP server. Generates images, video, audio, face-swaps and talking-avatars, and chats across 300+ models (Claude, GPT, Gemini, DeepSeek…) - all from one balance and one API key.16MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.