Skip to main content
Glama

ModelsLab

Server Details

Generate images, video, speech and music, and run LLM chat, across 10,000+ models on ModelsLab.

If you are the author of this connector, you can claim ownership by verifying the domain or GitHub account it belongs to. Claimed connector authors can inspect health checks, view analytics, and manage their listing.
Status
Healthy
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL

TDQS

A3.6/5.0

Scored across 23 tools

Disambiguation4/5

Most tools target distinct modalities and actions, but fetch-generation overlaps with the specialized fetch-image, fetch-video, and fetch-audio tools. Descriptions help clarify usage, but the redundancy could still cause confusion when an agent chooses which fetch tool to use.

Naming Consistency4/5

All tools use lowercase kebab-case and are readable, but the naming patterns vary (verb-noun, X-to-Y, noun-noun, bare noun). This is mostly consistent with minor deviations from a single verb_noun convention.

Tool Count4/5

With 23 tools covering text, image, video, audio, music, and model listing, the count is on the high side but reasonable given the platform's broad multi-modal scope. Each tool appears to target a distinct capability, though the fetch-generation redundancy suggests one tool could be removed.

Completeness4/5

The surface covers core generation, transformation, and retrieval workflows across all major modalities. Minor gaps exist, such as no explicit cancel-generation tool and limited specialized editing operations like upscaling or background removal, but these are not critical dead ends.

Available Tools

23 tools
chat-completionChat CompletionBInspect
Chat with AI language models.
Send messages to an LLM and receive AI-generated responses.
Supports various models and configuration options.
A message may carry a PDF as a {"type":"file","file":{"filename":...,"file_data":"data:application/pdf;base64,..."}} content part.
ParametersJSON Schema
NameRequiredDescriptionDefault
nNoNumber of completions to generate (1-10).
seedNoRandom seed for reproducible results.
stopNoUp to 4 sequences where the API will stop generating.
top_kNoTop-k sampling parameter.
top_pNoNucleus sampling parameter (0-1).
pluginsNoOpenRouter plugins. Send [{"id":"file-parser","pdf":{"engine":"native"}}] alongside a {"type":"file"} content part to have a PDF attachment parsed.
messagesYesArray of message objects with "role" (system/user/assistant) and "content" keys.
model_idYesThe LLM model ID to use (e.g., "gpt-4", "claude-3").
max_tokensNoMaximum number of tokens to generate.
temperatureNoSampling temperature (0-2). Higher values make output more random.
response_formatNoResponse format configuration (e.g., {"type": "json_object"}).
presence_penaltyNoPenalty for new topics (-2 to 2).
frequency_penaltyNoPenalty for frequent tokens (-2 to 2).

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide only openWorldHint=true, so the description carries most of the behavioral burden. It does add non-obvious context: PDF attachment handling via a typed file content part, tying into the plugins/file-parser mechanism. However, it says nothing about cost, streaming, rate limits, or error behavior, and 'Supports various models and configuration options' is filler.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and only four short sentences. The PDF detail earns its place, but 'Supports various models and configuration options' is a low-value sentence that could be cut.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 13 parameters, nested content objects, and no output schema, the description covers the essential action and the notable PDF attachment case but omits the shape of the response (choices, streaming), model selection cost implications, and error handling. Adequate but with clear gaps for a tool this complex.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 13 parameters are already documented, which sets the baseline at 3. The description supplements this by showing the exact file content-part shape for messages, adding real value beyond the schema, but it adds nothing for the sampling/knob parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: chat with AI language models, send messages and receive generated responses. It clearly identifies the text-LLM capability amid media-generation siblings, but never names a sibling or explicitly contrasts itself (e.g. against list-models), so differentiation is implied rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use/when-not guidance and no routing to alternatives. The only practical hint is that a message may carry a PDF, which is a capability note rather than selection guidance. Model choice is left entirely to the caller with just two example IDs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dubbingDubbingAInspect

Create dubbed audio content. Takes a video and translates/dubs the audio from one language to another. Returns a request ID that can be used with fetch-audio to retrieve results.

ParametersJSON Schema
NameRequiredDescriptionDefault
webhookNoURL to receive webhook notification when dubbing completes.
model_idYesThe model ID to use for dubbing.
track_idNoCustom tracking ID for the request.
init_videoYesURL or base64 string of the video to dub.
output_langYesTarget language code for the dubbed output.
source_langYesSource language code (e.g., "en", "es", "fr").

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare openWorldHint=true, so the description carries the rest. It usefully discloses the asynchronous job pattern: the call returns a request ID that must be redeemed via fetch-audio, which is meaningful behavioral context beyond the annotation. It does not mention cost, latency, or auth requirements, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, each earning its place: the action, the transformation, and the return value. The core purpose is front-loaded and nothing is redundant. Minor leading whitespace in the raw string is immaterial.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly explains the return value (a request ID) and how to consume it via fetch-audio, closing the biggest gap for an async tool. It does not cover failure modes, dubbing duration, or webhook behavior, which the webhook parameter implies but the description never elaborates.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all six parameters (model_id, init_video, source_lang, output_lang, webhook, track_id) are already documented in the schema. The description restates only that a video and two languages are involved and adds no format, enum, or constraint detail beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('Create dubbed audio content') and clarifies the transformation: a video's audio is translated/dubbed between languages. It also routes the agent to the follow-up tool 'fetch-audio', which helps separate it from pure text-to-speech or speech-to-speech siblings, though it never explicitly contrasts itself with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied — the agent must infer that this is the tool for dubbing a video's audio. The one concrete guideline is that the returned request ID should be used with fetch-audio, which tells the agent how to complete the workflow but not when to choose dubbing over alternatives like lip-sync or speech-to-speech.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fetch-audioFetch Audio ResultA
Read-only
Inspect

Retrieve the status and results of an audio generation request. Use the request ID returned from text-to-speech, speech-to-text, music-generation, and other audio tools.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe request ID returned from a previous audio generation call.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation already establishes this is a safe, non-mutating read. The description adds value by disclosing that the call returns both status and results, implying polling semantics, but it says nothing about what happens while a request is still pending or how failures are reported. Adequate but thin beyond the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler, and the core operation is front-loaded before the sourcing guidance. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read-only fetch tool with no output schema, the description covers the operation, the input source, and the existence of status/result data. The only gap is sibling routing against fetch-generation, which is minor given the audio scoping.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single parameter documented at 100% schema coverage, so the schema already carries the meaning. The description reinforces it by telling the agent the ID comes from an audio tool, but adds no format or constraint detail beyond that. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description pairs a specific verb ("Retrieve") with a specific resource ("status and results of an audio generation request"), which is far better than a tautology. It scopes itself to audio, which separates it from fetch-image and fetch-video, but it never explicitly distinguishes itself from the sibling fetch-generation tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The second sentence gives clear context: call this with the request ID produced by text-to-speech, speech-to-text, music-generation, and other audio tools. That is a real when-to-use signal, though there is no explicit exclusion telling the agent when to prefer fetch-generation instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fetch-generationFetch Generation StatusA
Read-only
Inspect
Check the status of a queued or processing generation request on the ModelsLab V7 API.

Use this tool when a generation tool returns a `status` of `"processing"`, and call it again
while the status stays `"processing"`. Image jobs usually finish in seconds; video and music
jobs can take several minutes, so tell the user the job is still running if it has not
finished after a few polls.

Pass the `id` from the original generation response and the `type` matching the category
of the original request (e.g. "images" for text-to-image, image-to-image, or inpaint-image).

The response will contain:
- `status`: "success", "processing", or "error"
- `output`: array of URLs when status is "success"
ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe numeric id returned by a generation tool when status was "processing".
typeYesThe generation category matching the original request: "images" for image tools, "videos" for video tools, "audios" for audio tools.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so safety is covered. The description goes beyond by disclosing polling behavior, expected latency per media type, and what the agent should tell the user while waiting — useful operational context. It stops short of noting rate limits or a max-polling cutoff.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose, then usage, then parameter guidance, then return shape — a logical order. Slightly verbose in the polling-timing paragraph, but every part earns its place given there is no output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by enumerating the return fields (`status`, `output`) and their values, and it covers the full call lifecycle including re-polling. An agent has everything needed to invoke and interpret this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description reinforces the type-to-category mapping with concrete examples ('images' for text-to-image/image-to-image/inpaint-image), which is helpful routing context, but it largely restates what the schema property descriptions already say.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Check the status of a queued or processing generation request.' Clear what it does. However, it never distinguishes itself from the fetch-image, fetch-video, and fetch-audio siblings, which an agent could easily confuse with this generic poller — the `type` parameter is the only hint at the difference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit trigger ('use this tool when a generation tool returns a status of "processing"'), explicit loop condition ('call it again while the status stays "processing"'), and explicit user-facing guidance about job duration. Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fetch-imageFetch Image ResultA
Read-only
Inspect

Retrieve the status and results of an image generation request. Use the request ID returned from text-to-image, image-to-image, or inpaint-image tools.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe request ID returned from a previous image generation call.

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true already tells the agent this is a safe, non-mutating read, so the safety burden is covered. The description adds that it returns 'status and results', hinting at an async polling pattern, but does not explain possible states (pending/failed/complete), whether polling is required, or how partial results are represented.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the primary action front-loaded and the prerequisite second. No filler, no redundancy with the title or schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description is the only place to learn what comes back, and 'status and results' is thin for a polling-style fetch tool: it omits the possible status values and what a caller should do when the generation is not yet ready. The input side is adequately covered via the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the sole 'id' parameter is already documented as the request ID from a previous image generation call. The description marginally enriches this by naming the specific tools that produce the ID, but adds no format or validation detail beyond the schema. Baseline 3 applies when the schema carries the semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (retrieve) and resource (status and results of an image generation request), which clearly separates it from generic fetchers like fetch-generation and from the other media fetchers (fetch-video, fetch-audio). It does not, however, explicitly name a sibling it is not, so the differentiation is implied by the 'image generation request' scope rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to use the request ID returned from text-to-image, image-to-image, or inpaint-image, giving the agent the prerequisite and provenance of the input. It stops short of when-not guidance (e.g., which fetch-* to use for non-image generations), so no explicit exclusions are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fetch-videoFetch Video ResultA
Read-only
Inspect

Retrieve the status and results of a video generation request. Use the request ID returned from text-to-video, image-to-video, video-to-video, lip-sync, or motion-control tools.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe request ID returned from a previous video generation call.

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true, and the description is consistent with that non-mutating profile. It notes that status and results are returned, but for an asynchronous job endpoint it omits important behavior such as in-progress/failed states or the need to poll until completion.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero padding, with the core action front-loaded before the sourcing guidance. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of describing return values, and 'status and results' is only a thin sketch of what comes back. It is adequate for invocation but incomplete about async state semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'id' parameter is fully documented as the request ID from a previous video generation call. The description largely restates that, adding only the list of generator tools that can source the ID, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Retrieve the status and results of a video generation request'), which clearly separates it from fetch-audio and fetch-image by resource type. It does not, however, distinguish itself from the sibling fetch-generation, leaving some ambiguity about which fetch endpoint applies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells the agent to call it with a request ID produced by text-to-video, image-to-video, video-to-video, lip-sync, or motion-control, which supplies clear triggering context. It stops short of stating when-not to use it or naming fetch-generation as the alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

image-to-imageImage to ImageAInspect

Transform existing images based on text prompts. Takes an input image and modifies it according to the provided prompt. Returns a request ID that can be used with fetch-image to retrieve results.

ParametersJSON Schema
NameRequiredDescriptionDefault
widthNoOutput image width in pixels (512-1024).
heightNoOutput image height in pixels (512-1024).
promptYesText description of how to transform the image.
samplesNoNumber of images to generate (1-4).
webhookNoURL to receive webhook notification when generation completes.
model_idYesThe model ID to use for image transformation.
strengthNoTransformation strength (0-1). Higher values mean more change from the original.
track_idNoCustom tracking ID for the request.
init_imageYesInput image URL or base64 string to transform.
aspect_ratioNoAspect ratio for the output image.
negative_promptNoText describing what to avoid in the output.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare openWorldHint=true, so the description carries most of the behavioral burden. It usefully discloses the async pattern (returns a request ID, results retrieved via fetch-image), but omits cost/latency, auth requirements, and what happens to the original image. This partial disclosure is adequate but thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with the core action front-loaded, followed by input semantics and the async return path. No filler or repetition of schema content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description appropriately explains that the immediate return is a request ID and that fetch-image retrieves the result, closing the main gap for an async tool. Coverage is nearly complete, missing only edge conditions such as failure or timeout behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% across 11 parameters, including ranges and defaults, so the schema does all parameter work. The description adds no extra meaning about strength, samples, or negative_prompt beyond what the schema already says, giving the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Transform existing images based on text prompts') and clarifies that an input image is modified, which implicitly separates it from text-to-image. It stops short of naming the closest siblings (text-to-image, inpaint-image, image-to-video), so an agent must infer which one fits.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The mention that the returned request ID is used with fetch-image gives a clear downstream routing hint, and 'takes an input image' implies usage. However, there is no explicit when-to-use vs when-not guidance relative to text-to-image or inpaint-image, which are the natural alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

image-to-videoImage to VideoAInspect

Animate static images into videos. Takes input image(s) and creates a video based on the prompt. Returns a request ID that can be used with fetch-video to retrieve results.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptNoText description of the video motion/transformation.
webhookNoURL to receive webhook notification when generation completes.
durationNoVideo duration in seconds (minimum 4).
model_idYesThe model ID to use for video generation.
portraitNoGenerate in portrait orientation.
track_idNoCustom tracking ID for the request.
init_audioNoURL of audio to sync with the video.
init_imageYesInput image URL or base64 string to animate.
resolutionNoOutput resolution preset.
aspect_ratioNoAspect ratio for the video.
negative_promptNoThings to avoid in the generated video.

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare openWorldHint=true, so the description carries most of the burden. It correctly discloses the asynchronous pattern: a request ID is returned and fetch-video must be used to retrieve results. It omits details on the webhook parameter's interplay with polling, cost, or timing, but the key async contract is disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with what the tool does, then input, then the return contract. No filler, though the leading whitespace/trailing structure is slightly loose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 11-parameter async generation tool with no output schema, the description covers the essential return contract (request ID + fetch-video). It leaves out how webhook interacts with polling, model discovery via list-models, and any notion of latency or cost, leaving the agent to infer the full workflow.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 11 parameters are already documented in the schema, establishing a baseline of 3. The description adds no parameter-level detail (e.g., model_id selection or duration constraints) beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: animate static images into videos, and clarifies the input is image(s) with a prompt. This implicitly separates it from text-to-video and video-to-video, though it never names those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The retrieval path is spelled out (use fetch-video with the returned request ID), which is genuinely useful workflow guidance. However, there is no guidance on when to choose this over sibling tools like image-to-image, video-to-video, or motion-control.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inpaint-imageInpaint ImageAInspect

Edit specific areas of images using masks. Provide an image, a mask indicating the area to edit, and a prompt describing the desired changes. Returns a request ID that can be used with fetch-image to retrieve results.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesText description of what to paint in the masked area.
webhookNoURL to receive webhook notification when generation completes.
model_idYesThe model ID to use for inpainting.
strengthNoStrength of the transformation (0-1). Higher values mean more change.
track_idNoCustom tracking ID for the request.
init_imageYesURL or base64 string of the input image to edit.
mask_imageYesURL or base64 string of the mask image (white areas will be edited, black areas preserved).
negative_promptNoThings to avoid in the generated content.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With only openWorldHint=true in annotations, the description carries most of the behavioral burden and delivers the key trait: this is asynchronous, returning a request ID rather than the image, retrieved later via fetch-image. It omits cost, rate limits, auth, and whether status can be polled instead of fetched.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short lines, front-loaded with the core action, then the required inputs, then the return-path. Every sentence earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter generation tool with full schema coverage and no output schema, the description covers the essential lifecycle: inputs required, async return, and retrieval tool. Minor gap in not mentioning the webhook alternative to fetching results, but nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 8 parameters are already documented in the schema, establishing a baseline of 3. The description restates the mask/prompt/image triad and confirms the mask semantics (edit vs preserve), adding little beyond what the schema already states verbatim.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Edit specific areas of images using masks') and immediately disambiguates from text-to-image/image-to-image by naming the mask-based workflow. An agent can identify this as the masked-region editing tool without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives the required operating context (supply image, mask, prompt) and names the follow-up tool fetch-image for retrieving results, which is clear temporal guidance. It stops short of stating when NOT to use it, e.g. versus image-to-image for whole-image edits or song-inpaint for audio.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lip-syncLip SyncAInspect

Sync video with audio for lip movements. Takes a video and audio file and syncs the lip movements to match the audio. Returns a request ID that can be used with fetch-video to retrieve results.

ParametersJSON Schema
NameRequiredDescriptionDefault
webhookNoURL to receive webhook notification when processing completes.
model_idYesThe model ID to use for lip sync.
track_idNoCustom tracking ID for the request.
init_audioYesURL or base64 string of the audio file to sync lips with.
init_videoYesURL or base64 string of the input video containing the face to sync.

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With only openWorldHint=true declared, the description usefully discloses the asynchronous execution pattern: a request ID is returned and must be redeemed via fetch-video, and the webhook parameter signals optional completion notification. It still omits processing duration, failure behavior, and any auth or cost considerations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action and the return/retrieval flow. Sentence one and two slightly overlap, but nothing is wasted and the async contract comes before supporting detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an async generation tool with no output schema, the description adequately covers the call-then-fetch lifecycle. It stops short of guidance on choosing model_id, input size or format limits, or what happens on failure, which an agent would need to call this reliably.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters (including model_id, init_video, init_audio) are already documented in the schema. The description adds no format, size, or model-selection detail beyond what the schema provides, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: syncing video lip movements to an audio track. It is clear what the tool produces, but it does not distinguish itself from the closely related 'dubbing' sibling, which an agent could easily confuse it with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: it says the tool takes a video plus audio and returns a request ID for fetch-video, which routes the agent to the retrieval sibling. There is no explicit when-to-use vs when-not guidance or any mention of alternatives such as dubbing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-modelsList ModelsA
Read-only
Inspect

List available AI models on the ModelsLab platform. Filter by category (imagen, video, audio, llm, 3d), provider, tags, and more. Returns model IDs that can be used with generation tools.

ParametersJSON Schema
NameRequiredDescriptionDefault
nsfwNoSet to false to exclude NSFW models, true to include them. Defaults to user preference.
sortNoSort order: "recommended" (default), "latest", "most-used".recommended
limitNoMaximum number of models to return (1-100).
searchNoSearch models by name, ID, description, or tags.
featureNoFilter by product feature: "imagen" (images), "videofusion" (videos), "audiogen" (audio/voice), "llmaster" (LLMs), "threedverse" (3D).
categoryNoFilter by model category (e.g., "stable_diffusion", "stable_diffusion_xl", "flux", "llm", "video", "voice_cloning").
providerNoFilter by model provider (e.g., "modelslab", "civitai").
subcategoryNoFilter by model subcategory (e.g., "lora", "controlnet", "embeddings", "checkpoint").

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true already establishes this as a safe read, so the description's main added value is that it returns model IDs intended for downstream generation tools. It does not disclose pagination/limit interaction or whether the result set is exhaustive, so the added context is modest.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the action and resource, then filters, then return value. No filler, though the middle sentence's filter list is loose ('and more').

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With readOnlyHint covering the safety profile, 100% schema coverage for all 8 optional params, and the description explaining what the call yields, an agent has enough to invoke it correctly. The absence of an output schema is partly offset by the description naming the returned model IDs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 8 parameters and the baseline is 3. The description's 'category (imagen, video, audio, llm, 3d)' actually mirrors the feature parameter's domain rather than the category parameter's values (stable_diffusion, flux, etc.), which introduces mild ambiguity instead of adding clarifying meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List available AI models on the ModelsLab platform') and clarifies the value of the output ('model IDs that can be used with generation tools'). It does not explicitly distinguish itself from the sibling list-providers, leaving that differentiation to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by noting the returned IDs feed generation tools, and it gestures at the available filters, but it never states when to prefer this tool over list-providers or how to chain it into a generation call. Usage is implied rather than instructed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-providersList ProvidersA
Read-only
Inspect

List all available model providers on the ModelsLab platform. Returns provider names with model counts for each. Use provider names to filter models in the list-models tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
featureNoFilter providers by product feature: "imagen" (images), "videofusion" (videos), "audiogen" (audio/voice), "llmaster" (LLMs), "threedverse" (3D).
categoryNoFilter providers by model category (e.g., "stable_diffusion", "flux", "llm", "video").

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover only readOnlyHint=true; the description adds the return shape (provider names plus per-provider model counts), which is beyond what the annotation or schema declares. It says nothing about pagination or rate limits, but for a simple read-only enumeration those are minor gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action and followed by the return content and the integration hint. Every sentence carries distinct information with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-required-parameter read-only listing tool with no output schema, the description covers action, return content, and usage chain adequately. The only omission is any acknowledgement of the optional filtering params, which the schema handles.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description does not mention the 'feature' or 'category' filters at all and even says 'List all available', so it adds no semantic guidance beyond the schema's own parameter documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List all available model providers'), names the platform, and describes the payload ('provider names with model counts'). It also implicitly separates itself from the sibling list-models by framing providers as the input to that tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit downstream use case: 'Use provider names to filter models in the list-models tool,' which tells the agent when to reach for this tool (before filtering models). It stops short of naming when not to use it or how it relates to the generation siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

motion-controlMotion ControlBInspect

Control motion in video generation. Uses an image and video to create motion-controlled output. Returns a request ID that can be used with fetch-video to retrieve results.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoProcessing mode: std (standard) or pro (professional).
promptNoOptional text prompt (max 2500 characters).
webhookNoURL to receive webhook notification when processing completes.
model_idYesThe model ID to use for motion control.
track_idNoCustom tracking ID for the request.
init_imageYesURL or base64 string of the input image (character/subject).
init_videoYesURL or base64 string of the video for motion reference.
keep_original_soundNoKeep original sound from the video.
character_orientationYesUse character orientation from image or video.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With only openWorldHint declared, the description carries most of the behavioral burden. It usefully discloses the async nature (returns a request ID) and the fetch-video handoff, but says nothing about cost, generation latency, permissions, or what happens if the optional webhook is omitted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with the core purpose front-loaded and the return/consumption detail last. There is a small amount of wasted leading whitespace and no further padding, so it is efficient without being exemplary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter, 4-required generation tool with no output schema, the description covers the essential async return contract via the request ID and fetch-video pointer. It remains thin on the required-input expectations and mode/model_id selection, leaving gaps that annotations (openWorldHint only) cannot fill.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3 — every parameter, including enums for mode, keep_original_sound, and character_orientation, is already documented in the schema. The description only restates the image/video roles, adding no syntax or format detail beyond it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Control motion in video generation') and clarifies the input combination (image + video) that produces motion-controlled output. It does not explicitly distinguish itself from close siblings like video-to-video or image-to-video, so it lands at clear-but-undifferentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides one useful workflow cue — the returned request ID is consumed by fetch-video — which implies the async two-step pattern. However, it gives no guidance on when to choose this over video-to-video, image-to-video, or lip-sync, nor any prerequisites for the required image/video inputs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

music-generationMusic GenerationBInspect

Create music from text prompts. Generates original music based on your description, optional tags, and lyrics. Returns a request ID that can be used with fetch-audio to retrieve results.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagsNoArray of genre/style tags for the music.
lyricsNoLyrics to include in the generated song.
promptYesText description of the music to generate.
webhookNoURL to receive webhook notification when generation completes.
model_idYesThe model ID to use for music generation.
track_idNoCustom tracking ID for the request.
music_length_msNoDuration of the music in milliseconds.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only include openWorldHint, so the description must carry behavioral context. It discloses that generation is asynchronous by returning a request ID to be used with fetch-audio, which is useful. However, it omits other traits such as authentication needs, rate limits, or handling of webhooks, leaving gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three concise sentences with no waste, and the core purpose is front-loaded. Every sentence adds value: what it does, what it works with, and how to retrieve results.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a music generation tool with 7 parameters and no output schema, the description covers the essential workflow: creation from prompts and retrieval via request ID. It could better highlight the asynchronous nature and webhook option, but the schema already documents those parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all parameters are already documented. The description references prompt, tags, and lyrics but adds no syntax, format, or constraint details beyond what the schema provides. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Create music from text prompts') and clarifies that it generates original music based on description, tags, and lyrics. It distinguishes itself from retrieval tools by mentioning fetch-audio, but does not explicitly differentiate from other generation siblings like sound-generation or song-extender.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no explicit when-to-use or when-not-to-use guidance. It mentions fetch-audio for retrieving results, but does not compare to alternatives or state prerequisites. Only implied usage: generate music when you have a text prompt.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

song-extenderSong ExtenderAInspect

Extend existing music tracks. Takes an existing audio file and extends it from either the beginning or end. Returns a request ID that can be used with fetch-audio to retrieve results.

ParametersJSON Schema
NameRequiredDescriptionDefault
sideYesWhich side to extend: left (beginning) or right (end).
tagsNoArray of genre/style tags.
lyricsNoLyrics for the extended portion.
promptNoOptional text description for the extended portion.
webhookNoURL to receive webhook notification when extension completes.
model_idYesThe model ID to use for song extension.
track_idNoCustom tracking ID for the request.
init_audioYesURL or base64 string of the audio file to extend.
crop_durationNoDuration to crop from the original.
extend_durationNoDuration to extend in seconds.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare openWorldHint=true, leaving the description to carry most of the behavioral burden. It usefully discloses the asynchronous contract (returns a request ID, results retrieved via fetch-audio), which an agent could not infer from the schema. It does not mention auth, cost, or duration limits, so it stops short of full transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with no filler, and the core action and its constraint are front-loaded before the retrieval note. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly compensates by explaining that the return is a request ID consumed by fetch-audio. Combined with 100% schema coverage this is nearly complete; only error/limit behavior is missing, which is minor for a generation task.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all ten parameters are already documented in the schema, including the side enum meanings. The description adds no syntax, format, or default information beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Extend existing music tracks') plus the scope constraint of extending from either the beginning or end. It is reasonably distinguishable from siblings like song-inpaint or music-generation, though it never explicitly contrasts itself with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a clear input precondition ('takes an existing audio file') and an explicit follow-up route ('returns a request ID that can be used with fetch-audio'), which is genuine workflow guidance. However, it gives no guidance on when to choose this over song-inpaint or music-generation, and no prerequisites for model_id.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

song-inpaintSong InpaintBInspect

Edit specific sections of songs. Takes an audio file and regenerates a specific section defined by start and end times. Returns a request ID that can be used with fetch-audio to retrieve results.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagsNoArray of genre/style tags.
lyricsNoLyrics for the regenerated section.
promptNoOptional text description for the regenerated section.
webhookNoURL to receive webhook notification when inpainting completes.
model_idYesThe model ID to use for song inpainting.
sectionsYesArray of 2 numbers: [start_time, end_time] in seconds for the section to regenerate.
track_idNoCustom tracking ID for the request.
init_audioYesURL or base64 string of the audio file to edit.
instrumentalNoGenerate instrumental only (no vocals).
selection_cropNoReturn only the regenerated section.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide only openWorldHint=true, so the description carries most of the burden. It does disclose a genuinely useful behavioral trait — that this is asynchronous, returning a request ID retrieved later via fetch-audio — but says nothing about cost, latency, or whether regeneration is reversible/destructive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the purpose before the mechanism and the retrieval path. No filler; only slight redundancy between 'edit specific sections' and 'regenerates a specific section'.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter async generation tool with no output schema and thin annotations, the description covers the core workflow and the fetch-audio handoff, which is the most important missing piece. It leaves the webhook callback, selection_crop, and instrumental behavior to the schema, which is acceptable but not rich.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 10 parameters including sections as [start_time, end_time]. The description only restates the section semantics, adding no syntax, format, or constraint detail beyond the structured field.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('edit specific sections of songs') and clarifies the mechanism (regenerates a section defined by start/end times from an audio file). An agent can distinguish it from inpaint-image and song-extender by domain noun alone, though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies a clear use case (regenerating a time-bounded section of existing audio) and usefully routes the agent to fetch-audio for results. But it gives no explicit when-not guidance or comparison against close siblings like song-extender or music-generation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sound-generationSound GenerationBInspect

Generate sound effects from text descriptions. Creates audio sound effects based on your prompt. Returns a request ID that can be used with fetch-audio to retrieve results.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesText description of the sound effect to generate.
webhookNoURL to receive webhook notification when generation completes.
durationNoDuration of the sound effect in seconds.
model_idYesThe model ID to use for sound generation.
track_idNoCustom tracking ID for the request.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only provide openWorldHint=true, so the description carries most of the burden. It does add real behavioral context: generation is asynchronous and yields a request ID, and a webhook parameter exists for completion notification. It does not disclose cost, latency, permissions, or whether generation can be cancelled.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The first two sentences are near-duplicate restatements of the same fact ("Generate sound effects from text descriptions" / "Creates audio sound effects based on your prompt"), so only the third sentence carries unique information. Front-loading is fine, but roughly half the text is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter tool with full schema coverage and no output schema, the async request-ID return and fetch-audio handoff are the key missing pieces the description supplies, and it supplies them. It is silent, however, on model selection guidance and completion timing, leaving some agent-facing questions open.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and every parameter (prompt, model_id, webhook, duration, track_id) is documented in the schema, so the baseline is 3. The description adds no syntax, format, or constraint detail beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ("Generate sound effects from text descriptions") that distinguishes it from music-generation and text-to-speech siblings. However, the second sentence restates the same fact rather than sharpening the differentiation, so the boundary against music-generation is left implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use guidance or exclusions relative to siblings like music-generation. The one useful workflow cue is that the result must be retrieved via fetch-audio, which routes the agent to the follow-up tool but says nothing about when this tool is the right choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

speech-to-speechSpeech to SpeechAInspect

Voice conversion and transformation. Takes an audio file and converts it to a different voice. Returns a request ID that can be used with fetch-audio to retrieve results.

ParametersJSON Schema
NameRequiredDescriptionDefault
webhookNoURL to receive webhook notification when conversion completes.
model_idYesThe model ID to use for voice conversion.
track_idNoCustom tracking ID for the request.
voice_idYesThe target voice ID to convert to.
init_audioYesURL or base64 string of the audio file to transform.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only carry openWorldHint=true, so the description bears most of the disclosure burden. It usefully reveals the async/request-ID retrieval pattern and, implicitly via the schema, webhook support, but says nothing about auth requirements, rate limits, latency, or what happens on failure. The async behavior is the main value added.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short, front-loaded sentences with no filler; purpose comes first, then mechanism, then the return/retrieval contract. Minor leading whitespace/indentation is cosmetic noise but nothing is bloated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly supplies the return contract (a request ID retrievable via fetch-audio), which is exactly the missing piece an agent needs. It is nearly complete for a 5-param async tool, lacking only error/timeout and permission context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters are already documented in the schema, and the baseline is 3. The description adds only the general notion that init_audio is transformed to a target voice, no format or constraint detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('voice conversion', 'takes an audio file and converts it to a different voice'), which distinguishes it from text-to-speech by making clear the input is audio, not text. It does not explicitly name or contrast with the closest siblings (text-to-speech, dubbing), so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the async workflow — 'returns a request ID that can be used with fetch-audio to retrieve results' — which tells the agent it must pair this with fetch-audio. However, it gives no guidance on when to choose this over siblings like dubbing or text-to-speech, and states no exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

speech-to-textSpeech to TextAInspect

Transcribe audio to text. Takes an audio file and converts it to text transcription. Returns a request ID that can be used with fetch-audio to retrieve results.

ParametersJSON Schema
NameRequiredDescriptionDefault
webhookNoURL to receive webhook notification when transcription completes.
model_idYesThe model ID to use for speech-to-text.
track_idNoCustom tracking ID for the request.
init_audioYesURL or base64 string of the audio file to transcribe.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare openWorldHint, so the description carries most of the burden, and it usefully discloses that this is an asynchronous operation returning a request ID rather than a transcript. It does not mention auth needs, rate limits, or format/length constraints on the audio, leaving some gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short, front-loaded sentences with the core purpose first. The second sentence ('Takes an audio file and converts it to text transcription') largely restates the first, so there is minor redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly explains the return value (a request ID) and the webhook/async flow, which is the key thing an agent needs. It is adequate for a straightforward transcription tool, though it omits supported audio formats and model selection context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters are already documented in the input schema. The description adds no format, size, or model-choice guidance beyond what the schema states, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Transcribe audio to text'), which is clearly distinct from sibling synthesis tools like text-to-speech or speech-to-speech. However, it never names those siblings explicitly, so differentiation relies on the reader's inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use or when-not-to-use statement relative to alternatives such as speech-to-speech. The final sentence implies a workflow ('use with fetch-audio to retrieve results'), which gives implicit guidance on the follow-up step but not on tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

text-to-imageText to ImageAInspect

Generate images from text prompts using AI models. Returns a request ID that can be used with fetch-image to retrieve results. Supports various AI image generation models.

ParametersJSON Schema
NameRequiredDescriptionDefault
widthNoImage width in pixels (512-1024).
heightNoImage height in pixels (512-1024).
promptYesText description of the image to generate.
samplesNoNumber of images to generate (1-4).
webhookNoURL to receive webhook notification when generation completes.
model_idYesThe model ID to use for image generation (e.g., "flux-dev", "sdxl").
track_idNoCustom tracking ID for the request.
aspect_ratioNoAspect ratio for the image (e.g., "1:1", "16:9", "9:16").
negative_promptNoText describing what to avoid in the image.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare openWorldHint, so the description carries most of the burden — and it delivers the key behavioral trait: the call is asynchronous and returns a request ID rather than an image, retrieved later via fetch-image. It does not cover auth, rate limits, or what happens if the request fails, keeping it below a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with the core action front-loaded and the async retrieval flow second. The closing sentence ('Supports various AI image generation models') is mild filler since model_id already implies this, so it is not maximally tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully compensates by explaining that the return value is a request ID consumed by fetch-image. Combined with 100% schema coverage and the openWorldHint annotation, an agent has enough to call it correctly; the webhook and failure behavior remain unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter (width, height, samples, webhook, negative_prompt, aspect_ratio, track_id, model_id) is already documented with examples and ranges. The description adds no format, default, or interaction detail beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Generate images from text prompts') and names the mechanism ('using AI models'). This cleanly separates it from siblings like image-to-image, text-to-video, and inpaint-image, which take different inputs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description routes the agent downstream by naming fetch-image as the retrieval tool and explaining the request-ID handoff, which is genuine usage guidance for an async tool. It stops short of stating when to prefer this over alternatives such as image-to-image or a specific model family, so it lacks explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

text-to-speechText to SpeechAInspect

Convert text to natural speech audio. Takes text and generates realistic speech using the specified voice. Returns a request ID that can be used with fetch-audio to retrieve results.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesThe text to convert to speech.
webhookNoURL to receive webhook notification when generation completes.
model_idYesThe model ID to use for text-to-speech.
track_idNoCustom tracking ID for the request.
voice_idYesThe voice ID to use for speech generation.
temperatureNoTemperature for voice variation (0-1).

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only supply openWorldHint, so the description carries most of the burden and it does disclose the key behavioral trait: generation is asynchronous and returns a request ID rather than audio. It omits permission/auth requirements, latency expectations, and whether the request is durable, but the async contract is the critical disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action, then the mechanism, then the retrieval path. No filler or restated name/title padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully explains that the return value is a request ID and how to use it, closing the biggest gap an agent would face. It stops short of covering error behavior or the webhook-vs-poll choice, but is sufficient to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all six parameters (prompt, voice_id, model_id, webhook, track_id, temperature) are already documented in the schema. The description adds no syntax, format, or constraint details beyond what the schema provides, which is the expected baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise verb+resource: 'Convert text to natural speech audio,' and specifies the inputs that drive generation ('using the specified voice'). This is unambiguously distinct from the nearby speech-to-speech, speech-to-text, and text-to-video siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear workflow context by naming fetch-audio as the follow-up tool for retrieving results, which tells the agent this is an asynchronous submit step. It does not, however, explain when to prefer this over alternatives such as speech-to-speech or when a webhook should be used instead of polling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

text-to-videoText to VideoAInspect

Generate videos from text descriptions. Creates AI-generated videos based on your text prompt. Returns a request ID that can be used with fetch-video to retrieve results.

ParametersJSON Schema
NameRequiredDescriptionDefault
fpsNoFrames per second for the output video.
widthNoVideo width in pixels (512-1024).
heightNoVideo height in pixels (512-1024).
promptYesText description of the video to generate.
webhookNoURL to receive webhook notification when generation completes.
durationNoVideo duration in seconds (minimum 4).
model_idYesThe model ID to use for video generation.
portraitNoGenerate in portrait orientation.
track_idNoCustom tracking ID for the request.
init_audioNoURL of audio to sync with the video.
resolutionNoOutput resolution preset.
aspect_ratioNoAspect ratio for the video (e.g., "16:9", "9:16").
camera_fixedNoKeep camera position fixed during generation.
enhance_promptNoUse AI to enhance the prompt.
generate_audioNoGenerate audio for the video.
negative_promptNoThings to avoid in the generated video.

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only supply openWorldHint=true, so the description carries most of the burden. It usefully discloses that generation is asynchronous and yields a request ID rather than a video directly, which is the single most important behavioral fact here. It omits cost, latency, and rate-limit behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The first two sentences are near-duplicates ('Generate videos from text descriptions' vs 'Creates AI-generated videos based on your text prompt'), wasting a line. The return-value/next-step sentence is the only one carrying new information, so the text is not front-loaded efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 16 parameters, no output schema, and only a weak openWorld annotation, the description does cover the one thing the schema cannot: that the call returns a request ID to be polled via fetch-video. But it says nothing about model choice, defaults for the 14 optional params (fps, resolution, aspect_ratio), or generation constraints, leaving real gaps for a high-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% across all 16 parameters, so the schema already documents fps, resolution, aspect_ratio, negative_prompt, etc. The description adds no parameter-level meaning beyond that, which is the expected baseline when the schema does the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Generate videos from text descriptions') and confirms it creates AI-generated video from a prompt, which cleanly separates it from siblings like image-to-video and video-to-video. It never explicitly names those alternatives, but the resource is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the async workflow by noting the result is a request ID consumed by fetch-video, which is genuinely useful context. However, it gives no guidance on when to pick this tool over image-to-video or video-to-video, nor any prerequisites (e.g., model selection via list-models).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

video-to-videoVideo to VideoBInspect

Transform existing videos with AI. Takes input video(s) and modifies them based on the prompt. Returns a request ID that can be used with fetch-video to retrieve results.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNoRandom seed for reproducible results (0-4294967295).
promptYesText description of how to transform the video.
webhookNoURL to receive webhook notification when generation completes.
durationNoVideo duration in seconds (minimum 4).
model_idYesThe model ID to use for video transformation.
track_idNoCustom tracking ID for the request.
init_imageNoOptional guidance image URLs applied in order across the input video.
init_videoYesInput video URL to transform.
aspect_ratioNoAspect ratio for the output video.
negative_promptNoThings to avoid in the generated video.
image_timestampsNoOptional seconds into the video where each init_image applies, in the same order. Images without a timestamp are spread evenly across the clip.
public_figure_thresholdNoThreshold for public figure detection.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only supply openWorldHint, so the description carries most of the burden. It usefully discloses the asynchronous contract (returns a request ID retrieved later via fetch-video), which is real added value. It omits cost, latency, auth, and what happens on failures, so it is adequate but incomplete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action and then the async retrieval path. No filler, though the middle sentence largely restates the first.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 12-parameter generation tool with no output schema, the description covers the essential input-modify-then-fetch loop. It does not note that three parameters are required, nor explain model selection or duration/aspect-ratio interactions, leaving gaps an agent must fill from the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every one of the 12 parameters is already documented in the schema. The description adds no syntax, format, or interaction detail (e.g. how init_image pairs with image_timestamps) beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Transform existing videos with AI') and clarifies that it takes an input video and modifies it per a prompt. This implicitly separates it from text-to-video and image-to-video siblings, though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'existing videos' and 'based on the prompt', and it hints at the async workflow by pointing to fetch-video. However there is no explicit when/when-not guidance or statement of prerequisites such as which model IDs are valid or the minimum duration.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 23 tool updates
    • First observedchat-completion
    • First observeddubbing
    • First observedfetch-audio
    • First observedfetch-generation
    • First observedfetch-image
    • First observedfetch-video
    • First observedimage-to-image
    • First observedimage-to-video
    • First observedinpaint-image
    • First observedlip-sync
    • First observedlist-models
    • First observedlist-providers
    • First observedmotion-control
    • First observedmusic-generation
    • First observedsong-extender
    • First observedsong-inpaint
    • First observedsound-generation
    • First observedspeech-to-speech
    • First observedspeech-to-text
    • First observedtext-to-image
    • First observedtext-to-speech
    • First observedtext-to-video
    • First observedvideo-to-video

Related MCP Connectors

Related MCP Servers

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources