Skip to main content
Glama

Server Details

Generate and edit images, videos, and audio with 150+ models from 20+ vendors.

Status
Healthy
Last Tested
Transport
Streamable HTTP
URL

Glama MCP Gateway

Connect through Glama MCP Gateway for full control over tool access and complete visibility into every call.

MCP client
Glama
MCP server

Full call logging

Every tool call is logged with complete inputs and outputs, so you can debug issues and audit what your agents are doing.

Tool access control

Enable or disable individual tools per connector, so you decide what your agents can and cannot do.

Managed credentials

Glama handles OAuth flows, token storage, and automatic rotation, so credentials never expire on your clients.

Usage analytics

See which tools your agents call, how often, and when, so you can understand usage patterns and catch anomalies.

100% free. Your data is private.
Tool DescriptionsA

Average 4.8/5 across 13 of 13 tools scored.

Server CoherenceA
Disambiguation5/5

Each tool has a clearly distinct purpose: dedicated tools for background removal/replacement, enhancement, vectorization, and generic generation; model discovery is split into UI vs data access; and support tools like credits, preflight, and job status are uniquely scoped. The descriptions explicitly cross-reference each other to prevent confusion.

Naming Consistency4/5

All tools share the `picsart_` prefix and snake_case style, but the second part mixes noun forms (credits, drive, job_status, model_catalog, model_params, music_studio) with verb forms (enhance, generate, preflight, vectorize) and verb_noun forms (change_bg, list_models, remove_bg). This is a minor inconsistency that does not hinder readability.

Tool Count5/5

13 tools is a well-scoped count for a full GenAI server covering image editing, generation, model browsing, metadata, cost preflight, credit tracking, drive management, and async job status. Each tool serves a distinct purpose and none feel redundant.

Completeness5/5

The tool surface covers the full lifecycle: model discovery, parameter inspection, preflight cost estimation, generation, background operations, enhancement, vectorization, job status, credit monitoring, and drive storage. The only possible gap is advanced editing (e.g., inpainting), but the generic generate tool can handle that, and the dedicated tools cover the most common workflows thoroughly.

Available Tools

52 tools
picsart_asset_reviewA
Read-onlyIdempotent
Inspect

Opens the Picsart Asset Review board: a side-by-side grid of candidate assets the user picks from, rejects, comments on, or asks for variations of. Call this right after generating SEVERAL candidate stills for one purpose — a character or product reference, a style frame, a thumbnail — instead of describing them in prose. Typical flow: run picsart_generate two to four times (or once with count > 1), then pass every resulting URL here in one call so the user can compare them at full size. Pass Picsart-hosted https URLs only; other hosts are blocked by the widget sandbox. The user's picks, per-candidate comments, and any variation requests come back as a JSON message in the conversation — read it and act on it: approve the winner as the reference for later steps, or regenerate the ones they flagged with their notes folded into the prompt. Returns the normalised board payload the widget renders. Read-only; spends no credits and works without authentication.

ParametersJSON Schema
NameRequiredDescriptionDefault
notesNoOne line of guidance shown above the grid.
purposeYesWhat the chosen asset is FOR, in a few words — e.g. "character reference for the ad", "product hero still". Shown as the board's heading so the user judges against the goal.
candidatesYesThe candidate assets to compare — 1 to 12 of them.
allowMultipleNotrue when several candidates can be kept together (e.g. a set of reference angles). Defaults to false: exactly one winner.

Output Schema

ParametersJSON Schema
NameRequiredDescription
notesNo
purposeYes
candidatesYes
allowMultipleYes
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description reinforces these by stating 'Read-only; spends no credits and works without authentication'. It adds behavioral context: the widget sandbox blocks non-Picsart URLs, returns a JSON message with user actions, and requires the agent to act on the feedback. This enriches the annotation-provided safety profile with practical interaction details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is approximately 150 words, well-structured with a clear lead sentence, followed by a usage flow, constraints, output description, and action. It is front-loaded with the core purpose. While every sentence is useful, a slight reduction in redundancy (e.g., the output description could be more concise) would improve conciseness. Overall, it earns its length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (4 parameters, 2 required, no nested objects) and the presence of annotations and an output schema, the description covers all necessary aspects: purpose, when to use, how to use, constraints, expected output, and post-invocation actions. The flow from generation to review to acting on feedback is fully described. No gaps remain for an agent to misuse the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is high. The description adds value by explaining the UI context of parameters: 'purpose' is 'Shown as the board's heading', 'notes' is 'One line of guidance', and 'candidates' must be Picsart-hosted URLs (max 12). It also clarifies the effect of 'allowMultiple' (default false, exactly one winner). This enhances the agent's understanding of how parameters affect the user experience.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Opens the Picsart Asset Review board: a side-by-side grid of candidate assets' and immediately distinguishes it from sibling tools like picsart_generate (which creates assets) and other review tools (e.g., picsart_scene_review, picsart_video_review). It clearly states the tool's function as a comparison grid for user feedback.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit guidelines are provided: 'Call this right after generating SEVERAL candidate stills for one purpose' and 'Typical flow: run picsart_generate two to four times... then pass every resulting URL here'. It also specifies URL constraints ('Pass Picsart-hosted https URLs only') and outlines the expected action after receiving the result ('read it and act on it'). This gives clear when-to-use and how-to-use instructions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_asset_sheetA
Read-onlyIdempotent
Inspect

Opens the Picsart Asset Sheet: one film asset's passport — its locked text descriptor as keyed lines, reference images by angle, state-variant tabs, and stress-test results — for the user to edit surgically and lock. Their decisions come back as a JSON message in the conversation (asset_sheet_feedback): line-level from→to descriptor edits, angle/re-render requests, new variants, and a lock/iterate/reject verdict. Call it whenever a character, location, or prop sheet needs review or locking — and never paste a descriptor at the user in chat; this board is how an asset is shown to a person. Apply descriptor edits to the named line only, keeping every other word unchanged, and treat lock as valid only when the stress-test has no failures. Pass Picsart-hosted https URLs only; other hosts are blocked by the widget sandbox. Returns the normalised passport payload the widget renders. Read-only; spends no credits and works without authentication (generation happens via picsart_generate from inside the widget).

ParametersJSON Schema
NameRequiredDescriptionDefault
tagYesThe asset's @tag, e.g. "@cal" — its id in every prompt and document.
voiceNoCharacter assets only: the voice block and its generated sample.
imagesYesReference images by angle. May be empty while the first render is pending.
scenesNoScene ids where this asset appears.
statusNoPassport status. Defaults to draft.
variantsNoState variants (wet, bloodied, costume change) — separate assets, listed as tabs.
assetTypeYes
descriptorYesThe descriptor as keyed lines — the board edits these surgically, one line at a time.
stressTestNoRepeatability battery results. Lock requires failures to be empty.
sheetVersionNoSheet version, e.g. "v1". Defaults to v1.

Output Schema

ParametersJSON Schema
NameRequiredDescription
tagYes
voiceNo
imagesYes
modelsNo
scenesNo
statusYes
variantsNo
assetTypeYes
descriptorYes
stressTestNo
sheetVersionYes
voicePresetsNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond annotations by detailing read-only behavior ('spends no credits and works without authentication'), return value ('normalised passport payload'), constraints ('Pass Picsart-hosted https URLs only'), and locking condition ('treat lock as valid only when the stress-test has no failures'). Annotations already mark readOnlyHint=true, idempotentHint=true, destructiveHint=false, and the description adds critical operational context without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively long but well-structured, front-loading the primary action and then systematically covering usage guidelines, specific instructions, constraints, and return information. Every sentence appears necessary given the tool's complexity (10 parameters, nested objects, output schema). It is appropriately sized but could be slightly tighter without losing clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's high complexity (10 parameters, nested objects, output schema, specific constraints), the description covers all essential aspects: when to call, how edits work, locking prerequisites, URL restrictions, return format, and cost/authentication status. With an output schema present and schema coverage at 90%, the description is complete enough for an agent to use the tool correctly without additional clarification.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 90%, so baseline is 3. The description adds some semantic value, such as explaining that descriptor edits are surgical ('Apply descriptor edits to the named line only') and the URL constraint, but it does not elaborate on most parameters beyond what the schema already provides. The added value is moderate but not substantial enough to raise the score above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool opens the Picsart Asset Sheet for review and locking of film assets, distinguishing it from chat-based presentation by noting 'never paste a descriptor at the user in chat; this board is how an asset is shown to a person.' It specifies the resource (asset sheet) and the action (opens for surgical editing and locking), with explicit context for character, location, or prop sheets.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to call the tool: 'Call it whenever a character, location, or prop sheet needs review or locking.' It also provides a clear negative guideline ('never paste a descriptor at the user in chat') but does not contrast directly with sibling tools like picsart_asset_review or picsart_save_asset. The guidance is clear but could be more specific about when not to use it versus alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_change_bgA
Destructive
Inspect

Replaces the background of an image with a new scene described by a prompt, keeping the foreground subject intact. Auto-picks the newest enabled Picsart change-bg model unless overridden via the model param — no need to call picsart_list_models first. Use this when the user wants to "change the background to X", "put this on a beach", "swap the background for a marble counter", or any compositing where the subject is kept and the backdrop changes. Do NOT use this to strip the background to transparency (use picsart_remove_bg), upscale or sharpen (use picsart_enhance), convert raster to SVG (use picsart_vectorize), or generate a brand-new image from scratch (use picsart_generate). Required inputs: image — a publicly-accessible URL, not a local file path — and prompt describing the new background. Optional: model to pin a specific change-bg model; preflight the explicit model id you plan to use (the default path may select recraftv3-replace-bg rather than the legacy picsart-change-bg). Example: { image: "https://example.com/product.jpg", prompt: "polished marble countertop with soft window light" }. Returns { assets, id, model, created_at, prompt, summary, why_relevant, url, results: [{ url, metadata? }], drive? } as a single JSON text block plus matching structuredContent (no resource_link blocks — the widget is the single source of visual truth, so result URLs are not duplicated as separate content blocks). id is the SDK's generation handle; metadata may include model-specific tags (e.g. exploreImageId for Recraft Explore models). Spends credits. Requires Authorization: Bearer .

ParametersJSON Schema
NameRequiredDescriptionDefault
imageYesInput image URL
modelNoOverride model ID (e.g. "recraftv3-replace-bg")
promptYesDescription of the new background

Output Schema

ParametersJSON Schema
NameRequiredDescription
idNo
urlNo
urlsNo
driveNo
modelNo
assetsYes
promptNo
videosNo
resultsNo
summaryNo
finalUrlNo
completedNo
created_atNo
iterationsNo
why_relevantNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate destructiveHint=true and readOnlyHint=false, but the description goes far beyond that by disclosing credit expenditure, required Authorization header, automatic model selection (recraftv3-replace-bg vs legacy), output structure, and the absence of resource_link blocks. It provides rich behavioral detail that annotations alone do not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is lengthy but every section serves a purpose: purpose, guidance, parameter details, output explanation, and warnings. It is front-loaded with the core function, then elaborates. Slight redundancy with the schema's required fields, but overall well-structured and dense with useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity and the presence of many sibling tools, the description is remarkably complete. It covers return format, model behavior, credits, auth, and exclusions. With an output schema present, it appropriately doesn't repeat all return fields but explains key nuances like the `id` semantic and `metadata` tags.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, giving baseline 3, but the description adds substantial meaning: image must be a publicly accessible URL (not local path), prompt describes the new background, and the model param lets you override the auto-picked newest model. It also includes a concrete example mapping parameter values, enriching semantics beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource statement: 'Replaces the background of an image with a new scene described by a prompt, keeping the foreground subject intact.' This clearly differentiates from sibling tools like picsart_remove_bg, picsart_enhance, and picsart_generate. The purpose is unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use the tool ('Use this when the user wants to "change the background to X"...') and provides a detailed list of when NOT to use it, naming exact alternatives ('use picsart_remove_bg', 'use picsart_enhance', etc.). It also mentions no need to call picsart_list_models first, adding practical usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_creditsA
Read-onlyIdempotent
Inspect

Returns the current Picsart credit balance for the authenticated user — balance (active credits available now) plus the breakdown into resettable (recurring monthly/period quota) and accumulative (top-ups and add-ons), total (active credits across both pools), and overdraftUsage (credits spent past the balance, if any). When the resettable pool has a scheduled reset, nextResetDate is the ISO timestamp of the next refill. Use this before expensive operations to warn the user when the balance is low, or after a 402 from picsart_generate to confirm the issue is credits and not something else. Do NOT use it to estimate the cost of a specific generation (use picsart_preflight); this tool only reports the balance, not per-call cost. Takes no input. Returns { balance, total, resettable, accumulative, overdraftUsage, nextResetDate? } where each number is non-negative. Requires Authorization: Bearer (per-user account data).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
totalYes
balanceYes
resettableYes
accumulativeYes
nextResetDateNo
overdraftUsageYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds value by explaining the returned fields in detail, noting the need for Authorization header, and clarifying that it takes no input. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is thorough but slightly long. It is front-loaded with the main purpose and then breaks down the fields. Every sentence is meaningful, but some technical details could be condensed slightly. Still, it is well-structured and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no parameters, a complete output schema, and rich annotations, the description leaves no gaps. It explains the meaning of each field, when to use the tool, and the authentication requirement. It is fully complete for an agent to select and invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, so the description naturally adds no parameter-level detail. With 0 parameters, the baseline is 4, and the description adequately compensates by explaining the return fields and usage context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it returns the Picsart credit balance and provides a detailed breakdown of fields (balance, resettable, accumulative, etc.). It also distinguishes from sibling tools like picsart_preflight by noting it should not be used for cost estimation, making its purpose precise and unique.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells when to use: before expensive operations to warn the user when balance is low, or after a 402 from picsart_generate. Also states when NOT to use it (for cost estimation) and points to the alternative tool (picsart_preflight). This provides clear context and exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_delete_assetA
Destructive
Inspect

INTERNAL — invoked by the Picsart Asset Sheet widget when the user deletes an asset from their library. Do NOT call this directly from chat: the assetId is a Drive folder uid that only the widget has after listing the library, and a wrong id would still be rejected by the server-side scope check. From chat, ask the user to open the library via picsart_list_assets and delete from there. On invocation it permanently removes the asset's subfolder (and the portrait, turnaround, and voice sample inside it) from Picsart Drive after verifying the target is a direct child of "Film Assets". Returns { deleted, assetId, message }. Writes to Picsart Drive; requires Authorization: Bearer .

ParametersJSON Schema
NameRequiredDescriptionDefault
assetIdYesThe asset subfolder uid (returned by list/save).

Output Schema

ParametersJSON Schema
NameRequiredDescription
assetIdYes
deletedYes
messageYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses that this is a permanent deletion ('permanently removes the asset's subfolder'), specifies what gets deleted (portrait, turnaround, voice sample), mentions server-side scope checking, and states that the operation writes to Picsart Drive. It also notes the required authorization format. Annotations already say destructiveHint=true and readOnlyHint=false, so the description adds rich behavioral context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the key information (INTERNAL, do not call directly) and provides all necessary details in a few well-structured sentences. It could be slightly more concise by removing the parenthetical 'and the portrait, turnaround, and voice sample inside it', but overall it's efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool is a single-parameter destructive action with a simple output schema, the description is complete. It covers purpose, limitations, prerequisites, and behavior. It could include a mention that the output schema shows return format ('deleted, assetId, message'), but that's already provided. The description leaves little ambiguity for an AI agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides full parameter documentation (100% coverage), so the baseline is 3. The description adds value by explaining that assetId is 'a Drive folder uid that only the widget has' and that a wrong id is rejected by scope checks, which provides contextual meaning beyond the schema's description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states this tool is for deleting an asset from the library, invoked by the Picsart Asset Sheet widget. It specifies the verb ('delete'), the resource ('asset'), and the context ('from their library'), distinguishing it well from siblings like picsart_list_assets or picsart_save_asset.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Do NOT call this directly from chat' and explains why (assetId is a Drive folder uid from the widget). It gives a clear alternative: ask the user to open the library via picsart_list_assets and delete from there. This provides both when-to-use and when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_driveA
Destructive
Inspect

Single entry point for the authenticated user's Picsart Drive. Pass action:

  • list: browse a folder. folderUid omitted = Drive root; set it to descend. Folders are returned first; files are paginated (page, pageSize<=128, optional sort, optional type filter). Set flat: true to list every file across all folders (folders are omitted in flat mode).

  • create_folder: requires name; folderUid = parent (omit for root). Optional description.

  • upload: save a file. Provide EITHER file (chat attachment) OR url+name (an HTTPS URL or an inline data: URI — data URIs are pushed to the Picsart CDN first). folderUid = destination, type = resource kind. result.url is the CDN-hosted URL of the saved file, ready to pass to picsart_generate reference params like imageUrls.

  • move: requires itemUids; targetFolderUid = destination (omit = root).

  • delete: requires itemUids; soft-deletes to trash unless permanent is true.

  • update: requires itemUid + attributes; sets custom key/value attributes on a file (e.g. { coverUrl }). Every action except upload returns the current folder listing (folders, files, page math) so the widget can render; upload skips the re-list (the destination was already explicit) and returns only result, so the browser panel does not reopen as a side effect of a plain save. Requires an authenticated Picsart session (per-user Drive content): the caller's identity is forwarded to the Drive service as a user-id header — no bearer token is sent onward.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNo
fileNo
flatNoList action: when true, returns every file across the whole Drive (all folders) instead of just the items directly inside `folderUid`. Folders are not returned in flat mode.
nameNo
pageNo
sortNo
typeNo
actionYes
itemUidNoUpdate action: uid of the file whose attributes to set.
itemUidsNo
pageSizeNo
folderUidNo
permanentNoDelete action: true permanently purges; false/omitted moves to trash (recoverable).
attributesNoUpdate action: custom key/value attributes to set on the file (e.g. `{ coverUrl }`).
descriptionNo
targetFolderUidNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
pageYes
filesYes
resultNo
foldersYes
hasNextYes
hasPrevYes
pageSizeYes
folderUidYes
breadcrumbNo
totalFilesNo
totalPagesNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint: true and readOnlyHint: false, but the description adds significant behavioral context beyond that: soft-deletes vs. permanent, authentication forwarding (user-id header, no bearer token), and side-effect differences (upload skips re-list to avoid panel reopening). The only minor gap is lack of explicit rate limit or quota info, but overall transparency is strong.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured: it opens with the tool's purpose, then uses a clear list for each action with consistent formatting. Every sentence adds value, covering action syntax, parameter options, behavioral notes, and a final authentication detail. No redundancy or filler—succinct yet comprehensive for a 6-action, 16-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (16 parameters, 6 actions, output schema present), the description addresses all major aspects: action selection, parameter usage, return behavior, authentication, and integration hints (e.g., url for picsart_generate). Annotations already provide readOnly/destructive hints, and output schema is separate, so the description adequately rounds out the context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25%, but the description compensates by explaining the role of key parameters per action (e.g., 'folderUid' for parent, 'file' vs 'url' + 'name' for upload, 'permanent' for delete). It adds context like 'flat' mode behavior and 'result.url' output, though some parameters (e.g., 'description', 'targetFolderUid') are only briefly mentioned. This goes well beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly defines 'picsart_drive' as a single entry point for the authenticated user's Drive, handling six distinct actions (list, create_folder, upload, move, delete, update). Each action is described with specific verbs and resources, effectively distinguishing the tool from siblings like 'picsart_list_assets' or 'picsart_media_upload', which target different resource types or workflows.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use each action, including syntax, optional parameters, and behavioral nuances (e.g., 'folderUid omitted = Drive root', 'flat: true' for flat listing). It also clarifies behavior for upload vs. other actions regarding re-listing, and notes that upload returns 'result.url' for use with 'picsart_generate', directly aiding tool selection among siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_enhanceA
Destructive
Inspect

Upscales and enhances an image — sharpens edges, denoises, and raises resolution by an optional scale factor. Auto-picks the newest enabled Picsart upscale / enhance model unless overridden via the model param. Use this when the user asks to "upscale", "enhance", "make it higher resolution", "sharpen", "clean up this photo", or "make this 4k". Do NOT use this to remove the background (use picsart_remove_bg), replace the background (use picsart_change_bg), convert raster to SVG (use picsart_vectorize), or generate a new image (use picsart_generate). Required input: image — a publicly-accessible URL, not a local file path. Optional: model to pin a specific enhance model, scaleFactor (e.g. 2 or 4) for upscale ratio. Example: { image: "https://example.com/photo.jpg", scaleFactor: 4 }. Returns { assets, id, model, created_at, summary, why_relevant, url, results: [{ url, metadata? }], drive? } as a single JSON text block plus matching structuredContent (no resource_link blocks — the widget is the single source of visual truth, so result URLs are not duplicated as separate content blocks). id is the SDK's generation handle; metadata may include model-specific tags. Spends credits. Requires Authorization: Bearer .

ParametersJSON Schema
NameRequiredDescriptionDefault
imageYesInput image URL
modelNoOverride model ID (e.g. "picsart-enhance")
scaleFactorNoUpscale factor (e.g. 2, 4)

Output Schema

ParametersJSON Schema
NameRequiredDescription
idNo
urlNo
urlsNo
driveNo
modelNo
assetsYes
promptNo
videosNo
resultsNo
summaryNo
finalUrlNo
completedNo
created_atNo
iterationsNo
why_relevantNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds substantial behavioral context beyond annotations: it spends credits, requires Authorization Bearer token, auto-picks the newest enabled model, returns a specific JSON structure with no duplicate resource_link blocks, and explains id and metadata semantics. Annotations already indicate destructive and non-read-only, but the description enriches this with actionable details without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but front-loaded with the core purpose, then usage triggers, exclusions, parameters, example, and output format. Every sentence adds distinct value. Slight length is justified by the tool's complexity, but it could be trimmed slightly by omitting the full output structure since an output schema exists.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 3 parameters, full schema coverage, annotations, and an output schema, the description is remarkably complete: it covers triggers, non-use cases, authentication, cost, model selection, input constraints, output shape, and special display behavior. No critical context is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema already covers 100% of parameters, but description adds crucial semantics: image must be a publicly-accessible URL not a local file path, model is optional to pin a specific model, scaleFactor examples (2, 4), and a concrete usage example. This goes beyond the schema's terse descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: "Upscales and enhances an image — sharpens edges, denoises, and raises resolution." It clearly distinguishes from siblings by naming alternative tools for background removal, background replacement, vectorization, and generation in the 'Do NOT use' section.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use: "Use this when the user asks to 'upscale', 'enhance', 'make it higher resolution', 'sharpen', 'clean up this photo', or 'make this 4k'." It also provides explicit exclusions with alternative tool names (picsart_remove_bg, picsart_change_bg, picsart_vectorize, picsart_generate), giving the agent clear decision criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_film_setupA
Read-onlyIdempotent
Inspect

Opens the Picsart Film Setup console: preset rows for genre, era, camera/film stock, lens character, texture, and tempo, plus the 60:30:10 palette fields. The widget compiles the picks client-side into the film's style prefix — the exact text pasted word for word into every prompt of that world — and the user's decision comes back as a JSON message in the conversation (film_setup_feedback: selections, stylePrefix, verdict locked|draft). Call it once per visual world at the visual-bible lock, and again with current when the user wants to revise a look. Before opening it, infer every category the script, logline, or moodboard already answers and pass those in suggested with a one-line why each — the console renders suggestions pre-selected for confirmation and expands only the genuinely open categories. Never open the console empty when the material has answers. Note the era category is the look of the IMAGE (film-stock decade), not the story's period — a period story usually suggests era "timeless" and carries its period in the location descriptors. Store the returned stylePrefix verbatim; never reword it. Returns the normalised console payload the widget renders. Read-only; spends no credits and works without authentication.

ParametersJSON Schema
NameRequiredDescriptionDefault
filmYesThe film title, shown as the console heading
worldNoWhich visual world this setup belongs to — 'main' unless the film deliberately has several worlds, each with its own style prefix.
currentNoPass the existing setup when revising, so the console opens pre-filled
suggestedNoYour inferences from the script, logline, and moodboard. ALWAYS fill this before opening the console: a category the material already answers must arrive pre-selected, not as an open question. The user confirms or corrects; only genuinely open categories render expanded.
paletteHintNoOne line of palette direction from the moodboard discussion — shown beside the palette fields as a hint, not applied automatically.

Output Schema

ParametersJSON Schema
NameRequiredDescription
filmYes
worldYes
currentNo
suggestedNo
paletteHintNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description fully discloses behavioral traits beyond annotations. It explains the read-only nature explicitly ('Read-only; spends no credits and works without authentication'), which aligns with the readOnlyHint annotation. It details the internal behavior (compiles client-side style prefix, returns JSON message), how the return value should be used ('Store the returned stylePrefix verbatim; never reword it'), and the idempotent nature implied by 'Call it once... and again with `current`...' (consistent with idempotentHint). No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose and quickly lists the console's categories. It then adds critical usage behavior, parameter guidance, and a special note on the 'era' category. While every sentence is valuable, the description is somewhat long (approximately 150 words) and could be tightened further, e.g., 'The widget compiles the picks client-side...' could be more direct. Still, it is well-structured and dense with actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of this tool (5 parameters, nested objects, a revision parameter, inference requirements), the description is remarkably complete. There is an output schema, so it need not explain return format. The description covers all critical aspects: purpose, usage timing, prerequisites, parameter semantics, behavioral notes, read-only nature, and post-call data handling. No gaps are evident for the agent to successfully invoke this tool in a correct workflow.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% coverage, so the baseline is 3. However, the description adds immense value beyond the schema. It explains the 'suggested' parameter's purpose in detail (why each category, must be pre-filled), the 'current' parameter's role in revision workflows, and the 'paletteHint' parameter's display behavior ('shown beside ... as a hint, not applied automatically'). It also gives semantic context for 'era' ('look of the IMAGE (film-stock decade), not the story's period'), which is far more than the schema's 'Preset ids per category' suggests.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states this tool opens the Picsart Film Setup console, a visual preset picker for genre, era, camera/film stock, lens character, texture, tempo, and palette. It distinguishes itself from siblings like picsart_grade_console (color grading) or picsart_shot_designer (shot composition) by focusing on establishing a film's visual world style prefix. The verb 'Opens' and the specific resource 'Picsart Film Setup console' are precise.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to call this tool ('once per visual world at the visual-bible lock'), when to revise ('again with current when the user wants to revise a look'), and crucial prerequisites ('Never open the console empty when the material has answers'). It tells the agent to infer categories from script, logline, or moodboard and pass them in 'suggested' with a 'why'. This clearly distinguishes this tool's usage from other setup or editing tools in the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_film_trackerA
Read-onlyIdempotent
Inspect

Opens the Picsart Film Tracker: the 11-stage film production pipeline as a living map — stage chips with gate status, each stage's exit checklist, and the scene×assets coverage matrix with its holes marked. The user checks items off, confirms the current gate, or overrides it with a reason; their decisions come back as a JSON message in the conversation. Call this when opening a film session ("you are here"), when judging a gate, and whenever the matrix changes — a hole in the matrix is a stop: no scene generates until its assets are locked. Pass all 11 stages every time so the map stays complete; pass the matrix during asset and generation phases. Record a confirmed gate in the film state, and record an override verbatim with its reason — an override is a debt, not a pass. Returns the normalised pipeline payload the widget renders. Read-only; spends no credits and works without authentication.

ParametersJSON Schema
NameRequiredDescriptionDefault
filmYesThe film's title
noteNoOne line of guidance shown under the header
phaseYesCurrent working phase: A development (stages 1–3), B assets (4–5), C scenes (6), D edit (7–8), E finishing (9–11)
matrixNoThe scene×assets coverage matrix (phases B–C)
stagesYesAll pipeline stages with their gate state — pass all 11 so the map is complete

Output Schema

ParametersJSON Schema
NameRequiredDescription
filmYes
noteNo
phaseYes
matrixNo
stagesYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint and destructiveHint false. The description adds value by stating 'Read-only; spends no credits and works without authentication', which expands on the annotations. It also clarifies that user decisions 'come back as a JSON message in the conversation' and that the tool 'Returns the normalised pipeline payload'. The rule 'a hole in the matrix is a stop' further explains behavioral implications. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence contributes value. It front-loads the purpose ('Opens the Picsart Film Tracker'), then covers visual elements, usage conditions, parameter instructions, and return value. It could be slightly more structured (e.g., bullet points) but remains clear and efficient for a complex tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (5 parameters with nested objects, output schema exists), the description covers purpose, when to use, parameter guidance, behavioral traits (read-only, no auth, no credits), and how to handle results. The output schema exists so return value details are not needed, but the description mentions the return type. The matrix rule and override debt concept add important contextual completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so each parameter already has a description. The description adds usage semantics beyond the schema: it instructs to 'Pass all 11 stages every time' for stages, and 'pass the matrix during asset and generation phases' for the matrix parameter. It also explains the significance of matrix holes. While not deeply enriching every parameter, it provides phase-specific context that improves correct usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Opens the Picsart Film Tracker: the 11-stage film production pipeline as a living map' and details what it shows (stage chips, gate status, checklists, matrix). This specific verb+resource combination distinguishes it from sibling tools like picsart_film_setup or picsart_generate. The purpose is unambiguous and well-differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly enumerates when to call the tool: 'when opening a film session, when judging a gate, and whenever the matrix changes'. It also gives parameter-specific guidance: 'Pass all 11 stages every time' and 'pass the matrix during asset and generation phases'. It instructs the agent on how to handle results: 'Record a confirmed gate in the film state, and record an override verbatim with its reason'. This provides comprehensive usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_generateA
Destructive
Inspect

Runs any Picsart AI model end-to-end to produce an image, video, audio, or text result. Spends credits. If you already have a model id/name in hand, skip straight to picsart_generate — no need to call picsart_list_models first. Optionally validate first via picsart_model_params (learn its inputs) and/or picsart_preflight (validate the payload and quote cost before spending credits); picsart_generate itself also rejects unsupported param values before charging. Only reach for picsart_list_models when you need to pick a model — e.g. no model was named, or the user wants to browse/compare visually via the model-picker widget. To browse or check model capabilities (e.g. supported aspect ratios) programmatically WITHOUT popping that widget, use picsart_model_catalog instead. Do NOT use this for editing operations that have dedicated tools — background removal (picsart_remove_bg), background replacement (picsart_change_bg), upscale / enhancement (picsart_enhance), or raster-to-SVG conversion (picsart_vectorize). Also do NOT use it to validate params, quote cost, or browse the catalog — those are separate tools above. Required inputs: model (id) and prompt. Model-dependent optional inputs: duration (video seconds), aspectRatio (e.g. "16:9", "9:16", "1:1"), resolution (e.g. "1080p", "4k"), count (1–10 outputs), quality, style, negativePrompt, imageUrls (for image-to-X models), videoUrl (for video-to-X), enhancePrompt, generateAudio, and extra — a free-form record for model-specific params (discover them via picsart_model_params). Example (image): { model: "flux-2-pro", prompt: "a cat in a hat", aspectRatio: "1:1", count: 1 }. Example (video): { model: "kling-v3-pro", prompt: "a cat skiing down a mountain", duration: 5, aspectRatio: "16:9" }. Returns { assets, id, model, created_at, prompt, summary, why_relevant, url, results: [{ url, metadata? }], drive? } as a single JSON text block plus matching structuredContent (no resource_link blocks — the widget is the single source of visual truth, so result URLs are not duplicated as separate content blocks). id is the SDK's generation handle; metadata may include model-specific tags (e.g. exploreImageId for Recraft Explore models). Text/LLM models (mode "text" in the catalog — e.g. gemini-3-pro, gpt-5.5, claude-*) run synchronously (async is ignored) and return the generated text as the text content block plus text in structured content. VIDEO models default to async: the call returns { job, status: "ACCEPTED" } immediately — poll picsart_job_status with the job handle until it completes (it then returns this same media payload). Never re-submit a video generation because a call seemed to hang or the host reported a timeout: the render is still running and already charged — poll instead. Pass async: false only for a video call you know finishes inside the host's window. ChatGPT renders images and videos with the Picsart media gallery UI; clients fetch the assets from URLs, never base64. Spends credits and writes to the user's Picsart Drive when the Drive option is enabled. Requires Authorization: Bearer .

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoSubmit the job and return immediately with `{ job, status: "ACCEPTED" }` instead of waiting for the result, then poll `picsart_job_status` with the returned handle. DEFAULTS TO TRUE FOR VIDEO MODELS: hosts cap a synchronous tool call at roughly a minute, and a severed sync call loses the result while the render still charges — a host "isn't responding" error on a sync video call is NOT a failed generation, the job is still running server-side with no handle left to recover it. Images and text finish fast; they stay synchronous unless you pass true.
countNoNumber of outputs
extraNoExtra model-specific params. The accepted shape varies per model — picsart_model_params returns the JSON schema for any model, and picsart_preflight can pre-check (and price) a candidate object before generation.
modelYesModel ID (e.g. "flux-2-pro", "kling-v3-pro")
styleNoStyle preset
promptYesGeneration prompt
qualityNoQuality preset
durationNoVideo duration in seconds
videoUrlNoInput video for video-to-X models
imageUrlsNoInput images for image-to-X models
resolutionNoResolution (e.g. "1080p", "4k")
aspectRatioNoAspect ratio (e.g. "16:9", "9:16", "1:1")
saveToDriveNoAuto-save the result into the user's Drive (default true). Set false when the caller persists the result itself (e.g. Music Studio saves into its own folder) to avoid a duplicate copy.
enhancePromptNoAI-enhance the prompt before generating
generateAudioNoGenerate audio track for video
negativePromptNoWhat to avoid

Output Schema

ParametersJSON Schema
NameRequiredDescription
idNo
urlNo
urlsNo
driveNo
modelNo
assetsYes
promptNo
videosNo
resultsNo
summaryNo
finalUrlNo
completedNo
created_atNo
iterationsNo
why_relevantNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite the annotations already marking `destructiveHint: true` (spends credits) and `readOnlyHint: false`, the description adds substantial behavioral context: it clarifies that the tool spends credits, writes to Picsart Drive when enabled, that video models default to async and require polling via `picsart_job_status`, and that the call will reject unsupported params before charging. It also warns not to re-submit a video generation that seems to hang, a critical transparency detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is quite long and comprehensive, but every sentence adds important context or guidance. However, it could be more front-loaded; the critical use-case guidance and examples are present, but the return structure details and async polling explanation could be condensed or structured as bullet points. At 500+ words, it is verbose though not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's high complexity (16 parameters, 2 required, nested objects, output schema exists, rich sibling set), the description is remarkably complete. It covers return format, async behavior, polling, credit spending, Drive integration, authentication, and provides both image and video examples. The output schema exists, but the description still explains the top-level keys and special cases (text models, video models).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema description coverage, the description adds moderate value beyond the schema. It groups parameters as 'required inputs' and 'model-dependent optional inputs', provides two concrete examples with different model IDs, and explains special fields like `extra` (free-form record for model-specific params). It does not explain every parameter's format or range beyond what the schema provides, but the examples and grouping significantly aid understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'runs any Picsart AI model end-to-end' and specifies the resources it produces: 'image, video, audio, or text result'. It differentiates this tool from many siblings by explicitly listing tools like `picsart_remove_bg`, `picsart_change_bg`, `picsart_enhance`, and `picsart_vectorize` that should not be used for their dedicated operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides extensive when-to-use and when-not-to-use guidance. It instructs to skip calling `picsart_list_models` if you already have a model id, and alternatively to use `picsart_model_params` or `picsart_preflight` for validation. It explicitly warns not to use this tool for editing operations with dedicated tools, and not to use it for validation, cost quoting, or catalog browsing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_grade_consoleA
Read-onlyIdempotent
Inspect

Opens the Picsart Grade Console: colour finishing for a picture-locked cut — a rail of look presets plus temperature/contrast/saturation/grain/highlights/exposure sliders. The user's decision comes back as a JSON message in the conversation. Resolve the looks with picsart_media_resolve_looks in this session first and pass them in — never a remembered look list — and pass your best-fit look for the bible's palette in suggested with a one-line why, so the console opens pre-set for confirmation instead of blank. Call with scope set to a scene id to unify that scene's shots, then once with scope "film" for the film-wide look. The widget performs no charged call itself: a "preview" verdict asks you to quote and run picsart_media_contact_sheet (charged per frame) and reopen this console with previewFrames; an "apply" verdict asks you to quote and run picsart_media_apply_look plus adjust patches on the scene. Returns the normalised console payload the widget renders. Read-only; spends no credits and works without authentication.

ParametersJSON Schema
NameRequiredDescriptionDefault
filmYesThe film title, for the header
bibleYesThe locked visual bible's palette line — shown as the reference the grade serves
looksYesThe looks the user may pick from. These MUST come from a live picsart_media_resolve_looks call in this session — never from memory; the catalog changes and a stale look id fails at apply time.
scopeNo"film" for the film-wide look (the default), or a scene id (e.g. "sc01") when unifying one scene's shots first — per-scene unification comes before the film-wide grade.
currentNoThe grade as it stands, when reopening the console after a preview
suggestedNoYour inferred starting grade — pick the resolved look that best serves the bible's palette line and pass it here so the console opens pre-set for confirmation instead of blank. `current` wins over `suggested`.
suggestedWhyNoOne line naming why the suggested look fits, e.g. "closest to the slate-blue dusk palette"
previewFramesNoContact-sheet frames rendered for the current grade, when a preview was requested

Output Schema

ParametersJSON Schema
NameRequiredDescription
filmYes
bibleYes
looksYes
scopeYes
currentNo
suggestedNo
suggestedWhyNo
previewFramesNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint and idempotentHint, but the description adds critical context: 'The widget performs no charged call itself', explains that 'preview' verdict triggers contact sheet (charged per frame) and 'apply' triggers apply_look, and notes it works without authentication. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is packed with essential workflow information, but it is somewhat long and could be more structured (e.g., separating prerequisites, steps, and return details). However, every sentence adds value, and it front-loads the core purpose. Minor loss for density over clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (8 parameters, nested objects, workflow dependencies, and an output schema), the description covers prerequisites, invocation order, verdict handling, and result payload. The output schema exists, so return value details are not needed in the description. Fully sufficient for an AI agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. However, the description adds significant value by explaining the workflow: why looks must come from live calls, how suggested/current interact, and why scope should be called per-scene then film. This exceeds what the schema alone provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states 'Opens the Picsart Grade Console: colour finishing for a picture-locked cut' and details what the console includes (look presets and sliders). It distinguishes from siblings by mentioning related tools like picsart_media_resolve_looks, picsart_media_contact_sheet, and picsart_media_apply_look, making its unique purpose clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit guidance on when and how to use the tool: must resolve looks via picsart_media_resolve_looks first, never use remembered look lists, pass a suggested look with a why line, call with scope scene id before film scope, and outlines consequences for 'preview' vs 'apply' verdicts. This clearly differentiates from alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_job_statusA
Read-onlyIdempotent
Inspect

Checks a generation job started by picsart_generate with async: true. Widget-facing: widgets poll this every few seconds with the returned job handle; assistants normally call picsart_generate synchronously and never need this tool. While running it returns { status: "ACCEPTED"|"IN_PROGRESS", progress?: { percent, estimatedSecondsLeft } }. Once finished it returns the same media payload picsart_generate would have returned ({ status: "COMPLETED", assets, results, url, ... }), or an error for FAILED/CANCELED jobs. Requires Authorization: Bearer .

ParametersJSON Schema
NameRequiredDescriptionDefault
jobYesThe job handle returned by `picsart_generate` when called with `async: true`.
modelYesModel ID the job was submitted with (e.g. "seedance-2.0")
promptNoOriginal prompt — echoed into the completed result metadata
saveToDriveNoMirror of the original call's saveToDrive flag; pass false when the caller persists the result itself.

Output Schema

ParametersJSON Schema
NameRequiredDescription
assetsYes
statusYesACCEPTED | IN_PROGRESS | COMPLETED (FAILED/CANCELED surface as tool errors)
progressNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds behavioral context beyond annotations, detailing the status objects (ACCEPTED, IN_PROGRESS with progress, COMPLETED with media, or error for FAILED/CANCELED), polling behavior, and the required Authorization header. Annotations only indicate readOnlyHint, etc., so this adds significant value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with 5 sentences, each earning its place: core purpose, usage context, return types, and authentication. It is well-structured and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (4 params, nested object, output schema exists), the description covers polling behavior, return types, relationships to other tools, and auth. It is complete for an agent to use correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with descriptions, but the description adds usage guidance such as 'pass false when the caller persists the result itself' for saveToDrive, and explains that prompt is echoed into metadata. This provides context beyond the schema, warranting a score above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it checks a generation job started by picsart_generate with async: true, specifying the resource (generation job) and the context (async polling). It distinguishes itself from siblings by noting that assistants normally use synchronous calls and don't need this tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use (widgets polling async jobs) and when not to (assistants using synchronous picsart_generate). It also mentions the return payload is the same as picsart_generate when complete, providing clear guidance on alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_list_assetsA
Read-onlyIdempotent
Inspect

Lists the user's saved Picsart film assets. Reads the "Film Assets" Drive folder, pairs each portrait file with its sibling turnaround file (when present), and parses each portrait's asset manifest. Use when the user wants to see the assets they built earlier ("show my film assets", "open the asset library"), pick one to reuse, or manage them. Do NOT use it to create or edit an asset (use picsart_asset_sheet). Takes no input. Returns the widget config (models, voicePresets) alongside { items, total } so the library can render and "Open" can drop into the sheet without a second round trip. Newest first. Read-only; spends no credits. Requires Authorization: Bearer .

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
itemsYes
totalYes
modelsYes
voicePresetsYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=true. The description adds value by stating it is 'Read-only; spends no credits' and specifies authentication needs ('Requires Authorization: Bearer <picsart_token>'). It also discloses internal behavior: reading a Drive folder, pairing files, and parsing manifests. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is 4-5 sentences, front-loaded with the main purpose. Each sentence serves a distinct role: purpose, internal behavior, usage guidance, behavioral summary, and return content. No redundant or extraneous information. Highly efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (no parameters, has output schema), the description covers all necessary aspects: purpose, when to use, behavioral traits (read-only, no credits, auth), and what it returns ('widget config alongside {items, total}'). The output schema handles return details, so no further explanation is needed. Complete for an agent to decide and invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and the input schema is empty. Schema description coverage is 100% (trivially). The description confirms 'Takes no input', which is the core semantic value. According to guidelines, 0 parameters earns a baseline of 4, and the description meets that without adding unnecessary detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states the verb 'Lists' and the resource 'user's saved Picsart film assets.' It details the reading of the 'Film Assets' Drive folder, pairing of portrait and turnaround files, and parsing of asset manifests. It also provides concrete example use cases ('show my film assets', 'open the asset library') that clearly distinguish it from siblings like `picsart_asset_sheet`.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description tells the agent exactly when to invoke this tool: 'when the user wants to see the assets they built earlier, pick one to reuse, or manage them.' It explicitly excludes creation/editing by stating 'Do NOT use it to create or edit an asset (use `picsart_asset_sheet`)'. This provides unambiguous guidance on when to use it and when not to.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_list_modelsA
Read-onlyIdempotent
Inspect

Lists Picsart AI models across ALL modes (image / video / audio / text) and renders the Picsart Studio model-picker widget so the USER can browse, compare, and pick a model visually. Each item carries id, name, mode, inputType, supportedAspectRatios/supportedResolutions (when the model declares an enum for that param) (and provider, badges, description when verbose is true). Use this when the user wants to SEE the available models or pick one themselves — especially when they have not committed to an output mode yet, or for cross-mode searches ("all flux models", "every model with image input"). To narrow to one output mode without a separate tool, pass the mode filter (image/video/audio/text) on this same tool. Ratio/resolution constraints ride along in supportedAspectRatios/supportedResolutions, so you rarely need picsart_model_params just to check whether a model supports a given aspect ratio or resolution. Do NOT use it to fetch a single model's FULL parameter schema (use picsart_model_params) or estimate per-call cost (use picsart_preflight). If you only need catalog knowledge for your own reasoning (no UI shown to the user), use picsart_model_catalog instead. Inputs (all optional): mode (filter to image/video/audio/text — text = LLM models that return generated text), provider (case-insensitive substring like "flux", "kling", "google"), acceptsImage (true → only models that take an image input — i2i, i2v, i2t), acceptsVideo (true → only models that take a video input — v2v, v2a, v2t), acceptsAudio (true → only models that take an audio input — a2v, sts), inputType (exact-match escape hatch; one of t2v/i2v/v2v/a2v/t2i/i2i/t2a/v2a/tts/sts/sfx/music/t2t/i2t/v2t), limit (1–100, default 20), verbose (default false; when true each item adds provider/badges/description). inputType codes — first letter is input modality, second is output: t2i (text→image), i2i (image→image), t2v (text→video), i2v (image→video), v2v (video→video), a2v (audio→video), t2a (text→audio), v2a (video→audio), tts (text-to-speech), sts (speech-to-speech), sfx (sound effects), music (music gen), t2t/i2t/v2t (LLM text output from text/image/video input). Example: { mode: "video", acceptsImage: true, limit: 10 } returns image-to-video models. Returns { items, total, truncated }truncated is true when more matched than were returned; refine filters or raise limit (max 100) to see more. Read-only; spends no credits and works without authentication.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoFilter by generation mode
limitNoMax items to return (1–100, default 20)
verboseNoWhen true, include provider/badges/description per item. Default false.
providerNoProvider substring (e.g. "kling", "flux", "google")
inputTypeNoExact inputType match (e.g. "i2v")
acceptsAudioNoOnly return models that accept an audio input (a2v, sts)
acceptsImageNoOnly return models that accept an image input (i2i, i2v)
acceptsVideoNoOnly return models that accept a video input (v2v, v2a)

Output Schema

ParametersJSON Schema
NameRequiredDescription
itemsYes
totalYes
truncatedYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint, idempotentHint, destructiveHint), the description discloses the tool renders a UI widget ('renders the Picsart Studio model-picker widget'), the fact it is read-only and costs no credits ('spends no credits and works without authentication'), and the truncation behavior in the return payload ('truncated is true when more matched than were returned'). This adds significant behavioral context not captured by annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but well-organized: purpose first, then usage, then parameters, then inputType legend, example, and return shape. Every sentence contributes unique information, though some parts (e.g., the detailed inputType codes) could arguably be condensed. Still, for a tool with 8 optional filters and cross-tool distinctions, the thoroughness is justified.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 8 parameters, a UI side-effect, and strong sibling differentiation, the description covers all necessary aspects: what it renders, how to filter, what the input codes mean, an example, and the truncation behavior. The output schema exists, so return values need no further explanation. The description is fully adequate for selecting and invoking this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema coverage is 100%, the description adds substantial meaning to parameters: it decodes the cryptic inputType codes (e.g., 't2v (text→video)'), explains the boolean filters in terms of model categories (acceptsImage → i2i, i2v, i2t), and provides a concrete example call. This goes well beyond the schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource+scope: 'Lists Picsart AI models across ALL modes' and explicitly distinguishes itself from siblings by naming alternatives like picsart_model_params, picsart_preflight, and picsart_model_catalog. This clearly identifies what the tool does and how it differs from related tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states precisely when to use the tool ('when the user wants to SEE the available models or pick one themselves... for cross-mode searches') and when not to use it, with explicit alternative tools for each case. It also explains how to narrow to a single mode using the `mode` filter, providing a clear decision framework.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_media_apply_effectA
Read-onlyIdempotent
Inspect

Apply a named effect (gaussian_blur, drop_shadow, stroke, ...) to a visual layer in an MP Scene, with catalog defaults filled for any omitted params. Returns the updated scene with the effect merged into the target layer's effects[] -- re-applying the same effect id REPLACES it (idempotent), so params can be tuned without stacking duplicate effects. Each param is validated by kind and numeric range; an unknown effect id, unknown param, or out-of-range value is rejected with a clear error. Discoverable effect ids and their parameter shapes (ranges, defaults) are listed under supports.effects in picsart_media_get_capabilities. Pure: takes the full scene by value, returns a new scene; no server-side state.

ParametersJSON Schema
NameRequiredDescriptionDefault
sceneYesThe MP Scene document to edit
effectYesEffect id (e.g. gaussian_blur, drop_shadow, stroke)
paramsNoOptional effect parameter overrides
layerIdYesId of the visual layer to apply the effect to
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, but the description adds substantial behavioral context: re-applying an effect replaces it in effects[] (explaining what idempotent means here), validation behavior with clear errors for unknown ids/params/ranges, and purity (no server-side state). This far exceeds the annotation baseline.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence in the description contributes meaningful information: purpose, return behavior, idempotency, validation, discoverability, and purity. There is no fluff, and the first sentence front-loads the core action. Dense but highly efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-param tool with nested objects and no output schema, the description covers return value (updated scene with effects[]), error cases (unknown id/param, out-of-range), idempotency, purity, and how to discover valid values. This gives an agent all necessary context to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% for all four parameters, but the description adds crucial semantics: omitted params get catalog defaults, params are validated by kind and numeric range, and effect ids/param shapes are discoverable via get_capabilities. This enriches the bare schema descriptions with actionable information.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Apply a named effect') to a specific resource ('a visual layer in an MP Scene') with concrete examples (gaussian_blur, drop_shadow, stroke). It clearly distinguishes from sibling tools like picsart_media_apply_look and picsart_media_apply_motion_preset by focusing on named effects and their parameter tuning.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context by explaining the effect application process, idempotent replacement, and discoverability via get_capabilities. It doesn't explicitly name alternative tools or when-not-to-use scenarios, but the context is clear enough for an agent to select this tool for named effect operations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_media_apply_lookA
Read-onlyIdempotent
Inspect

Apply a composite 'look' (e.g. vintage_bw, light_leak, shimmer) to a media OR scene_ref layer. A look wraps the layer's content into an isolated nested composition so the treatment hits the composited result as one image. TWO MODES: mode:'reference' (DEFAULT) just appends a thin entry to the layer's looks[] array ({look, params}) and KEEPS the layer's content as the look's subject — nothing is baked; translate/preview/query expand it just-in-time (resolveLooks), or call picsart_media_resolve_looks to bake a self-contained scene. This keeps the stored scene small and the look non-destructively editable (re-tune params / drop the entry / stack multiple looks). mode:'expand' eagerly bakes the full nested sub-scene now (the original behavior). Params have catalog defaults; omit params for the default look. Discoverable looks + params are under supports.looks in picsart_media_get_capabilities. NOTE: a layer carrying its own effects/animations/mask/blendMode can't take a look (the look's scene_ref wrapper can't preserve them). Pure: takes the full scene by value, returns a new scene.

ParametersJSON Schema
NameRequiredDescriptionDefault
lookYesLook id (e.g. vintage_bw, light_leak, shimmer)
modeNo"reference" (default) defers expansion; "expand" bakes the nested comp now
sceneYesThe MP Scene document to edit
paramsNoOptional look parameter overrides
layerIdYesId of the media or scene_ref layer to apply the look to
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite readOnlyHint/idempotentHint annotations, the description adds crucial behavioral context: reference mode appends to looks[] and does NOT bake, expand eagerly bakes, and the pure functional nature ('takes the full scene by value, returns a new scene'). It also discloses the limitation with layered effects. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer and dense, but it is well-structured with labeled sections (TWO MODES, NOTE, Pure) and front-loads the main purpose. Every sentence contributes meaningful information; however, the density makes it slightly heavy for quick scanning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (5 params, nested objects, no output schema), the description covers the essential aspects: modes, baking behavior, discoverability, purity, and limitations. It states the return value ('returns a new scene') even without an output schema. It does not discuss error conditions, but remains sufficient for a tool of this scope.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema already documents all 5 params with 100% coverage, but the description enriches semantics: explains mode enum values, notes that params have catalog defaults and can be omitted, and clarifies layerId targets media or scene_ref. This goes beyond the schema's standalone text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb+resource ('Apply a composite look ... to a media OR scene_ref layer') and includes concrete examples (vintage_bw, light_leak). It clearly distinguishes itself from sibling tools like apply_effect by explaining the isolated nested composition and referencing resolve_looks.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly describes two modes with clear use cases: reference (default, deferred, non-destructive) vs expand (eager bake). Provides a when-not condition (layers with effects/animations/mask/blendMode can't take a look) and points to get_capabilities for discoverable looks and resolve_looks as an alternative for baking.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_media_apply_motion_presetA
Read-onlyIdempotent
Inspect

Apply a named motion preset (glow_pulse, ken_burns, scale_pop, ...) to a visual layer in an MP Scene. Returns the updated scene with the preset's animations + effects merged into the target layer. Each preset declares an appliesTo set of layer content kinds (text/media/color) -- mismatches are rejected with a clear preset_kind_mismatch error. Discoverable presets and their parameter shapes are listed under supports.motionPresets in picsart_media_get_capabilities. Pure: takes the full scene by value, returns a new scene; no server-side state.

ParametersJSON Schema
NameRequiredDescriptionDefault
sceneYesThe MP Scene document to edit
paramsNoOptional preset parameter overrides
presetYesMotion preset id (e.g. glow_pulse, ken_burns, scale_pop)
layerIdYesId of the visual layer to animate
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint, idempotentHint, etc.), the description explicitly states the tool is 'Pure'—takes the full scene by value, returns a new scene, no server-side state. It also discloses error behavior (preset_kind_mismatch) and describes how animations+effects are merged. This adds substantial behavioral context beyond the structured hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact yet information-dense. It front-loads the main action and then covers return value, preset constraints, discovery method, and purity. Every sentence earns its place with no fluff or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (nested objects, no output schema, multiple apply-* siblings), the description is remarkably complete. It explains what the tool returns, how to discover presets, the error condition, and side-effect-free behavior. It references the capabilities tool for additional details, which is appropriate. No significant gaps remain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already covers all four parameters with descriptions (100% coverage), so the baseline is 3. The description adds value by explaining preset semantics (appliesTo kinds), error conditions, and that params are optional overrides. It also clarifies the scene is a full document, not a reference. This enriches the parameter meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool applies a named motion preset to a visual layer in an MP Scene, with specific verb ('apply'), resource ('motion preset'), and target ('visual layer'). It distinguishes itself from sibling 'apply_*' tools by focusing on motion presets and providing examples (glow_pulse, ken_burns, scale_pop). The return value (updated scene) is also stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context: the tool is for applying motion presets to layers, and it points to 'picsart_media_get_capabilities' for discovering available presets and parameter shapes. It also implies that presets have an 'appliesTo' constraint, warning about mismatches. However, it does not explicitly contrast this with sibling tools like 'apply_effect' or 'apply_look', so it lacks explicit when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_media_apply_scene_templateA
Read-onlyIdempotent
Inspect

Instantiate a scene template with the supplied parameter bindings. Two modes: reference (DEFAULT, canonical) packages an MpSceneRefContent snippet (kind:'scene_ref' with bindings + optional timeScale/trim/fit) to drop into a parent scene's layers[]; at translate/preview time the ref becomes its OWN nested composition (keeps its resolution/duration, fit/transform honoured — the template stays a reusable parameterized unit). For a STANDARD library template (mpscene://montage, …) the parent needs NO scenes entry; the translator resolves it from the registry. bootstrap returns a complete resolved standalone MpScene (substitutes every $param, synthesizes asset entries for asset-typed parameters, strips the parameters declaration) — use it to bake an editable starting scene. inline MERGES the template's resolved concrete layers + assets INTO a target scene you pass (no scene_ref) and returns it — the 'bring a preset's editable layers into my composition' op (ids prefixed so nothing collides, assetId refs rewired, brought-in layer starts offset by at). Use picsart_media_describe_scene_template first to learn what parameters the template accepts. Pure: returns either sceneRef or scene plus the applied parameter map. For the common video jobs — concatenating/merging N clips with per-clip trim and fit/align into one output — call picsart_media_quickstart with recipe:'concat_videos' FIRST. It returns a complete, validated parameters payload for mpscene://montage; you should not need picsart_media_describe_scene_template or picsart_media_get_scene_schema round-trips for that case.

ParametersJSON Schema
NameRequiredDescriptionDefault
atNoInline mode only: offset each brought-in layer's start by this many seconds
fitNoReference mode only: fit mode passed through to the scene_ref
uriYesTemplate URI, e.g. "mpscene://montage"
modeNo"reference" (default) packages a scene_ref; "bootstrap" bakes a standalone scene; "inline" merges into `scene`
trimNoReference mode only: trim window passed through to the scene_ref
sceneNoInline mode only: the scene to merge the template's editable layers into
idPrefixNoInline mode only: id-prefix for the brought-in layers/assets
timeScaleNoReference mode only: passed through to the scene_ref
parametersNoParameter bindings; keys must match the template's declared parameters
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description details the exact behavior of each mode: reference produces a scene_ref that becomes a nested composition, bootstrap strips parameters and synthesizes assets, inline merges with id-prefixing and asset rewiring. It also confirms the operation is pure and returns a sceneRef or scene, adding value beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively long (about 200 words) but every sentence conveys necessary information about modes, usage, and fallbacks. It could be slightly more concise with bullet points, but the dense content is justified for a tool with three distinct operational modes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite lacking an output schema, the description clearly states the return type ('sceneRef or scene plus the applied parameter map'), describes mode-specific side effects, and explains when to use this tool versus quickstart. It also covers required prerequisites (describing template parameters) and edge cases like template resolution from registry.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already covers all 9 parameters with descriptions (100% coverage), giving a baseline of 3. The description adds meaning by explaining which parameters apply to which mode (e.g., 'at' is inline-only, 'fit'/'trim'/'timeScale' are reference-only) and clarifying the default mode. This improves parameter understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with "Instantiate a scene template with the supplied parameter bindings," which is a specific verb+resource combination. It further clarifies the three modes (reference, bootstrap, inline) and contrasts with sibling tools like apply_effect, making it easy to distinguish.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit guidance is given: use picsart_media_describe_scene_template first to learn parameters, and for common video jobs use picsart_media_quickstart with recipe:'concat_videos' instead. It also states when no scenes entry is needed for standard library templates, covering when-to-use and alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_media_apply_text_animationA
Read-onlyIdempotent
Inspect

Apply a named text-animation preset (typewriter, fade_in_chars, slide_up_lines, ...) to a text layer in an MP Scene. Returns the updated scene with the animation appended to the layer's content.animations[]. Discoverable presets and their parameter shapes are listed under supports.textAnimationPresets in picsart_media_get_capabilities. Agents wanting custom shapes can write MpTextAnimation entries directly into a scene without going through this tool. Pure: takes the full scene by value, returns a new scene; no server-side state.

ParametersJSON Schema
NameRequiredDescriptionDefault
sceneYesThe MP Scene document to edit
paramsNoOptional preset parameter overrides
presetYesText-animation preset id (e.g. typewriter, fade_in_chars, slide_up_lines)
layerIdYesId of the text layer to animate
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint, idempotentHint, etc.), the description discloses the pure functional behavior: 'takes the full scene by value, returns a new scene; no server-side state.' It also specifies the exact mutation point (appended to content.animations[]) and the return type (updated scene), which are not implied by annotations alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is five sentences long, with each sentence serving a distinct purpose: main action, return behavior, preset discovery, alternative approach, and purity guarantee. It is front-loaded with the key action and avoids any wasteful repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (nested objects, no output schema) and rich annotations, the description thoroughly covers the essential information: what it does, how to find valid presets, what it returns, where changes occur, and when to use a different approach. No critical gaps that would prevent an agent from using it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema provides 100% coverage with descriptions for all parameters, but the description adds extra semantics by clarifying that 'params' are optional preset overrides, the 'scene' parameter is the full scene passed by value, and providing examples for the 'preset' parameter. This goes beyond mere schema repetition.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool applies a named text-animation preset to a text layer in an MP Scene, with concrete examples like 'typewriter' and 'fade_in_chars'. It distinguishes itself from siblings like picsart_media_apply_motion_preset by focusing specifically on text animations and also mentions an alternative for custom shapes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance: presets and their parameter shapes are discoverable via picsart_media_get_capabilities. It also explicitly states when not to use the tool—when custom animation shapes are needed, agents can write MpTextAnimation entries directly—thus giving a clear alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_media_contact_sheetA
Destructive
Inspect

Render a cheap multi-frame OVERVIEW of a scene as low-res jpeg thumbnails YOU CAN ACTUALLY SEE: each sampled frame comes back as an inline image content block (labeled with its scene time), alongside the machine-readable { frames: [{time, url, inline}] }. Two modes: pass times (PREFERRED — you usually know the interesting moments: clip seams, animation midpoints, entrance ends) to get an EXACT thumbnail per requested time, rendered concurrently; or pass frames (default 8, max 24) for evenly-spaced sampling across the whole timeline (one image-sequence render at 1fps — integer-second granularity only). Sits between a single full-res frame check and picsart_media_export (full encode): use it to eyeball pacing, seams, and content presence across the WHOLE timeline before exporting, instead of checking single frames repeatedly. Auth is handled by the platform automatically — no token setup needed on your end. Validate the scene first. Inline images are best-effort under a total size/count budget: a frame that fails to download, is oversized, or falls outside the budget still comes back with its url and inline: false in the JSON — a partially-inlined sheet is a SUCCESS, not an error, because the render has already happened and been charged. Do not retry this call just to get the missing pictures (that re-charges the render) — open those urls instead. Scope: this samples a scene you have already authored, to look at it. It is NOT a metadata probe — use picsart_media_probe_media first for a source clip's duration/dimensions/frame rate. Sparse-sampling frames here to discover duration is only the fallback when picsart_media_probe_media genuinely cannot determine it (every frame is a charged GPU render). Use picsart_media_query_layout to check layer geometry, overlap and stacking order without rendering. At most 8 frames come back inline per call regardless of how many you request.

ParametersJSON Schema
NameRequiredDescriptionDefault
sceneYesThe MP Scene document to sample, or an http(s):// URL to fetch it from
timesNoEXACT scene times (seconds) to thumbnail. At most 8 frames come back as inline images; requesting more no longer renders/charges the extras (see above). When set, `frames` is ignored
framesNoHow many evenly-spaced thumbnails to return (default 8, max 8 shown inline)
resolutionNoThumbnail box override (default: composition aspect scaled to width 480)

Output Schema

ParametersJSON Schema
NameRequiredDescription
idNo
urlNo
urlsNo
driveNo
modelNo
assetsYes
promptNo
videosNo
resultsNo
summaryNo
finalUrlNo
completedNo
created_atNo
iterationsNo
why_relevantNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations by disclosing auth is handled automatically, inline images are best-effort under a size/count budget, partially-inlined sheets are a success, and retrying re-charges the render. It also clarifies that at most 8 frames come back inline regardless of request count, adding crucial behavioral nuance.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but well-organized, front-loading the purpose and then systematically covering modes, usage context, auth, failure semantics, and alternatives. While each sentence adds necessary information, the length could be trimmed slightly without losing value, so it earns a 4 rather than a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers return format, charging behavior, inline vs url fallback, prerequisites (validate scene), scope, and exclusions. Despite having an output schema, the description still explains the inline vs url response and success semantics, making it fully complete for an agent to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Adds significant meaning beyond the schema: `times` yields exact thumbnails per requested time and is preferred, `frames` does evenly-spaced sampling with 1fps integer-second granularity, `times` overrides `frames`, and `resolution` defaults to width 480. This is far more than the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states it renders a multi-frame overview of a scene as low-res jpeg thumbnails with inline images and a machine-readable frames JSON. It explicitly situates itself between a single full-res frame check and full export, distinguishing it from sibling tools like picsart_media_probe_media and picsart_media_query_layout.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit guidance: prefer `times` for exact moments, use for eyeballing pacing/seams before export, and not as a metadata probe (use picsart_media_probe_media first). Also instructs to validate the scene first and names alternative tools for layout checks, giving clear when/when-not context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_media_describe_scene_templateA
Read-onlyIdempotent
Inspect

Describe a single MP Scene template by URI -- returns its declared parameters (with types, defaults, ranges, required flags, descriptions) and composition dimensions. Use this before picsart_media_apply_scene_template so you know what to pass. Accepts mpscene://<id> URIs in v1; other schemes are deferred. Returns { id?, uri, composition, parameters } where parameters is the full declarations map.

ParametersJSON Schema
NameRequiredDescriptionDefault
uriYesTemplate URI, e.g. "mpscene://montage"
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare `readOnlyHint=true`, `idempotentHint=true`, and `destructiveHint=false`. The description adds valuable behavioral constraints beyond that: the URI scheme version restriction ('in v1; other schemes are deferred') and the exact return shape (`{ id?, uri, composition, parameters }`). This is more than baseline but not exhaustive (e.g., error behavior is not described).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each dense with information: purpose, usage context, and return details. No filler or repetition. The key information is front-loaded in the first sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter describe tool with strong annotations, the description is very complete. It states what it does, when to use it (with an explicit sibling reference), what URI schemes are accepted, and what the response contains. There is no output schema, so the explicit return format `{ id?, uri, composition, parameters }` covers that gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 100% coverage with one parameter (`uri`) already described as 'Template URI, e.g. "mpscene://montage"'. The description adds further meaning by specifying the accepted scheme ('mpscene://<id>') and the version limitation ('in v1; other schemes are deferred'), which goes beyond the schema's basic example.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Describe a single MP Scene template by URI' and clearly states what it returns ('declared parameters... and composition dimensions'). It also distinguishes from the sibling tool by explicitly connecting to 'picsart_media_apply_scene_template'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit usage context: 'Use this before `picsart_media_apply_scene_template` so you know what to pass.' It also specifies URI scheme support ('Accepts `mpscene://<id>` URIs in v1; other schemes are deferred'), which tells the agent when the tool is applicable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_media_expand_scene_refA
Read-onlyIdempotent
Inspect

DETACH a scene_ref into editable layers: replace a resolvable scene_ref layer (or every resolvable one, if no layerId) with the referenced template's concrete resolved layers, inlined IN PLACE (prefixed by the ref layer's id, started at its start, assetId refs rewired). The result has no scene_ref for the expanded layers, so you can edit the brought-in layers directly. Inverse of the by-reference model: 'I referenced a template, now bake this instance to customize it beyond its parameters.' One level in v1 (a brought-in layer that is itself a scene_ref stays a reference — run again to go deeper). This is a SYNC op that does NOT fetch: it resolves only in-memory mpscene:// refs (local scenes{} / inherited / registry). A REMOTE https:///http:// ref (e.g. a CDN-hosted scene) is left as-is here — those ARE resolved automatically at translate / render-preview / query-layout time (and by the deck compiler picsart_media_overview when called over MCP, though its bare runner does not self-fetch), so detach after a translate/preview/query has inlined it, or inline the piece locally. Unresolvable refs (unfetched remote URLs / unknown ids) are left as-is in bulk mode; in single-layer mode they error. Pure.

ParametersJSON Schema
NameRequiredDescriptionDefault
sceneYesThe MP Scene document containing the scene_ref(s) to detach
layerIdNoExpand only this scene_ref layer; omit to expand every resolvable one
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond annotations by disclosing that it does NOT fetch, only resolves in-memory mpscene:// refs, remote refs remain untouched, unresolvable behavior differs between bulk mode (left as-is) and single-layer mode (error), and the operation is one-level deep. Ends with 'Pure', confirming no side effects. This richly supplements the readOnlyHint and idempotentHint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is lengthy but dense and front-loaded with the core action. Every sentence provides meaningful behavioral detail, including sync/async, remote vs local refs, and error cases. Slightly verbose with some technical asides (e.g., 'prefixed by the ref layer's id'), but no wasted content warrants a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of handling scene_refs, the description fully covers operation semantics, limitations (one level, only in-memory), remote ref handling, error behavior, and practical timing advice. With no output schema, the description still leaves the agent well-informed about what happens and when to call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both parameters have descriptions in the schema. The tool description adds no new parameter-level syntax or meaning beyond what the schema already provides; it mostly explains behavior that depends on the parameters rather than their individual semantics. Meets the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb 'DETACH a scene_ref into editable layers' and clearly explains the resource and action. It also distinguishes itself from siblings by explicitly framing this as the inverse of the by-reference model, which differentiates it from related tools like picsart_media_apply_scene_template.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit usage guidance is provided: 'I referenced a template, now bake this instance to customize it beyond its parameters' gives a concrete use case. It also tells when NOT to use it by explaining remote refs are left as-is and recommending detach after a translate/preview/query has inlined them, and mentions alternative automatic resolution at translate/render/query-layout.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_media_exportA
Destructive
Inspect

Render a fully-authored MP Scene to a final file via the server-side media-platform/v3/export workflow, returning the output URL(s). format:"video" (default) renders an MP4; format:"image" renders a single PNG still. Unlike a pure scene→scene deck compiler, this performs the ACTUAL render: it translates the scene to a V3/Replay project and dispatches it to the export service over an authenticated HTTP call — no local renderer. Auth is handled by the platform automatically — no token setup needed on your end. Validate the scene (picsart_media_validate_scene) and check a frame (picsart_media_contact_sheet with times:[t] — sample one frame near the point you want to verify; it returns that frame as an inline image you can look at) BEFORE exporting. If you do not yet have a scene, get the whole sequence from picsart_media_quickstart rather than reverse-engineering the document shape from picsart_media_get_scene_schema.

ParametersJSON Schema
NameRequiredDescriptionDefault
fpsNoOutput frame rate (the export service requires fps >= 1)
sceneYesThe MP Scene document to render, or an http(s):// URL to fetch it from
archiveNoimage-sequence only: false -> per-frame urls[]; true -> one zip url
durationNoCap on rendered duration (seconds) from startTime
mediaTypeNoOutput media type. Omit for the workflow default (mp4, full video)
startTimeNoScene time (seconds) of the rendered frame, for single-still media types
resolutionNoOutput resolution override in px (server-side downscale)
imageFormatNoPer-frame still codec for image-sequence (server default: png)
jpegQualityNoJPEG quality 1-100 (jpeg outputs only)

Output Schema

ParametersJSON Schema
NameRequiredDescription
idNo
urlNo
urlsNo
driveNo
modelNo
assetsYes
promptNo
videosNo
resultsNo
summaryNo
finalUrlNo
completedNo
created_atNo
iterationsNo
why_relevantNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the tool as mutating (readOnlyHint: false) and destructive (destructiveHint: true). The description adds valuable behavioral context: 'Auth is handled by the platform automatically — no token setup needed,' no local renderer, and that it performs the 'ACTUAL render' via server-side dispatch. It doesn't detail potential destructive effects, but the annotations already cover that baseline, and the description adds context beyond them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense with valuable information, but every sentence earns its place: purpose, format differentiation, auth note, and workflow guidance. It's not overly verbose, though it could be slightly tightened by removing the 'no local renderer' redundancy with 'server-side' mentioned earlier. Overall well-structured and informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (9 params, nested objects, output schema present), the description is complete: it explains the output (URLs), differentiates major modes, provides prerequisites and alternatives, and mentions auth. It doesn't need to detail return values since an output schema exists. The description leaves no major gaps for an agent to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with rich param descriptions, giving a strong baseline. The description adds meaning by mapping `format:"video"` / `format:"image"` to the `mediaType` parameter and explaining the default behavior. Minor confusion arises because the description uses 'format' while the schema field is 'mediaType,' but it still clarifies the output type semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'Render a fully-authored MP Scene to a final file via the server-side `media-platform/v3/export` workflow, returning the output URL(s).' It specifies output formats and explicitly differentiates from a 'pure scene→scene deck compiler,' which helps distinguish it from sibling tools like `picsart_media_video_render`.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit when-to-use guidance: instructs to validate the scene first (`picsart_media_validate_scene`), check a frame (`picsart_media_contact_sheet`), and points users to `picsart_media_quickstart` if they don't yet have a scene. It also clarifies when the tool is not the right choice, e.g., avoiding reverse-engineering from schema. This goes beyond basic usage to offer a clear workflow.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_media_get_capabilitiesA
Read-onlyIdempotent
Inspect

Returns the MP SDK tool layer's capabilities: supported layer content kinds, animatable properties, effect ids, transition ids, and operational limits (max duration, max layers, max resolution). Call this first when planning a composition so subsequent calls stay within supported bounds. Omitting sections returns the section INDEX (names only); the full document is ~135KB, so pass sections to fetch only the slices you need, or ["all"] for the whole thing — PREFER the granular supports sub-keys over the whole supports block (e.g. ["effects","limits"] or ["generativeTemplates","limits"] for a montage plan). looks alone is ~62KB (full prose + full param schemas for all 17 looks) and can overflow a result cap even as the ONE section requested — ask for looksCompact (id/summary/params only, ~8KB) first to see what exists, then looks for the id you actually need. Top-level: supports, limits, engine, featureMatrix; supports sub-keys: effects, transitions, easing, blendModes, animatableProperties, textAnimationPresets, motionPresets, looks, looksCompact, layerContentKinds, sceneRefs, mask, generativeTemplates, presentation, export, templateModes, expressionAnimation. Idempotent and dependency-free; safe to cache.

ParametersJSON Schema
NameRequiredDescriptionDefault
sectionsNoOptional subset of capability sections to return. Omit to get the section INDEX (names only) — the full document is ~135KB, so fetch just the slices you need; pass ["all"] for the whole document. "looks" alone is ~62KB and can overflow a result cap even as the only section requested — ask for "looksCompact" (id/summary/params only, ~8KB) first, then "looks" for the one id you need in full
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds valuable behavioral context beyond annotations: the size implications (~135KB full, ~62KB looks, ~8KB looksCompact), the result-cap overflow risk, the index-only default, and an explicit 'dependency-free; safe to cache' note. This enhances transparency without contradicting annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence earns its place: purpose, usage timing, parameter behavior, a critical size warning, a key list, and cacheability. It is front-loaded with the most important information. Slightly verbose, but the density is justified by the complexity of the operation and the need to prevent invocation errors.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the full burden of explaining return behavior. It explicitly states the default return (section index names only), lists top-level keys (supports, limits, engine, featureMatrix) and all sub-keys, and warns about response sizes. This is complete enough for an agent to know what to expect and how to avoid overflow, fulfilling all necessary contextual information.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and already explains the `sections` parameter well. The description adds further meaning by enumerating the exact top-level and sub-keys available (effects, transitions, looksCompact, etc.) and providing usage patterns like PREFER granular sub-keys. This goes beyond the schema's summary and helps the agent select valid sections, though the schema already carried much of the burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: 'Returns the MP SDK tool layer's capabilities' and enumerates the exact content (layer content kinds, animatable properties, effect ids, transition ids, operational limits). It distinguishes itself from siblings by explicitly positioning it as 'Call this first when planning a composition,' making its role clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs when to use: 'Call this first when planning a composition so subsequent calls stay within supported bounds.' It also provides detailed how-to guidance on section selection, including the preference for granular sub-keys and the specific warning to use 'looksCompact' before 'looks' to avoid overflow. This is actionable and distinguishes usage patterns.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_media_get_scene_schemaA
Read-onlyIdempotent
Inspect

Returns the JSON Schema for an MP Scene document. Use this to construct valid scenes from scratch or to remind yourself of the exact shape of layers, animations, effects, and transitions before calling picsart_media_validate_scene or checking a frame with picsart_media_contact_sheet. The returned schema is authoritative; any document that validates against it is accepted by the renderer. Pass name (e.g. "MpMediaContent") to get back just that one definition instead of the full ~66 KB schema — its internal $refs point into the full schema's #/definitions.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoOptional single definition name (e.g. "MpMediaContent") to return instead of the full schema
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only and idempotent hints, but the description adds meaningful behavior: the schema is 'authoritative' and any document validating against it is accepted. It also discloses the ~66 KB size and that name returns a single definition with internal $refs pointing to #/definitions, which is not visible in annotations or schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: purpose, usage guidance with sibling mentions, and parameter details. Front-loaded with the core action, no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only schema getter with one optional parameter, the description covers purpose, when to use, behavior of the returned schema, and parameter semantics. The absence of an output schema is compensated by the description explaining what it returns. This is complete for the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already describes the 'name' parameter well, so baseline is 3. The description adds value by explaining the size difference (~66 KB full schema vs single definition) and the internal $ref behavior, which helps the agent decide whether to pass 'name'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: 'Returns the JSON Schema for an MP Scene document.' It clearly distinguishes this tool from siblings like picsart_media_validate_scene and picsart_media_contact_sheet by framing it as a schema-retrieval utility.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use: 'Use this to construct valid scenes from scratch or to remind yourself... before calling picsart_media_validate_scene or checking a frame with picsart_media_contact_sheet.' This names alternatives and provides workflow context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_media_list_fontsA
Read-onlyIdempotent
Inspect

Returns the curated font catalog. Use this before authoring any text layer so font.family resolves to a real font: every entry has a stable key (passable directly in MpFont.family) and a resolved .otf/.ttf URL. The renderer has no system-font fallback — passing CSS family names like Inter, Arial, or Helvetica produces empty text and a unknown_font_family validation error. Agents may also pass a direct font URL matching accepted_url_pattern. The tool is dependency-free and idempotent; safe to cache.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, non-destructive. Description adds critical context: no system-font fallback, stable key, resolved URL, accepted_url_pattern, dependency-free, safe to cache. No contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured, front-loaded with purpose, then usage, then error context. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple zero-param tool, the description covers purpose, usage, output format, error behavior, and caching properties. Complete given no output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Tool has zero parameters, so baseline 4 applies. The description mentions no parameter details, but none are needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource: 'Returns the curated font catalog.' It clearly distinguishes from sibling tools as the font listing tool, with no ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states 'Use this before authoring any text layer' and explains the consequence of not using it (unknown_font_family error). No direct alternatives exist, but context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_media_list_scene_templatesA
Read-onlyIdempotent
Inspect

Enumerate the curated MP Scene template catalog. Each entry summarises a reusable, parameterized scene (title cards, lower thirds, product cards, ...). Use this to discover templates by id/category/aspect before describing or applying one. Returns { templates: [{ id, displayName, description, category?, width, height, parameterCount }, ...] }. Pure: no inputs, no state.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark readOnly, idempotent, non-destructive. Description adds 'Pure: no inputs, no state', which reinforces no-side-effect behavior and adds a stateful claim beyond annotations. It also documents return payload shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, front-loaded with the core action. No filler; each sentence adds a distinct type of info: operation, data content, usage context, return/purity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given it's a zero-input enumeration tool, the description covers purpose, usage, return format, and behavioral guarantees. Sibling context confirms this sits next to describe/apply tools; the description explicitly places itself 'before' those.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters; schema is empty with 100% coverage. Baseline is 4 per rubric; description appropriately notes 'no inputs' and explains that discovery is by template attributes, not inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description opens with 'Enumerate the curated MP Scene template catalog' – a specific verb ('enumerate') and resource ('MP Scene template catalog'). It also distinguishes from siblings by noting use 'before describing or applying one', clearly differentiating from describe/apply tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use: 'Use this to discover templates by id/category/aspect before describing or applying one.' This provides timing relative to sibling operations, effectively giving usage guidance and alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_media_overviewA
Read-onlyIdempotent
Inspect

Compile a presentation DECK (an MP Scene with one top-level track layer of scene_ref slide clips) into an OVERVIEW / 'badges' board: a single MP Scene that lays the SAME slide scene_refs into a grid of shrunken thumbnails — same content, a spatial projection. Each cell is the page positioned at its grid-cell centre with fit:"none" and a baked contain-scale (the compiler resolves each page to read its real composition size, because scene_ref.fit fits against the whole composition, not a cell). Defaults: a near-square grid (ceil(sqrt(n)) columns), the deck format for board size, and a duration long enough for every slide's entrance to settle. Carries the deck's scenes/assets so the refs still resolve. Override cols/gap/width/height/duration. Pure scene→scene; validate / preview / translate the output. NOT a capability index or a getting-started tool despite the name — it compiles thumbnails for an existing deck scene you already authored. If you are looking for 'what can this server do' or 'how do I do X', call picsart_media_quickstart instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
gapNoGap between/around cells in board pixels; default 2% of board width
colsNoColumns; default a near-square grid
sceneYesThe presentation DECK (an MP Scene) to lay out as a grid of thumbnails
widthNoBoard width; defaults to the deck format width
heightNoBoard height; defaults to the deck format height
durationNoBoard duration; defaults to the longest slide's content duration
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide readOnly, openWorld, idempotent, and non-destructive hints. The description adds substantial behavioral detail beyond these: it creates a new MP Scene with a grid layout, explains how scene_ref.fit is handled, notes that it carries the deck's scenes/assets, and describes default behavior for grid, size, and duration. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-organized, front-loaded with the core purpose, then technical details, defaults, and final exclusions. Every sentence adds useful information without fluff or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, a rich input schema, and no output schema, the description sufficiently covers the transformation process, defaults, constraints, and expected output. It also provides enough context for a user to decide whether to call this tool over siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema covers 100% of parameters with meaningful descriptions, so baseline is 3. The description adds value by emphasizing which parameters are overridable ('Override cols/gap/width/height/duration') and providing technical context about scene_ref.fit and board sizing, going slightly beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action: 'Compile a presentation DECK ... into an OVERVIEW / badges board' with a grid of shrunken thumbnails. This clearly distinguishes it from sibling tools, and it explicitly clarifies what it is NOT: 'NOT a capability index or a getting-started tool despite the name.'

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear usage context: it is for an 'existing deck scene you already authored' and is a 'Pure scene→scene' operation. It also explicitly names the alternative for capability queries: 'If you are looking for what can this server do ... call picsart_media_quickstart instead.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_media_patch_sceneA
Read-onlyIdempotent
Inspect

Apply a batch of incremental, ID-ANCHORED edit ops to an MP Scene and return the patched result — the token-cheap alternative to re-emitting the entire scene on every edit. Three verbs: set (create-or-replace a value; last-write-wins, idempotent; replace is accepted as a first-class alias), remove (delete a field/array element, or an entire anchored entity when no path is given), and add (insert a new id-keyed layer/asset/audio/marker; requires kind and a value carrying the new entity's id). Address the target with exactly one ANCHOR + id — layer, asset, audio, or marker — never a root index path: {op:"set",path:"layers/3/..."} is rejected with use_id_anchor and a corrected anchored form embedded in the message; omit the anchor only for true document-root fields like composition/*, version, or scenes/<id>. path is a slash-separated RFC-6901-style pointer relative to the anchored entity (e.g. content/text, effects/0/params/amount) — ~0/~1 escapes are honored, and dot-separated paths fail strict with a slash-form hint. All ops in one call are ATOMIC: they apply sequentially against a single clone and the first failing op aborts the WHOLE batch with no partial result (a typed error carrying code and opIndex); the input scene is never mutated, and later ops see earlier ops' results. Validation runs INSIDE this tool: the patched result goes through the exact picsart_media_validate_scene pipeline (JSON-Schema structural pass + semantic rules) and any severity:error fails the whole batch loudly with the diagnostics — a success has already passed that pipeline, so do NOT call picsart_media_validate_scene again right after a successful patch; only validate once you've made further hand-edits outside this tool. (Validation proves conformance to the scene contract, not that every referenced remote asset or scene_ref is fetchable at render time.) Id-less collections — track clips, transitions, animations, keyframes, text content/animations — carry no per-element id, so edit them with a whole-array set on the owning anchor's array path (e.g. {op:"set",layer:"main_track",path:"content/layers",value:[...]}) rather than addressing individual elements by index across calls.

ParametersJSON Schema
NameRequiredDescriptionDefault
opsYes1-64 id-anchored patch ops (set/remove/add) applied atomically
sceneYesThe MP Scene document to patch, or an http(s):// URL to fetch it from
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint, idempotentHint, etc.), the description discloses critical behavioral traits: atomicity of the entire batch ('first failing op aborts the WHOLE batch'), non-mutation of the input scene ('input scene is never mutated'), internal validation via the picsart_media_validate_scene pipeline, and error structure ('typed error carrying code and opIndex'). It also clarifies which collections are id-less and how to handle them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence delivers an operational rule or clarification. It is front-loaded with the purpose, then systematically covers verbs, addressing, path syntax, atomicity, validation, and edge cases. There is no fluff—each detail like 'replace is accepted as a first-class alias' or the id-less collection paragraph earns its place for a complex tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity and absence of an output schema, the description covers all necessary aspects: how to construct ops, how to address anchors vs root fields, path escaping, atomic failure semantics, internal validation and its implications, and handling of id-less collections. It even explains the interplay with picsart_media_validate_scene, making it self-sufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

While the schema covers the two parameters with basic descriptions, the description adds extensive meaning for the 'ops' parameter: verb semantics (set/remove/add), anchor types (layer/asset/audio/marker), path syntax with ~0/~1 escapes, and the requirement for a value carrying an id. For the 'scene' parameter, it clarifies it can be a URL. This far exceeds the schema's bare descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb+resource: 'Apply a batch of incremental, ID-ANCHORED edit ops to an MP Scene and return the patched result'. It clearly distinguishes this tool from siblings by positioning it as 'the token-cheap alternative to re-emitting the entire scene on every edit' and enumerates three specific verbs (set/remove/add).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to use the tool (token-cheap incremental edits) and when not to: 'do NOT call picsart_media_validate_scene again right after a successful patch; only validate once you've made further hand-edits outside this tool.' It also provides rules for addressing targets (anchor vs root index) and id-less collections, giving clear conditional guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_media_probe_mediaA
Read-onlyIdempotent
Inspect

Returns a remote media URL's metadata WITHOUT downloading the file: kind (image/video/audio), as-displayed width/height (EXIF orientation applied, plus the raw orientation), durationSeconds (mp4/mov/m4a), contentType, and total bytes — from at most two small ranged fetches (≤256KiB). Use this BEFORE authoring a scene to size compositions, pick fit, set clip windows to the real source duration, and reject wrong assets early — instead of shelling out to curl/ffprobe or guessing. Covers PNG/JPEG/GIF/WebP/HEIC/AVIF/SVG/TIFF dimensions and ISO-BMFF (mp4/mov/m4a) dimensions+duration, including moov-at-end files. Unrecognized containers still return kind/contentType/bytes plus a notes[] entry — dimensions are never invented. URLs are SSRF-guarded (public http(s) only; every redirect hop re-checked). Marked stochastic only because a URL's content can change between calls — probing an immutable asset is stable. Skip this call if you already know the asset's exact width/height/duration from a prior tool result (e.g. a generation tool's own response) — this is a cost-saving lookup, not a required gate. A missing durationSeconds means the duration is unknown — it is never emitted as 0 for a real zero-length clip. Fragmented MP4 / streaming uploads carry no duration in their container header, but this tool now recovers it for most of them by walking the file's fragment timeline, so a duration usually comes back anyway; when the field is still absent it genuinely could not be determined. That pass costs a few more small ranged reads than the two mentioned above — still bounded, still free, still no download. Do not assume 0; resolve it another way before computing trims. The response may also carry fps, videoCodec, audioCodec and hasAudio when they are derivable from the container — these are best-effort extras and are omitted rather than guessed.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesPublic http(s) URL of the media to probe (SSRF-guarded ranged fetch, file is not downloaded)
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations (readOnlyHint, idempotentHint, destructiveHint) already establish safety, but the description adds a wealth of behavioral context: SSRF-guarding, bounded ranged fetches, 'dimensions are never invented' for unrecognized containers, the semantics of missing durationSeconds (never 0 for real zero-length), and the stochastic annotation rationale ('a URL's content can change between calls'). This goes far beyond the structured hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but each paragraph serves a purpose: summary, usage, security, duration semantics, and best-effort extras. It is front-loaded with the core return fields and usage context. While some sentences could be tightened, the complexity of the tool (edge cases for duration, fragmented MP4, SSRF) justifies the length. It earns a 4 rather than a 5 due to slight redundancy in the stochastic explanation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description takes full responsibility for explaining return values, edge cases, and failure modes. It covers what is returned for recognized and unrecognized formats, how missing fields should be interpreted, security behavior, and cost-savings usage. This is a complete picture for an AI agent to decide when and how to invoke the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already has 100% coverage for the single 'url' parameter, including 'Public http(s) URL... SSRF-guarded ranged fetch, file is not downloaded'. The description does not add new parameter-level details beyond what the schema states; it reinforces the same constraints but doesn't introduce novel semantics, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb ('Returns') and resource ('remote media URL's metadata'), listing specific fields ('kind', 'width'/'height', 'durationSeconds', 'contentType', 'bytes'). It explicitly contrasts with downloading, making the tool's purpose unmistakable and distinct from siblings like picsart_media_export or picsart_media_validate_scene.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit when-to-use guidance ('Use this BEFORE authoring a scene to size compositions, pick fit, set clip windows...') and when-not-to-use ('Skip this call if you already know the asset's exact width/height/duration from a prior tool result'). It also names alternatives ('instead of shelling out to curl/ffprobe or guessing'), covering both positive and negative usage cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_media_query_layoutA
Read-onlyIdempotent
Inspect

Resolves where every layer ACTUALLY LANDS in the root composition at a given time — without rendering pixels. Runs the engine's Scene Query System and returns each layer's final geometry in root-composition space: box {x,y,width,height} (top-left pixels), center, rotationDeg, paint-order zIndex, visible/active, an onCanvas coverage flag (full/partial/off), timing, and composition/parent/nesting. This is the cheap, structured answer to 'what does my composition look like, spatially?' — use it instead of (or before) checking a rendered frame (picsart_media_contact_sheet) to verify positions, sizes, overlap, off-canvas layers, and stacking order, since reasoning over numbers beats eyeballing a PNG. Coordinates are ALWAYS top-left pixels in root space for BOTH engines — the query layer unifies Jet's pixel coords and V3's internal normalized coords, so a layer WITH RESOLVED GEOMETRY reports the same box on jet or v3 (the engine arg only changes the translation path). Text is the exception (see limitations). Nested children (scene_ref/look comps) are returned too, namespaced (e.g. a__bg) with nestingLevel and composition; inactive layers at the queried time come back marked active:false with their box omitted. Set verbose:true to also get raw 4x4 transform matrices and the per-component transform chain (for transform-editing tools); omit it for layout reasoning. Determinism is deterministic — identical (scene,time,engine) returns an identical report. Limitations: (1) text layers — the query system does not resolve laid-out glyph bounds, so a text layer's POSITION/center is accurate but its box size is a placeholder (Jet 0x0, V3 ~2000x1); a notes[] entry flags any layer with a degenerate dimension. (2) multi-group V3 scenes may need a sceneContext (e.g. { canvasGroupId }); single-composition scenes do not. (3) GEOMETRY ONLY — this is NOT a render/export gate. The query is GL-free and never resolves effects, assets, codecs, or engine-specific support, so a clean layout here does NOT guarantee the scene renders or exports. A clean report carries this caveat in its scope field — confirm renderability with picsart_media_contact_sheet / picsart_media_export.

ParametersJSON Schema
NameRequiredDescriptionDefault
timeYesScene time (seconds) at which to resolve layer geometry
sceneYesThe MP Scene document to query, or an http(s):// URL to fetch it from
engineNoTarget engine translator: "jet" or "v3" (coordinate convention is identical either way)
verboseNoInclude raw transform matrices + component chain per layer (default false)
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and idempotentHint, but the description adds substantial behavioral context: coordinate unification across both engines, deterministic output, text-layer placeholder boxes, the need for `sceneContext` in multi-group V3 scenes, and the caveat that a clean layout does not guarantee renderability. No contradiction with annotations exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but well-structured: it front-loads the core purpose, then systematically details return fields, limitations, and caveats. Each section is necessary for a complex tool. Minor redundancy exists (e.g., 'Determinism is deterministic') and the repeated references to contact_sheet add a slight amount of extra length, but overall the length is justified.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Even without an output schema, the description previews the major return fields (`box`, `center`, `rotationDeg`, `zIndex`, `active/visible`, `onCanvas`, `timing`, `nesting`), explains prerequisites like `sceneContext` for multi-group V3, and details text-layer limitations and the `scope` caveat. It is exceptionally complete for a tool of this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, giving a baseline of 3. The description adds meaning beyond the schema by explaining that `engine` only changes the translation path while coordinates are unified, that `verbose` yields raw 4x4 matrices and chain details, and that `scene` can be a URL. This is valuable extra context, though the schema already documents types and requirements.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Resolves where every layer ACTUALLY LANDS in the root composition at a given time — without rendering pixels.' It clearly differentiates from sibling `picsart_media_contact_sheet` by emphasizing that it returns geometry rather than a rendered frame, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly instructs when to use the tool: 'use it instead of (or before) checking a rendered frame (`picsart_media_contact_sheet`)' and when not to use it: 'GEOMETRY ONLY — this is NOT a render/export gate.' It also directs users to confirm renderability with `picsart_media_contact_sheet` / `picsart_media_export`, providing clear alternatives and exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_media_quickstartA
Read-onlyIdempotent
Inspect

Call FIRST when the user asks how to accomplish a media task (merge/concat videos, make a contact sheet, or export/render a scene). Returns ready-to-run picsart_media_* tool-call sequences. Omit recipe for the index.

ParametersJSON Schema
NameRequiredDescriptionDefault
recipeNoOmit for the recipe index; pass one for its step-by-step tool-call sequence
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint, idempotentHint, etc. The description adds behavioral details: it returns sequences (not executing them) and that omitting recipe gives an index. No contradiction with annotations. Slightly less than perfect because it doesn't mention any potential limitations or response structure, but adequate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, each earning its place: first states purpose and timing, second gives parameter guidance. No fluff, front-loaded with the key instruction 'Call FIRST'.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with one optional parameter and no output schema, the description fully covers what the tool does, when to use it, how to invoke it, and what to expect in return. No gaps given the context signals.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with enum values. The description adds meaning by explaining the two modes: omitting for the index or passing a value for step-by-step sequences. This clarifies usage beyond the schema's enum list.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool provides 'ready-to-run picsart_media_* tool-call sequences' for specific media tasks. It uses a specific verb ('Call FIRST') and resource ('how-to recipes'), and distinguishes itself from sibling tools that perform the actual tasks.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Call FIRST when the user asks how to accomplish a media task' and lists example tasks. It also tells when to omit the parameter ('Omit `recipe` for the index'), providing clear usage context and alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_media_resolve_looksA
Read-onlyIdempotent
Inspect

Expand every by-reference look (each layer's looks[] annotation) in an MP Scene into its concrete nested composition, returning a SELF-CONTAINED scene — the preset logic baked in, no look-catalog dependency at render time. The explicit, on-demand counterpart of the just-in-time resolution that picsart_media_translate_scene / picsart_media_contact_sheet / picsart_media_query_layout already do internally. Use it to 'flatten' a thin look-annotated scene into a portable one (e.g. to hand off, archive, or edit the expanded layers directly). Idempotent: a scene with no looks[] returns unchanged. Pure: takes the full scene by value, returns a new scene.

ParametersJSON Schema
NameRequiredDescriptionDefault
sceneYesThe MP Scene document whose by-reference looks should be expanded
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false. The description adds critical behavioral traits beyond those: it is pure (takes full scene by value, returns new scene), idempotent even without looks, and bakes preset logic to eliminate look-catalog dependency. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but efficient—four sentences each carrying essential information: action/result, sibling context, usage directive, and idempotence/purity. No filler or repetition of structured annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 1-parameter tool with no output schema and strong annotations, the description covers what, why, when, and edge-case behavior. It explains the return value (self-contained scene) and what it means for rendering dependencies, making it adequately complete for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the parameter description already explains the scene role. The description adds semantics beyond the schema by clarifying the exact transformation (expand by-reference looks into concrete nested composition), purity ('takes the full scene by value'), and the no-op case when no looks[] are present.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Expand') and resource ('by-reference look in an MP Scene'), clearly stating the transformation into a self-contained scene. It also distinguishes itself from sibling tools by framing it as the explicit, on-demand counterpart to internal just-in-time resolution in tools like picsart_media_translate_scene.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Use it to flatten...' and provides concrete use cases (hand off, archive, edit expanded layers), while contrasting with internal resolution in siblings. It lacks an explicit 'when not to use' but the implied alternative guidance is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_media_translate_sceneA
Read-onlyIdempotent
Inspect

Translate an MP Scene document into an engine's project format, WRITE it to a content-addressed file, and return { path, cached, summary }. The engine project itself is NOT inlined — a large scene's project can be ~100 KB of JSON and would overflow this tool's result token cap, so it goes to disk (under MP_AI_OUTPUT_DIR/projects/..json) and you get a path plus a compact summary. The target engine is chosen by engine (default v3 → a full, openable Replay file { meta, context:{layers,…}, actions, settings }; jet → a Jet { compositions, activeCompositionID, ColorSpace } project). Read the file when you need the full project. Validate the scene first; the translator does not re-validate. Deterministic: identical (scene, engine) hash to the same path (cached:true on a repeat).

ParametersJSON Schema
NameRequiredDescriptionDefault
sceneYesThe MP Scene document to translate, or an http(s):// URL to fetch it from
engineNoTarget engine: "v3" (default) or "jet"
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotation contradiction: The description explicitly says 'WRITE it to a content-addressed file' and 'it goes to disk', which is a write operation, yet the annotations declare readOnlyHint: true. This directly contradicts the read-only hint. While the description discloses other behaviors (caching, determinism, no re-validation), the contradiction forces a score of 1.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is somewhat long but every sentence provides necessary information: the action, the reason for writing to disk, engine specifics, output file location, and determinism. It is front-loaded with the core purpose and structured logically. It could be slightly trimmed, but the length is justified given the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully covers return values (`{ path, cached, summary }`) and explains the external side effects (file path, caching). It also addresses edge cases like token overflow, engine-specific outputs, and the need to validate first. Given the tool's complexity, the description is remarkably complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description adds significant meaning beyond the schema. It explains the `engine` parameter's effect in detail (`v3` → Replay file structure, `jet` → Jet project format), the default value, and the `scene` parameter's URL option. This goes well beyond the bare schema descriptions, making the parameters fully understandable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's verb and resource: 'Translate an MP Scene document into an engine's project format, WRITE it to a content-addressed file, and return { path, cached, summary }.' It distinguishes this from sibling tools by emphasizing the translation/writing behavior and the specific output format, making its purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context on when to use the tool and how to handle outputs: it explains that the file should be read when the full project is needed, and explicitly recommends validation first ('Validate the scene first; the translator does not re-validate'), which implies using the validation sibling tool. It doesn't name alternative tools directly but gives practical usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_media_uploadA
Read-onlyIdempotent
Inspect

Opens a drag-and-drop upload widget so the user can get a local image/video/audio file into this conversation as a URL — no filesystem access on your side, and no external CLI needed. Call this whenever a user wants to use a local file with any picsart_media_* tool, or after a local-path argument was refused with a local_source_not_supported error. Call with no arguments to just open it. Optional purpose labels the dropzone; accept restricts file kind; detected_files is a BEST-EFFORT, MODEL-SUPPLIED label hint ONLY — e.g. filenames you can see attached in this conversation but have no other way to reach — the widget shows it purely as a suggestion ("Drop clip.mp4 here") and NEVER filters or rejects what the user actually drops, since this hint can be wrong or hallucinated; known_urls seeds pickable chips in a "Use existing" tab from URLs already produced earlier in this conversation. The upload itself happens directly in the user's browser to Picsart's CDN — no bytes pass through this tool. Once the user finishes (or picks existing URLs), the widget reports the resulting URL(s) back into the conversation. This is a two-turn handshake: this call only OPENS the widget and returns no URL(s) itself — they arrive on the user's NEXT message, not in this call's result.

ParametersJSON Schema
NameRequiredDescriptionDefault
acceptNoRestricts the dropzone to one media kind. Omit to accept anything.
purposeNoShort label for what the file is for, shown in the widget's dropzone, e.g. "the video you want to upscale".
known_urlsNoURLs already produced earlier in this conversation (by this tool or another), offered in the widget as pickable chips under "Use existing".
detected_filesNoFilenames you can see attached in this conversation (e.g. a chat attachment you have no other way to reach). BEST-EFFORT LABEL HINT ONLY, never a filter: this is model-supplied and can be hallucinated — the widget only ever displays it as a suggestion and never rejects or gates on what the user actually drops.

Output Schema

ParametersJSON Schema
NameRequiredDescription
driveNo
toolsNo
acceptNo
purposeNo
ingestedNo
knownUrlsNo
seedVersionNo
capabilitiesNo
detectedFilesNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true, but the description adds critical behavioral details beyond that: the upload happens in the user's browser directly to Picsart's CDN without passing through the tool, the two-turn handshake, and the fact that detected_files is a best-effort hint that never filters. These go well beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively long but every sentence serves a purpose: purpose, usage, parameter details, and the handshake explanation. It is front-loaded with the core action. Could be slightly more concise, but it earns its length for the complexity involved.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's two-turn handshake and user-interaction nature, the description is fully complete. It explains the asynchronous result flow, the role of each parameter, and the non-filtering behavior of detected_files. With a rich annotation set and output schema implied, nothing is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds meaningful extra context for detected_files (warning about hallucination and that it's a suggestion only) and known_urls (seeds pickable chips). This goes slightly beyond the schema's own descriptions, justifying a 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool 'Opens a drag-and-drop upload widget' to get a local media file into the conversation as a URL. The verb 'opens' and specific resource 'upload widget' distinguish it from sibling picsart tools, none of which handle local file upload directly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly specifies when to call: 'whenever a user wants to use a local file with any picsart_media_* tool, or after a local-path argument was refused with a local_source_not_supported error.' Also notes the two-turn handshake, advising that results come on the next message. No alternative usage is needed; the context is fully covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_media_validate_sceneA
Read-onlyIdempotent
Inspect

Validates an MP Scene document and returns a list of structured diagnostics. Use this whenever you've assembled or modified a scene, especially before rendering. Each diagnostic carries a JSON path, a stable machine-readable code, a human-readable message, and optional hints. A scene is considered valid when no diagnostic has severity error. The tool is pure: it takes the full scene in and returns diagnostics; if you keep editing, call again with the updated scene.

ParametersJSON Schema
NameRequiredDescriptionDefault
sceneYesThe MP Scene document to validate, or an http(s):// URL to fetch it from
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, and non-destructive hints. The description adds meaningful behavioral details beyond that: it states the tool is pure, takes the full scene, and that results are only valid for the current scene state, advising to call again after edits. It also explains the diagnostic structure and validity condition, enriching the annotation context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences with no filler. It front-loads the verb and resource, then provides concise, useful details about the output format and usage scenario. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity, the description covers the key aspects: what it validates, when to use, what the output looks like, how to interpret validity, and its pure behavior. With annotations already handling safety, no output schema exists but the diagnostic structure is described. This is complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description already fully covers the single parameter ('The MP Scene document to validate, or an http(s):// URL to fetch it from'). The description adds the notion of passing the 'full scene' and mentions re-calling after edits, but this is more usage guidance than parameter semantics. At 100% schema coverage, the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Validates') with an explicit resource ('an MP Scene document') and states the main output ('a list of structured diagnostics'). This clearly distinguishes it from sibling tools like picsart_media_patch_scene or picsart_media_export, which perform different actions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit guidance on when to use it: 'Use this whenever you've assembled or modified a scene, especially before rendering.' It does not mention alternatives or when not to use it, but given the tool's unique validation role among siblings, this is clear enough context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_media_video_createA
Destructive
Inspect

Generate a NEW short motion-graphics video (title cards, animated logos, ambient loops, kinetic-typography clips) from a natural-language brief, at up to 3840x2160 — use this for output above 1920x1920, since the picsart_media_* Scene tools (picsart_media_export, picsart_media_apply_scene_template, etc.) cap at 1920x1920. NOT for template-based slideshows/decks (picsart_media_apply_scene_template), montage/concat of existing clips (picsart_media_apply_scene_template with the montage template), contact sheets (picsart_media_contact_sheet), or captions burned over existing user media — those are all picsart_media_* Scene tools and are cheaper and deterministic; this tool writes NEW code from a brief every time. Returns a code_url (never renders — call picsart_media_video_render next) plus a thumbnail_url; THIS tool returns the thumbnail as a url only, not as something you can look at (picsart_media_contact_sheet is the tool on this surface that returns inline images), so describe the result to the user as unverified until it has actually been rendered. Costs 25 credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
briefYesNatural-language description of the video to generate: subject, mood, color, pacing, story arc. Self-contained — write the full brief here; there is no conversation memory carried between separate picsart_media_video_create/picsart_media_video_revise calls.
durationNoVideo duration in seconds (5-600), or omit/'auto' to let the LLM decide from the brief. Note: picsart_media_video_render currently caps duration_seconds conservatively (15s for hd/full_hd, 10s for ultra_hd) pending calibration of longer renders — a much longer video generated here may need trimming before it can be rendered.
attachmentsNoReference images/video/audio (e.g. a brand logo or product photo) to incorporate
aspect_ratioNoVideo aspect ratio. Omit for the workflow default (16:9)
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds significant context beyond annotations: it notes that the tool never renders, returns a code_url that must be rendered separately, the thumbnail is a url only and unverified until rendered, the tool costs 25 credits, and it writes new code from a brief every time (non-idempotent). It also mentions render duration limitations. Annotations already indicate destructiveHint=true, and this description is consistent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with the main action first, but it is somewhat long with multiple clauses and alternative listings. Each sentence adds value, but it could be slightly more concise without losing essential details. The front-loading is good.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no output schema, the description covers the return values (code_url, thumbnail_url), explains the thumbnail is a url only and not viewable inline, provides credits cost, differentiates from siblings, and gives usage guidelines. It is complete for the agent to understand the tool's behavior and integration.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so baseline is 3. The description adds extra meaning for the 'brief' parameter (self-contained, no conversation memory) and the 'duration' parameter (render cap limitation, may need trimming). This improves the agent's understanding beyond the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states 'Generate a NEW short motion-graphics video ... from a natural-language brief', providing a specific verb and resource. It distinguishes from sibling tools by noting that Scene tools cap at 1920x1920 and listing excluded use cases (template-based slideshows, montage, contact sheets, captions).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clear when-to-use guidance: 'use this for output above 1920x1920, since the picsart_media_* Scene tools cap at 1920x1920'. Also explicit when-not-to-use: 'NOT for template-based slideshows/decks ... montage/concat ... contact sheets ... captions burned over existing user media — those are all picsart_media_* Scene tools and are cheaper and deterministic'. Additionally, outlines the next step (call picsart_media_video_render).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_media_video_renderA
Destructive
Inspect

Render a picsart_media_video_create/picsart_media_video_revise code_url to a final MP4, up to 3840x2160 — use for output above 1920x1920 (the picsart_media_* Scene tools, e.g. picsart_media_export, cap at 1920x1920). Resolutions: hd (1280x720), full_hd (1920x1080, default), ultra_hd (3840x2160, the one output size no picsart_media_* tool can produce). Cost: 5 credits (full_hd), 10 credits (ultra_hd); hd not separately measured, expected no higher than full_hd. IMPORTANT: duration_seconds is currently capped conservatively — 15s for hd/full_hd, 10s for ultra_hd — because a 30s render measured ~32 minutes of real wall time in testing regardless of resolution; requests above the cap are refused up front rather than risking a charged, uncompleted render. If the clip will be composited into an MP Scene afterward (picsart_media_patch_scene, picsart_media_export, etc.), render at hd/full_hd — MP composition itself caps at 1920x1920, so a 4K render only matters as a final, standalone output, never as a scene input.

ParametersJSON Schema
NameRequiredDescriptionDefault
code_urlYescode_url from a picsart_media_video_create/picsart_media_video_revise result
resolutionNoOutput resolution. Omit for the workflow default (full_hd, 1920x1080)
duration_secondsYesCopy metadata.duration from the picsart_media_video_create/picsart_media_video_revise result that produced this code_url. Used ONLY by this tool's own pre-render safety check — see the cap noted in this tool's description.

Output Schema

ParametersJSON Schema
NameRequiredDescription
idNo
urlNo
urlsNo
driveNo
modelNo
assetsYes
promptNo
videosNo
resultsNo
summaryNo
finalUrlNo
completedNo
created_atNo
iterationsNo
why_relevantNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate destructiveHint=true and readOnlyHint=false. Description adds cost per resolution (5-10 credits), duration caps (15s/10s), and a real-time warning (30s render ~32 min). This justifies the conservative cap and aligns with annotations. No contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is structured with clear sections: purpose, resolution options, cost, caps, scene composition note. It's front-loaded with the core purpose. Slightly verbose but every sentence adds necessary context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers all essential aspects: input source, output format/max resolution, cost, duration caps with rationale, and guidance for scene integration. Output schema exists, so return values need not be described. Complete for a complex tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, baseline 3. Description adds meaning for duration_seconds: tells agent to copy metadata.duration from the source result and explains it's used for a pre-render safety check. For resolution, it clarifies defaults and scene composition implications.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the tool renders a code_url from video_create/revise to a final MP4 up to 3840x2160. It distinguishes from sibling tools like export by noting they cap at 1920x1920, so this tool handles higher resolutions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to use for output above 1920x1920. Provides guidance when compositing into an MP Scene: use hd/full_hd because MP composition caps at 1920x1920. Also details cost and duration caps, helping the agent decide when to invoke.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_media_video_reviseA
Destructive
Inspect

Edit or fix an existing picsart_media_video_create/picsart_media_video_revise output (code_url). Use for output above 1920x1920 — same routing rule as picsart_media_video_create: for template slideshows, montage/concat, contact sheets, or captions over user media, use the picsart_media_* Scene tools instead. Pass error (a compilation error/stack trace) only when fixing a broken render from picsart_media_video_render; omit it for a normal content edit. Stateless — describe the FULL desired change in instruction each call; no conversation memory is kept between calls. Returns a new code_url (never renders — call picsart_media_video_render next). Costs 25 credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
errorNoCompilation error/stack trace from a failed picsart_media_video_render call (fix mode only)
code_urlYescode_url from a prior picsart_media_video_create/picsart_media_video_revise result
attachmentsNoReference images/video/audio to incorporate into the revision
instructionYesThe FULL desired change, in prose. Restate context — nothing prior is remembered.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate destructiveHint=true and readOnlyHint=false. The description adds critical behavioral context: statelessness (no memory between calls), that it returns a new code_url but does not render (call picsart_media_video_render next), and the cost of 25 credits. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single paragraph of about 5 sentences, covering purpose, usage, behavior, parameter nuance, and cost. It is efficient with no wasted words, but could be slightly more structured (e.g., bullet points for the key rules). Still, it earns its length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 4 parameters, no output schema, and annotations present, the description addresses purpose, usage alternatives, parameter behavior, statelessness, next step, and cost. It does not explicitly describe the return value format beyond 'new code_url', but that is sufficient for a revision tool that returns a code_url. The mention of 'same routing rule' is a bit vague but acceptable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% (all 4 parameters described). The description adds value beyond the schema by clarifying when to use the `error` parameter ('fix mode only'), emphasizing that `instruction` must be a full description because of statelessness, and reinforcing that `code_url` comes from a prior create/revise result. This contextual guidance helps the agent use parameters correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb+resource ('Edit or fix an existing picsart_media_video_create/picsart_media_video_revise output') and explicitly distinguishes from sibling tools by naming the picsart_media_* Scene tools for alternative use cases. It also clarifies the tool's stateless nature and what it returns, making its role unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit when-to-use guidance ('Use for output above 1920x1920...') and when-not-to-use ('for template slideshows... use the picsart_media_* Scene tools instead'). It also explains the conditional use of the `error` parameter (only for fixing a broken render) and the stateless requirement to describe the full change each call. This gives the agent clear decision-making rules.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_model_catalogA
Read-onlyIdempotent
Inspect

Returns the Picsart AI model catalog as plain data — renders NO widget or UI. Use this when YOU (the assistant) need catalog knowledge for your own reasoning: picking a model before picsart_generate, answering "which models support X", or comparing options — without pushing a model-picker widget into the conversation. When the user wants to SEE or browse models visually, use picsart_list_models instead (it renders the Picsart Studio picker). Same filters and result shape as picsart_list_models, but every item is rich by default: id, name, mode, inputType, provider, badges, description, plus supportedAspectRatios/supportedResolutions when the model declares an enum for that param — enough to answer "which models support 16:9" without picsart_model_params. Do NOT use it to fetch a single model's FULL parameter schema (use picsart_model_params) or estimate per-call cost (use picsart_preflight). Inputs (all optional): mode (filter to image/video/audio/text — text = LLM models that return generated text), provider (case-insensitive substring like "flux", "kling", "google"), acceptsImage (true → only models that take an image input — i2i, i2v, i2t), acceptsVideo (true → only models that take a video input — v2v, v2a, v2t), acceptsAudio (true → only models that take an audio input — a2v, sts), inputType (exact-match escape hatch; one of t2v/i2v/v2v/a2v/t2i/i2i/t2a/v2a/tts/sts/sfx/music/t2t/i2t/v2t), limit (1–100, default 20), concise (default false; when true items carry only id/name/mode/inputType plus the ratio/resolution fields, to save tokens). inputType codes — first letter is input modality, second is output: t2i (text→image), i2i (image→image), t2v (text→video), i2v (image→video), v2v (video→video), a2v (audio→video), t2a (text→audio), v2a (video→audio), tts (text-to-speech), sts (speech-to-speech), sfx (sound effects), music (music gen), t2t/i2t/v2t (LLM text output from text/image/video input). Example: { mode: "audio", inputType: "music" } returns music-generation models. Returns { items, total, truncated }truncated is true when more matched than were returned; refine filters or raise limit (max 100) to see more. Read-only; spends no credits and works without authentication.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoFilter by generation mode
limitNoMax items to return (1–100, default 20)
conciseNoWhen true, items carry only id/name/mode/inputType. Default false (rich items).
providerNoProvider substring (e.g. "kling", "flux", "google")
inputTypeNoExact inputType match (e.g. "i2v")
acceptsAudioNoOnly return models that accept an audio input (a2v, sts)
acceptsImageNoOnly return models that accept an image input (i2i, i2v)
acceptsVideoNoOnly return models that accept a video input (v2v, v2a)

Output Schema

ParametersJSON Schema
NameRequiredDescription
itemsYes
totalYes
truncatedYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, idempotentHint=true; description adds context by noting 'Read-only; spends no credits and works without authentication.' It also discloses return behavior: 'Returns `{ items, total, truncated }`' and explains `truncated` semantics, plus the token-saving `concise` mode. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Long but structured: purpose, usage vs sibling, parameter semantics, inputType code legend, return shape. Front-loaded with the key distinction. Every sentence earns its place; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 8 optional params, an output schema, and rich annotations. The description is fully self-contained: it covers all params, return shape, pricing/credit behavior, authentication, and explicit exclusions. Sibling tools are referenced for complementary needs. Extremely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% but the description adds meaning beyond enums: explains `mode` values (e.g., 'text = LLM models that return generated text'), `provider` is 'case-insensitive substring', `inputType` is an 'exact-match escape hatch', and defines all 15 inputType codes. It also clarifies `concise` changes item shape. This goes well beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description opens with a specific verb+resource: 'Returns the Picsart AI model catalog as plain data — renders NO widget or UI.' It immediately distinguishes from sibling `picsart_list_models` by specifying that tool renders a picker. This removes ambiguity about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use: 'when YOU (the assistant) need catalog knowledge for your own reasoning...' and when not to: 'When the user wants to SEE or browse models visually, use `picsart_list_models` instead.' It also names alternatives for related tasks: `picsart_model_params` for full schema, `picsart_preflight` for cost. Clear when/when-not.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_model_choiceA
Read-onlyIdempotent
Inspect

Opens the Picsart Model Choice board: a short list of candidate models as cards, each with what it is best for and its key limits (max clip length, resolutions, aspect ratios, audio, cost note). The user taps one and their pick comes back as a JSON message in the conversation. Call this BEFORE the first charged generation of a project, after narrowing the catalog with picsart_model_catalog: pass 2 to 6 candidates that genuinely fit the brief, write each bestFor line for THIS project rather than as a spec sheet, and mark at most one as recommended. Do not pick a model silently when candidates differ in trade-offs the user would care about (duration ceiling vs resolution ceiling, cost, audio) — that trade-off is theirs to make. For free-form catalog browsing use picsart_list_models instead; this board is a decision, not a catalog. Returns the normalised board payload the widget renders. Read-only; spends no credits and works without authentication.

ParametersJSON Schema
NameRequiredDescriptionDefault
notesNoOne line of guidance shown above the cards.
purposeYesWhat the model will be used FOR, in a few words — e.g. "9:16 hero video, 4 shots, character must stay consistent". Shown as the board's heading.
candidatesYesThe shortlist to choose from — 2 to 6 models you already vetted against the project.

Output Schema

ParametersJSON Schema
NameRequiredDescription
notesNo
purposeYes
candidatesYes
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint. The description adds valuable context: 'Read-only; spends no credits and works without authentication' and describes the user interaction flow (user taps, pick returns as JSON). No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured: first sentence defines purpose, then usage guidelines, then constraints, then alternatives. Every sentence adds value. No redundancy or verbosity. Length is appropriate for the complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's role as a decision board with rich input schema and annotations, the description covers all necessary aspects: when to call, how to prepare candidates, constraints, what happens (user picks, JSON in conversation), read-only nature, and return of normalized board payload. No gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with detailed field descriptions. The description adds meaning beyond schema by explaining how to use parameters: 'write each bestFor line for THIS project rather than as a spec sheet', 'mark at most one as recommended', and provides example for purpose. This guidance helps the agent fill parameters correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool opens a Model Choice board for user decision, with a specific verb ('Opens'), resource ('Picsart Model Choice board'), and outcome. It distinguishes from siblings: 'For free-form catalog browsing use picsart_list_models instead; this board is a decision, not a catalog.'

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to call ('BEFORE the first charged generation of a project, after narrowing the catalog with picsart_model_catalog'), how to prepare candidates (2-6, bestFor per project, at most one recommended), and when not to use it ('Do not pick a model silently...'). Also provides alternative: picsart_list_models for catalog browsing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_model_paramsA
Read-onlyIdempotent
Inspect

Returns the parameter schema for a specific Picsart model — a map of param name to descriptor ({ type, required, default, enum, min, max, step, label, accept }). Use this once you have a model id and need to construct the params payload for picsart_generate or feed picsart_preflight with a candidate object. Do NOT use it to discover which models exist (use picsart_list_models) or estimate cost (use picsart_preflight). Required input: model id. Example: { model: "flux-2-pro" }. Returns { model, schema: { <paramName>: { type: "string"|"number"|"boolean"|"file", required?, default?, enum?, min?, max?, step?, label?, accept? } } }. Read-only; spends no credits and works without authentication.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelYesModel ID (e.g. "flux-2-pro")

Output Schema

ParametersJSON Schema
NameRequiredDescription
modelYes
schemaYes
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false. The description adds value by stating 'Read-only; spends no credits and works without authentication', which is congruent with annotations. It also describes the return structure, providing additional behavioral context beyond what annotations offer.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured. It starts with the primary purpose, then provides usage guidelines, required input, an example, and a note on credits/auth. Every sentence adds value, though it could be slightly more terse without losing clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple tool (1 param, output schema exists), the description is complete. It covers input, output structure, usage context, what the tool does and does not do, and behavioral notes (read-only, no credits). With an output schema present, there's no need to detail return values further, but the description still outlines the schema shape.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with one parameter 'model' fully described. The description adds an example usage '{ model: "flux-2-pro" }' and mentions the type, which slightly enhances understanding beyond the schema alone. The baseline for high coverage is 3, and the example justifies a 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it 'Returns the parameter schema for a specific Picsart model' and distinguishes from siblings by explicitly saying 'Do NOT use it to discover which models exist (use picsart_list_models) or estimate cost (use picsart_preflight)'. The verb 'Returns' and resource 'parameter schema' are specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit when-to-use guidance: 'once you have a model id and need to construct the params payload for picsart_generate or feed picsart_preflight with a candidate object'. It also gives when-not-to-use: 'Do NOT use it to discover which models exist (use picsart_list_models) or estimate cost (use picsart_preflight)'. Includes required input and an example.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_music_studioA
Read-onlyIdempotent
Inspect

Opens the Picsart Music Studio: browse music/audio models, compose with a guided prompt builder, generate and play tracks, create AI album-cover art, revisit previously generated tracks, and save everything into a "Music Studio" folder in the user's Picsart Drive. Use when the user wants to MAKE music, a song, a soundtrack, a jingle, or sound effects. Covers text-to-music (MiniMax Music v2, Google Lyria 3 Pro/Clip, ElevenLabs Music v2), short audio clips (Kling T2A), and sound effects (ElevenLabs SFX). Does NOT edit existing audio (no trimming, remixing, or stem work), and is not for text-to-speech / voice cloning or image/video generation. Takes no input. Returns { items, total, truncated } — the curated music catalog the widget renders. Read-only; spends no credits and works without authentication.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
itemsYes
totalYes
truncatedYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint, idempotentHint, destructiveHint), the description adds valuable context: 'Read-only; spends no credits and works without authentication' and the return shape '{ items, total, truncated }'. This explains the tool's side-effect profile and runtime behavior clearly. The 'create/save' language refers to the widget's UI actions, and the read-only clarification resolves any ambiguity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a dense, front-loaded paragraph where every sentence earns its place. It moves from core purpose to use cases, exclusions, input, return value, and behavioral characteristics in logical order. Despite its length, there is negligible waste and the structure makes it easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (many embedded capabilities), the description covers purpose, use cases, boundaries, parameter count, return type, authentication, and credit impact. The presence of an output schema means return-value details are already structured, and the description complements it rather than repeating it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and the schema coverage is 100%, so there is no parameter syntax to clarify. Per the calibration baseline, 0 params warrants a 4. The description appropriately notes 'Takes no input,' which adds useful confirmation without redundancy.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb ('Opens the Picsart Music Studio') and enumerates concrete capabilities (browse models, compose, generate/play, create album art, revisit tracks, save to folder). It strongly distinguishes itself by explicitly listing what it does NOT do (editing audio, TTS, image/video generation), which separates it from sibling tools like picsart_media_apply_effect or picsart_generate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use: 'Use when the user wants to MAKE music, a song, a soundtrack, a jingle, or sound effects.' It also gives clear exclusions: 'Does NOT edit existing audio... not for text-to-speech / voice cloning or image/video generation.' This provides strong when-to-use and when-not-to-use guidance, even without naming alternative tools directly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_preflightA
Read-onlyIdempotent
Inspect

Free pre-flight check before picsart_generate: in ONE call it (1) validates a candidate params object against the model's parameter schema + inter-parameter constraints, and (2) quotes the credit cost — without running the model or charging the user. Use this after assembling params (user input, derived defaults, model swaps) and before generating, to surface bad arguments and show cost. Do NOT use it to look up which params a model accepts (use picsart_model_params) or to actually generate (use picsart_generate). Required inputs: model id and a params object (put the prompt inside params). Example: { model: "flux-2-pro", params: { prompt: "a cat in a hat", aspectRatio: "1:1", count: 1 } }. Returns { model, valid, errors?, credits }: valid/errors are from local validation (always present, no auth needed; errors only when invalid); credits is the dry-run cost (a number), or null when pricing is unavailable or the request is unauthenticated.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelYesModel ID
paramsYesCandidate params (include the prompt). Validated locally and priced.

Output Schema

ParametersJSON Schema
NameRequiredDescription
modelYes
validYes
errorsNo
creditsYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, and the description aligns with these. It adds non-obvious behavioral details: the validation is local and requires no auth, the credits quote is a dry-run cost, and it returns null when pricing is unavailable or unauthenticated. This goes beyond the annotation signal to disclose actual side-effect-free behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact yet information-dense, using numbered flow, an example, and a return format spec. Every sentence adds value: purpose, usage timing, exclusions, required inputs, example, output semantics. No fluff or repetition of schema fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (nested params, optional errors, conditional credits) and the presence of an output schema, the description covers all necessary context: what valid/errors mean, when credits is null, and auth implications. It is complete enough for an agent to invoke correctly without further research.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already describes both parameters and notes that params should include the prompt, so baseline is 3. The description adds a concrete example (`{ model: "flux-2-pro", params: { prompt: "a cat in a hat", aspectRatio: "1:1", count: 1 } }`) and clarifies that the prompt goes inside the params object, which is useful for correct invocation. It doesn't over-explain since schema coverage is 100%.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-plus-resource statement: 'Free pre-flight check before `picsart_generate`'. It clearly enumerates the tool's two functions (validate params, quote cost) and distinguishes it from siblings by explicitly naming alternatives (`picsart_model_params`, `picsart_generate`). No ambiguity remains about what this tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit timing: 'Use this after assembling params ... and before generating'. It also provides direct exclusions with named alternatives: 'Do NOT use it to look up which params a model accepts (use `picsart_model_params`) or to actually generate (use `picsart_generate`)'. This is textbook usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_remove_bgA
Destructive
Inspect

Removes the background from an image, returning a transparent cutout of the foreground subject. Auto-picks the newest enabled Picsart remove-bg model unless overridden via the model param — no need to call picsart_list_models first. Use this when the user asks to "remove the background", "cut out the subject", or "make the background transparent". Do NOT use this to replace the background with a new scene (use picsart_change_bg), upscale or sharpen the result (use picsart_enhance), convert raster to SVG (use picsart_vectorize), or generate a new image from scratch (use picsart_generate). Required input: image — a publicly-accessible URL. Local files are not supported; if you only have a local file, first make it available as a public or app-authorized URL. Optional: model to pin a specific remove-bg model, outputFormat (e.g. "png"). Example: { image: "https://example.com/portrait.jpg" }. Returns { assets, id, model, created_at, summary, why_relevant, url, results: [{ url, metadata? }], drive? } as a single JSON text block plus matching structuredContent (no resource_link blocks — the widget is the single source of visual truth, so result URLs are not duplicated as separate content blocks). id is the SDK's generation handle; metadata may include model-specific tags. Spends credits. Requires Authorization: Bearer .

ParametersJSON Schema
NameRequiredDescriptionDefault
imageYesInput image URL
modelNoOverride model ID (e.g. "picsart-sod-v8-2")
outputFormatNoOutput format (e.g. "png")

Output Schema

ParametersJSON Schema
NameRequiredDescription
idNo
urlNo
urlsNo
driveNo
modelNo
assetsYes
promptNo
videosNo
resultsNo
summaryNo
finalUrlNo
completedNo
created_atNo
iterationsNo
why_relevantNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

While annotations indicate readOnlyHint=false and destructiveHint=true, the description adds significant context: auto-picks model, requires Authorization, spends credits, handles local files, and explains the return format (single JSON text block, no resource_link blocks, widget as source of truth). No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is detailed but every sentence adds value: purpose, usage, exclusions, parameters, example, return structure, and auth. Well-organized and front-loaded with the core action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with auth, credit costs, model selection, and output quirks, the description covers all necessary context. Even with an output schema, it explains the return structure and critical integration details (single JSON block, no duplicate result blocks).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but description adds crucial semantics: 'image' must be publicly accessible, 'model' is an optional override with auto-pick behavior, and 'outputFormat' is exemplified. This goes well beyond the schema baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'Removes the background from an image, returning a transparent cutout of the foreground subject.' It uses a specific verb and resource, and explicitly distinguishes itself from sibling tools (e.g., 'use picsart_change_bg' for replacement).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit usage guidance: 'Use this when the user asks to remove the background...' and explicitly lists exclusions with alternative tools. Also notes no need to call picsart_list_models first, and covers required input constraints (public URL, no local files).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_save_assetAInspect

INTERNAL — invoked by the Picsart Asset Sheet widget when the user saves an asset. Do NOT call this directly from chat: the required arguments (descriptor lines, portraitUrl, sheet version) are only meaningful after the user has worked the sheet. From chat, use picsart_asset_sheet to open the board instead. On invocation it persists the passport to Picsart Drive — creates (or reuses) "Film Assets", adds a per-asset subfolder, and saves the portrait (carrying the manifest in attributes) plus an optional turnaround sheet and voice sample. Returns { saved, assetId, folder, portraitUid, voiceUid?, manifest, message }. Writes to Picsart Drive; requires Authorization: Bearer .

ParametersJSON Schema
NameRequiredDescriptionDefault
tagYesThe asset's @tag.
nameYesAsset display name; also the Drive subfolder name.
voiceNoCharacter assets only: voice configuration and generated sample URL.
scenesNoScene ids where this asset appears.
statusYes
assetTypeYes
descriptorYes
portraitUrlYesPicsart-hosted HTTPS URL of the asset's primary reference image.
sheetVersionYese.g. "v1" — a new version is a new save, never an overwrite.
turnaroundUrlNoHTTPS URL of the optional passport/turnaround sheet, when one was generated.
referenceImagesNoAdditional angle reference URLs (HTTPS only).

Output Schema

ParametersJSON Schema
NameRequiredDescription
tagNo
savedYes
folderYes
assetIdYes
messageYes
manifestYes
voiceUidNo
portraitUidNo
turnaroundUidNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (which indicate readOnlyHint=false, destructiveHint=false, openWorldHint=true), the description adds crucial behavioral traits: it persists data to Picsart Drive, creates or reuses folders, and requires a Picsart token. It also describes side effects (creating subfolders, saving optional files) that are not captured by annotations. The contradiction flag is false because annotations and description are consistent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured: it starts with a core restriction ('INTERNAL'), then provides usage guidance, followed by what the tool does, its return value, and side effects. Every sentence adds value, but the description is relatively long (4 sentences). It earns a 4 because it packs essential context without verbosity, though some details could be slightly tighter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (11 params, nested objects, internal usage, side effects on Drive), the description covers all critical aspects: invocation context, forbidden direct usage, behavior (create/reuse folders), persistence location, return shape, and auth requirement. The output schema exists, so the description doesn't need to detail return fields, but it still lists them briefly. It is fully complete for an AI agent to determine when and how to use this tool safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 73%, which provides decent baseline. The description adds context that `name` becomes the Drive subfolder name and `sheetVersion` implies versioning ('a new version is a new save, never an overwrite'). However, it does not elaborate on the meaning of `tag` or `scenes` beyond what the schema provides. For 11 parameters with nested objects, the description could add more value, but overall it enhances understanding sufficiently.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it is an internal tool that saves an asset to Picsart Drive when invoked by the Picsart Asset Sheet widget. It specifies the created resources (folder, portrait, voice sample) and explicitly distinguishes it from the sibling `picsart_asset_sheet`, ensuring the agent understands this is not a direct chat tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states 'Do NOT call this directly from chat' and provides the contextual prerequisite ('after the user has worked the sheet'). It also directs the agent to use the sibling `picsart_asset_sheet` for opening the board from chat, giving clear when-to-use and alternatives guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_scene_reviewA
Read-onlyIdempotent
Inspect

Opens the Picsart Scene Review storyboard: the assembled cut as a list of clips the user can reorder, drop, and comment on, plus the seam transition to pick. Their decisions come back as a JSON message in the conversation. Call this after assembling a scene (picsart_media_apply_scene_template + picsart_media_validate_scene) and BEFORE any charged render — and never paste the scene JSON into the conversation; this board is how a scene is shown to a person. Translate the scene into the storyboard args yourself: one entry per clip in playback order with its played duration, trim window, and the prompt that produced it. On feedback, apply the changes by re-running the montage bootstrap with the new clip order, removals, and transition (a free operation), then re-validate. Pass Picsart-hosted https URLs only; other hosts are blocked by the widget sandbox. Returns the normalised storyboard payload the widget renders. Read-only; spends no credits and works without authentication.

ParametersJSON Schema
NameRequiredDescriptionDefault
clipsYesThe clips in playback order — this IS the storyboard.
notesNoOne line of guidance shown above the storyboard.
titleNoHeading for the board, e.g. "Hero montage — cut 1"
canvasNoThe composition size the scene renders at
transitionNoThe transition currently applied at every seam (montage has one for all seams). Empty string or omitted = hard cut.
allowedTransitionsNoTransition ids the user may switch to (from picsart_media_get_capabilities). Always include "" for a hard cut. Defaults to a safe basic set.
transitionDurationNoSeconds the transition takes, e.g. 0.5

Output Schema

ParametersJSON Schema
NameRequiredDescription
clipsYes
notesNo
titleNo
canvasNo
transitionYes
totalDurationYes
allowedTransitionsYes
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint: true, idempotentHint: true, and destructiveHint: false, so the bar for behavioral transparency is lower. The description adds value by stating 'Read-only; spends no credits and works without authentication' and explains that the tool returns a 'JSON message in the conversation' and 'normalised storyboard payload'. However, it doesn't detail the exact structure of the returned JSON or the format of the user feedback message, which would be helpful but not mandatory given the output schema exists. Score 4.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long (over 150 words) and front-loaded with a clear purpose sentence, but it meanders into multiple workflow instructions ('On feedback, apply the changes...', 'Pass Picsart-hosted https URLs only...') that could be split or shortened. It earns its space by conveying essential guidelines, but the density of information could be better organized for quick scanning. Score 3.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (7 parameters, nested objects, output schema) and the richness of sibling tools, the description covers the core workflow, prerequisites, restrictions, and post-call actions. It references sibling tools and capabilities (picsart_media_get_capabilities) which helps with context. The only gap is not describing what happens in the feedback loop after board interaction, but the return of a JSON message is mentioned. Score 4.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters. The description adds meaning by explaining the intended workflow context ('one entry per clip in playback order', 'Translate the scene into the storyboard args yourself') and provides semantic guidance for constructing the clips array. It also clarifies that the tool expects Picsart-hosted URLs and translates the scene automatically. There's still room to explain the canvas or transition usage more, but this is strong context beyond the schema. Score 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear, specific verb ('Opens the Picsart Scene Review storyboard') and describes the concrete resource ('assembled cut as a list of clips... plus the seam transition'). It explicitly distinguishes this tool from siblings by detailing its unique interactive purpose (user review, reorder, comment, pick transition) and contrasts it with tools like picsart_asset_review and picsart_video_review in the context of scene editing. This is a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-call instructions: 'Call this after assembling a scene (picsart_media_apply_scene_template + picsart_media_validate_scene) and BEFORE any charged render'. It also states a critical when-not ('never paste the scene JSON into the conversation; this board is how a scene is shown to a person'), provides a sibling-like alternative workflow ('On feedback, apply the changes by re-running the montage bootstrap... then re-validate'), and warns about blocked URL hosts. This is a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_shot_designerA
Read-onlyIdempotent
Inspect

Opens the Picsart Shot Designer: one film shot designed with presets instead of prose — camera movement, shot size, field of view, aperture and angle, a two-source lighting console, per-character acting (emotion wheel, gaze, hidden intention), and format. Every selection compiles to an exact prompt-block text; the user's design, the compiled optics/camera/lighting/acting/format blocks, and a generate-or-hold verdict come back as a JSON message in the conversation — paste the blocks into the shot's fifteen-block prompt verbatim, never reworded. The widget dispatches its own draft takes via picsart_generate (charged, confirmed in-widget) and saves them to a "Shot Designer" Drive folder. Call it per shot during scene generation, passing the shot card's action, the world's style prefix and tempo, and everyone in frame — and seed suggested from what the shot card and world defaults already decide (size, move, FOV, duration, register), so the console opens pre-set and only deviations need touching. Returns the normalised shot context the widget renders. Read-only; this call spends no credits and works without authentication.

ParametersJSON Schema
NameRequiredDescriptionDefault
filmYesFilm title — also the Drive folder draft takes save into
shotYesShot id within the scene, e.g. "12B"
draftNoThe shot's latest draft take, if one exists — shown on the stage
sceneYesScene id, e.g. "sc01"
actionYesThe shot card's action line — what physically happens, one action
currentNoThe shot's existing design, when revising
modelIdNoThe film's video model. Known ids: seedance-2.0, kling-v3 — anything else falls back to seedance-2.0
presetsNoSaved signature setups from film.json.presets
suggestedNoYour inferred starting design — from the shot card's lanes (size, move, FOV, duration), the world defaults (tempo → camera register), and the scene's light. ALWAYS fill what the card already answers: the console opens pre-set and the user only touches deviations. `current` wins over `suggested` field by field.
aspectRatioNoThe film's aspect ratio, e.g. "16:9" (the default) or "9:16"
suggestedWhyNoOne line per suggested field naming what decided it, e.g. { "size": "shot card 12B says MCU" }
worldDefaultsYesDefaults inherited from the film-setup console
charactersInFrameYesEveryone in this shot, by asset tag — drives the Acting tab

Output Schema

ParametersJSON Schema
NameRequiredDescription
filmYes
noteNo
shotYes
draftNo
modelYes
sceneYes
actionYes
currentNo
presetsNo
suggestedNo
aspectRatioYes
suggestedWhyNo
worldDefaultsYes
charactersInFrameYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds valuable context: 'Read-only; this call spends no credits and works without authentication,' and discloses that the widget dispatches draft takes via picsart_generate (a charged operation), which goes beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but well-organized, front-loading the purpose and then detailing behavior, usage, and output. Every sentence serves a purpose, and the structure is clear with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 13 parameters, nested objects, and an output schema, the description covers the tool's role in scene generation, the console interaction, output format (JSON message), and integration with other tools. It is complete enough for an agent to understand the tool's context and correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description adds no significant parameter-specific semantics beyond the schema's own descriptions. It does explain the high-level usage of action, worldDefaults, charactersInFrame, and suggested, but this is more about workflow than parameter detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly defines the tool as opening the Picsart Shot Designer for shot design with presets, detailing components like camera, lighting, acting. It distinguishes itself from sibling tools by referencing picsart_generate and the Drive folder, and explains its role in scene generation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit guidance is provided: 'Call it per shot during scene generation, passing the shot card's action, the world's style prefix and tempo, and everyone in frame — and seed suggested from what the shot card and world defaults already decide.' This tells the agent when and how to invoke the tool, and includes post-use instructions for handling the output blocks.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_shotlist_boardA
Read-onlyIdempotent
Inspect

Opens the Picsart Shotlist Board: a scene's shot cards as an editable list — sizes, moves, FOV, durations, and asset tags per shot — with reorder and cut controls, and in conveyor mode a status chip plus take tiles to accept a take or order a reshoot. Their decisions come back as a JSON message in the conversation (type shotlist_board_feedback). Use it twice in a film: to review the preliminary shotlist before assets are made (mode "plan"), and per scene during generation (mode "conveyor") — and never paste the shotlist as raw JSON into the conversation; this board is how a shotlist is shown to a person. On feedback, apply card edits to the shot cards and recompile only the owning prompt block per changed key, promote accepted takes to selects, and fold reshoot notes in surgically. The board performs no generation itself. Pass Picsart-hosted https URLs only for takes; other hosts are blocked by the widget sandbox. Returns the normalised shotlist payload the widget renders. Read-only; spends no credits and works without authentication.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoplan = preliminary shot cards (no takes yet); conveyor = generation tracking with statuses and take tiles. Defaults to plan.
noteNoOne line of guidance shown above the board.
slugYesSlug line, e.g. "EXT. GARAGE — DAY"
sceneYesScene id, e.g. "sc01"
shotsYesThe scene's shots in intended order — this IS the shotlist.

Output Schema

ParametersJSON Schema
NameRequiredDescription
modeYes
noteNo
slugYes
sceneYes
shotsYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds valuable context: 'Read-only; spends no credits and works without authentication', 'Returns the normalised shotlist payload', 'The board performs no generation itself', and explains that human decisions come back as a JSON feedback message. Also warns about widget sandbox blocking non-Picsart URLs. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is comprehensive but not overly verbose. It front-loads the main action, then lists features, usage, and behavioral notes in a logical order. Every sentence contributes unique information. Could be slightly tighter, but for the complexity of the tool it is well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (5 parameters, nested objects, two modes, feedback loop) and the presence of an output schema, the description covers all key aspects: what the board shows, mode usage, feedback handling, URL restrictions, read-only nature, and expected post-processing. The only minor omission is an explicit statement that the board waits for human input, but that is implied by 'Their decisions come back as a JSON message'.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds extra value by explaining the two modes in context, reinforcing the URL restriction for takes, and clarifying that the `shots` array IS the shotlist. This goes beyond the schema's per-field descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool opens a specific UI board ('Picsart Shotlist Board') for editing scene shot cards with reorder/cut controls and take tiles in two modes. It distinguishes itself from sibling tools by emphasizing that this is the proper way to present a shotlist to a human, not as raw JSON, and that it performs no generation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use: 'Use it twice in a film: to review the preliminary shotlist before assets are made (mode "plan"), and per scene during generation (mode "conveyor")'. Also gives a clear negative instruction: 'never paste the shotlist as raw JSON into the conversation; this board is how a shotlist is shown to a person'. Additionally mentions URL restrictions ('Pass Picsart-hosted https URLs only').

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_vectorizeA
Destructive
Inspect

Converts a raster image (PNG, JPG) into an SVG vector. Auto-picks the newest enabled Picsart vectorize model unless overridden via the model param. Use this when the user asks to "vectorize", "convert to SVG", "make this a vector", or wants a scalable version of a logo or icon. Best results on logos, icons, and simple graphics — photographic images vectorize poorly and the user should be warned. Do NOT use this to remove the background (use picsart_remove_bg), replace the background (use picsart_change_bg), upscale a raster image (use picsart_enhance), or generate a new image (use picsart_generate). Required input: image — a publicly-accessible URL to a PNG or JPG (not a local file path). Optional: model to pin a specific vectorize model. Example: { image: "https://example.com/logo.png" }. Returns { assets, id, model, created_at, summary, why_relevant, url, results: [{ url, metadata? }], drive? } as a single JSON text block plus matching structuredContent (no resource_link block for the SVG URL — the widget is the single source of visual truth, so it is not duplicated as a separate content block). id is the SDK's generation handle. Clients fetch the SVG from that URL. Spends credits. Requires Authorization: Bearer .

ParametersJSON Schema
NameRequiredDescriptionDefault
imageYesInput image URL
modelNoOverride model ID

Output Schema

ParametersJSON Schema
NameRequiredDescription
idNo
urlNo
urlsNo
driveNo
modelNo
assetsYes
promptNo
videosNo
resultsNo
summaryNo
finalUrlNo
completedNo
created_atNo
iterationsNo
why_relevantNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description explains that the newest model is auto-picked, that photographic images vectorize poorly and the user should be warned, that it spends credits, and requires Bearer auth. It also clarifies the output format (single JSON text block plus structuredContent, no separate resource_link). Annotations do not cover these details, and there is no contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than average but every section earns its place: function, model selection, use cases, warnings, exclusions, input requirements, output format, side effects, and auth. It is front-loaded with the core purpose and structured with clear examples.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the two-parameter tool with an output schema, the description covers all operational aspects: input format, model selection, expected output structure, client fetching behavior, cost, and authentication. It also explicitly compares to five sibling tools, leaving no ambiguity about when to invoke it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description adds key constraints: image must be a publicly-accessible URL (not a local file) to a PNG or JPG, and model is optional to pin a specific vectorize model. This goes beyond the schema's bare 'Input image URL' and 'Override model ID'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Converts a raster image (PNG, JPG) into an SVG vector,' giving a specific verb and resource. It clearly differentiates from siblings by naming alternatives for background removal, enhancement, and generation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit when-to-use triggers ('when the user asks to "vectorize"') and explicit exclusions ('Do NOT use this to remove the background... use picsart_remove_bg'). This is comprehensive guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_video_compareA
Read-onlyIdempotent
Inspect

Opens the Picsart Video Compare board: two versions of a video playing in lockstep, with the stretches you changed marked on a shared timeline, so the user can see what moved and say which version wins. Call this whenever you have produced a NEW version of a video the user already saw — after regenerating a flagged segment, after changing a transition, after a resolution pass — rather than presenting the new one alone and hoping they remember the old. Pass the pair oldest-first and list what you changed in changedRanges. Pass Picsart-hosted https URLs only; other hosts are blocked by the widget sandbox. The verdict comes back as a JSON message in the conversation: the winning version, how strong the preference is, per-range notes, and a comment. Act on it — if the older version wins, restore its URL as the current cut rather than keeping the new one. For reviewing a single video on its own timeline use picsart_video_review instead. Returns the normalised board payload the widget renders. Read-only; spends no credits and works without authentication.

ParametersJSON Schema
NameRequiredDescriptionDefault
notesNoOne line summarising what you changed and why.
videosYesExactly two versions to compare, oldest first: [previous, new].
changedRangesNoThe stretches you actually changed between the two versions. These are marked on the shared timeline so the user looks where it matters instead of hunting the whole clip.

Output Schema

ParametersJSON Schema
NameRequiredDescription
notesNo
videosYes
changedRangesYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, idempotentHint=true. The description adds further behavioral context: 'Read-only; spends no credits and works without authentication.' It also explains that the verdict comes back as a JSON message in the conversation and instructs the agent on what to do next ('Act on it'). No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense paragraph, but every sentence earns its place. It is front-loaded with the primary purpose, then usage context, parameter details, behavioral notes, and sibling reference. While it covers everything well, it could be slightly more structured (e.g., separate sections) to aid quick scanning. However, no extraneous content exists.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (video comparison with verdict), full schema coverage, annotations, and an output schema, the description covers all necessary aspects: purpose, when to use, parameter expectations, behavioral traits (read-only, no auth), post-call actions, and explicit sibling differentiation. The return payload is mentioned ('normalised board payload') and the verdict format is described. No gaps remain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, but the description adds significant meaning beyond the schema: it specifies that videos must be Picsart-hosted https URLs (other hosts blocked), that the pair must be oldest-first, and that `changedRanges` marks the stretches actually changed. It explains that `generationMeta` is provenance shown under the player. `notes` is described as 'One line summarising what you changed and why.' This provides context the schema alone does not convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool opens a Picsart Video Compare board and explains its function: two video versions playing in lockstep with changed stretches on a shared timeline. It distinguishes from the sibling tool `picsart_video_review` by specifying that this is for comparing two versions while that is for single video review. The verb 'opens' plus resource 'Video Compare board' is specific.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use guidance is provided: 'Call this whenever you have produced a NEW version of a video the user already saw' with concrete examples like regenerating a segment or changing a transition. It also states when not to use: 'For reviewing a single video on its own timeline use picsart_video_review instead.' This directly tells the agent when to pick this tool over a sibling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

picsart_video_reviewA
Read-onlyIdempotent
Inspect

Opens the Picsart Video Review board for one video: a player with a thumbnail filmstrip timeline where the user selects the stretches that do not work, comments on each one, and marks manual cut points. Call this after every video draft — a picsart_generate result, a picsart_media_export assembly, a picsart_media_video_render output — instead of asking "how does it look?" in prose. Pass the shot list in segments when the video is a multi-segment assembly so each complaint lands on a known segment and its prompt. Pass Picsart-hosted https URLs only; other hosts are blocked by the widget sandbox. The user's verdict comes back as a JSON message in the conversation: per-range actions (keep, cut, regenerate) with comments, plus any manual cut in/out points. Act on it — regenerate only the ranges they flagged, folding their comment into that segment's prompt, and turn cut ranges into trim windows on the assembly rather than re-rendering the whole thing. Returns the normalised board payload the widget renders. Read-only; spends no credits and works without authentication.

ParametersJSON Schema
NameRequiredDescriptionDefault
notesNoOne line of guidance shown above the player.
titleNoHeading for the board, e.g. "Boot ad — draft 2"
segmentsNoThe shots this video was assembled from, in order. Supply these whenever the video is a multi-segment assembly: they become pre-marked regions, so a complaint maps straight back to the segment (and prompt) that needs regenerating.
videoUrlYesPicsart-hosted https URL of the video to review

Output Schema

ParametersJSON Schema
NameRequiredDescription
notesNo
videoYes
segmentsYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds that it's 'Read-only; spends no credits and works without authentication,' confirming the read-only nature. It details the interactive flow (user selects/comments, verdict comes back as JSON) and what the agent should do with results. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is comprehensive but slightly verbose—covers purpose, usage, parameters, behavior, and post-call actions in multiple paragraphs. It is well-structured and front-loaded with the core action, but could be tightened without losing clarity. Still very effective.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (interactive review board, multi-step response handling) and the presence of annotations, output schema (not shown but referenced), and full parameter descriptions, the description fully equips an agent: it explains the UI, when to call, parameter nuances, read-only nature, response format, and recommended follow-up actions. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description adds significant value: for 'title' it gives a concrete example; for 'segments' it explains why to supply them (maps complaints back to segments/prompts for regeneration); for 'videoUrl' it warns about host blocking. Only 'notes' is not extended, but overall the description meaningfully supplements the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it 'Opens the Picsart Video Review board for one video' with a specific role: a player with thumbnail filmstrip timeline for selecting non-working stretches, commenting, and marking cut points. It distinguishes from siblings like picsart_asset_review (image-focused) or picsart_video_compare by specifying the video review workflow and when to call it after generation/export/render.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Call this after every video draft' listing specific predecessors (picsart_generate, picsart_media_export, picsart_media_video_render) and advises against asking 'how does it look?' in prose. It also guides when to pass segments, restricts URLs to Picsart-hosted, and instructs how to act on the verdict (regenerate flagged ranges only, turn cut ranges into trim windows). This is complete when-to and how-to guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Discussions

No comments yet. Be the first to start the discussion!

Related MCP Servers

View all MCP Servers

Try in Browser

Your Connectors

Sign in to create a connector for this server.

Resources