higgsfield-mcp-unified
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@higgsfield-mcp-unifiedcan you generate a photorealistic image of a cat in a spacesuit?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
higgsfield-mcp-unified
Unified Model Context Protocol (MCP) server for Higgsfield AI. It puts 43 Higgsfield image and video models — Sora 2, Veo 3.x, Kling 3.0, Seedance 2.0, Wan, Soul, DOP, Nano Banana, FLUX, and more — plus Soul character training, account/history, and talking-head speech behind a single, typed tool surface for any MCP client (Claude Desktop, Claude Code, Cursor, …).
Status: alpha. Local/self-hosted and not yet on PyPI — install from source (below). The cloud web backend is opt-in and experimental.
Why this over the hosted MCP
The official hosted MCP (mcp.higgsfield.ai) is great but runs on Higgsfield's servers behind OAuth. This one is local-first and adds things a hosted server can't:
Runs on your machine — prompts and media go straight to Higgsfield, no intermediary proxy.
Dual backend under one surface — routes per model across the official REST API and the cloud web app.
Typed structured output — every tool returns a schema'd result (
outputSchema+structuredContent), not an opaque blob.Discovery without burning a generation —
recommend_model,validate_params,preflight_check, and an MCP resource catalog.Reliability built in — retries with backoff + jitter (honoring
Retry-After), a circuit breaker, idempotency keys, and a structured error taxonomy.MCP prompt templates — reusable cinematic/product scaffolds.
Related MCP server: Higgsfield Unlimited MCP
Backends
Official REST API (
platform.higgsfield.ai) — stable,KEY:SECRETauth. Default; rock-solid.Cloud web app (
fnf.higgsfield.ai) — the newest models, reachable only via a Clerk cookie. Opt-in and experimental (see warning below).
Install (from source)
Not published to PyPI yet. Clone and run with uv:
git clone https://github.com/Hikhakk/higgsfield-mcp-unified
cd higgsfield-mcp-unified
uv sync
uv run higgsfield-mcp # starts the stdio MCP serverConfigure
The official backend needs two environment variables:
export HIGGSFIELD_API_KEY=... # from the platform.higgsfield.ai dashboard
export HIGGSFIELD_SECRET=...To unlock the cloud-only models (Sora 2 / Veo 3.x / Kling 3.0 / …), opt in to the web backend:
export HIGGSFIELD_ENABLE_WEB_BACKEND=1
export HIGGSFIELD_CLERK_CLIENT=... # __client cookie from cloud.higgsfield.ai (lasts ~7 days)
# or, for one-shot tests only:
export HIGGSFIELD_JWT=... # __session cookie (expires in ~1 minute)Run preflight_check from your client to confirm both backends are reachable before generating.
Web backend is experimental and will likely break
The
cloud.higgsfield.ai/fnf.higgsfield.aisurface is not a public API — it is the consumer web app's private backend, integrated by reverse-engineering. Expect:
Auth churn. The Clerk JWT lives ~1 minute. The server refreshes it from a long-lived
__clientcookie, but Clerk rotates that cookie too (~7 days). When it rotates, paste a new one.Bot protection.
fnf.higgsfield.aisits behind Cloudflare's managed challenge (TLS fingerprinting). Browser-impersonating TLS (curl_cffi) clears it today, but it is brittle.Schema drift. Endpoint paths, body keys, and slugs are undocumented and change without notice. Several of the newest model endpoints in this server are marked
inferred(best-guess, hidden by default) until verified against live traffic.Probably against ToS. Driving the consumer app programmatically is almost certainly unsupported. Review the terms before enabling it.
Off by default. Without
HIGGSFIELD_ENABLE_WEB_BACKEND=1, only the official-API models are reachable, and that path is solid.
Claude Desktop
~/Library/Application Support/Claude/claude_desktop_config.json:
{
"mcpServers": {
"higgsfield": {
"command": "uv",
"args": ["run", "--directory", "/path/to/higgsfield-mcp-unified", "higgsfield-mcp"],
"env": {
"HIGGSFIELD_API_KEY": "...",
"HIGGSFIELD_SECRET": "..."
}
}
}
}Cursor / Claude Code / any stdio MCP client: point at the same uv run … higgsfield-mcp command and pass env vars. See examples/.
Tools
Tool | What it does |
| The model registry. Inferred (unverified) models are hidden unless |
| Rank models for a described goal (local, no API call). |
| Check params against a model's supported set before submitting. |
| Validate auth + reachability for both backends without spending a generation. |
| Submit a text-to-image / image-edit job (supports Soul character refs). |
| Submit a text-to-video / image-to-video job. |
| Fan out multiple image/video submits concurrently. |
| Talking-head video from a face image + WAV audio. |
| Poll, or long-poll until terminal, with output URLs. |
| Cancel a queued/in-progress job. |
| Upload a local image and get a hosted URL. |
| Train and manage reusable Soul characters. |
| Soul style and DOP motion presets, by name. |
| Available credits + plan (official backend). |
| Recent generations (history). |
Resources & prompts
Resources:
higgsfield://models(full catalog) andhiggsfield://models/{kind}— subscribe to the catalog instead of repeatedly callinglist_models.Prompts:
cinematic_shot,product_360,animate_portrait,action_sequence,b_roll— parameterized scaffolds to feed into the generation tools.
Models
Run list_models() (or read the higgsfield://models resource) for the live catalog. The registry carries a confidence tier:
verified— endpoint confirmed against SDK source / live probe (shown by default).inferred— slug confirmed real but the endpoint is a best-guess pending live verification (hidden unlessinclude_unverified=true, and noted as such).
Official backend (verified): Soul, Reve, Seedream v4 (text-to-image + edit), FLUX.1 Kontext Max, DOP (preview/standard), Seedance v1 Pro, Kling v2.1 Pro.
Cloud backend: Kling 3.0 / O3 FLF, Seedance 2.0 / 1.5, Wan 2.6, Veo 3, Grok, Sora 2, Nano Banana, Soul v2, OpenAI Hazel, and newest-tier inferred entries (Veo 3.1, Wan 2.7, Kling 3.0 Turbo, Hailuo 02, FLUX.2, Z-Image, Cinematic Studio, …).
Development
uv sync --extra dev
uv run --extra dev pytest -q
uv run --extra dev ruff check . && uv run --extra dev ruff format --check .
uv run --extra dev mypy srcDesign specs and phased implementation plans live in docs/superpowers/.
Contributing
PRs welcome. See CONTRIBUTING.md. Merges to main require code-owner approval (@Hikhakk) and green CI.
Credits
Built on two earlier community efforts:
geopopos/geo_higgsfield_ai_mcp— first Python MCP for the official API.jfikrat/higgsfield-mcp— first MCP to expose the cloud web models; source of the original model registry.
License
MIT — see LICENSE.
Available Tools
20 toolscancel_job_toolC
Cancel a queued or in-progress job.
| Name | Required | Description | Default |
|---|---|---|---|
| job_handle | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| cancelled | Yes | |
| job_handle | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it discloses almost nothing beyond the one-line action. It does not state that cancellation is irreversible, whether in-progress work is stopped or merely flagged, permission requirements, or what happens to jobs that have already completed. For a state-mutating tool this is a substantial gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the action and object front-loaded and zero filler. It is appropriately sized, though its brevity is partly the cause of the missing behavioral detail rather than pure efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. However, for a destructive, unannotated mutation with one entirely undocumented parameter, the description omits the preconditions, irreversibility, and handle provenance an agent needs to invoke it safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single parameter 'job_handle', and the description adds nothing about its origin or format (e.g., where to obtain it from list_jobs_tool or a prior generate call). The parameter is documented only by its name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a clear verb+resource ('Cancel a ... job') and even scopes it to queued or in-progress states, which is more specific than the bare tool name. It does not, however, name or differentiate itself from siblings such as delete_character_tool or list_jobs_tool, so the agent gets no explicit routing signal.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this versus alternatives (e.g., list_jobs_tool to find the handle, get_status_tool to check state first). Nothing is said about whether already-completed jobs can be cancelled or what to do if the job is finished, leaving the agent to infer all preconditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_character_toolC
Train a reusable Soul character from reference images. Results best-effort until verified.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| image_urls | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| id | Yes | |
| raw | No | |
| name | Yes | |
| status | Yes | |
| image_count | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose one real behavioral trait – the training is best-effort and results are unverified until checked – which is useful beyond the schema. It says nothing about long-running/async behavior, whether a job id is returned (despite list_jobs_tool and get_status_tool siblings), auth requirements, or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler, and the core action is front-loaded. The caveat sentence is terse and earns its place, though the description is arguably too lean for the gaps it leaves.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-shape explanation is not required. However, for a training tool with two undocumented required parameters, no annotations, and asynchronous siblings in the toolset, the description omits the parameter meaning, workflow prerequisites, and job-tracking context an agent needs to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there are two required parameters. 'Reference images' loosely maps to image_urls but gives no format, count, or source constraints (e.g., must come from upload_image_tool), and the required 'name' parameter is never explained at all.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource combination ('Train a reusable Soul character from reference images') that is distinguishable from siblings like generate_image_tool or upload_image_tool. It does not explicitly contrast itself with any named sibling, which keeps it short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus alternatives such as generate_image_tool (create a one-off image), upload_image_tool (stage a reference image), or preflight_check_tool (validate first). 'Results best-effort until verified' hints at a caveat but never states prerequisites or the intended workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_character_toolC
Delete a trained Soul character.
| Name | Required | Description | Default |
|---|---|---|---|
| character_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| deleted | Yes | |
| character_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure, yet it says nothing about permanence, reversibility, authorization requirements, or side effects of deleting a trained character. 'Delete' implies destruction but the description never confirms whether it is irreversible.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with zero filler, front-loading the verb and resource. Nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but for a destructive tool with no annotations the description omits the essential behavioral context an agent needs (permanence, side effects, required permissions). It is under-specified for the operation's risk level.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one required parameter (character_id) with 0% schema description coverage, and the description adds no meaning: no format, no source (e.g., from list_characters_tool), no constraints. With low coverage the description needed to compensate and does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Delete') and resource ('trained Soul character'), which clearly separates it from sibling tools like create_character_tool and get_character_tool. It stops short of 5 because it does not name an alternative or clarify scope (e.g., whether the character must be untrained/uncited first).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this versus alternatives, no preconditions, and no mention of confirmation or prerequisites. Usage is only implied by the verb.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_batch_toolB
Submit multiple generations at once. Each request: {kind, model_id, prompt, ...params}.
| Name | Required | Description | Default |
|---|---|---|---|
| requests | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| count | Yes | |
| results | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It says 'submit' but not whether this is asynchronous (job IDs, list_jobs/get_status siblings suggest it likely is), what happens on partial failure, or any auth/rate-limit constraints. Only bare minimum is conveyed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and followed by the per-item shape. Nothing wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. Still, for a batch mutation with no annotations and a single opaque array parameter, the description omits partial-failure behavior, whether batches are homogeneous, and sync/async expectations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the schema only declares an untyped array of free-form objects, so the description's '{kind, model_id, prompt, ...params}' is a genuine value add. However, the '...params' tail is vague and the semantics of 'kind' (which values map to which generation type) are not explained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: submit multiple generations at once. This distinguishes it from the single-generation siblings (generate_image_tool, generate_video_tool) by the word 'multiple', but it never names those siblings or clarifies the batch-versus-single split explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to choose batch submission over individual generation calls, no note on ordering, limits, or whether mixing kinds in one batch is allowed. The agent is left to infer usage entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_image_toolC
Submit an image-generation request. Returns a job_handle to poll.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | ||
| prompt | Yes | ||
| quality | No | ||
| soul_id | No | ||
| model_id | Yes | ||
| image_url | No | ||
| batch_size | No | ||
| resolution | No | ||
| aspect_ratio | No | ||
| soul_strength | No | ||
| enhance_prompt | No | ||
| input_image_urls | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| backend | Yes | |
| model_id | Yes | |
| job_handle | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations the description carries the full burden, but it does disclose the important async pattern ('Returns a job_handle to poll'), telling the agent the result is deferred and must be polled via get_status_tool. It omits cost/credit implications, auth requirements, and failure behavior for a 12-parameter generation call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight, front-loaded sentences with no filler. It is efficient, though its brevity is arguably under-specification rather than disciplined conciseness given the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values are partly covered, but for a 12-parameter generation tool with no annotations and 0% param coverage the description is far too thin. It should at minimum explain the model/soul/image input modes and the async polling workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 12 parameters (seed, quality, soul_id, soul_strength, enhance_prompt, input_image_urls, etc.), and the description adds zero parameter meaning. The agent has no guidance on model_id format, soul vs input-image modes, or how optional fields interact.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Submit an image-generation request'), which distinguishes it from the video/speech generation siblings at a high level. However, it never explicitly differentiates itself from the many other generate_* tools (generate_video_tool, generate_batch_tool) beyond the resource noun.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no mention of related siblings like preflight_check_tool, validate_params_tool, recommend_model_tool, or list_models_tool that clearly gate this call. The agent must infer the entire usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_speech_video_toolC
Talking-head video from a face image + WAV audio (official backend).
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | No | ||
| audio_url | Yes | ||
| image_url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| backend | Yes | |
| model_id | Yes | |
| job_handle | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, yet it discloses almost nothing: it does not say whether generation is async (siblings like get_status_tool and cancel_job_tool imply it is), whether it costs credits, what the output artifact is, or what constraints apply to inputs. '(official backend)' is the only behavioral hint and it is unexplained.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler, which is structurally clean. However, at this length for a 3-parameter generation tool it reads as under-specified rather than efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. But with zero annotations, zero schema descriptions, and an undocumented prompt parameter, the definition leaves an agent without the async/job, cost, or input-format context needed to invoke it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It loosely maps 'face image' to image_url and 'WAV audio' to audio_url, but says nothing about the third parameter, prompt, or about URL formats, accepted image types, or audio length limits.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: generating a talking-head video from a face image plus WAV audio. That clearly separates it from generate_video_tool and generate_image_tool. The only gap is the unexplained '(official backend)' qualifier, which hints at an alternative backend without naming it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to choose this tool over generate_video_tool or another speech-video path, and no prerequisites or exclusions. The required image+audio pairing is inferable only from the schema, not the description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_video_toolC
Submit a video-generation request. Returns a job_handle to poll.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | ||
| sound | No | ||
| prompt | Yes | ||
| duration | No | ||
| model_id | Yes | ||
| image_url | No | ||
| resolution | No | ||
| end_image_url | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| backend | Yes | |
| model_id | Yes | |
| job_handle | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses one behavioral trait—that the operation is asynchronous and returns a job_handle to poll—but says nothing about authentication, cost, side effects, or what happens to submitted jobs. This is insufficient for a submission tool with eight parameters.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and followed by the return behavior. There is no wasted text, though the extreme brevity may under-serve such a parameter-heavy tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, 0% schema coverage, and eight parameters, the description is far too thin. It mentions the async return but omits parameter semantics, usage context, and behavioral details. The existence of an output schema excuses return-value explanation but not the other gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description adds no meaning for any of the eight parameters. Required fields like model_id and prompt, as well as optional fields like seed, sound, and duration, are left completely unexplained beyond their names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Submit') and resource ('video-generation request'), and mentions the return value ('job_handle'). It distinguishes itself from image or speech siblings by naming video generation, though it doesn't explicitly differentiate from generate_speech_video_tool or generate_batch_tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives like generate_image_tool or generate_speech_video_tool. It only implies usage through the tool name and purpose statement, with no prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_balance_toolA
Get available credits + plan (best-effort; official API key required).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| raw | No | |
| plan | No | |
| credits | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose two real behavioral traits: it is 'best-effort' (results may be incomplete or fail) and it needs an official API key. It does not mention rate limits or what happens on failure, but the auth and reliability caveats are substantive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the resource named first and constraints in a compact parenthetical. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the simple zero-parameter read is adequately covered by the purpose, auth requirement, and best-effort caveat. Only the concrete meaning of 'best-effort' (stale data? partial failure?) is left undefined.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters and the schema is trivially empty, so there is nothing to document; the baseline for 0 params applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb + resource: 'Get available credits + plan' tells the agent exactly what it returns, and no sibling tool in the list overlaps with billing/credits. The only slight dilution is the vague parenthetical qualifier, which adds caveats rather than precision.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states a prerequisite ('official API key required'), which is useful usage context, but it never says when to call this versus other tools or how to react to the credits/plan values. Usage is implied only.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_character_toolC
Check a Soul character's training status.
| Name | Required | Description | Default |
|---|---|---|---|
| character_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| id | Yes | |
| raw | No | |
| name | Yes | |
| status | Yes | |
| image_count | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing: no read-only confirmation, no polling/retry behavior for a training job, no auth or rate-limit context. The only behavioral signal is the implied non-mutating 'Check', which is weak.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short, front-loaded sentence with no wasted words. It is efficient, though arguably too terse given the gaps elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read tool with an output schema that documents the return shape, the description is minimally sufficient. The missing piece is any identification of the character_id parameter or when polling is appropriate, which leaves the definition adequate but incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single required character_id has 0% schema description coverage, and the description says nothing about its format, where to obtain it, or whether it is a Soul character ID versus some other identifier. With one undocumented parameter, the description fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Check) and resource (a Soul character's training status), narrowing the generic tool name get_character_tool to a status lookup. It is clear what the tool does, though it never distinguishes itself from siblings like get_status_tool or list_characters_tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to call this versus get_status_tool, list_characters_tool, or create_character_tool. There are no prerequisites, no exclusions, and no indication of when a character is in a state worth polling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_status_toolC
Check job state and return output URLs when ready.
| Name | Required | Description | Default |
|---|---|---|---|
| job_handle | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | |
| state | Yes | |
| images | No | |
| timeout | No | |
| progress | No | |
| video_url | No | |
| job_handle | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full behavioral disclosure. It notes that output URLs appear 'when ready,' but does not describe failed-job behavior, polling cadence, rate limits, or whether the call blocks. For a status-check tool with zero annotation coverage, this is a notable gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no wasted words. It is appropriately sized, though its brevity contributes to the gaps in behavioral and parameter detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. However, for a status-check tool with no annotations and an undocumented parameter, the description should cover job states, error handling, and how to interpret the 'ready' condition. It remains too thin for the tool's role.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single required parameter job_handle has 0% schema description coverage, and the description does not explain what a job handle is, where it comes from, or its expected format. The description adds no semantic value beyond the parameter name itself.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Check') and resource ('job state'), and mentions the output URLs returned when ready. It distinguishes itself from list_jobs_tool by focusing on a single job handle, though it does not explicitly name siblings or scope limitations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: poll a job after submission to see if output URLs are ready. There is no explicit guidance on when to use this versus list_jobs_tool, cancel_job_tool, or other job-related siblings, and no exclusions or prerequisites are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_characters_toolC
List trained Soul characters.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | ||
| page_size | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| count | Yes | |
| characters | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden. A single sentence says nothing about pagination behavior, auth requirements, ordering, or result limits — meaningful gaps for a listing tool with no annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short sentence with no wasted words and the purpose is front-loaded, but the extreme terseness leaves essential context unstated — this reads as under-specification rather than deliberate concision.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and a simple paginated list does not demand much prose. Still, pagination behavior and the scope of 'trained Soul characters' (ownership, filtering) are left ambiguous.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description adds nothing about the two parameters (page, page_size). It does not explain pagination semantics or the meaning of page defaults, so it fails to compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb (List) and a specific resource (trained Soul characters), which distinguishes it from list_jobs_tool, list_soul_styles_tool, and list_motions_tool. However, it never explicitly names a sibling or scope condition, so the differentiation is only implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus get_character_tool or the other list_* tools, and no mention of prerequisites or context. The agent must infer usage purely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_jobs_toolC
List recent generations (history) from the official backend.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | ||
| page_size | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| jobs | Yes | |
| count | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, yet it only says 'list recent generations'. It omits pagination behavior, default ordering/recency window, what 'official backend' implies, and any auth assumptions for a mutation-adjacent history tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no padding, which is appropriate. The parenthetical '(history)' is slightly redundant, and 'from the official backend' is filler that adds little actionable meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, but with zero annotation coverage and zero parameter documentation, the description should have covered read-only intent and pagination. As written it is too thin for a 2-param listing tool nested among many similar siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so both page and page_size are undocumented in the schema. The description does not mention paging at all, leaving the agent with no idea that results are paginated or what the defaults (1 / 20) imply.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb (List) and resource (recent generations/history), which is unambiguous on its own. However it does nothing to distinguish itself from siblings like get_status_tool or cancel_job_tool, all of which operate on the same job/generation domain.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus get_status_tool (single job) or the other list_* tools. Nothing states prerequisites, that it is read-only, or how to page through results, so the agent must guess.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_models_toolC
List every supported Higgsfield model.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | Filter by 'image', 'video', or 'speech'. | |
| backend | No | Filter by 'official' or 'web'. | |
| include_unverified | No | Include models whose endpoint is suspected wrong upstream (currently: nano-banana-1). |
Output Schema
| Name | Required | Description |
|---|---|---|
| count | Yes | |
| models | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It only says what is listed; it says nothing about auth requirements, rate limits, or that include_unverified defaults to false and exposes suspect endpoints. For a read-only listing tool this is low risk, but the disclosure is still thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient, front-loaded sentence with no waste. It is arguably too terse for the tool's filtering capability, but nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the schema covers all three parameters. The description is still minimal for an agent deciding between this and recommend_model_tool, leaving a routing gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents kind, backend, and include_unverified fully. The description adds nothing about filter semantics, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb (List) and resource (Higgsfield models) with scope ('every supported'). It does not distinguish itself from the sibling recommend_model_tool, which also concerns model selection, and 'every' sits awkwardly against the kind/backend filters in the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this versus alternatives like recommend_model_tool or validate_params_tool, and no note that the filters exist as a way to narrow results. The agent must infer usage entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_motions_toolA
List DOP motion presets for image-to-video.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| raw | No | |
| count | Yes | |
| names | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It is a zero-parameter read-only listing, so destructive risk is minimal, but the description says nothing about result size, pagination, or filtering behavior. Adequate for a simple enumeration but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence that front-loads the verb and resource with no filler. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the description need not explain return values, and with zero parameters there is no schema semantics gap to fill. The one remaining ambiguity is what 'DOP' means and how presets relate to generate_video_tool, but for a simple list tool this is close to complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate. The baseline for a no-parameter tool is 4, and nothing here creates confusion.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (List) and a specific resource (DOP motion presets) scoped to image-to-video, which is enough to tell it apart from list_models_tool or list_soul_styles_tool. It stops short of explicitly naming a sibling, so it lands at 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use statement, but the 'for image-to-video' scope implies this is a discovery step in the image-to-video workflow. Usage is inferable rather than stated, which matches a 3.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_soul_styles_toolA
List Soul image style presets to pick by name.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| raw | No | |
| count | Yes | |
| names | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full burden. It does not state that the operation is read-only, has no side effects, or returns names, relying solely on the verb 'List' to imply safe behavior. This is minimal but not misleading.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no wasted words. It efficiently conveys the tool's function.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and zero parameters, the description only needs to state the tool's purpose, which it does. It is complete for a simple list operation, though it could briefly note that it returns all available presets.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so parameter semantics are not applicable; the baseline score of 4 is appropriate. The schema is empty and requires no further explanation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('Soul image style presets'), and adds a purpose ('to pick by name'). It is distinguishable from sibling list tools by its unique resource, though it does not explicitly contrast with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'To pick by name' implies the tool is used before selecting a style, but there is no explicit when-to-use, when-not-to-use, or mention of alternative list tools (e.g., list_motions_tool). Guidance is minimal.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preflight_check_toolA
Check auth + config for both backends before submitting a generation.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| web | Yes | |
| official | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and it discloses the check's scope (auth and config for two backends), implying a non-mutating diagnostic. It does not say whether a failed check throws an error or returns a status object, nor whether it incurs cost or requires credentials to already be present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler: verb, resources, and timing all in one clause. Nothing could be removed without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and with zero parameters the schema burden is nil. The description covers what is checked and when; only the failure-handling behavior and the identity of 'both backends' remain unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline of 4 applies; there is nothing for the description to disambiguate. It correctly does not pad with parameter talk.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (check) and the resources inspected (auth + config) plus the temporal context (before submitting a generation). It is distinguishable from validate_params_tool and get_balance_tool, though the phrase 'both backends' is never explained, leaving a small ambiguity about scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly names the trigger condition — run this before submitting a generation — which gives an agent a concrete moment to invoke it. It stops short of naming alternatives (e.g., validate_params_tool) or stating what to do if the check fails, so it is not a full when/when-not/alternatives treatment.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recommend_model_toolB
Suggest models for a described goal. Ranks by keyword overlap; verified-only by default.
| Name | Required | Description | Default |
|---|---|---|---|
| top | No | ||
| kind | No | ||
| intent | Yes | ||
| include_unverified | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| intent | Yes | |
| recommendations | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It usefully discloses the ranking mechanism ('keyword overlap') and the default filtering behavior ('verified-only'), which are non-obvious traits. However, it says nothing about read-only nature, whether 'verified' means quality or availability, or the shape of the ranking output.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core action and followed by the two most decision-relevant behavioral facts. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be explained. But with 4 parameters at 0% schema coverage, the description leaves top and kind unexplained, and an agent must infer their meaning from names alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 parameters. The description only loosely maps to two of them: 'described goal' corresponds to intent, and 'verified-only by default' corresponds to include_unverified. The top and kind parameters are entirely undocumented in both schema and description, leaving a real gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Suggest models for a described goal.' This distinguishes it from the enumerating sibling list_models_tool, since the action is goal-driven recommendation rather than listing. It stops short of explicitly naming that sibling, so a 5 isn't warranted.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage (use when you have a goal description rather than a known model id), but never states when to prefer this over list_models_tool or the generate_* tools. The 'verified-only by default' clause hints at a default mode but gives no guidance on when to override it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
subscribe_toolB
Long-poll until the job reaches a terminal state (completed/failed/etc).
| Name | Required | Description | Default |
|---|---|---|---|
| job_handle | Yes | ||
| poll_interval | No | ||
| timeout_seconds | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | |
| state | Yes | |
| images | No | |
| timeout | No | |
| progress | No | |
| video_url | No | |
| job_handle | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose the key behavioral trait: this is a blocking long-poll that returns only at a terminal state (completed/failed/etc). However it says nothing about timeout handling, whether the default 600s timeout raises an error, or the cost of holding the call open.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that states the blocking nature and the terminal states with zero filler. Nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. But with no annotations and three undocumented parameters, the description is thin for a blocking call whose timeout and poll-interval semantics materially affect how an agent should invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 3 parameters, so the description must compensate, and it does not. job_handle, poll_interval, and timeout_seconds are neither explained nor given format/units guidance (e.g. that timeout_seconds bounds the wait).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific behavior (long-poll until terminal state) plus the resource (job) and enumerates terminal states. It lets an agent distinguish this blocking-wait tool from a one-shot check, though it never explicitly names get_status_tool or list_jobs_tool as alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: an agent can infer this is the tool to call when it wants to wait for a job to finish rather than polling manually. There is no explicit 'use this instead of get_status_tool' guidance or statement of when NOT to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upload_image_toolC
Upload a local image and get a hosted URL to use as image_url.
| Name | Required | Description | Default |
|---|---|---|---|
| mime | No | image/png | |
| path | No | ||
| backend | No | web | |
| data_base64 | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | Yes | |
| backend | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it only discloses the return shape (hosted URL). It says nothing about auth requirements, size/format limits, which backend is used, or that the upload is a mutating external side effect.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the action and result front-loaded and no filler. It is efficient, though almost too terse given the undocumented parameter set.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, but the description is incomplete for a 4-parameter tool with 0% schema coverage and no annotations, notably omitting that path and data_base64 are alternative input modes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 parameters, so the description must compensate and does not. 'Local image' vaguely gestures at path but leaves the path vs. data_base64 alternative, the mime default, and the official/web backend enum entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Upload a local image') and the outcome ('get a hosted URL'), which cleanly separates it from generate_image_tool. It does not explicitly name that sibling, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The clause 'to use as image_url' implies the downstream context where this tool is appropriate, which is useful inferred guidance. However, there is no explicit when-to-use versus generate_image_tool, no prerequisites, and no exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_params_toolB
Pre-flight check params against a model's supported set (no generation).
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes | ||
| model_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| valid | Yes | |
| model_id | Yes | |
| supported | Yes | |
| constraints | Yes | |
| known_model | Yes | |
| unsupported | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose the key behavioral trait that nothing is generated (no side effects). It does not state what happens on failure, whether it raises errors or returns them, or any auth/rate constraints, so the disclosure is partial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the operative constraint front-loaded. Nothing is wasted, though the parenthetical could arguably be folded into a clearer statement of scope.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but for a validation tool with two required parameters and a nested untyped object, the description omits what 'supported set' means, what the validation result looks like conceptually, and how it relates to the near-identical sibling preflight_check_tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for both parameters, and 'params' is a free-form nested object with additionalProperties=true whose keys are never enumerated. The description tells the agent the object holds generation parameters checked against a model, but adds no key names, formats, or valid value ranges, so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: validating 'params' against a model's supported set, with a clarifying '(no generation)' to mark it as a non-generative operation. However, it does not differentiate itself from the sibling 'preflight_check_tool', which appears to cover very similar territory, leaving the agent to guess which of the two to call.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'pre-flight check' implies the tool is used before invoking a generation tool, and '(no generation)' signals it is a dry-run. But there is no explicit statement of when to prefer this over preflight_check_tool or how it fits into a generate workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
20 tool updates
v0.1.0- First observed
cancel_job_tool - First observed
create_character_tool - First observed
delete_character_tool - First observed
generate_batch_tool - First observed
generate_image_tool - First observed
generate_speech_video_tool - First observed
generate_video_tool - First observed
get_balance_tool - First observed
get_character_tool - First observed
get_status_tool - First observed
list_characters_tool - First observed
list_jobs_tool - First observed
list_models_tool - First observed
list_motions_tool - First observed
list_soul_styles_tool - First observed
preflight_check_tool - First observed
recommend_model_tool - First observed
subscribe_tool - First observed
upload_image_tool - First observed
validate_params_tool
TDQS
Scored across 20 tools
Most tools target distinct resource+action pairs (image vs video vs speech-video generation, character CRUD, model/job listing). The only mild overlap is get_status_tool vs subscribe_tool (both inspect job state) and the three pre-flight helpers (preflight_check, validate_params, get_balance), but descriptions make their separate targets clear.
Every tool follows the same verb_noun_tool convention (generate_image_tool, list_characters_tool, get_status_tool, cancel_job_tool), giving a fully predictable, uniform naming pattern with no deviations.
20 tools is on the heavier side but each covers a plausibly needed capability across generation, character lifecycle, presets, job control, and utilities. It leans verbose (several discovery/validation helpers) but nothing is clearly redundant.
Strong coverage: generation (image/video/speech/batch), full character CRUD, job status/cancel/subscribe, upload, model/style/motion discovery, balance, and pre-flight validation. Minor gaps like job history deletion or character update are absent but not blocking for the core workflow.
Maintenance
Related MCP Connectors
Multi-model AI image and video generator. 14 models behind one OAuth-secured MCP endpoint.
Generate images, video, audio and short films with 140+ AI models from any MCP client.
LLM chat, text tools, image generation, editing, batch image jobs, and asynchronous video generation
Generate images, video, music and voice from your CLI or AI agent. On-brand AI media toolkit.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables AI-powered image and video generation using Higgsfield AI models through MCP-compatible clients like Claude Desktop and Perplexity.346 npmMIT
- AlicenseBqualityAmaintenanceMCP server for Higgsfield AI that enables unlimited-mode image, video, audio generation, uploads, and job management via multiple parallel accounts, using browser session authentication.4354MIT
- FlicenseNot gradedqualityCmaintenanceMCP server that combines Higgsfield AI image/video generation with direct native posting to X, YouTube, Instagram, Facebook, and Bluesky, including a composite generate_and_post tool.-
- AlicenseNot gradedqualityCmaintenanceEnables MCP-compatible clients to generate and upscale AI images, videos, music, and sound effects, manage generation jobs, access model catalogs, and track credits.3MIT