Skip to main content
Glama

🎬 omnicinema-mcp

A local asset-creation engine exposed as a Model Context Protocol (MCP) server. Repository: https://github.com/joeshwoa/omnicinema-mcp

One local server that plans, designs, generates and edits production assets:

  • 🧠 Persona consultation β€” a Director of Photography, Graphic Designer, Voice Director and Music Producer compile a prompt strategy (positive/negative prompt, palette, lens, BPM/key/structure, pacing) before anything is generated.

  • πŸ–ΌοΈ Offline design engine β€” logos (emblem, monogram, wordmark, horizontal lockup; with reversed and icon variants), flat vector illustrations, and device-framed UI mockups as clean SVG (+ PNG), with no key and no network. Photos/textures via image APIs with your own key.

  • πŸŽ™οΈ Audio β€” narration via your system's TTS offline (or an HF TTS model), composed + synthesized soundtracks in seven genres (stereo WAV + editable multi-track MIDI), and procedural sound effects.

  • 🎞️ Video pipeline β€” one-line idea β†’ coherent multi-scene screenplay with shot-to-shot continuity β†’ stock footage (Pexels / Pixabay / Unsplash) or an offline storyboard animatic β†’ a frame-accurate edit with dissolves, Ken Burns motion, captions, a ducked soundtrack β†’ MP4 via Remotion or ffmpeg.

  • πŸ›‘οΈ Budget guard β€” persistent free-tier tracking; generative calls halt with a cost breakdown until the user approves.

  • πŸ”Œ Local REST API β€” token-authed, 127.0.0.1-only, so other local tools can request assets.

  • πŸ”­ Review-only discovery β€” finds candidate providers on public catalogs and queues them for a human; never integrates anything by itself.


What you get without any keys (honest version)

Capability

Offline result

Quality notes

Logo / vector art / UI mockup

Real SVG + PNG designs

Deterministic, brief-driven, professional-looking flat design. Text uses system fonts (Poppins β†’ Helvetica β†’ DejaVu fallbacks), so letterforms vary slightly by machine; convert text to outlines in a vector editor for final brand use.

Cinematic photo / texture

A clearly labelled placeholder

Not a photo. Needs an image API (HF_IMAGE_MODEL / REPLICATE_IMAGE_MODEL) + generative:true.

Voiceover

Real speech via say (macOS), pico2wave, ffmpeg's built-in flite, or espeak-ng

Intelligible (verified by transcribing it with Whisper) but clearly synthetic. With no TTS engine installed you get a timing tone that the tool labels NOT SPEECH. For a natural voice use HF_TTS_MODEL + generative:true.

Soundtrack

Composed & synthesized stereo WAV + multi-track MIDI

Real arrangement (chords, bass, drums, melody, section dynamics), mastered to about βˆ’14 LUFS with no clipping. It is a synth demo, not a studio recording; the MIDI is there to re-voice in a DAW.

Sound effects

13 procedural recipes (whoosh, rain, thunder, wind, ocean, impact, riser, click, ding, beep, fire, footsteps, heartbeat)

Good for placeholders/transitions; real recordings via Freesound with a key.

Video

An animatic: storyboard frames per shot with camera-matched motion, title card, captions, music, narration β†’ MP4

Real footage needs a stock key (free) or a generative video model (paid/quota).

Related MCP server: imagine-mcp

Scope & ethics

Everything uses official, documented APIs with your own keys, within each service's Terms. It does not reuse browser sessions, scrape web UIs (Seedance, Higgsfield, Suno, Udio…), or auto-integrate code found online. If a service has no first-party API you can get a key for, it is out of scope.


Requirements

  • Node.js β‰₯ 18.17.

  • Recommended system tools (all optional, each unlocks something):

    • ffmpeg β€” renders MP4 without Remotion, measures media, normalizes TTS loudness, rasterizes SVG (if built with librsvg) and provides the flite voice.

    • rsvg-convert (librsvg2-bin / brew install librsvg) β€” PNG export of every SVG.

    • espeak-ng (Linux) β€” offline narration if neither say nor ffmpeg-flite is available.

    • Remotion toolchain β€” npm run setup:render (downloads a headless Chrome on first render).

Install

git clone https://github.com/joeshwoa/omnicinema-mcp.git
cd omnicinema-mcp
npm install            # core server (installs optional Remotion too unless you pass --omit=optional)
npm run build          # compile to dist/
npm run setup:render   # optional: install the pinned Remotion render toolchain

All runtime output lives under CINEMA_ROOT (default: the repo folder). Point it at an external drive to keep caches and renders off your system disk.

Register as an MCP server

Claude Desktop β€” see examples/claude_desktop_config.json:

{
  "mcpServers": {
    "omnicinema": {
      "command": "node",
      "args": ["/absolute/path/to/omnicinema-mcp/dist/index.js"],
      "env": { "CINEMA_ROOT": "/absolute/path/to/omnicinema-mcp" }
    }
  }
}

Claude Code: claude mcp add omnicinema -- node /absolute/path/to/omnicinema-mcp/dist/index.js

Keys (all optional)

cp .env.example .env   # fill in ONLY the keys you have

Variable

Enables

Network / cost

PEXELS_API_KEY / PIXABAY_API_KEY / UNSPLASH_ACCESS_KEY

Stock video/photos in run_cinema_pipeline

Free APIs; used automatically when set (stock:false to stay offline)

FREESOUND_API_KEY

Real SFX recordings (generate_sfx + generative:true)

Free API

HUGGINGFACE_API_TOKEN + HF_IMAGE_MODEL / HF_TTS_MODEL / HF_MUSIC_MODEL / HF_VIDEO_MODEL

Photos, natural voice, MusicGen, video

Quota / paid, only with generative:true

REPLICATE_API_TOKEN + REPLICATE_IMAGE_MODEL / REPLICATE_MUSIC_MODEL / REPLICATE_VIDEO_MODEL

Same, via Replicate

Paid, only with generative:true

FAL_API_KEY + FAL_VIDEO_MODEL

Video via fal.ai

Paid, only with generative:true

ANTHROPIC_API_KEY

Optional screenplay prose polish (1 call per run; enrich:false to skip)

Paid

Model ids are yours to choose, so new models work without code changes. Behaviour toggles: CINEMA_TTS=off|say|pico2wave|flite|espeak-ng, CINEMA_DISABLE_REMOTION=1 (force the ffmpeg renderer), CINEMA_DISABLE_RASTER=1, LIMIT_<PROVIDER>_<DAILY|WEEKLY|MONTHLY>=n.

Tools (15)

Every tool's text output ends with the files it wrote and next steps; a JSON block follows for programs.

Tool

Use it to…

Network / money

run_cinema_pipeline

Turn an idea into a video. Start with dry_run:true to get the screenplay, shot list, providers and a cost estimate instantly.

Stock APIs if keys set; generative:true spends quota

compile_montage

Finish an interactive_montage project or re-cut any project (reorder via montage-order.json, swap clip files).

None

generate_image

Logo, vector art, UI mockup (offline SVG+PNG), photo/texture (API).

Only photo/texture with generative:true

generate_voiceover

Speak a script (offline TTS or HF TTS).

Only with generative:true

generate_soundtrack

Compose music in a genre, optionally to a length (durationSeconds).

Only with generative:true

generate_sfx

Sound effect / ambience.

Freesound with generative:true

consult_personas

Preview the compiled brief without generating.

None

check_limits

See free-tier usage per provider.

None

list_providers

See which providers are configured.

None

install_dependencies

Preview (consent:false) or run (consent:true) ffmpeg/Blender/Remotion installs.

Installs only with consent

discover_providers / approve_suggestion

Queue candidate providers for review / catalogue one (added disabled).

Public catalog search

ipc_start / ipc_status / ipc_stop

Local REST API for other tools (for a long-running service use npm run ipc; devuniverse-mcp fetches assets from its POST /generate).

Localhost only

Typical flow

  1. run_cinema_pipeline { prompt, dry_run: true } β†’ show the user the shot list and spend (usually "Nothing").

  2. run_cinema_pipeline { prompt, narration, soundtrack: true, musicStyle } β†’ MP4 path + screenplay + timeline.

  3. Optional: edit montage-order.json or replace clip files in the project folder β†’ compile_montage { projectId }.

How the video is built

  • Screenplay: the idea is parsed into subject / action / setting / story element; the film gets one lighting look (e.g. dusk β†’ storm night β†’ grey dawn), scenes follow an arc (setup β†’ inciting β†’ climax β†’ resolution), shots follow film grammar (establishing β†’ medium β†’ close). Each shot's opening frame is identical to the previous shot's closing frame, camera moves are derived from the framing change, and screen direction is held. screenplay.md is human-readable.

  • Footage: per shot, the configured stock providers are searched with short, specific queries; hits are ranked by orientation, duration and resolution, never reused, downloaded, and licensed in attributions.txt. Unfilled shots get storyboard frames.

  • Edit: hard cuts inside a scene, true crossfades between scenes, Ken Burns motion matching the camera move, title card, captions from the narration, soundtrack fitted to the cut and ducked under the voice, fade out. Rendered by Remotion (remotion/compositions/CinemaTimeline.tsx) or, without it, by an equivalent ffmpeg filter graph.

Output layout

assets/                      # standalone assets (<subject>_<kind>.svg/.png/.wav/.mid)
projects/<id>/               # screenplay.{md,json}, timeline.json, clips, audio, attributions.txt, manifest.json
output/<id>.mp4              # rendered video
data/usage-limits.json       # budget guard database   (gitignored)
data/review-queue.json       # discovery queue         (gitignored)
data/ipc-token.txt           # IPC bearer token, 0600  (gitignored)

Tests

npm test     # 76 tests, offline, against a throwaway CINEMA_ROOT (your data/ is never touched)

Covers the designer (valid, deterministic SVG; librsvg render when installed), stock clients against a mocked fetch (no key β†’ no request), screenplay quality rules, timeline/editorial rules, music/MIDI/SFX checks (loudness, clipping, structure), the MCP tool surface, and an offline end-to-end run that renders a real MP4 when ffmpeg or Remotion is present (skips cleanly otherwise). OMNICINEMA_TEST_REMOTION=1 npm test exercises Remotion instead of ffmpeg in that test.

Known limits

  • Offline photos are placeholders; offline video is an animatic, not footage.

  • Offline voices are robotic; offline music is synthesized (no sampled instruments, no vocals).

  • Logo/UI text depends on installed fonts; there is no font embedding or outlining.

  • The screenplay engine is template-based (deterministic). ANTHROPIC_API_KEY adds an optional prose rewrite.

  • Generative providers (Replicate/fal/HF) and stock APIs were unit-tested against mocked responses; they were not exercised against the live services in this release.

  • Remotion needs its headless Chrome download on first render; if Remotion fails, the pipeline falls back to ffmpeg.

Remotion licensing: free for individuals and small teams; a company license applies above a threshold β€” see https://remotion.dev/license.

License

MIT. Downloaded/generated assets keep their own licenses (see each project's attributions.txt). Please keep the scope boundary intact: official APIs + user-owned keys only; no scraping, token reuse, or auto-integration of untrusted code.

Available Tools

15 tools
approve_suggestionCatalogue a Reviewed SuggestionA

Use only after the user approves a discovered candidate. Adds it to tools-registry.json as DISABLED and unimplemented (a human must write an adapter and enable it). No network. Requires approve:true; otherwise does nothing.

ParametersJSON Schema
NameRequiredDescriptionDefault
approveNoMust be true to make any change.
suggestionIdYesThe suggestion id from discover_providers.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the write target, that the entry lands in a DISABLED/unimplemented state, that a human adapter is still required, that there is no network access, and that the operation is a no-op unless approve:true. It stops short of stating whether the write is idempotent, reversible, or what happens on an unknown suggestionId.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, each earning its place: precondition first, then side effect, then network/safety note, then the guard flag. No filler and nothing buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter mutation with no output schema, the description covers the trigger, the effect, and the enablement guard, which is enough to call it correctly. Error behavior on an invalid suggestionId is the only notable omission.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented, and the description's 'Requires approve:true; otherwise does nothing' largely restates the schema's 'Must be true to make any change.' It adds the no-op consequence slightly more explicitly but no format or validation detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource and the concrete side effect: the candidate is written into tools-registry.json as DISABLED and unimplemented. An agent can distinguish this from discover_providers (which finds candidates) without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit precondition ('Use only after the user approves a discovered candidate'), which simultaneously sets the trigger and rules out premature invocation. The 'only' framing plus the named source of the candidate (discovery) leaves no ambiguity about the alternative path.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_limitsCheck Free-Tier UsageA

Use before a generative run or when the user asks about remaining quota. Reports per-provider usage vs the budget guard's daily/weekly/monthly limits from data/usage-limits.json. Offline.

ParametersJSON Schema
NameRequiredDescriptionDefault
providerNoA single provider slug; omit for all tracked providers.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden: it usefully discloses that the operation is offline and that the limits come from data/usage-limits.json, implying a local read-only query. It does not state whether anything is written, what the report looks like, or failure behavior, leaving real gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, all front-loaded: the usage trigger comes first, then what is reported, then the source and the offline trait. Nothing is redundant with the name or schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-optional-param read tool with no output schema, the description covers when to call it, what it reports, where the data lives, and that it runs offline. Only a hint at the shape of the returned usage report is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single provider parameter is fully documented in the schema (single slug, omit for all). The description adds no syntax or format detail beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (check/report) and resource (per-provider usage against budget guard limits), names the data source, and is clearly distinguishable from siblings like list_providers or ipc_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Use before a generative run or when the user asks about remaining quota" gives explicit triggering conditions, which is exactly the routing an agent needs among the generation siblings. No when-not guidance or named alternative, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compile_montageCompile / Re-render a Video ProjectA

Use after run_cinema_pipeline in interactive_montage mode, or to re-cut any project: rebuilds the timeline from the project folder (honoring an optional montage-order.json and any clip files you replaced), validates it (no gaps, no missing files) and renders the MP4. Offline; no network. Returns the MP4 path, validation warnings and next steps.

ParametersJSON Schema
NameRequiredDescriptionDefault
renderNoRender after compiling (default true).
projectIdYesThe projectId returned by run_cinema_pipeline.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does well: it discloses that the operation is offline ('no network'), that it rebuilds rather than merely copies, that it validates for gaps and missing files, and that it renders an MP4. It does not state whether an existing render is overwritten or what permissions are needed, but for a local compile step the key traits are disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the usage trigger, then the behavior, then the offline note and return values. Three tight sentences with no filler; every clause adds actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema and no annotations, the description states the return payload (MP4 path, validation warnings, next steps), the offline constraint, and the prerequisites. Nothing an agent needs to call this correctly appears to be missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the two parameters are already documented, making 3 the baseline. The description adds related context about implicit inputs (montage-order.json and replaced clip files) but does not explain the declared 'render' flag or 'projectId' beyond what the schema says.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses specific verbs and resources: 'rebuilds the timeline from the project folder... validates it... and renders the MP4.' This clearly separates it from the pipeline sibling (run_cinema_pipeline) which is named explicitly as the upstream step, so an agent can tell the two apart without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states exactly when to use it ('after run_cinema_pipeline in interactive_montage mode, or to re-cut any project'), which implicitly names the alternative workflow for initial production. The trigger conditions and the sibling tool are both explicit, leaving little to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

consult_personasPreview the Creative Brief (no generation)A

Use before generating to show/tune the strategy. Runs the persona consultation (Director of Photography, Graphic Designer, Voice Director, Music Producer) for an asset kind and returns the compiled brief: positive/negative prompt, technical params (palette, lens, BPM/key/structure, pacing) and the debate transcript. Instant, offline, writes nothing.

ParametersJSON Schema
NameRequiredDescriptionDefault
styleNoStyle hint, e.g. 'noir', 'lo-fi', 'brutalist'.
subjectYesThe subject/description.
assetKindYesWhat to design.
aspectRatioNoe.g. '16:9', '1:1', '9:16'.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the return contents (positive/negative prompt, technical params, debate transcript) and key behavioral traits β€” 'instant, offline, writes nothing' β€” which tells the agent this is a safe, side-effect-free read. It stops short of noting any limits or auth considerations, but the core behavior profile is covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the use case ('Use before generating'), followed by tightly packed specifics about what it runs and returns, then a closing behavioral clause. Every sentence carries information with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although no output schema exists, the description enumerates the brief's components (prompt, technical params, transcript), so an agent knows what comes back. Combined with the side-effect disclosure, nothing needed to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all four parameters including the assetKind enum, and the description only loosely references 'an asset kind'. Baseline 3 applies since the schema does the heavy lifting and the description adds no format or constraint detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource β€” runs a persona consultation and returns a compiled creative brief β€” and the title reinforces the 'preview, no generation' framing. It is clearly distinguishable from the generator siblings (generate_image, generate_voiceover, etc.) because it explicitly produces a brief rather than an asset.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use before generating to show/tune the strategy' gives explicit timing and intent, which implicitly contrasts with the generate_* siblings that actually produce assets. It lacks an explicit exclusion or a named alternative, but the context for when to reach for it is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

discover_providersDiscover New Providers (review-only)A

Use when the user asks to find new video/model tools. Searches the public Hugging Face Hub and GitHub Search APIs (network, no key needed; GITHUB_TOKEN optional) and QUEUES candidates in data/review-queue.json. Never installs, enables or runs anything. Returns the top candidates and the queue size.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesSearch term, e.g. 'text to video'.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and mostly does: it discloses network access, that no key is required but GITHUB_TOKEN is optional, that it writes to data/review-queue.json, and that it never installs, enables, or runs anything. It omits rate-limit behavior and how many candidates are returned or whether the queue is appended to versus overwritten.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the trigger condition, then mechanism, then side-effect boundary and return value. No filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description compensates by stating the return (top candidates and queue size) and the persistent artifact it creates. For a one-parameter discovery tool, an agent has everything needed to call it and understand the consequences.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single `query` parameter has 100% schema description coverage, so the baseline is 3. The description implies a search term but adds no syntax, scoping, or qualifier guidance beyond what the schema's example already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (discover/queue provider candidates) and names the two external sources searched (Hugging Face Hub, GitHub Search). It is clearly separable from siblings like list_providers (which presumably lists existing ones) and approve_suggestion (which acts on the queue this tool fills).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Use when the user asks to find new video/model tools" gives a clear triggering condition, and "Never installs, enables or runs anything" implicitly routes install/run intent to siblings such as install_dependencies. It stops short of explicitly naming those alternatives or stating when NOT to use it (e.g., if the queue already has candidates).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_imageDesign an Image (logo, vector art, UI mockup, photo)A

Use for logos, illustrations, app/web mockups, photos and textures.

  • logo / vector-art / ui-mockup: designed OFFLINE as clean SVG (plus a PNG export when rsvg-convert or ffmpeg is installed). Logos come with -reversed (dark backgrounds) and -icon variants. Free, instant, no network.

  • cinematic-photo / texture: need an image API. Without generative:true you get a clearly LABELLED placeholder, not a photo. generative:true calls Replicate/Hugging Face with your key (uses quota; budget-guarded β†’ returns a halt with a cost breakdown; retry with approveOverBudget:true only if the user agrees). Style words steer the result: palette ('vibrant', 'warm', 'earth', 'mono'), logo layout ('monogram', 'wordmark', 'horizontal', 'emblem'), type voice ('luxury', 'tech', 'playful'), 'sharp' corners. Returns file paths, size, license.

ParametersJSON Schema
NameRequiredDescriptionDefault
styleNoStyle keywords (see description).
subjectYesWhat to create, e.g. 'Nova Labs' (logo), 'fitness mobile app' (ui-mockup), 'mountain sunset'.
assetKindYesType of visual asset.
generativeNoUSES QUOTA: photo/texture via your image API key. Ignored for vector kinds.
aspectRatioNoe.g. '1:1', '16:9', '3:1' (wide logo lockup).
approveOverBudgetNoOnly after the user agrees: proceed past the budget guard.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so richly: offline/no-network generation, PNG export contingent on rsvg-convert or ffmpeg, logo -reversed and -icon variants, quota consumption via the user's API key, budget guard with cost breakdown and halt behavior, and the return payload (paths, size, license).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with a one-line purpose then tight bullets grouped by asset family, with cost/quota caveats placed next to the feature that triggers them. Dense but every clause carries operational information; only marginal trimming is possible.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite a 6-parameter schema with no annotations and no output schema, the description covers the decision-critical facts: offline vs API path, placeholder-vs-real output, budget guard behavior, and what is returned. An agent has everything needed to call it correctly and set expectations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds real meaning beyond the schema: it names style-word categories (palette, logo layout, type voice, 'sharp' corners), gives subject examples per asset kind, and clarifies generative quota use and approveOverBudget consent. The schema even defers to the description for style semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (generate/design) plus the resource (visual assets) and enumerates the exact asset kinds it produces: logo, vector-art, ui-mockup, cinematic-photo, texture. It maps cleanly onto the assetKind enum and cannot be confused with the audio siblings (generate_voiceover, generate_soundtrack, generate_sfx).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: logos/vector/UI are produced offline and free, while photo/texture require generative:true, otherwise 'you get a clearly LABELLED placeholder, not a photo.' It also states the budget-guard halt condition and that approveOverBudget should be used 'only if the user agrees,' giving clear when/and-when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_sfxCreate a Sound EffectA

Use for a sound effect or ambience. Offline by default: procedural synthesis matched to the request (whoosh, rain, thunder, wind, ocean, impact, riser, click, ding, beep, fire, footsteps, heartbeat). generative:true searches Freesound with your FREESOUND_API_KEY (network, CC-licensed results, budget-guarded). Returns the file path, exact duration and license.

ParametersJSON Schema
NameRequiredDescriptionDefault
styleNo
subjectYesThe effect, e.g. 'whoosh transition', 'rain ambience'.
generativeNoNETWORK: fetch a real recording from Freesound.
approveOverBudgetNoOnly after the user agrees: proceed past the budget guard.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses offline-by-default procedural synthesis, that generative:true is a networked call requiring FREESOUND_API_KEY, that results are CC-licensed, that cost is budget-guarded, and what the call returns. The budget guard implies the approval gate but doesn't spell out the abort behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose, then the offline behavior, then the network branch and return contract. The long parenthetical of effect names is dense but earns its place by anchoring the valid subject space; otherwise tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema and no annotations, yet the description covers the return contract (file path, exact duration, license) and the safety/cost profile (network, API key, budget guard). Sufficient for correct invocation, though the approveOverBudget flow could be more explicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, and the description adds meaning beyond it: generative is characterized as network + Freesound + CC-licensed + budget-guarded (schema only says 'fetch a real recording'), and the style/subject intent is illustrated by the effect list. It doesn't explain what approveOverBudget does when triggered, leaving that to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('sound effect or ambience') and enumerates concrete effect types, distinguishing it from siblings like generate_soundtrack and generate_voiceover. An agent can tell at a glance this produces SFX files, not music or speech.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear when-to-use context ('Use for a sound effect or ambience') and explains the branch condition that selects offline vs network mode via generative:true. It stops short of naming an explicit alternative to prefer for non-SFX audio, so it's clear but not fully routed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_soundtrackCompose a Soundtrack / BeatA

Use for background music or a beat in a genre: hip-hop, trap/rap, cinematic orchestral, rock, lo-fi, electronic, ambient. Offline by default: the Music Producer plans tempo/key/structure/instruments and the engine composes and synthesizes a stereo WAV (voice-led chords, bass, drums, a melody, section dynamics, mastered to about -14 LUFS) plus an editable multi-track MIDI. It is a synthesized demo, not a studio recording. generative:true uses your MusicGen model on Replicate/HF (quota, budget-guarded). Returns WAV + MIDI paths, exact duration, and the arrangement.

ParametersJSON Schema
NameRequiredDescriptionDefault
styleNoGenre, e.g. 'hip-hop', 'cinematic orchestral', 'lo-fi', 'electronic'.
subjectYesTheme/mood, e.g. 'rainy midnight city'.
generativeNoUSES QUOTA: render via your MusicGen API model.
durationSecondsNoTarget length; the arrangement is re-flowed to fit (whole bars).
approveOverBudgetNoOnly after the user agrees: proceed past the budget guard.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the offline-by-default behavior, the Music Producer/engine pipeline, the exact synthesis contents and loudness target, that output is a synthesized demo rather than a studio recording, the quota/budget guard on generative:true, and the returned artifacts. Missing auth/permission and error/failure behavior keeps it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the use case and genre list before diving into pipeline detail, and every clause is substantive. It is dense and runs long, but no sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a multi-parameter generative tool with no output schema and no annotations, the description covers the pipeline, mode tradeoffs, quality caveats, and return values well. It stops short of error/limit handling, but an agent has enough to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents genre enums, subject, quota, duration re-flow, and budget approval. The description mostly restates these (genres, generative quota, WAV/MIDI), adding little that the schema does not already convey, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (compose a soundtrack/beat) and immediately scopes it to background music in named genres. It clearly distinguishes from audio siblings like generate_sfx and generate_voiceover by framing itself as music/beat composition.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Opens with explicit when-to-use ('Use for background music or a beat in a genre') and explains the default offline mode versus the generative:true mode. It does not explicitly compare itself against generate_sfx or generate_voiceover, so the sibling boundary is left implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_voiceoverGenerate Voiceover / NarrationA

Use to turn a script into spoken audio (WAV). Offline by default via the system TTS engine (macOS 'say', pico2wave, ffmpeg's flite, or espeak-ng): intelligible but synthetic-sounding, loudness-normalized to -16 LUFS. If no engine exists you get a timing placeholder TONE that is explicitly labelled NOT SPEECH. generative:true uses your Hugging Face TTS model (HF_TTS_MODEL; quota, budget-guarded) for a natural voice. Returns the WAV path and exact duration in ms.

ParametersJSON Schema
NameRequiredDescriptionDefault
styleNoDelivery: 'dramatic' (slower, deeper), 'calm documentary', 'energetic'.
scriptNoExplicit script (overrides subject as the spoken text).
subjectYesThe narration text (or a topic if you also pass script).
generativeNoUSES QUOTA: natural voice via HF_TTS_MODEL.
approveOverBudgetNoOnly after the user agrees: proceed past the budget guard.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so: it discloses the default engine chain, the loudness normalization target (-16 LUFS), the explicit failure-mode fallback (a timing TONE labelled NOT SPEECH), quota/budget-guarding for the generative path, and the return payload (WAV path plus exact duration in ms). This is unusually complete behavioral disclosure for an un-annotated tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose and the default mode before the generative option, and every sentence carries information (engine fallback, normalization, budget guard, return values). It is dense with parenthetical lists and runs slightly long, but contains no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema and no annotations exist, so the description must cover behavior and returns, and it does: it names both the offline and generative paths, the degenerate fallback case, the budget approval flow, and the returned fields. An agent has everything needed to invoke it correctly and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline would be 3; the description earns above that by explaining the runtime consequence of 'generative' (HF model, quota, budget guard) and the ordering constraint on 'approveOverBudget' (only after user agreement), which the schema texts state only tersely. 'style' and the subject/script interaction are left to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('turn a script into spoken audio (WAV)') and immediately disambiguates the two modes of production. An agent can distinguish it from generate_soundtrack/generate_sfx/generate_image without opening any schema, since the output artifact is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear conditional guidance: offline is the default, and 'generative:true' switches to the HF model with quota and budget guard. It also explains the approveOverBudget precondition ('Only after the user agrees'). It does not name sibling alternatives or state when NOT to use this tool, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

install_dependenciesInstall Local Dependencies (consent required)A

Use when rendering fails because ffmpeg, Blender or the Remotion toolchain is missing. With consent:false (default) it only detects the OS and returns the exact commands it WOULD run β€” nothing is installed. Only call with consent:true after the user explicitly agrees; it then runs the package manager (may need sudo) and npm (network).

ParametersJSON Schema
NameRequiredDescriptionDefault
consentNotrue runs the install commands. Only set after explicit user approval.
targetsNoWhich dependencies (default all).

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and discharges it: it discloses that consent:false is a safe dry-run that installs nothing, that consent:true runs the package manager (may need sudo) and npm (network), i.e. privilege escalation and network side effects. This is exactly the behavioral context a mutating install tool needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero waste. The trigger context is front-loaded and the behavioral consequences of each consent value follow immediately, so a reader gets the actionable facts in the first pass.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-param install tool with no output schema and no annotations, the description covers purpose, trigger, mode selection, safety gate, privilege and network implications. Nothing an agent needs in order to decide whether and how to call it is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so the baseline is 3, and targets is fully specified by its enum plus the 'default all' description. The description goes beyond the schema for consent by tying it to the dry-run/real-run distinction and the user-approval gate, adding genuine safety semantics rather than restating 'true runs the install commands'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (install) and resource (local dependencies), and frames the exact trigger: rendering fails because ffmpeg, Blender or the Remotion toolchain is missing. An agent can distinguish this from generation/pipeline siblings like run_cinema_pipeline or generate_image without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('when rendering fails because X is missing') and, crucially, a when-to-use-which-mode rule: consent:false for detection only, consent:true only after explicit user agreement. This is a decision rule an agent can act on directly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ipc_startStart Local REST APIA

Use when another local program needs to request assets from this engine over HTTP. Starts a server bound to 127.0.0.1 with a bearer token (stored in data/ipc-token.txt, mode 0600). Returns the URL and token. Stop it with ipc_stop.

ParametersJSON Schema
NameRequiredDescriptionDefault
portNoPort (default OMNICINEMA_IPC_PORT or 8787; 0 = any free port).

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and discloses important security context: binding to 127.0.0.1, bearer-token auth, token storage path, and 0600 permissions. It also states return values and how to stop the server. It does not cover edge behaviors such as what happens if the port is already in use or whether the process blocks, leaving some operational uncertainty.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short, front-loaded with the usage condition, and then covers behavior, return values, and lifecycle in four efficient sentences. Every sentence contributes useful information for invoking or managing the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a start-server tool with one schema-documented parameter and no output schema, the description covers the core contract: purpose, bind address, auth token, return values, and stop counterpart. It lacks some lifecycle/error detail, but the essential information an agent needs to start the service correctly is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single port parameter is fully documented in the schema, including default environment-variable fallback and the meaning of 0. The description adds no parameter-level guidance, so the baseline of 3 is appropriate when schema coverage is 100%.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: starts a local REST API server bound to 127.0.0.1. It also distinguishes from the sibling ipc_stop by directing the agent to use that tool for stopping, so the agent can tell the pair apart without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit condition for use: 'Use when another local program needs to request assets from this engine over HTTP.' It also names the complementary lifecycle tool, ipc_stop, so the agent has a clear alternative action for stopping rather than inferring it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ipc_statusLocal REST API StatusB

Use to check whether the local REST API is running and where. Offline.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full behavioral burden. It only says the tool checks status and location, offering no details about side effects, return format, authentication, or the meaning of 'Offline.'β€”leaving key behavioral traits undisclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The main sentence is front-loaded and efficient, but the trailing fragment 'Offline.' is cryptic and adds no clear value, slightly undermining the otherwise concise structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple zero-parameter status tool with no output schema, the description conveys that it returns whether the API is running and where. However, it does not describe the return format or clarify the 'Offline.' fragment, leaving some ambiguity for an agent trying to interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the description has no parameter semantics to add. The baseline score for a zero-parameter tool is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('check') and resource ('local REST API'), and specifies what is checked (running status and location). It does not explicitly distinguish itself from siblings like ipc_start or ipc_stop, but the purpose is clear without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It begins with 'Use to check...', providing a clear context for when to invoke the tool. It does not mention when not to use it or name alternative tools, but no misleading guidance is present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ipc_stopStop Local REST APIA

Use to shut down the local REST API started by ipc_start.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses the core action and the prerequisite that the API must have been started by ipc_start, but it omits idempotency, failure behavior if the API is not running, and whether active connections or tasks are interrupted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the action and target front-loaded. It contains no filler and is appropriately sized for a zero-parameter lifecycle operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple no-parameter stop tool with no output schema and no annotations, the description supplies the essential action and prerequisite. Remaining gaps around error states and idempotency are minor, though not fully covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, and the schema description coverage is 100%. The description adds no parameter detail, but none is needed, so the baseline of 4 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('shut down') and resource ('local REST API'), and explicitly ties it to ipc_start. This clearly distinguishes it from ipc_start and ipc_status as the opposite lifecycle action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Use to shut down... started by ipc_start' gives a clear condition for use and an implicit prerequisite. It does not name when-not-to-use cases or alternatives, but there is no obvious sibling alternative for stopping the API.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_providersList Providers & What Is ConfiguredA

Use to answer 'what can this server do with my keys?'. Lists every catalogued provider (stock, generative video/image, audio) with whether its key/model env vars are set. Offline; reads tools-registry.json and env only.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does meaningful work: it declares the tool is offline and reads only tools-registry.json and env, telling the agent there is no network call, no provider invocation, and no credit consumption. It does not describe cost of calling or output shape beyond the key/env status, but the safety-relevant behavior is well disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the question it answers, then scope, then behavioral constraints. Every clause carries information and nothing is repeated from the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must convey what comes back, and it does: a list of every catalogued provider with whether its key/model env vars are set. The main remaining gap is the unresolved relationship to the sibling 'discover_providers'.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, which is the baseline-4 case. The description adds that filtering/status is determined from env vars rather than arguments, reinforcing that no input is expected.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Lists every catalogued provider') and enumerates the catalog scope (stock, generative video/image, audio). It does not distinguish itself from the sibling 'discover_providers', which appears to be a closely related catalogue tool, so an agent cannot fully disambiguate the two from the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Frames a concrete use case ('Use to answer what can this server do with my keys?') and notes the offline context, which implies when it is appropriate. However, it never names an alternative or a when-not condition, and the closely named sibling 'discover_providers' is left unaddressed, so routing between the two remains inferential.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_cinema_pipelineMake a Short Video (screenplay β†’ footage β†’ edit β†’ MP4)A

Use when the user wants a short video from a one-line idea. Writes a multi-scene screenplay with shot-to-shot continuity, fills each shot with stock footage (Pexels/Pixabay/Unsplash, only if their keys are set) or an offline storyboard frame, optionally adds narration (offline system TTS + captions) and a soundtrack fitted to the cut, edits a frame-accurate timeline (hard cuts in a scene, dissolves between scenes) and renders an MP4 with Remotion or ffmpeg. TIP: call with dry_run:true first β€” it returns the screenplay, shot list, provider choice and a cost/network estimate instantly without writing files, so you can show the user before producing. Network/money: stock APIs are used automatically when keys exist (free; set stock:false to stay offline). generative:true calls paid-or-quota video APIs (budget-guarded; approveOverBudget:true to exceed). With no keys everything runs offline. Returns: projectId, rendered MP4 path (if rendered), screenplay.md, timeline.json, attributions, warnings, next steps. interactive_montage pauses before the edit so a human can reorder/replace clips, then compile_montage finishes.

ParametersJSON Schema
NameRequiredDescriptionDefault
fpsNoFrames per second (default 30).
stockNoUse stock APIs when keys are configured (default true). false = fully offline storyboard/animatic.
styleNoVisual style, e.g. 'noir', 'documentary', 'neon cyberpunk'. Also steers lighting.
widthNoFrame width (default 1920). Use 1080 with height 1920 for vertical.
enrichNoSPENDS MONEY: true + ANTHROPIC_API_KEY polishes the screenplay prose with one Anthropic API call. Default false (offline template prose).
heightNoFrame height (default 1080).
preferNoPrefer motion b-roll ('video', default) or stills ('image') from stock.
promptYesThe video idea, e.g. 'a lone lighthouse keeper watching a storm roll in'.
renderNoRender the MP4 now (default true in fully_automated). Uses Remotion if installed, else ffmpeg.
dry_runNotrue = plan only: screenplay, shot list, providers, cost estimate. No files, no network, no quota.
captionsNoBurn narration captions into the video (default true).
narrationNoNarration script. Spoken offline by the system TTS (say/flite/espeak-ng), captioned, locked to the timeline; the last shot is held if the narration runs long.
generativeNoSPENDS QUOTA/MONEY: generate clips with your Replicate/fal/HF video model (default false).
musicStyleNoSoundtrack genre: 'lo-fi', 'cinematic orchestral', 'hip-hop', 'trap', 'rock', 'electronic', 'ambient'.
sceneCountNoNumber of scenes (default 3–5 from the prompt). Scenes follow a story arc.
soundtrackNoAdd an offline-synthesized soundtrack fitted to the video length, ducked under narration.
shotsPerSceneNoShots per scene (default 2): establishing β†’ medium β†’ close.
workflow_modeNofully_automated renders in one call; interactive_montage stops after gathering clips (then use compile_montage).fully_automated
approveOverBudgetNoOnly after the user agrees: proceed past the free-tier budget guard.
shotDurationSecondsNoSeconds per shot (default 4).

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses that stock APIs only engage when keys exist, that everything runs offline otherwise, that generative:true spends money/quota behind a budget guard, and that approveOverBudget must be user-approved. It also enumerates the return payload and the warnings category.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the trigger, then the TIP, then cost/network caveats, then returns β€” a sensible priority order. It is dense and a little long, but each block (usage, tip, cost, returns) earns its place for a 20-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 20-parameter, no-output-schema tool, the description covers the essentials an agent needs: the trigger, the cost/network decision tree, the offline fallback, the interactive versus automated paths, and the explicit return fields. Nothing critical is left to inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description still adds meaning: it explains the dry_run shortcut, the cost implications of generative/enrich, the stock:false offline mode, and the interactive_montage handoff. It does not contradict or merely restate the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+output: produces a short video from a one-line idea via screenplay, footage, edit, and MP4 render. It names the sibling relationship (interactive_montage pauses, then compile_montage finishes), so an agent can distinguish it from generate_image or generate_voiceover in the same family.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Opens with an explicit trigger ('Use when the user wants a short video from a one-line idea') and adds an actionable TIP to call with dry_run:true first to show the plan before producing. It also routes the agent between workflow_mode values and names compile_montage as the follow-up tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 15 tool updatesv0.3.0
    • First observedapprove_suggestion
    • First observedcheck_limits
    • First observedcompile_montage
    • First observedconsult_personas
    • First observeddiscover_providers
    • First observedgenerate_image
    • First observedgenerate_sfx
    • First observedgenerate_soundtrack
    • First observedgenerate_voiceover
    • First observedinstall_dependencies
    • First observedipc_start
    • First observedipc_status
    • First observedipc_stop
    • First observedlist_providers
    • First observedrun_cinema_pipeline

TDQS

A4.1/5.0

Scored across 15 tools

Disambiguation4/5

Most tools have clearly distinct purposes (IPC trio, generation tools per asset kind, pipeline vs montage). The provider-management cluster (list_providers, discover_providers, check_limits, approve_suggestion) overlaps slightly, but descriptions differentiate 'what's available', 'find new', 'approve', and 'quota' well enough. generate_image bundles several sub-modes (logo/vector/ui-mockup vs cinematic-photo/texture) but is still a single coherent generation tool.

Naming Consistency5/5

All 15 tools use consistent snake_case with a verb_noun or verb_noun_noun pattern (generate_image, run_cinema_pipeline, ipc_start). The ipc_ prefix trio is a clean, predictable convention. No mixed camelCase or style deviations.

Tool Count5/5

15 tools sit at the top of the healthy 3-15 range and each earns its place: distinct asset generators (image/voice/soundtrack/sfx), a project pipeline, dependency install, provider registry management, and an IPC lifecycle. Nothing feels padded or redundant.

Completeness4/5

Covers image, voiceover, music, SFX generation, a full pipeline, provider discovery/approval, quota checks, and IPC serving. Minor gaps: no standalone create/list/delete project tools despite projectId being central, and no explicit generate_video tool although the pipeline references generative video APIs β€” agents can partly work around via run_cinema_pipeline and compile_montage.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers