omnicinema-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| CINEMA_ROOT | No | Root directory for the project (default /Volumes/PortableSSD/omnicinema-mcp) | |
| FAL_API_KEY | No | Enables fal.ai video | |
| HF_TTS_MODEL | No | Model ID for HuggingFace TTS | |
| HF_IMAGE_MODEL | No | Model ID for HuggingFace image generation | |
| HF_MUSIC_MODEL | No | Model ID for HuggingFace music generation | |
| HF_VIDEO_MODEL | No | Model ID for HuggingFace video generation | |
| PEXELS_API_KEY | No | Enables stock video/photos/stills from pexels.com/api | |
| FAL_VIDEO_MODEL | No | Model ID for fal.ai video generation | |
| PIXABAY_API_KEY | No | Enables stock video/photos/stills from pixabay.com/api/docs | |
| ANTHROPIC_API_KEY | No | Optional screenplay enrichment from console.anthropic.com | |
| FREESOUND_API_KEY | No | Enables stock SFX/ambience from freesound.org/apiv2/apply | |
| REPLICATE_API_TOKEN | No | Enables Replicate image/music/video | |
| UNSPLASH_ACCESS_KEY | No | Enables stock video/photos/stills from unsplash.com/developers | |
| HUGGINGFACE_API_TOKEN | No | Enables HF Inference image/TTS/music/video | |
| REPLICATE_IMAGE_MODEL | No | Model ID for Replicate image generation | |
| REPLICATE_MUSIC_MODEL | No | Model ID for Replicate music generation | |
| REPLICATE_VIDEO_MODEL | No | Model ID for Replicate video generation |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| run_cinema_pipelineA | Use when the user wants a short video from a one-line idea. Writes a multi-scene screenplay with shot-to-shot continuity, fills each shot with stock footage (Pexels/Pixabay/Unsplash, only if their keys are set) or an offline storyboard frame, optionally adds narration (offline system TTS + captions) and a soundtrack fitted to the cut, edits a frame-accurate timeline (hard cuts in a scene, dissolves between scenes) and renders an MP4 with Remotion or ffmpeg. TIP: call with dry_run:true first — it returns the screenplay, shot list, provider choice and a cost/network estimate instantly without writing files, so you can show the user before producing. Network/money: stock APIs are used automatically when keys exist (free; set stock:false to stay offline). generative:true calls paid-or-quota video APIs (budget-guarded; approveOverBudget:true to exceed). With no keys everything runs offline. Returns: projectId, rendered MP4 path (if rendered), screenplay.md, timeline.json, attributions, warnings, next steps. interactive_montage pauses before the edit so a human can reorder/replace clips, then compile_montage finishes. |
| compile_montageA | Use after run_cinema_pipeline in interactive_montage mode, or to re-cut any project: rebuilds the timeline from the project folder (honoring an optional montage-order.json and any clip files you replaced), validates it (no gaps, no missing files) and renders the MP4. Offline; no network. Returns the MP4 path, validation warnings and next steps. |
| install_dependenciesA | Use when rendering fails because ffmpeg, Blender or the Remotion toolchain is missing. With consent:false (default) it only detects the OS and returns the exact commands it WOULD run — nothing is installed. Only call with consent:true after the user explicitly agrees; it then runs the package manager (may need sudo) and npm (network). |
| list_providersA | Use to answer 'what can this server do with my keys?'. Lists every catalogued provider (stock, generative video/image, audio) with whether its key/model env vars are set. Offline; reads tools-registry.json and env only. |
| discover_providersA | Use when the user asks to find new video/model tools. Searches the public Hugging Face Hub and GitHub Search APIs (network, no key needed; GITHUB_TOKEN optional) and QUEUES candidates in data/review-queue.json. Never installs, enables or runs anything. Returns the top candidates and the queue size. |
| approve_suggestionA | Use only after the user approves a discovered candidate. Adds it to tools-registry.json as DISABLED and unimplemented (a human must write an adapter and enable it). No network. Requires approve:true; otherwise does nothing. |
| consult_personasA | Use before generating to show/tune the strategy. Runs the persona consultation (Director of Photography, Graphic Designer, Voice Director, Music Producer) for an asset kind and returns the compiled brief: positive/negative prompt, technical params (palette, lens, BPM/key/structure, pacing) and the debate transcript. Instant, offline, writes nothing. |
| generate_imageA | Use for logos, illustrations, app/web mockups, photos and textures.
|
| generate_voiceoverA | Use to turn a script into spoken audio (WAV). Offline by default via the system TTS engine (macOS 'say', pico2wave, ffmpeg's flite, or espeak-ng): intelligible but synthetic-sounding, loudness-normalized to -16 LUFS. If no engine exists you get a timing placeholder TONE that is explicitly labelled NOT SPEECH. generative:true uses your Hugging Face TTS model (HF_TTS_MODEL; quota, budget-guarded) for a natural voice. Returns the WAV path and exact duration in ms. |
| generate_soundtrackA | Use for background music or a beat in a genre: hip-hop, trap/rap, cinematic orchestral, rock, lo-fi, electronic, ambient. Offline by default: the Music Producer plans tempo/key/structure/instruments and the engine composes and synthesizes a stereo WAV (voice-led chords, bass, drums, a melody, section dynamics, mastered to about -14 LUFS) plus an editable multi-track MIDI. It is a synthesized demo, not a studio recording. generative:true uses your MusicGen model on Replicate/HF (quota, budget-guarded). Returns WAV + MIDI paths, exact duration, and the arrangement. |
| generate_sfxA | Use for a sound effect or ambience. Offline by default: procedural synthesis matched to the request (whoosh, rain, thunder, wind, ocean, impact, riser, click, ding, beep, fire, footsteps, heartbeat). generative:true searches Freesound with your FREESOUND_API_KEY (network, CC-licensed results, budget-guarded). Returns the file path, exact duration and license. |
| check_limitsA | Use before a generative run or when the user asks about remaining quota. Reports per-provider usage vs the budget guard's daily/weekly/monthly limits from data/usage-limits.json. Offline. |
| ipc_startA | Use when another local program needs to request assets from this engine over HTTP. Starts a server bound to 127.0.0.1 with a bearer token (stored in data/ipc-token.txt, mode 0600). Returns the URL and token. Stop it with ipc_stop. |
| ipc_statusB | Use to check whether the local REST API is running and where. Offline. |
| ipc_stopA | Use to shut down the local REST API started by ipc_start. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 15 tools
Most tools have clearly distinct purposes (IPC trio, generation tools per asset kind, pipeline vs montage). The provider-management cluster (list_providers, discover_providers, check_limits, approve_suggestion) overlaps slightly, but descriptions differentiate 'what's available', 'find new', 'approve', and 'quota' well enough. generate_image bundles several sub-modes (logo/vector/ui-mockup vs cinematic-photo/texture) but is still a single coherent generation tool.
All 15 tools use consistent snake_case with a verb_noun or verb_noun_noun pattern (generate_image, run_cinema_pipeline, ipc_start). The ipc_ prefix trio is a clean, predictable convention. No mixed camelCase or style deviations.
15 tools sit at the top of the healthy 3-15 range and each earns its place: distinct asset generators (image/voice/soundtrack/sfx), a project pipeline, dependency install, provider registry management, and an IPC lifecycle. Nothing feels padded or redundant.
Covers image, voiceover, music, SFX generation, a full pipeline, provider discovery/approval, quota checks, and IPC serving. Minor gaps: no standalone create/list/delete project tools despite projectId being central, and no explicit generate_video tool although the pipeline references generative video APIs — agents can partly work around via run_cinema_pipeline and compile_montage.