Skip to main content
Glama
README.md
# studio-mcp

[![CI](https://github.com/rishbjain1/studio-mcp/actions/workflows/ci.yml/badge.svg)](https://github.com/rishbjain1/studio-mcp/actions/workflows/ci.yml)

**The block-method film pipeline as agent-callable MCP tools.**

An MCP server that lets an LLM agent drive a full `brief → shots → render → QC → cut`
workflow by calling tools — instead of clicking through generation UIs by hand.
Plan-then-generate, never reverse: every shot is planned (type, move, duration)
before any pixel is made, and every still is QC'd against a locked look so the
piece stays on-model.

![MARÉA shot 1 — Higgsfield Soul Cinema, 21:9, QC-passed](docs/marea_shot1.png)

> *Above: one shot from the brief "MARÉA — wordless coastal slow-burn" — planned,
> rendered on Higgsfield Soul Cinema at 21:9, and passed by the style-drift QC
> (intent 80 / look 84 / character 90). The agent re-rolls anything that fails.*

## Why

Single-tool MCP servers (one model, one call) are common. This is the layer above:
it **orchestrates** across your generation stack with a block-method planner and a
style-drift QC gate — the part nobody ships.

## Three ways to drive it

- **As an MCP server** — connect to Claude Desktop / Claude Code and the LLM calls
  the 14 tools directly (stdio, or `--transport streamable-http` for web clients).
- **The web console** — a cinematic "grading bay" UI over the server: browse
  instruments, invoke them from schema-driven forms, stream long renders, and brief
  **the Director** (an LLM that drives the instruments for you). Live demo:
  **https://console-pied-eight.vercel.app** · [`console/`](console/).
- **The orchestration graph** — Planner → Generation ⇄ QC → Assembly as an agent
  graph with retries, human gates, and replayable per-run traces
  ([`studio_mcp/orchestration/`](studio_mcp/orchestration/)).

Plus **Langfuse tracing + an offline benchmark suite** for evaluating the pipeline
([`eval/`](eval/)). CI runs tests on 3.10–3.12 and builds the package on every push.

## Tools (14)

**Plan & lock**
| Tool | What it does |
|------|--------------|
| `plan_shots(brief, project, n_shots)` | brief → block-method shot plan (type · move · duration · lighting · lens · time · hold · vibe) |
| `lock_campaign(project, aspect, camera, day_stock, hex_palette, elements, …)` | lock the look once — every shot's prompt inherits it |
| `palette_from_image(image)` | extract a HEX palette (dominant/secondary/accent) from a moodboard/still |
| `reference_prompt(reference, swap_subject)` | break a reference image into a ready 6-layer prompt (build *from* a ref) |

**Render & QC** *(via [Higgsfield CLI](https://higgsfield.ai/cli))*
| Tool | What it does |
|------|--------------|
| `gen_still(project, shot_id, note, model)` | 6-layer Soul prompt → render; `note` re-rolls with a QC fix; per-shot `model` |
| `qc_still(project, image, shot_id, threshold)` | vision **style-drift QC** vs shot intent + lock; pass/fail + fix_suggestion |
| `animate(project, shot_id, still, model, direct, hero)` | img2vid — `direct` = DP persona reads the frame & directs the move; `hero` = full Seedance timecoded/lip-sync prompt |
| `train_character(project, name, photos)` | soul-id self-clone from 3–5 photos |
| `upscale(media, kind, model)` | final-polish image/video upscale |

**Assemble & utility**
| Tool | What it does |
|------|--------------|
| `cut(project)` | ffmpeg-concat the rendered clips into one `<project>_cut.mp4` (offline, free) |
| `assemble(project, clips)` | cut manifest — order, durations, diegetic-audio notes |
| `list_models(kind)` | list available image/video models so an agent can route per shot |
| `project_status(project)` | what stages exist for a project |

**Ground** *(via [creative-rag](https://github.com/rishbjain1/creative-rag))*
| Tool | What it does |
|------|--------------|
| `craft_lookup(question, top_k)` | query the craft knowledge base for a **grounded, cited, verified** answer (stocks/lenses/lighting/prompt structure) — use while planning/locking so prompts trace to the real library, not generic guesses. Needs creative-rag running (`CRAG_URL`, default `http://127.0.0.1:8000`). |

## Integration — the studio trio

studio-mcp is one of three interlocking pieces; see [INTEGRATION.md](INTEGRATION.md).

- **`ai-content-pipeline` skill** — the *method* (block plan → lock → stills → animate → cut). It maps each stage to the studio-mcp tool that executes it and calls `craft_lookup` to ground prompts.
- **studio-mcp** (this repo) — the *tools* that execute the method.
- **[creative-rag](https://github.com/rishbjain1/creative-rag)** — the *cited craft KB* behind `craft_lookup`.

Smoke-test the full chain (skill method → `craft_lookup` → creative-rag):

```bash
python scripts/smoke_chain.py        # needs creative-rag on :8000
```

## Provider-agnostic

The LLM layer talks to **any OpenAI-compatible endpoint** — Anthropic, OpenRouter,
OpenAI, or a local server — chosen entirely through env config. No provider is
hard-coded.

```bash
cp .env.example .env   # set STUDIO_LLM_BASE_URL / _MODEL / _API_KEY
```

## Install

```bash
python3 -m venv .venv && source .venv/bin/activate
pip install -e .
```

## Run

```bash
studio-mcp        # stdio MCP server
```

Register with an MCP client (e.g. Claude Code / Claude Desktop):

```json
{
  "mcpServers": {
    "studio": {
      "command": "/path/to/studio-mcp/.venv/bin/studio-mcp",
      "env": { "STUDIO_LLM_API_KEY": "sk-..." }
    }
  }
}
```

## Example flow

```
plan_shots("MARÉA — wordless coastal slow-burn, 90s", project="marea")
lock_campaign("marea", aspect="21:9", day_stock="Kodak 500T, soft handheld",
              hex_palette=["#1b2a3a","#c8a15a"], elements=["the woman in grey"],
              audio="diegetic SFX only, no music")
# render a still (v1.1) → then:
qc_still("marea", image="assets/shot1.png", shot_id=1)   # pass/fail + fix
assemble("marea")                                         # cut manifest
```

State lives under `STUDIO_ROOT` (default `~/studio-projects/<project>/`) as plain
JSON — human-inspectable, and a clean contract a future console can read.

## License

MIT

TDQS

A3.8/5.0

Scored across 5 tools

Disambiguation5/5

Each tool has a distinct purpose: assemble builds the cut manifest, lock_campaign sets the visual look, plan_shots creates the shot plan, project_status reports stage completion, and qc_still performs quality control. There is no overlap or ambiguity between them.

Naming Consistency3/5

Tool names are not perfectly consistent: 'assemble' is a bare verb, while others follow verb_noun (lock_campaign, plan_shots) or noun_noun (project_status) patterns, and 'qc_still' uses an abbreviation. The naming is understandable but lacks a uniform convention.

Tool Count4/5

With 5 tools, the count is appropriate for the domain of film production planning and quality control. It covers the main workflow stages without being too heavy or too sparse, though a few more specialized tools could be added later.

Completeness4/5

The tool set covers the core stages: planning, look locking, assembly, QC, and status tracking. Notable gaps like actual rendering or exporting are intentionally omitted per v1 scope. The surface is fairly complete for its intended purpose.

Maintenance

ActivityStale
ResponsivenessNo issues