aidemo
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| OPENAI_API_KEY | No | API key for OpenAI TTS (required if AIDEMO_TTS_PROVIDER is not 'local') | |
| OPENAI_BASE_URL | No | Optional base URL for OpenAI-compatible TTS endpoint (e.g., a local server). | |
| AIDEMO_TTS_PROVIDER | No | TTS provider; set to 'local' to use in-process voice (no API key needed). Default is 'openai' if key is set, else 'local'. |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
| logging | {} |
| resources | {
"listChanged": true
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| get_authoring_guideA | The canonical guide to authoring demos: storyboard schema, action vocabulary, demo-director principles, ChatGPT-app recording facts. Call this FIRST before authoring or editing a storyboard. Pass |
| get_storyboard_schemaA | JSON Schema for storyboard.json, generated from the engine's own zod schema — the exact contract the engine validates against. |
| validate_storyboardA | Validate a storyboard against the engine schema without running anything. Pass exactly one of: dir (uses generated/storyboard.json), path (a storyboard file), or json (storyboard JSON as a string). relaxed makes narration optional (probe semantics). |
| lint_storyboardA | Browser-free preflight over a storyboard: a per-scene pacing forecast (which scenes compose will mostly freeze-hold or cut because the narration and the recorded action don't match), selector/wait pitfalls (type+Enter with no wait, :text-is on nested labels, focus without a zoom block, anchored waitForChange regexes), and no-op keys. Run it after every storyboard edit, before probe/render. Pass exactly one of dir / path / json. lang lints a narrations[lang] translation at that language's speaking rate. |
| init_demoA | Create demos// with a starter brief + storyboard. dir is the repo to scaffold into (absolute path recommended; default: server cwd). With fromUrl the page is inspected first (no LLM): its headings become scenes and its unique selectors become the beats + a |
| doctorA | Check prereqs: node, ffmpeg, Chrome, TTS/STT endpoint (flags LLM-only servers like Ollama), API key, playwright, installed skill. |
| embedA | Ready-to-paste embed snippets (markdown GIF, markdown still, HTML ) for a demo, using stable raw.githubusercontent URLs on the 'demo-media' branch. Pure string generation — no network, no render. Owner/repo come from the repo's origin remote; the demo name from the directory basename. Pair with the demo-publish workflow so CI keeps the URLs fresh. See docs/EMBEDS.md. |
| feedbackA | File a bug, surprise, or workaround you hit this session as a GitHub issue on the aidemo engine repo (github.com/tandryukha/aidemo) — the in-band equivalent of |
| job_statusA | Status, stage, per-scene progress (scenesTotal/scenesDone/currentScene — populated for probe/record/render/voice/captions/compose), log tail, and final result or error (with failure-artifact paths and, on a genuine failure, a feedbackHint pointing at the feedback tool) of a pipeline job. |
| job_listA | All jobs this server session, newest last. |
| job_cancelA | Best-effort cancel (checked between actions/scenes/stages). A mid-record cancel salvages the partial timeline + footage. The job settles asynchronously — poll job_status for the final state. |
| probeA | Record-only dry run to verify selectors/timing (narration optional). With updateGolden, writes golden/probe.json (the regression baseline); with golden, deep-compares against it and returns match + field-level diffs. Returns a jobId immediately — poll job_status. Rejects if another job is running. |
| import_traceA | Turn a Playwright trace.zip (context.tracing / |
| recordA | Drive the storyboard in Chrome and record raw video + timeline.json. Returns a jobId immediately — poll job_status. Rejects if another job is running. |
| renderA | Full pipeline: voice → record → captions → compose (the CLI |
| voiceA | Generate per-scene TTS narration (skips unchanged scenes). Returns a jobId immediately — poll job_status. Rejects if another job is running. |
| captionsA | Transcribe narration.mp3 to captions with word timing. Returns a jobId immediately — poll job_status. Rejects if another job is running. |
| composeA | Trim, sync, mux and caption into output/final-demo.mp4. Returns a jobId immediately — poll job_status. Rejects if another job is running. |
| qaA | Post-render checks on the FINISHED MP4 — blank opening, blank poster, master loudness, music-bed level, per-scene hold shares + compose warnings, caption readability, frame/chapter hygiene. Run it after compose/render: it sees what a viewer sees, which neither lint (pre-take) nor report.json (structural) can. Returns a jobId immediately — poll job_status. Rejects if another job is running. |
| gifA | Convert output/final-demo.mp4 to a README-ready GIF. Returns a jobId immediately — poll job_status. Rejects if another job is running. |
| stillsA | Extract named stills (screenshot mode) from the recorded take into output/stills/ — one PNG per |
| framesA | Dump evenly spaced PNG frames from the final video (or the raw take) into output/frames/ for review — look at them instead of hand-running |
| walkthroughA | Export output/walkthrough/ from the final video: index.html (one card per scene — payoff frame, title, narration, jump-to-time; ← → keyboard nav), guide.md (README/SOP-ready), per-scene PNGs, SRT/VTT and a JSON manifest. Needs only the rendered video + report.json (no key, no browser). Returns a jobId immediately — poll job_status. Rejects if another job is running. |
| inspectA | Open a URL in the recording profile (logged-in state included) and list every visible interactive element with UNIQUE selectors ranked data-testid → id → aria-label → role/text → name/placeholder → class → path, plus headings and iframes. Use it BEFORE writing targets — no selector guessing, no wasted probe. Writes logs/inspect-.json and a screenshot in the demo dir. Returns a jobId immediately — poll job_status. Rejects if another job is running. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
| authoring-guide | Canonical agent-neutral demo-authoring guide |
| storyboard-schema | JSON Schema generated from the engine's zod schema |
TDQS
Scored across 24 tools
Most tools have clearly distinct purposes—job management, authoring, pipeline stages, and output helpers are well separated. The only mild ambiguity is between validate_storyboard and lint_storyboard (schema vs. preflight pacing), and between probe/record/render, but descriptions clarify these boundaries.
The dominant pattern is snake_case verb_noun (job_list, get_storyboard_schema, validate_storyboard), which is highly readable. However, several tools use bare verbs (probe, record, render, voice, compose, embed, feedback) and two use abbreviations (qa, gif), creating minor inconsistency.
At 24 tools this is on the heavy side, fitting the '16-25 feels heavy' band. The breadth is defensible given the full pipeline (authoring, validation, rendering, QA, output formats, feedback), but the count is borderline and could be tightened by merging some job-status helpers or utility commands.
The toolset covers the entire demo lifecycle: import/init, storyboard authoring aids, validation/linting, run pipeline (probe/record/voice/captions/compose), QA, output formats (gif/stills/frames/walkthrough), embedding, and feedback. Minor gaps like no direct demo listing or deletion tools exist, but the core workflows have no dead ends.