mcp-openshorts
Provides AI voiceover and dubbing with support for 30+ languages and voice cloning.
Provides AI capabilities for viral moment detection, script generation, web research, and video effects.
Enables direct publishing and scheduling of videos to Instagram Reels.
Enables direct publishing and scheduling of videos to TikTok.
Provides lip-sync technology for AI-generated talking head videos.
Enables direct publishing to YouTube, along with AI-generated thumbnails, titles, and descriptions.
Supports the creation and publishing of YouTube Shorts.
OpenShorts.app
Open source AI video platform with 3 tools in one: Clip Generator, AI Shorts (UGC videos with AI actors), and YouTube Studio.

Two people on camera? OpenShorts stacks them instead of shrinking the wide shot, puts the captions on the seam where they cover nobody, and switches back to a face-tracked crop when the cut goes to one person. The AI picks the layout per video; nothing to configure.
Two ways to run it, same software either way:
Self-hosted (this repo) | Hosted on openshorts.app | |
Price | Free forever, MIT | Free plan, paid from $12/mo |
Speed | 5 to 8 min per 8-min video on CPU | About 50s on our NVIDIA GPU |
API keys | Bring your own Gemini, ElevenLabs, fal.ai | Gemini included, nothing to set up |
Watermark / limits | None, ever | Watermark and 20 min/mo on the free plan, neither on paid |
Setup | Docker, 8GB+ RAM, model downloads | Sign in and paste a link |
MCP / API for agents | Same | Always-on endpoint at mcp.openshorts.app, API keys in one click |
Your data | Your server | Ours |
Self-hosting is genuinely free and always will be. It costs you a machine, your own API keys and the time to keep it running. The hosted plans exist to cover that hardware and those keys, not to unlock features.
https://github.com/user-attachments/assets/b45fa983-16b4-48b5-ac5b-a267836b9ad9
Video Tutorial: How it works

Click the image above to watch the full walkthrough.
3 Tools in 1 Platform
1. Clip Generator
Turn your long-form videos — podcasts, webinars, livestreams, vlogs, interviews — into viral-ready 9:16 shorts for TikTok, Instagram Reels, and YouTube Shorts.

2. AI Shorts (UGC Video Creator)
Generate marketing videos with AI actors for any product or business. No camera, no studio, no influencer budget. Just describe your product or paste a URL.

Two cost modes: Low Cost (
$0.65/video) and Premium ($2/video)Works for any business: SaaS, restaurants, e-commerce, coaching, local businesses
AI-generated actors with lip-sync, voiceover, b-roll, and TikTok-style subtitles
Choose from a shared avatar gallery or upload your own photo
Publish directly to TikTok, Instagram, and YouTube
3. YouTube Studio
Complete free AI YouTube toolkit: thumbnails, titles, descriptions, and direct publishing.

AI thumbnail generator with face overlay
10 viral title suggestions with refinement chat
Auto-generated descriptions with chapter timestamps
One-click publish to YouTube
UGC Video Gallery
All generated videos and avatars are saved to a public gallery with SEO pages for each video.

Public gallery page with hover-to-play (
/gallery)Individual SEO video pages with og:video meta tags (
/video/{id})JSON-LD structured data for search engines
Avatar gallery with prompt history
Related MCP server: ssemble-mcp-server
Key Features
Clip Generator
Viral Moment Detection: Google Gemini 3.1 Flash-Lite analyzes transcripts and scene boundaries to detect 3-15 high-potential moments
Runs fully local if you want: point
LLM_BASE_URLat Ollama, LM Studio, vLLM or any OpenAI-compatible server and the moment picker runs on your own model, no Google key needed (see Run without a Google key)Smart 9:16 Cropping: AI reframing per scene — TRACK mode (MediaPipe + YOLOv8 face tracking), GENERAL mode (blurred background), SPLIT mode (two speakers stacked, captions on the seam) and SCREENCAST mode (screen over presenter); the layout is picked per video by Gemini or forced from the dashboard
Auto Subtitles: faster-whisper with word-level timestamps, styled and burned into clips
AI Voice Dubbing: ElevenLabs integration for 30+ languages with voice cloning
Hook Text Overlays: AI-generated attention-grabbing text overlays
AI Video Effects: Gemini-generated FFmpeg filters for professional effects
AI Shorts Pipeline
Analyze: Scrape website URL + web research, or generate from manual description
Script: AI writes viral scripts (hook - problem - solution - CTA format)
Actor: Generate AI actors with Flux 2 Pro or select from shared gallery
Voice: ElevenLabs TTS voiceover (English/Spanish, male/female)
Video: Talking head generation (Hailuo 2.3 Fast img2video + VEED Lipsync)
B-roll: AI-generated visuals with Ken Burns effect
Composite: FFmpeg final assembly with subtitles and hook overlays
Publish: Direct posting to TikTok, Instagram Reels, YouTube Shorts via Upload-Post
YouTube Studio
AI-powered title generation with 10 viral options
Interactive refinement chat for titles
AI thumbnail generation with custom face + background
Auto descriptions with chapter timestamps from Whisper transcript
Direct YouTube publishing via Upload-Post
Social Auto-Publishing
One-click posting to TikTok, Instagram Reels, and YouTube Shorts simultaneously
Schedule uploads for any date and time — plan your content calendar and let OpenShorts publish automatically
Multi-platform distribution — publish to all your social networks at once from a single interface
Upload-Post integration with async uploads
Infrastructure
S3 cloud backup (private bucket for clips, public bucket for gallery/avatars)
SEO gallery pages served by FastAPI with JSON-LD structured data
Shared avatar gallery across all users
Async job queue with configurable concurrency
Who Is This For?
Content creators — Turn long videos into shorts automatically, publish to all platforms at once
Marketing agencies — Generate UGC videos for clients at scale, no actors or studios needed
SaaS founders — Create product demos and marketing shorts from just a URL
E-commerce brands — Product videos with AI actors for TikTok Shop, Instagram, YouTube
Local businesses — Restaurants, gyms, real estate, coaching — affordable video marketing
Developers — Self-host, customize the pipeline, integrate via API
AI Shorts Showcase
Videos generated with OpenShorts AI Shorts — no camera, no studio, no actors:
|
|
|
Biohacking for Investors · LOW COST | Secret Weapon for Devs · LOW COST | El Secreto de los Agentes de IA · PREMIUM |
Browse all videos at openshorts.app/gallery
OpenShorts vs Competitors
Feature | OpenShorts | Opus Clip | CapCut | Vizard | Klap | Descript |
Price | Free self-hostedfrom $12/mo hosted | $15-29/mo | $8/mo | $15-20/mo | $23-63/mo | $24-65/mo |
Self-hosted | Yes | No | No | No | No | No |
Open source | Yes | No | No | No | No | No |
Watermark | Never self-hostedfree plan only when hosted | Free tier | Some | Free tier | Free tier | Free tier |
Upload limits | None self-hostedby plan when hosted | 10-30GB | Credit-based | 60min-10hr | 10-100 vids/mo | 60min-40hr |
AI clip detection | Yes | Yes | Yes | Yes | Yes | Yes |
Smart 9:16 reframing | Yes | Yes | Yes | Yes | Yes | No |
Auto subtitles | Yes | Yes | Yes | Yes | Yes | Yes |
Voice dubbing (30+ langs) | Yes | No | Pro only | No | Pro only | Business only |
AI UGC actors | Yes | No | No | No | No | No |
AI video effects | Yes | No | Yes | No | No | No |
Hook text overlays | Yes | No | No | No | No | No |
YouTube Studio (titles, thumbnails) | Yes | No | No | No | No | No |
Social auto-publishing | Yes | Pro only | TikTok only | Paid only | Paid only | No |
Schedule uploads | Yes | Pro only | No | Paid only | Paid only | No |
Data privacy | Your server | Their cloud | Their cloud | Their cloud | Their cloud | Their cloud |
Works with a local LLM (Ollama) | Yes | No | No | No | No | No |
How Much Does It Cost?
Self-hosting OpenShorts is free. You provide the machine and you only pay for the AI APIs you use, and most have generous free tiers:
Service | Free Tier | Paid Cost | Used For |
Google Gemini | Free trial with generous limits | < $0.01 per 10-min video | Viral moment detection, script generation, web research |
Local LLM (Ollama, LM Studio, vLLM...) | Free, your hardware | $0 | Viral moment detection instead of Gemini ( |
fal.ai | Pay-per-use | ~$0.50-1.50 per AI Short | Actor generation, talking head video, lip-sync |
ElevenLabs | Free tier available | Pay-per-use | Voiceover, voice dubbing |
Upload-Post | 10 free uploads/month to all networks (no credit card) | Pay-per-use | Auto-publishing to TikTok, Instagram, YouTube |
AWS S3 | Optional | ~$0.023/GB | Cloud backup for clips and gallery |
Bottom line: You can clip videos for practically free with Gemini, and publish 10 videos/month to all social networks at zero cost with Upload-Post.
Don't want to run any of that? openshorts.app is the same software on our hardware: our NVIDIA GPU clips an 8-minute video in about 50 seconds instead of the 5 to 8 minutes it takes on a typical CPU, the Gemini key is included, and auto-publishing is already wired up. Free plan is 20 minutes a month with a watermark and no credit card; paid plans start at $12/mo for 100 minutes without watermark.
Requirements
Docker & Docker Compose
Google Gemini API Key (Free — get it here) — required for all AI features
fal.ai API Key (Pay-per-use) — required for AI Shorts (actor generation, video, lip-sync)
ElevenLabs API Key (Free tier) — required for voiceover/dubbing
Upload-Post API Key (free tier) — required for direct social posting
Getting Started
1. Clone
git clone https://github.com/mutonby/openshorts.git
cd OpenShorts2. Configure (optional)
cp .env.example .env
# Edit .env with your AWS keys for S3 backup3. Launch
docker compose up --build4. Open Dashboard
Navigate to http://localhost:5175
Go to Settings and enter your API keys (Gemini, fal.ai, ElevenLabs, Upload-Post)
Clip Generator: Upload a long-form video to generate viral shorts
AI Shorts: Describe your product or paste a URL to generate UGC marketing videos
YouTube Studio: Generate thumbnails, titles, and descriptions for YouTube
UGC Gallery: Browse all generated videos and avatars
5. GPU acceleration (optional, NVIDIA)
The default image is CPU-only. With an NVIDIA card (any card with NVENC, e.g. RTX 4060) an 8-minute video clips in about a minute instead of 5 to 8. Nothing is passed through in the VM sense — the container just gets access to the host GPU.
Host: install the NVIDIA driver (nvidia-smi must work) and the NVIDIA Container Toolkit:
sudo nvidia-ctk runtime configure --runtime=docker && sudo systemctl restart docker
docker run --rm --gpus all nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi # sanity checkOn Windows use Docker Desktop with the WSL2 backend and the Windows NVIDIA driver; no driver inside WSL.
Compose: create docker-compose.override.yml next to docker-compose.yml (picked up automatically). GPU: "1" adds cuBLAS/cuDNN and onnxruntime-gpu to the image (~2 GB); video is required for NVENC.
services:
backend:
build:
context: .
args:
GPU: "1"
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu, video].env:
WHISPER_MODEL=large-v3-turbo
WHISPER_DEVICE=cuda
WHISPER_COMPUTE=float16
FFMPEG_ENCODER=auto # probes h264_nvenc at startup, falls back to x264
TRANSCRIBE_BACKEND=parakeet # optional: ~2x faster than whisper, 25 European languages, auto-falls back to whisper
ASR_GPU_CONCURRENCY=1Verify:
docker compose up --build -d
docker exec openshorts-backend nvidia-smi -L
docker exec openshorts-backend ffmpeg -hide_banner -f lavfi -i testsrc=size=256x256:rate=1 -frames:v 1 -c:v h264_nvenc -f null -The backend log on the first job reports the chosen encoder and transcription device. A CUDA error in whisper (e.g. VRAM exhausted) retries once on CPU automatically. 8 GB of VRAM is enough for large-v3-turbo fp16 plus the detection models.
6. Run without a Google key (local LLM, optional)
The only cloud call in the clip pipeline is the moment picker: it sends the transcript (never the video) to Gemini. Point it at any OpenAI-compatible server instead and the whole pipeline stays on your box:
# .env
LLM_BASE_URL=http://host.docker.internal:11434/v1 # Ollama on the host
LLM_MODEL=qwen2.5:14b # any chat model that follows instructions
# LLM_API_KEY=... # only if your server checks one (vLLM --api-key, OpenRouter)Works with Ollama, LM Studio, vLLM, llama.cpp server, LocalAI and OpenRouter. The dashboard stops asking for a Gemini key when this is set. Two things to know:
Context length. A scoring call carries three transcript windows (~2-3k tokens) and the detail call up to ten (~5k on a long podcast). Ollama defaults to a 4096-token context and truncates silently, so run it with
OLLAMA_CONTEXT_LENGTH=16384(or setnum_ctxin a Modelfile); raiseLLM_SCORE_BATCHabove 3 only if your context allows it. 7-8B models return valid JSON reliably, 3B ones do not.What still needs Gemini. Anything that has to look at frames: the automatic layout picker (
AUTO_LAYOUT), the on-screen content detector and silent videos (no speech to clip by). Without a Gemini key those fall back to the plain face-tracking crop, and a silent video fails with a message that says so. Add a key alongsideLLM_BASE_URLand you get both.
Technical Pipeline
Clip Generator
Ingest — Local video upload (or self-hosted URL ingest via yt-dlp)
Transcribe — faster-whisper with word-level timestamps
Detect — PySceneDetect for scene boundaries
Analyze — Gemini identifies 3-15 viral moments (15-60s each)
Extract — FFmpeg precise clip cutting
Reframe — AI vertical cropping with subject tracking
Effects — Subtitles, hooks, AI video effects
Publish — S3 backup + Upload-Post social distribution
AI Shorts
Analyze — Website scraping + Gemini web research (or manual description)
Script — Gemini generates viral scripts with segments
Actor — Flux 2 Pro portrait generation (or gallery/upload)
Voice — ElevenLabs TTS voiceover
Video — Hailuo 2.3 Fast img2video + VEED Lipsync (Low Cost) or Kling Avatar v2 (Premium)
B-roll — Flux 2 Pro image generation + Ken Burns effect
Composite — FFmpeg assembly with ASS subtitles and hook overlays
Gallery — Upload to public S3 with metadata for SEO pages
Publish — Upload-Post to TikTok, Instagram, YouTube
Automate It: MCP Server, REST API and Webhooks
You don't need the dashboard. The whole pipeline is callable by AI agents and scripts.
MCP server (/mcp)
OpenShorts ships a built-in MCP server, so Claude, ChatGPT, Cursor or any MCP client can clip and publish videos for you:
claude.ai and ChatGPT: paste https://mcp.openshorts.app/mcp as a custom connector (Settings → Connectors) and approve the access on openshorts.app. The server does OAuth 2.1 with dynamic client registration, so there is no key to copy; the connection shows up under Account → API keys, where revoking it disconnects the app.
# Claude Code / Cursor / n8n (hosted): create an API key in your account page
claude mcp add --transport http openshorts https://mcp.openshorts.app/mcp \
--header "Authorization: Bearer osk_..."
# Self-hosted (no key needed, BYOK rules apply):
claude mcp add --transport http openshorts http://localhost:8000/mcp
# Self-hosted without running the web server: same tools over stdio
claude mcp add openshorts -- python mcp_stdio.pyTools: process_video (URL or upload_id; captions: false when the source already has subtitles, auto_hook: false to skip the hook line, burned by default like the dashboard), create_upload (hand the agent a local file: PUT the bytes, then process), get_job_status, list_clips, get_quota, add_subtitles, recut_clip, publish_clip. A prompt like "clip this podcast and schedule the best 3 to TikTok" is now a one-liner in your agent of choice.
REST API + API keys
Hosted accounts can mint osk_... API keys (account page). A key authenticates as you everywhere — same plan, same minutes, same job ownership:
curl -X POST https://api.openshorts.app/api/process \
-H "Authorization: Bearer osk_..." -H "Content-Type: application/json" \
-d '{"url": "https://youtube.com/watch?v=...", "acknowledged": true,
"webhook_url": "https://your-server.com/hooks/openshorts"}'Interactive docs at /docs (OpenAPI) on any instance.
Completion webhooks
Pass webhook_url (and optionally webhook_secret) to POST /api/process and you get exactly one POST when the job reaches a terminal state — no polling loops in your n8n / Zapier / cron pipelines:
{"event": "job.completed", "job_id": "…",
"clips": [{"index": 0, "title": "…", "video_url": "…", "download_url": "…"}]}With a secret, the body is signed: X-OpenShorts-Signature: sha256=<hmac-sha256(body)>.
CLI
The same API from the terminal, zero dependencies (cli/):
pip install openshorts # or: uvx openshorts
export OPENSHORTS_API_KEY=osk_... # hosted
# export OPENSHORTS_API_URL=http://localhost:8000 # self-hosted, no key
openshorts process "https://youtube.com/watch?v=..." --wait
openshorts clips <job_id>
openshorts publish <job_id> 0 --platforms tiktok,youtubeAgent skill
skills/openshorts/SKILL.md follows the open
Agent Skills standard, so it works in any
skill-capable agent:
# Claude Code (and most agents): copy the folder into the skills directory
cp -r skills/openshorts ~/.claude/skills/
# Hermes Agent: install straight from this repo
hermes skills install mutonby/openshorts/skills/openshorts
# OpenClaw: from ClawHub
openclaw skills install @mutonby/openshortsn8n
An importable workflow (video URL in, published-ready clips out, no polling)
lives in examples/n8n/.
Tech Stack
Layer | Technology |
Backend | Python 3.11, FastAPI, google-genai, faster-whisper, ultralytics (YOLOv8), mediapipe, opencv-python, yt-dlp, FFmpeg, httpx |
Frontend | React 18, Vite 4, Tailwind CSS 3.4 |
AI APIs | Google Gemini, fal.ai (Flux, Hailuo, VEED, Kling), ElevenLabs |
Infrastructure | Docker + Docker Compose, AWS S3 |
Publishing | Upload-Post API (TikTok, Instagram, YouTube) |
Environment Variables
Server-side (.env):
Variable | Description |
| AWS access key for S3 |
| AWS secret key |
| AWS region (default: us-east-1) |
| Private bucket for clip backup |
| Public bucket for gallery/avatars |
| Concurrent processing limit (default: 5) |
| OpenAI-compatible server for the moment picker (Ollama, vLLM, LM Studio...). Set it and the Gemini key becomes optional |
| Model name on that server (default |
| Bearer token for that server, if it checks one |
| Transcript windows per scoring call (default 3 local, 8 Gemini) |
Client-side (encrypted in localStorage):
Key | Description |
| Google Gemini — required unless |
| fal.ai — required for AI Shorts |
| ElevenLabs — required for voiceover/dubbing |
| Upload-Post — required, for social posting |
Security & Performance
Non-Root Execution: Containers run as dedicated
appuserConcurrency Control: Semaphore-based job queue (
MAX_CONCURRENT_JOBS)Auto-Cleanup: Automatic purging of old jobs (1h retention)
Encrypted Keys: API keys encrypted client-side, never stored server-side
Upload Validation: Image uploads validated for format and minimum size
File Limits: 2GB upload limit protection
Social Media Setup (Upload-Post)
Register: app.upload-post.com/login
Create Profile: Go to Manage Users
Connect Accounts: Link TikTok, Instagram, and/or YouTube
Get API Key: Navigate to API Keys
Use in OpenShorts: Paste the key in Settings
Star History
Contributions
Contributions are welcome! Whether it's adding new AI models, improving the lip-sync pipeline, or building new features — feel free to open a PR.
License
MIT License for the core application — OpenShorts is yours to use, modify, and scale.
Exception: the cloud/ directory (billing, managed keys, and the hosted-service infrastructure behind the optional BILLING_ENABLED flag) is source-available under the OpenShorts Commercial License. You can read it, modify it, and self-host it for personal or internal use, but you can't offer it to third parties as a paid/hosted service. Self-hosting the core app never requires this directory.
Available Tools
8 toolsadd_subtitlesBurn styled captions onto a clipBInspect
Re-style the captions of one clip (clips already ship with default captions). style 'karaoke' highlights the active word.
| Name | Required | Description | Default |
|---|---|---|---|
| style | No | ||
| job_id | Yes | ||
| position | No | ||
| font_name | No | ||
| font_size | No | ||
| uppercase | No | ||
| clip_index | Yes | 0-based index from list_clips. | |
| font_color | No | Hex color, e.g. #FFFFFF. | |
| highlight_color | No | Karaoke active-word color. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions that the karaoke style highlights the active word, but it does not disclose side effects (e.g., whether it overwrites the clip), reversibility, output format, or any prerequisites. For a mutation tool, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with zero redundancy, front-loading the main purpose and then adding one key behavioral detail about the karaoke style. Every word earns its place, and the structure is clean and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 9 parameters, no output schema, and no annotations, the description is too sparse. It doesn't explain the role of most parameters, doesn't mention that it appears to be job-based (given the required job_id), and doesn't clarify what happens to the clip after styling. The description leaves too much for the agent to infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%, and the description adds little beyond that. It mentions the karaoke style behavior, but that's already captured by the highlight_color parameter description and the style enum. Other parameters like position, font_name, font_size, and uppercase are left unexplained, and the description does not compensate for the low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 're-style' and the resource 'captions of one clip', making it clear the tool modifies existing captions rather than adding new ones. The parenthetical 'clips already ship with default captions' clarifies scope and distinguishes it from any caption-adding tool among the siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool (when you want to change the style of existing captions) by noting that clips already have captions, but it doesn't explicitly state when not to use it or name alternative tools like recut_clip or process_video. The guidance is implicit, not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_uploadReserve an upload slot for a local video fileAInspect
Use when the user hands you a video FILE instead of a link. Returns an upload_url: send the file's raw bytes to it with an HTTP PUT (for example curl -T video.mp4 <upload_url>), then call process_video with the returned upload_id. Unused slots are deleted after expires_in seconds (6 h); max size is max_mb. FALLBACK if your sandbox cannot reach upload_url: upload the file to tmpfiles.org (curl -F file=@video.mp4 https://tmpfiles.org/api/v1/upload), and pass the returned url to process_video as source_url right away (their files expire after 60 minutes; the server refreshes the signed download link itself). Otherwise ask the user to run the curl command or to share a public link.
| Name | Required | Description | Default |
|---|---|---|---|
| filename | No | Original file name (optional, for the extension). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does so thoroughly: it reveals the returned upload_url and upload_id, requires raw-byte HTTP PUT, states expiry and size limits, and explains the tmpfiles.org fallback lifecycle including the 60-minute expiry and signed-link refresh.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the primary use condition and workflow, then adds constraints and fallback instructions in a logical order. Every sentence carries operational value, including exact curl examples and timing details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must explain return values and next steps, and it does: upload_url, upload_id, expires_in, max_mb, and the required call to process_video. It also covers failure handling, making it complete for an agent to execute the workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single optional filename parameter is fully documented in the schema as 'Original file name (optional, for the extension).' The description does not add parameter-specific meaning, but the schema coverage is 100%, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and resource: reserve an upload slot for a local video file. It clearly distinguishes this from processing a video by explicitly contrasting file input with a link and by naming process_video as the follow-up step.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It opens with 'Use when the user hands you a video FILE instead of a link,' giving a direct condition for use. It also provides a fallback path for when the sandbox cannot reach the upload URL, and a final fallback to ask the user, so an agent knows exactly how to route the task.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_job_statusGet processing job statusAInspect
Status of a processing job: 'queued', 'processing', 'completed' or 'failed', with recent log lines and, once completed, the clips.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must convey behavioral expectations. It implies a read-only operation (status retrieval) and mentions return contents (status, logs, clips), but does not explicitly state that it has no side effects or require any special permissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence that efficiently conveys the tool's purpose and output. There is no redundancy or extraneous information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple status-check tool, the description is largely complete. It mentions the statuses returned, log lines, and clips, which gives a good sense of the output. It does not cover error scenarios or edge cases, but these are not critical for such a straightforward tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema only defines job_id as a string with no description (0% coverage). The description does not compensate by explaining the parameter, though the name is self-explanatory. Given the low schema coverage, the lack of any elaboration in the description reduces the score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: retrieving the status of a processing job, including possible statuses and additional data. It distinguishes itself from sibling tools like process_video or list_clips by focusing specifically on job status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The usage context is implied rather than explicit. It is clear that this tool should be used when checking job status, but it does not explicitly state when to use it over alternatives or provide any exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_quotaGet plan and remaining minutesAInspect
The authenticated user's plan and remaining processing minutes. Call before large jobs; process_video fails with quota_exceeded when the balance is insufficient.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the tool requires an authenticated user and that it returns quota information. It also mentions the failure mode of process_video, which is useful context. It does not explicitly state read-only behavior, but 'get' implies it, and there are no side effects to disclose.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences: the first states what the tool does, the second provides actionable usage guidance. No filler or redundancy. The information is front-loaded for quick comprehension.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple getter with no output schema, the description adequately explains what is returned (plan and remaining minutes) and when to invoke it. It also ties into the process_video failure mode, covering the main usage scenario. The tool is straightforward, so nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and an empty input schema, so the baseline is 4. The description correctly implies there is nothing to configure, and the focus is on the return value rather than inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: it retrieves the authenticated user's plan and remaining processing minutes. It uses a specific verb (get) and a specific resource (quota), distinguishing it from sibling tools like process_video or list_clips. No ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs when to use: 'Call before large jobs', and provides the consequence of not doing so: 'process_video fails with quota_exceeded when the balance is insufficient'. This names the alternative tool and the condition, giving the agent clear decision criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_clipsList a job's clipsCInspect
The clips of a completed job, with titles, platform-ready descriptions and download URLs. In MCP Apps-capable clients the result also renders as an interactive clip picker.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does state the output includes titles, descriptions, download URLs, and interactive rendering in MCP Apps clients. It does not mention side effects, error conditions, or read-only behavior, but the operation is inherently a listing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, directly states the key output fields, and includes a notable client rendering behavior without unnecessary detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple and the description covers return content, but it omits job_id semantics and any error or prerequisite details. It is adequate for a basic understanding but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter, job_id, has no schema description and the tool description does not explain its meaning, format, or requirements. Since schema description coverage is 0%, the description fails to compensate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The title 'List' and description 'The clips of a completed job' clearly identify the resource and action. However, it does not distinguish this tool from sibling tools like get_job_status or recut_clip, so it misses the differentiation expected for a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use list_clips versus alternatives, nor about prerequisites such as job completion. The description is not misleading but provides no usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
process_videoProcess a video into short clipsAInspect
Start clipping a video from its URL. OpenShorts downloads the source itself, transcribes it, finds the most viral moments with AI and renders vertical (9:16) clips. Captions and the AI hook line are burned by default; pass captions=false or auto_hook=false to skip either. Call this directly with the URL the user gave you; do not fetch, search or inspect the URL yourself first (you cannot access the video, and it is not needed). Returns a job_id immediately — the work takes minutes; poll get_job_status or pass webhook_url to be called back. The caller must own the content or hold the rights to process it (confirm_rights).
| Name | Required | Description | Default |
|---|---|---|---|
| layouts | No | Optional extra reframe layouts. 'auto' lets AI pick per video. | |
| captions | No | Default true: burn word-level captions on every clip. Set false when the source already has subtitles burned in (they would stack) or the user wants clean clips; add_subtitles can still caption a clip later. | |
| auto_hook | No | Burn the AI-written hook line (the clip's title) over the first seconds of each clip, as the dashboard does. Default true; set false for clean clips. | |
| upload_id | No | Instead of source_url: the id from create_upload after the file was PUT to its upload_url. Use when the user gave you a video file rather than a link. | |
| hook_style | No | Look of the hook text (with auto_hook). Default classic. | |
| source_url | No | Public video URL, passed through exactly as the user gave it (YouTube watch/short/live URL, a direct video file URL, or a tmpfiles.org link from the create_upload fallback). The server does the downloading. Omit when using upload_id. | |
| webhook_url | No | Optional public HTTPS URL POSTed once when the job finishes or fails. | |
| target_clips | No | How many clips to aim for. A target, not a guarantee: fewer come back when the material doesn't hold them. Default: the AI decides (usually 2-6). | |
| output_format | No | Clip aspect. Default auto (vertical). | |
| confirm_rights | Yes | Must be true: the user owns the content or has rights to process it. | |
| webhook_secret | No | Optional secret; the webhook body is then HMAC-SHA256 signed (X-OpenShorts-Signature). | |
| clip_max_seconds | No | Maximum clip length in seconds (default 60). Must be ≥ 5s above the minimum. | |
| clip_min_seconds | No | Minimum clip length in seconds (default 15). | |
| force_low_quality | No | Set true to proceed after a needs_confirmation low-resolution warning. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Fully discloses the tool's behavior: it returns a job_id immediately, takes minutes, and can notify via webhook on success or failure. It also transparently explains defaults for captions and hooks, plus the force_low_quality override. No annotations were provided, so the description bears the full burden and meets it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is thorough and logically organized, covering input, defaults, async flow, and rights. It is somewhat verbose, with several clauses that could be tightened, but every sentence contributes meaningful detail, so the length is justified given the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description provides all necessary context for a complex asynchronous tool: input variants, default behavior, output (job_id), completion mechanisms, and required rights confirmation. Without an output schema, it still covers what the caller should expect. No gaps are apparent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds significant context beyond the schema: it explains interactions (captions stacking, hook text, target_clips as a non-guarantee, clip length constraints, source_url vs upload_id). With 100% schema coverage already, the description enriches parameter understanding rather than repeating it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Start clipping a video') and identifies the resource (a video from a URL or upload). It differentiates from sibling tools by explicitly covering the asynchronous workflow and alternative input methods, making the tool's purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit instructions on when to call directly (pass the user's URL without fetching it) and when to use upload_id instead. It also explains the async behavior (poll or webhook) and how to handle low-resolution warnings, giving clear usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
publish_clipPublish a clip to social platformsAInspect
Post one clip to the user's connected accounts (TikTok lands as a draft in the app; Instagram and YouTube publish directly). Requires a connected social profile (cloud) or an Upload-Post key (self-host). Optionally schedule with an ISO-8601 scheduled_date.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | ||
| job_id | Yes | ||
| timezone | No | ||
| platforms | Yes | ||
| clip_index | Yes | ||
| description | No | ||
| scheduled_date | No | ISO-8601; omit to post now. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It does reveal important behavior: TikTok posts land as drafts, Instagram and YouTube publish directly, and scheduling is optional. However, it does not disclose potential side effects (e.g., irreversible actions, rate limits, error handling) or the response format, leaving gaps for a mutating operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, front-loaded with the core action and then providing crucial details about platform behavior, requirements, and scheduling. Every sentence adds value without redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 7 parameters and no output schema, the description is only partially complete. It covers prerequisites and scheduling but omits details on how parameters like job_id and clip_index are used, what happens after posting (return values), and potential error conditions. The description provides a good starting point but leaves an agent guessing on several operational aspects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 14% (only scheduled_date has a schema description). The description adds context for scheduled_date (ISO-8601) and mentions platforms implicitly, but fails to explain job_id, clip_index, title, description, timezone, or how they relate to the publishing process. Given the low schema coverage, the description should have compensated with parameter explanations but did not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: posting one clip to connected social accounts. It specifies the resource (clip) and action (post), and distinguishes from siblings by focusing on publication rather than processing, listing, or quota operations. The platform-specific behavior (TikTok draft vs direct) adds clarity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear prerequisites (connected social profile or Upload-Post key) and mentions optional scheduling, which helps the agent decide when to use it. However, it does not explicitly state when not to use this tool or compare it to alternatives like process_video or recut_clip, though the name and purpose make the context obvious.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recut_clipRe-cut a clip from an edited segment listAInspect
Re-render one clip from a new list of source-video segments (the same engine behind the dashboard's clip editor). Times are seconds in the ORIGINAL source video; segments are concatenated in the given order, so you can trim, extend, drop a dead moment in the middle, or reorder. Segments inside the clip's original range re-render in seconds; going outside needs the retained source video and re-runs the reframe engine.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | ||
| framing | No | Layout override: 'full' shows the whole source frame (no side-cropping), 'track' forces the subject-tracking crop, 'auto' resets to the AI classifier. Omit to keep the clip's current framing. Non-auto values re-run the reframe engine and need the retained source video. | |
| segments | Yes | Ordered source segments the new clip is made of. | |
| clip_index | Yes | 0-based index from list_clips. | |
| snap_to_words | No | Snap each boundary onto transcript word boundaries (recommended). | |
| reapply_captions | No | Burn default captions back on after the recut (default true). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must disclose behavior. It does explain that times are in the ORIGINAL source video, that segments are concatenated in order, and that going outside the clip's original range requires retained source video and re-runs the reframe engine. However, it does not disclose side effects (e.g., whether the original clip is replaced or a new one created), potential errors, or any permissions or rate limits. It covers the key mechanics but leaves significant behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, tightly written, and front-loads the purpose before explaining key constraints. Every clause adds value: the engine reference, time units, concatenation order, and the range/reframe condition. No fluff or repetition. This is a model of concise, structured documentation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 6 parameters, no output schema, and no annotations, so the description must carry a lot. It covers the core behavior and the critical constraint about retained source video. However, it does not explain what the tool returns after re-rendering, how to access the result, or what happens to the original clip. An agent would not know whether the operation is destructive or returns a new clip object. Given the complexity and missing output schema, this is a significant gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 83% (5 of 6 params described). The description adds meaning beyond the schema: it clarifies that segments are concatenated in the given order, that times are in the source video's seconds, and that reordering/dropping is possible. This goes beyond the simple 'Seconds in the source video' schema descriptions. For the one undocumented param (job_id), it's an ID that is self-explanatory. The description enhances understanding of segments and the reframe behavior, justifying a score above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Re-render one clip from a new list of source-video segments.' It names the exact operation and scope, and adds useful context ('the same engine behind the dashboard's clip editor'). This clearly distinguishes it from siblings like list_clips or add_subtitles, which are about different operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus alternatives. The description explains what it does but does not mention when it is appropriate or when to prefer a sibling (e.g., 'use process_video for initial processing, not here'). An agent must infer from the sibling names that this is for re-cutting an existing clip, but no direct comparison or exclusion is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v1.0.0- First observed
add_subtitles - First observed
create_upload - First observed
get_job_status - First observed
get_quota - First observed
list_clips - First observed
process_video - First observed
publish_clip - First observed
recut_clip
TDQS
Scored across 8 tools
Each tool has a distinct purpose: quota checking, upload creation, video processing, job status polling, clip listing, subtitle styling, clip recutting, and publishing. No two tools appear to overlap in functionality.
All tool names follow a consistent verb_noun snake_case pattern (e.g., list_clips, get_quota, process_video, publish_clip). The naming is predictable and clear.
With 8 tools, the set is well-scoped for a video clipping service, covering the main workflow from upload/processing through editing and publishing without being bloated.
The tool surface covers the full core lifecycle: create upload, process video, monitor jobs, list clips, edit clips, and publish. Minor gaps like deleting clips or retrieving individual clip details exist but are not critical to the primary workflow.
Maintenance
Related MCP Connectors
Turn long videos into AI-curated short clips: caption, reframe, thumbnail, schedule, and publish.
Turns long videos into captioned vertical clips, cut at the moments that stand on their own.
AI video editing + publishing: turn clips into vertical shorts, post to TikTok/Instagram/YouTube.
Turn any video or livestream into scored, captioned, ready-to-post vertical clips.
Related MCP Servers
- AlicenseAqualityFmaintenanceTurn YouTube videos into short clips — from Claude, Cursor, or any AI assistant that supports MCP. You give it a YouTube link. It finds the best moments, reframes them for vertical video, adds subtitles, and gives you download links. All from a chat.6642MIT
- AlicenseAqualityFmaintenanceCreate AI-powered short-form video clips from YouTube videos using any AI assistant. 9 tools for creating shorts, browsing caption templates, music, gameplay overlays, and meme hooks.91476MIT
- FlicenseNot gradedqualityDmaintenanceTurns long-form videos into short-form clips (TikTok/Reels) by reasoning over word-timestamped transcripts, with silence-aware rendering, STT-based validation, and optional reframing/captions.-

OpusClip MCPofficial
AlicenseNot gradedqualityBmaintenanceTurn long videos into AI-curated short clips. Submit a video file or URL and OpusClip finds the best moments, adds captions, reframes to vertical, and returns ready-to-post clips.2223MIT


