MCP-Powered Video RAG
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
| logging | {} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| extensions | {
"io.modelcontextprotocol/ui": {}
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| ingest_videoA | Ingest a video file into the RAG system. This transcribes the video using Whisper (locally, for free), splits the transcript into timed chunks, embeds them, and stores them in a local ChromaDB vector database for semantic search. |
| search_videoB | Semantically search across all indexed video transcripts. Returns the most relevant timestamped transcript chunks for a given query. You can optionally filter to a specific video by name. |
| ask_videoA | Ask a natural language question about your videos and get an AI-generated answer. This tool uses RAG: it retrieves the most relevant transcript chunks from ChromaDB, then sends them as context to a Groq LLM (free tier) to generate a precise answer with timestamps so you can jump directly to the relevant moment in the video. |
| list_videosB | List all videos currently indexed in the RAG system. Returns a summary of each ingested video including its name, path, and chunk count. Returns: JSON list of indexed videos with metadata. |
| delete_videoA | Remove a video and all its indexed data from the RAG system. This deletes all transcript chunks for the specified video from ChromaDB. The original video file is NOT deleted — only the indexed data is removed. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 5 tools
Each tool maps to a distinct action (ingest, search, ask, list, delete). The only potential overlap is search_video vs ask_video, since both retrieve transcript chunks, but the descriptions clearly differentiate raw chunk retrieval from LLM-generated answers with context.
All five tools follow a strict verb_noun snake_case pattern (ingest_video, search_video, ask_video, list_videos, delete_video). The convention is predictable and uniform.
Five tools is well-scoped for a video RAG system. Each tool covers a meaningful, non-redundant part of the workflow with no filler.
The surface covers the full lifecycle: ingestion, two retrieval modes (search and Q&A), listing indexed content, and deletion. Re-ingesting a video handles the update case, so there are no obvious dead ends.