Video RAG MCP Server
The Video RAG MCP Server enables video-based retrieval-augmented generation through semantic search and video processing capabilities:
Ingest video data: Load and process video files from a specified directory into the Ragie index, including audio-video analysis for indexing
Perform semantic search: Query indexed video content using natural language to retrieve relevant text excerpts with timestamps and source file information
Extract video chunks: Create and save specific video segments from original videos based on document name, start time, and end time
AI assistant integration: Provide these capabilities as MCP tools (
ingest_data_tool,retrieve_data_tool,show_video_tool) for use in Cursor IDE and other MCP-compatible environments
Provides environment variable management for storing and accessing API keys securely.
Allows cloning and management of the video RAG project repository.
Enables issue tracking and repository management for the video RAG project.
Enables interactive development and experimentation with video RAG capabilities through notebook support.
Integrates with Python 3.12 or higher for video processing and analysis capabilities.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Video RAG MCP Serverfind the part where they explain quantum computing concepts"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Video RAG (Retrieval-Augmented Generation)
A powerful video retrieval and analysis system that uses the Ragie API to process, index, and query video content with natural language. This project enables semantic search through video content, extracting relevant video chunks based on text queries.
šÆ MCP-powered video-RAG using Ragie
This project demonstrates how to build a video-based Retrieval Augmented Generation (RAG) system powered by the Model Context Protocol (MCP). It uses Ragie's video ingestion and retrieval capabilities to enable semantic search and Q&A over video content and integrate them as MCP tools via Cursor IDE.
Tech Stack
Ragie for video ingestion + retrieval (video-RAG)
Cursor as the MCP host
Model Context Protocol (MCP) for AI assistant integration
Related MCP server: YouTube MCP Server
šÆ Features
Video Processing: Upload and process video files with audio-video analysis
Semantic Search: Query video content using natural language
Video Chunking: Extract specific video segments based on search results
MCP Integration: Model Context Protocol (MCP) server for AI assistant integration
Jupyter Notebook Support: Interactive development and experimentation
Automatic Indexing: Clear and rebuild video indexes as needed
š Quick Start
Prerequisites
Python 3.12 or higher
Ragie API key
Video files to process
Cursor IDE (for MCP integration)
Setup and Installation
1. Install uv
First, let's install uv and set up our Python project and environment:
MacOS/Linux:
curl -LsSf https://astral.sh/uv/install.sh | shWindows:
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"2. Clone and Setup Project
# Clone the repository
git clone https://github.com/FiloHany/Video_RAG_MCP.git
cd video_rag
# Create virtual environment and activate it
uv venv
source .venv/bin/activate # MacOS/Linux
# OR
.venv\Scripts\activate # Windows
# Install dependencies
uv sync3. Configure Environment Variables
Create a .env file in the project root:
RAGIE_API_KEY=your_ragie_api_key_here4. Add Your Video Files
Place your video files in the video/ directory.
MCP Server Setup with Cursor IDE
1. Configure MCP Server in Cursor
Go to Cursor settings
Select MCP Tools
Add new global MCP server
In the JSON configuration, add:
{
"mcpServers": {
"ragie": {
"command": "uv",
"args": [
"--directory",
"/absolute/path/to/project_root",
"run",
"server.py"
],
"env": {
"RAGIE_API_KEY": "YOUR_RAGIE_API_KEY"
}
}
}
}Note: Replace /absolute/path/to/project_root with the actual absolute path to your project directory.
2. Connect MCP Server
In Cursor MCP settings, make sure to toggle the button to connect the server to the host
You should now see the MCP server listed in the MCP settings
3. Available MCP Tools
Your custom MCP server provides 3 tools:
ingest_data_tool: Ingests the video data to the Ragie indexretrieve_data_tool: Retrieves relevant data from the video based on user queryshow_video_tool: Creates a short video chunk from the specified segment from the original video
You can now ingest your videos, retrieve relevant data and query it all using the Cursor Agent. The agent can even create the desired chunks from your video just with a single query!
š Usage
Basic Usage
Run the main script to process videos and perform queries:
python main.pyThis will:
Clear the existing index
Ingest all videos from the
video/directoryPerform a sample query
Interactive Development
Use the Jupyter notebook for interactive development:
jupyter notebook video_rag.ipynbMCP Server
Start the MCP server for AI assistant integration:
python server.pyš§ API Reference
Core Functions
clear_index()
Removes all documents from the Ragie index.
ingest_data(directory: str)
Processes and uploads all video files from the specified directory to the Ragie index.
Parameters:
directory(str): Path to the directory containing video files
retrieve_data(query: str)
Performs semantic search on the indexed video content.
Parameters:
query(str): Natural language query to search for in video content
Returns:
List of dictionaries containing:
text: The retrieved text contentdocument_name: Name of the source video filestart_time: Start timestamp of the video segmentend_time: End timestamp of the video segment
chunk_video(document_name: str, start_time: float, end_time: float, directory: str = "videos")
Extracts a specific video segment and saves it as a new file.
Parameters:
document_name(str): Name of the source video filestart_time(float): Start time in secondsend_time(float): End time in secondsdirectory(str): Directory containing the source video (default: "videos")
Returns:
Path to the created video chunk file
MCP Tools
The project includes an MCP server with the following tools:
ingest_data_tool(directory: str)
MCP wrapper for the ingest_data function.
retrieve_data_tool(query: str)
MCP wrapper for the retrieve_data function.
show_video_tool(document_name: str, start_time: float, end_time: float)
MCP wrapper for the chunk_video function.
š Examples
Example 1: Basic Video Processing and Query
from main import clear_index, ingest_data, retrieve_data
# Clear existing index
clear_index()
# Ingest videos from directory
ingest_data("video")
# Query the video content
results = retrieve_data("What is the main topic of the video?")
print(results)Example 2: Extract Video Chunks
from main import retrieve_data, chunk_video
# Get search results
results = retrieve_data("Show me the goal scoring moments")
# Extract video chunks for each result
for result in results:
if result['start_time'] and result['end_time']:
chunk_path = chunk_video(
result['document_name'],
result['start_time'],
result['end_time']
)
print(f"Created chunk: {chunk_path}")Example 3: Jupyter Notebook Workflow
# Load environment and initialize Ragie
import os
from dotenv import load_dotenv
from ragie import Ragie
load_dotenv()
ragie = Ragie(auth=os.getenv('RAGIE_API_KEY'))
# Upload a video
file_path = "video/messi-goals.mp4"
result = ragie.documents.create(request={
"file": {
"file_name": "messi-goals.mp4",
"content": open(file_path, "rb"),
},
"mode": {
"video": "audio_video"
}
})
# Query the video
response = ragie.retrievals.retrieve(request={
"query": "Give detailed description of the video with timestamp of the events"
})
# Process results
for chunk in response.scored_chunks:
print(f"Time: {chunk.metadata.get('start_time')} - {chunk.metadata.get('end_time')}")
print(f"Content: {chunk.text}")
print("-" * 50)šļø Project Structure
video_rag/
āāā main.py # Core functionality and main script
āāā server.py # MCP server implementation
āāā video_rag.ipynb # Jupyter notebook for development
āāā pyproject.toml # Project configuration and dependencies
āāā README.md # This file
āāā video/ # Directory for video files
ā āāā messi-goals.mp4 # Example video file
āāā video_chunks/ # Output directory for video chunks (created automatically)š Environment Variables
Variable | Description | Required |
| Your Ragie API authentication key | Yes |
š¦ Dependencies
ragie: Video processing and retrieval API
moviepy: Video editing and manipulation
python-dotenv: Environment variable management
mcp: Model Context Protocol implementation
ipykernel: Jupyter notebook support
š¤ Contributing
Fork the repository
Create a feature branch (
git checkout -b feature/amazing-feature)Commit your changes (
git commit -m 'Add some amazing feature')Push to the branch (
git push origin feature/amazing-feature)Open a Pull Request
š License
This project is licensed under the MIT License - see the LICENSE file for details.
š Acknowledgments
Ragie for providing the video processing API
MoviePy for video manipulation capabilities
MCP for AI assistant integration
š Support
If you encounter any issues or have questions:
Check the Issues page
Create a new issue with detailed information
Include your Python version, error messages, and steps to reproduce
š Changelog
v0.1.0
Initial release
Basic video processing and retrieval functionality
MCP server integration
Jupyter notebook support
Video chunking capabilities
Available Tools
3 toolsingest_data_toolB
Loads data from a directory into the Ragie index. Wait until the data is fully ingested before continuing.
Args:
directory (str): The directory to load data from.
Returns:
str: A message indicating that the data was loaded successfully.
| Name | Required | Description | Default |
|---|---|---|---|
| directory | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals that the operation is blocking ('Wait until the data is fully ingested') and indicates success/failure through a return message. However, it doesn't disclose critical behavioral traits such as what types of data are supported, whether the operation is idempotent, what happens to existing data in the index, error handling, or performance characteristics like rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized and front-loaded, with the core purpose stated first followed by parameter and return details. The three-sentence structure is efficient, though the 'Args' and 'Returns' sections could be integrated more seamlessly into the narrative flow. There's no wasted text, but minor improvements in cohesion are possible.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of a data ingestion tool with no annotations, no output schema, and low schema description coverage (0%), the description is incomplete. It covers the basic operation and parameter but lacks details on data formats, indexing behavior, error scenarios, and what 'fully ingested' entails. For a tool that modifies an index, more comprehensive guidance is needed to ensure safe and effective use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds minimal semantic context beyond the input schema. It specifies that the 'directory' parameter is 'The directory to load data from,' which slightly clarifies the purpose but doesn't provide format requirements (e.g., local path, network path), supported directory structures, or examples. With 0% schema description coverage and only one parameter, this is adequate but leaves gaps in practical usage details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('Loads') and resource ('data from a directory into the Ragie index'). It distinguishes from sibling tools like 'retrieve_data_tool' and 'show_video_tool' by focusing on ingestion rather than retrieval or display. However, it doesn't explicitly differentiate from potential similar ingestion tools beyond the named siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides implied usage guidance through the instruction 'Wait until the data is fully ingested before continuing,' suggesting this is a blocking operation that should be used when immediate continuation isn't needed. However, it lacks explicit guidance on when to use this tool versus alternatives (e.g., batch vs. streaming ingestion) or any prerequisites for the directory structure.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
retrieve_data_toolB
Retrieves data from the Ragie index based on the query. The data is returned as a list of dictionaries, each containing the following keys:
- text: The text of the retrieved chunk
- document_name: The name of the document the chunk belongs to
- start_time: The start time of the chunk
- end_time: The end time of the chunk
Args:
query (str): The query to retrieve data from the Ragie index.
Returns:
list[dict]: The retrieved data.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that data is returned as a list of dictionaries with specific keys, which adds some behavioral context. However, it lacks details on permissions, rate limits, error handling, or whether the operation is read-only or has side effects. For a retrieval tool with zero annotation coverage, this is insufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized and front-loaded, starting with the core purpose. The structure includes sections for Args and Returns, which is efficient. However, the 'Returns' section partially repeats information from the description body, slightly reducing conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity (1 parameter, no output schema, no annotations), the description is somewhat complete. It covers the purpose, parameter semantics, and return format. However, it lacks usage guidelines and sufficient behavioral transparency, making it adequate but with clear gaps for effective agent use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds significant meaning beyond the input schema, which has 0% description coverage. It explains that the 'query' parameter is 'The query to retrieve data from the Ragie index,' clarifying its purpose and usage. With only one parameter, this compensation is effective, though it could be more detailed (e.g., query format).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Retrieves data from the Ragie index based on the query.' It specifies the verb ('retrieves'), resource ('data from the Ragie index'), and mechanism ('based on the query'). However, it doesn't explicitly differentiate from sibling tools like 'ingest_data_tool' or 'show_video_tool', which would require a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention sibling tools, prerequisites, or specific contexts for usage. The only implied usage is for retrieving data from the Ragie index, but this is basic and lacks explicit when/when-not instructions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
show_video_toolB
Creates and saves a video chunk based on the document name, start time, and end time of the chunk.
Returns a message indicating that the video chunk was created successfully.
Args:
document_name (str): The name of the document the chunk belongs to
start_time (float): The start time of the chunk
end_time (float): The end time of the chunk
Returns:
str: A message indicating that the video chunk was created successfully
| Name | Required | Description | Default |
|---|---|---|---|
| document_name | Yes | ||
| start_time | Yes | ||
| end_time | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It discloses that the tool creates and saves a video chunk (implying a write/mutation operation) and returns a success message, which is basic behavioral context. However, it lacks details on permissions, side effects, error handling, or rate limits, leaving significant gaps for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with separate sections for Args and Returns, and each sentence adds value (purpose, parameters, return). It's appropriately sized but could be slightly more front-loaded by moving the purpose statement earlier without the parameter details inline.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, 0% schema coverage, and no output schema, the description provides basic purpose and parameter info but is incomplete. It doesn't cover error cases, output format beyond a string message, or how the video chunk is stored/accessed, which are important for a creation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explicitly documents all three parameters (document_name, start_time, end_time) with brief semantics, adding meaningful context beyond the bare schema. However, it doesn't specify units for times (e.g., seconds) or document name format, leaving some ambiguity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool 'creates and saves a video chunk' with specific resources (document name, start time, end time), providing a concrete verb+resource combination. However, it doesn't distinguish from sibling tools (ingest_data_tool, retrieve_data_tool) which appear to handle different operations, so it misses full differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like ingest_data_tool or retrieve_data_tool. The description mentions creating video chunks but doesn't specify prerequisites, constraints, or when-not-to-use scenarios, leaving usage context implied at best.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
Each tool has a clearly distinct purpose with no overlap: ingest_data_tool loads data into an index, retrieve_data_tool queries the index for information, and show_video_tool generates video chunks from retrieved metadata. The descriptions clearly differentiate their functions, making misselection unlikely.
All tools follow a consistent verb_noun pattern with snake_case naming: ingest_data_tool, retrieve_data_tool, and show_video_tool. This predictable naming convention makes the tool set easy to understand and navigate.
With only 3 tools, the set feels thin for a video RAG system. While the core operations (ingest, retrieve, show) are covered, typical RAG workflows might benefit from additional tools like index management, query refinement, or batch processing. The count is borderline but functional.
The tools cover the essential RAG lifecycle: ingestion, retrieval, and video generation. However, there are minor gaps such as missing update/delete operations for the index, query history, or configuration tools. Agents can work around these, but the surface is not fully comprehensive.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
YouTube transcripts, search, channels, playlists and bulk transcript jobs for AI agents. 14 tools.
Understand your videos with Reka AI ā search, ask questions, and extract insights.
15 media & data tools for AI agents: search, transcribe, subtitles, voiceover, translate & more.
Video analysis AI: transcripts, summaries, visual scenes/shots, clips, answers in natural language.
Related MCP Servers
AlicenseNot gradedqualityDmaintenanceEnables AI agents to upload, index, search, and analyze videos through the Reka Vision API, supporting natural language search, visual question answering, and extraction of transcripts and captions.1Apache 2.0- FlicenseNot gradedqualityDmaintenanceEnables LLMs to interact with YouTube videos by fetching transcripts, summarizing content, and answering questions based on video context.
- AlicenseNot gradedqualityAmaintenanceEnables agents to analyze long videos by downloading them, extracting transcripts and storyboards, and zooming into specific moments with high-resolution frames and OCR.MIT
- AlicenseNot gradedqualityDmaintenanceProvides tools for searching YouTube videos, retrieving transcripts, and performing semantic search over video content.MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/FiloHany/Video_RAG_MCP'
If you have feedback or need assistance with the MCP directory API, please join our Discord server