Skip to main content
Glama
FiloHany

Video RAG MCP Server

by FiloHany

Video RAG (Retrieval-Augmented Generation)

A powerful video retrieval and analysis system that uses the Ragie API to process, index, and query video content with natural language. This project enables semantic search through video content, extracting relevant video chunks based on text queries.

šŸŽÆ MCP-powered video-RAG using Ragie

This project demonstrates how to build a video-based Retrieval Augmented Generation (RAG) system powered by the Model Context Protocol (MCP). It uses Ragie's video ingestion and retrieval capabilities to enable semantic search and Q&A over video content and integrate them as MCP tools via Cursor IDE.

Tech Stack

  • Ragie for video ingestion + retrieval (video-RAG)

  • Cursor as the MCP host

  • Model Context Protocol (MCP) for AI assistant integration

Related MCP server: YouTube MCP Server

šŸŽÆ Features

  • Video Processing: Upload and process video files with audio-video analysis

  • Semantic Search: Query video content using natural language

  • Video Chunking: Extract specific video segments based on search results

  • MCP Integration: Model Context Protocol (MCP) server for AI assistant integration

  • Jupyter Notebook Support: Interactive development and experimentation

  • Automatic Indexing: Clear and rebuild video indexes as needed

šŸš€ Quick Start

Prerequisites

  • Python 3.12 or higher

  • Ragie API key

  • Video files to process

  • Cursor IDE (for MCP integration)

Setup and Installation

1. Install uv

First, let's install uv and set up our Python project and environment:

MacOS/Linux:

curl -LsSf https://astral.sh/uv/install.sh | sh

Windows:

powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"

2. Clone and Setup Project

# Clone the repository
git clone https://github.com/FiloHany/Video_RAG_MCP.git
cd video_rag

# Create virtual environment and activate it
uv venv
source .venv/bin/activate  # MacOS/Linux
# OR
.venv\Scripts\activate     # Windows

# Install dependencies
uv sync

3. Configure Environment Variables

Create a .env file in the project root:

RAGIE_API_KEY=your_ragie_api_key_here

4. Add Your Video Files

Place your video files in the video/ directory.

MCP Server Setup with Cursor IDE

1. Configure MCP Server in Cursor

  1. Go to Cursor settings

  2. Select MCP Tools

  3. Add new global MCP server

  4. In the JSON configuration, add:

{
    "mcpServers": {
        "ragie": {
            "command": "uv",
            "args": [
                "--directory",
                "/absolute/path/to/project_root",
                "run",
                "server.py"
            ],
            "env": {
                "RAGIE_API_KEY": "YOUR_RAGIE_API_KEY"
            }
        }
    }
}

Note: Replace /absolute/path/to/project_root with the actual absolute path to your project directory.

2. Connect MCP Server

  1. In Cursor MCP settings, make sure to toggle the button to connect the server to the host

  2. You should now see the MCP server listed in the MCP settings

3. Available MCP Tools

Your custom MCP server provides 3 tools:

  • ingest_data_tool: Ingests the video data to the Ragie index

  • retrieve_data_tool: Retrieves relevant data from the video based on user query

  • show_video_tool: Creates a short video chunk from the specified segment from the original video

You can now ingest your videos, retrieve relevant data and query it all using the Cursor Agent. The agent can even create the desired chunks from your video just with a single query!

šŸ“– Usage

Basic Usage

Run the main script to process videos and perform queries:

python main.py

This will:

  1. Clear the existing index

  2. Ingest all videos from the video/ directory

  3. Perform a sample query

Interactive Development

Use the Jupyter notebook for interactive development:

jupyter notebook video_rag.ipynb

MCP Server

Start the MCP server for AI assistant integration:

python server.py

šŸ”§ API Reference

Core Functions

clear_index()

Removes all documents from the Ragie index.

ingest_data(directory: str)

Processes and uploads all video files from the specified directory to the Ragie index.

Parameters:

  • directory (str): Path to the directory containing video files

retrieve_data(query: str)

Performs semantic search on the indexed video content.

Parameters:

  • query (str): Natural language query to search for in video content

Returns:

  • List of dictionaries containing:

    • text: The retrieved text content

    • document_name: Name of the source video file

    • start_time: Start timestamp of the video segment

    • end_time: End timestamp of the video segment

chunk_video(document_name: str, start_time: float, end_time: float, directory: str = "videos")

Extracts a specific video segment and saves it as a new file.

Parameters:

  • document_name (str): Name of the source video file

  • start_time (float): Start time in seconds

  • end_time (float): End time in seconds

  • directory (str): Directory containing the source video (default: "videos")

Returns:

  • Path to the created video chunk file

MCP Tools

The project includes an MCP server with the following tools:

ingest_data_tool(directory: str)

MCP wrapper for the ingest_data function.

retrieve_data_tool(query: str)

MCP wrapper for the retrieve_data function.

show_video_tool(document_name: str, start_time: float, end_time: float)

MCP wrapper for the chunk_video function.

šŸ“ Examples

Example 1: Basic Video Processing and Query

from main import clear_index, ingest_data, retrieve_data

# Clear existing index
clear_index()

# Ingest videos from directory
ingest_data("video")

# Query the video content
results = retrieve_data("What is the main topic of the video?")
print(results)

Example 2: Extract Video Chunks

from main import retrieve_data, chunk_video

# Get search results
results = retrieve_data("Show me the goal scoring moments")

# Extract video chunks for each result
for result in results:
    if result['start_time'] and result['end_time']:
        chunk_path = chunk_video(
            result['document_name'],
            result['start_time'],
            result['end_time']
        )
        print(f"Created chunk: {chunk_path}")

Example 3: Jupyter Notebook Workflow

# Load environment and initialize Ragie
import os
from dotenv import load_dotenv
from ragie import Ragie

load_dotenv()
ragie = Ragie(auth=os.getenv('RAGIE_API_KEY'))

# Upload a video
file_path = "video/messi-goals.mp4"
result = ragie.documents.create(request={
    "file": {
        "file_name": "messi-goals.mp4",
        "content": open(file_path, "rb"),
    },
    "mode": {
        "video": "audio_video"
    }
})

# Query the video
response = ragie.retrievals.retrieve(request={
    "query": "Give detailed description of the video with timestamp of the events"
})

# Process results
for chunk in response.scored_chunks:
    print(f"Time: {chunk.metadata.get('start_time')} - {chunk.metadata.get('end_time')}")
    print(f"Content: {chunk.text}")
    print("-" * 50)

šŸ—ļø Project Structure

video_rag/
ā”œā”€ā”€ main.py              # Core functionality and main script
ā”œā”€ā”€ server.py            # MCP server implementation
ā”œā”€ā”€ video_rag.ipynb      # Jupyter notebook for development
ā”œā”€ā”€ pyproject.toml       # Project configuration and dependencies
ā”œā”€ā”€ README.md           # This file
ā”œā”€ā”€ video/              # Directory for video files
│   └── messi-goals.mp4 # Example video file
└── video_chunks/       # Output directory for video chunks (created automatically)

šŸ”‘ Environment Variables

Variable

Description

Required

RAGIE_API_KEY

Your Ragie API authentication key

Yes

šŸ“¦ Dependencies

  • ragie: Video processing and retrieval API

  • moviepy: Video editing and manipulation

  • python-dotenv: Environment variable management

  • mcp: Model Context Protocol implementation

  • ipykernel: Jupyter notebook support

šŸ¤ Contributing

  1. Fork the repository

  2. Create a feature branch (git checkout -b feature/amazing-feature)

  3. Commit your changes (git commit -m 'Add some amazing feature')

  4. Push to the branch (git push origin feature/amazing-feature)

  5. Open a Pull Request

šŸ“„ License

This project is licensed under the MIT License - see the LICENSE file for details.

šŸ™ Acknowledgments

  • Ragie for providing the video processing API

  • MoviePy for video manipulation capabilities

  • MCP for AI assistant integration

šŸ“ž Support

If you encounter any issues or have questions:

  1. Check the Issues page

  2. Create a new issue with detailed information

  3. Include your Python version, error messages, and steps to reproduce

šŸ”„ Changelog

v0.1.0

  • Initial release

  • Basic video processing and retrieval functionality

  • MCP server integration

  • Jupyter notebook support

  • Video chunking capabilities

Available Tools

3 tools
ingest_data_toolB
Loads data from a directory into the Ragie index. Wait until the data is fully ingested before continuing.

Args:
    directory (str): The directory to load data from.

Returns:
    str: A message indicating that the data was loaded successfully.
ParametersJSON Schema
NameRequiredDescriptionDefault
directoryYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals that the operation is blocking ('Wait until the data is fully ingested') and indicates success/failure through a return message. However, it doesn't disclose critical behavioral traits such as what types of data are supported, whether the operation is idempotent, what happens to existing data in the index, error handling, or performance characteristics like rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded, with the core purpose stated first followed by parameter and return details. The three-sentence structure is efficient, though the 'Args' and 'Returns' sections could be integrated more seamlessly into the narrative flow. There's no wasted text, but minor improvements in cohesion are possible.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of a data ingestion tool with no annotations, no output schema, and low schema description coverage (0%), the description is incomplete. It covers the basic operation and parameter but lacks details on data formats, indexing behavior, error scenarios, and what 'fully ingested' entails. For a tool that modifies an index, more comprehensive guidance is needed to ensure safe and effective use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds minimal semantic context beyond the input schema. It specifies that the 'directory' parameter is 'The directory to load data from,' which slightly clarifies the purpose but doesn't provide format requirements (e.g., local path, network path), supported directory structures, or examples. With 0% schema description coverage and only one parameter, this is adequate but leaves gaps in practical usage details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with a specific verb ('Loads') and resource ('data from a directory into the Ragie index'). It distinguishes from sibling tools like 'retrieve_data_tool' and 'show_video_tool' by focusing on ingestion rather than retrieval or display. However, it doesn't explicitly differentiate from potential similar ingestion tools beyond the named siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides implied usage guidance through the instruction 'Wait until the data is fully ingested before continuing,' suggesting this is a blocking operation that should be used when immediate continuation isn't needed. However, it lacks explicit guidance on when to use this tool versus alternatives (e.g., batch vs. streaming ingestion) or any prerequisites for the directory structure.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

retrieve_data_toolB
Retrieves data from the Ragie index based on the query. The data is returned as a list of dictionaries, each containing the following keys:
- text: The text of the retrieved chunk
- document_name: The name of the document the chunk belongs to
- start_time: The start time of the chunk
- end_time: The end time of the chunk

Args:
    query (str): The query to retrieve data from the Ragie index.

Returns:
    list[dict]: The retrieved data.
ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that data is returned as a list of dictionaries with specific keys, which adds some behavioral context. However, it lacks details on permissions, rate limits, error handling, or whether the operation is read-only or has side effects. For a retrieval tool with zero annotation coverage, this is insufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded, starting with the core purpose. The structure includes sections for Args and Returns, which is efficient. However, the 'Returns' section partially repeats information from the description body, slightly reducing conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's low complexity (1 parameter, no output schema, no annotations), the description is somewhat complete. It covers the purpose, parameter semantics, and return format. However, it lacks usage guidelines and sufficient behavioral transparency, making it adequate but with clear gaps for effective agent use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds significant meaning beyond the input schema, which has 0% description coverage. It explains that the 'query' parameter is 'The query to retrieve data from the Ragie index,' clarifying its purpose and usage. With only one parameter, this compensation is effective, though it could be more detailed (e.g., query format).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Retrieves data from the Ragie index based on the query.' It specifies the verb ('retrieves'), resource ('data from the Ragie index'), and mechanism ('based on the query'). However, it doesn't explicitly differentiate from sibling tools like 'ingest_data_tool' or 'show_video_tool', which would require a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention sibling tools, prerequisites, or specific contexts for usage. The only implied usage is for retrieving data from the Ragie index, but this is basic and lacks explicit when/when-not instructions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

show_video_toolB
Creates and saves a video chunk based on the document name, start time, and end time of the chunk.
Returns a message indicating that the video chunk was created successfully.

Args:
    document_name (str): The name of the document the chunk belongs to
    start_time (float): The start time of the chunk
    end_time (float): The end time of the chunk

Returns:
    str: A message indicating that the video chunk was created successfully
ParametersJSON Schema
NameRequiredDescriptionDefault
document_nameYes
start_timeYes
end_timeYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden. It discloses that the tool creates and saves a video chunk (implying a write/mutation operation) and returns a success message, which is basic behavioral context. However, it lacks details on permissions, side effects, error handling, or rate limits, leaving significant gaps for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with separate sections for Args and Returns, and each sentence adds value (purpose, parameters, return). It's appropriately sized but could be slightly more front-loaded by moving the purpose statement earlier without the parameter details inline.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, 0% schema coverage, and no output schema, the description provides basic purpose and parameter info but is incomplete. It doesn't cover error cases, output format beyond a string message, or how the video chunk is stored/accessed, which are important for a creation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explicitly documents all three parameters (document_name, start_time, end_time) with brief semantics, adding meaningful context beyond the bare schema. However, it doesn't specify units for times (e.g., seconds) or document name format, leaving some ambiguity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool 'creates and saves a video chunk' with specific resources (document name, start time, end time), providing a concrete verb+resource combination. However, it doesn't distinguish from sibling tools (ingest_data_tool, retrieve_data_tool) which appear to handle different operations, so it misses full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like ingest_data_tool or retrieve_data_tool. The description mentions creating video chunks but doesn't specify prerequisites, constraints, or when-not-to-use scenarios, leaving usage context implied at best.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

A3.5/5.0
Disambiguation5/5

Each tool has a clearly distinct purpose with no overlap: ingest_data_tool loads data into an index, retrieve_data_tool queries the index for information, and show_video_tool generates video chunks from retrieved metadata. The descriptions clearly differentiate their functions, making misselection unlikely.

Naming Consistency5/5

All tools follow a consistent verb_noun pattern with snake_case naming: ingest_data_tool, retrieve_data_tool, and show_video_tool. This predictable naming convention makes the tool set easy to understand and navigate.

Tool Count3/5

With only 3 tools, the set feels thin for a video RAG system. While the core operations (ingest, retrieve, show) are covered, typical RAG workflows might benefit from additional tools like index management, query refinement, or batch processing. The count is borderline but functional.

Completeness4/5

The tools cover the essential RAG lifecycle: ingestion, retrieval, and video generation. However, there are minor gaps such as missing update/delete operations for the index, query history, or configuration tools. Agents can work around these, but the surface is not fully comprehensive.

Maintenance

ActivityInactive
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/FiloHany/Video_RAG_MCP'

If you have feedback or need assistance with the MCP directory API, please join our Discord server