Skip to main content
Glama
mario-andreschak

MCP Video Generation with Veo2

Veo2를 사용한 MCP 비디오 생성

대장간 배지

이 프로젝트는 Google의 Veo2 비디오 생성 기능을 제공하는 모델 컨텍스트 프로토콜(MCP) 서버를 구현합니다. 클라이언트가 텍스트 프롬프트나 이미지에서 비디오를 생성하고, MCP 리소스를 통해 생성된 비디오에 접근할 수 있도록 지원합니다.

특징

  • 텍스트 프롬프트에서 비디오 생성

  • 이미지에서 비디오 생성

  • MCP 리소스를 통해 생성된 비디오에 액세스하세요

  • 비디오 생성 템플릿 예시

  • stdio 및 SSE 전송 모두 지원

Related MCP server: hyper-video-service

예시 이미지

1dec9c71-07dc-4a6e-9e17-8da355d72ba1

예시 이미지를 비디오로

이미지에서 비디오로 - Grok에서 생성된 강아지

이미지를 비디오로 - 실제 고양이로부터

필수 조건

  • Node.js 18 이상

  • Gemini API와 Veo2 모델에 액세스할 수 있는 Google API 키(= API 키로 신용 카드를 설정해야 합니다! -> aistudio.google.com으로 이동)

설치

FLUJO 에 설치하기

  1. 서버 추가를 클릭하세요

  2. Github URL을 복사하여 FLUJO에 붙여넣기

  3. 분석, 복제, 설치, 빌드 및 저장을 클릭합니다.

Smithery를 통해 설치

Smithery를 통해 Claude Desktop에 mcp-video-generation-veo2를 자동으로 설치하려면 다음을 수행합니다.

지엑스피1

수동 설치

  1. 저장소를 복제합니다.

    git clone https://github.com/yourusername/mcp-video-generation-veo2.git
    cd mcp-video-generation-veo2
  2. 종속성 설치:

    npm install
  3. Google API 키로 .env 파일을 만듭니다.

    cp .env.example .env
    # Edit .env and add your Google API key

    .env 파일은 다음 변수를 지원합니다.

    • GOOGLE_API_KEY : Google API 키(필수)

    • PORT : 서버 포트(기본값: 3000)

    • STORAGE_DIR : 생성된 비디오를 저장하는 디렉토리(기본값: ./generated-videos)

    • LOG_LEVEL : 로깅 수준(기본값: 치명적)

      • 사용 가능한 레벨: 자세한 정보, 디버그, 정보, 경고, 오류, 치명적, 없음

      • 개발의 경우 더 자세한 로그를 보려면 debug 또는 info 로 설정하세요.

      • 생산을 위해 콘솔 출력을 최소화하기 위해 fatal 이라고 유지하십시오.

  4. 프로젝트를 빌드하세요:

    npm run build

용법

서버 시작

stdio 또는 SSE 전송을 사용하여 서버를 시작할 수 있습니다.

stdio 전송(기본값)

npm start
# or
npm start stdio

SSE 운송

npm start sse

이렇게 하면 서버가 포트 3000(또는 .env 파일에 지정된 포트)에서 시작됩니다.

MCP 도구

서버는 다음과 같은 MCP 도구를 제공합니다.

텍스트에서 비디오 생성

텍스트 프롬프트에서 비디오를 생성합니다.

매개변수:

  • prompt (문자열): 비디오 생성을 위한 텍스트 프롬프트

  • config (객체, 선택 사항): 구성 옵션

    • aspectRatio (문자열, 선택 사항): "16:9" 또는 "9:16"

    • personGeneration (문자열, 선택 사항): "dont_allow" 또는 "allow_adult"

    • numberOfVideos (숫자, 선택 사항): 1 또는 2

    • durationSeconds (숫자, 선택 사항): 5~8 사이

    • enhancePrompt (boolean, 선택 사항): 프롬프트를 향상시킬지 여부

    • negativePrompt (문자열, 선택 사항): 생성하지 않을 내용을 설명하는 텍스트

예:

{
  "prompt": "Panning wide shot of a serene forest with sunlight filtering through the trees, cinematic quality",
  "config": {
    "aspectRatio": "16:9",
    "personGeneration": "dont_allow",
    "durationSeconds": 8
  }
}

이미지에서 비디오 생성

이미지에서 비디오를 생성합니다.

매개변수:

  • image (문자열): Base64로 인코딩된 이미지 데이터

  • prompt (문자열, 선택 사항): 비디오 생성을 안내하는 텍스트 프롬프트

  • config (객체, 선택 사항): 구성 옵션(위와 동일하지만 personGeneration은 "dont_allow"만 지원함)

목록 생성 비디오

생성된 모든 비디오를 나열합니다.

MCP 리소스

서버는 다음과 같은 MCP 리소스를 제공합니다.

비디오://{id}

ID로 생성된 비디오에 접근합니다.

비디오://템플릿

비디오 생성 템플릿의 예를 살펴보세요.

개발

프로젝트 구조

  • src/ : 소스 코드

    • index.ts : 메인 진입점

    • server.ts : MCP 서버 구성

    • config.ts : 구성 처리

    • tools/ : MCP 도구 구현

    • resources/ : MCP 리소스 구현

    • services/ : 외부 서비스 통합

    • utils/ : 유틸리티 함수

건물

npm run build

개발 모드

npm run dev

특허

MIT

Available Tools

7 tools
generateImageC

Generate an image from a text prompt using Google Imagen

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYes
numberOfImagesNo
includeFullDataNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavioral traits. It does not mention side effects, authentication needs, rate limits, or whether the image is stored or returned directly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise. However, it sacrifices valuable information for brevity. Not every sentence earns its place when it omits critical details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 3 parameters, no output schema, and no annotations, the description is severely incomplete. It does not explain return values or parameter behavior, leaving the agent without enough context to use the tool effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%. The description adds no meaning beyond the parameter names and types in the input schema. It fails to explain prompt constraints, the number of images, or the includeFullData field.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb ('Generate'), the resource ('an image'), and the input ('from a text prompt using Google Imagen'). It effectively distinguishes from sibling tools focused on video generation or listing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives like getImage or listGeneratedImages. The description does not mention prerequisites, exclusions, or context for invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generateVideoFromGeneratedImageC

Generate a video from a generated image (one-step process)

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYes
aspectRatioNo16:9
videoPromptNo
autoDownloadNo
enhancePromptNo
negativePromptNo
numberOfImagesNo
numberOfVideosNo
durationSecondsNo
includeFullDataNo
personGenerationNodont_allow

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, and the description only mentions 'one-step process'. Important behavioral traits such as cost, async behavior, output format, and side effects are not disclosed. The description fails to compensate for the lack of annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is too brief given the tool's complexity (11 parameters). It is under-specified and does not effectively convey necessary information despite being concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is minimal given the large parameter count and lack of output schema. It does not explain the overall workflow (e.g., whether the image is generated first or input is needed) or how the tool fits with siblings. Important context like return values and constraints is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description must explain parameters but does not. None of the 11 parameters are described in the description, and the schema itself lacks descriptions. The agent cannot understand what parameters like 'autoDownload' or 'personGeneration' do.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the action (generate video) and the source (generated image). However, it does not differentiate from the sibling tool 'generateVideoFromImage', which could lead to confusion about when to use this tool vs that one.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like generateVideoFromImage or generateVideoFromText. There is no explanation of prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generateVideoFromImageC

Generate a video from an image

ParametersJSON Schema
NameRequiredDescriptionDefault
imageYes
promptNoGenerate a video from this image
aspectRatioNo16:9
autoDownloadNo
enhancePromptNo
negativePromptNo
numberOfVideosNo
durationSecondsNo
includeFullDataNo
personGenerationNodont_allow

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, and the description only states the basic operation. It does not disclose any behavioral traits such as processing time, image format requirements, or potential limitations, leaving the agent without important context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is under-specified for a tool with 10 parameters. It lacks structure and important details, making it minimally informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (10 parameters, no output schema, no annotations), the description is grossly incomplete. It fails to provide context for the output, parameter behavior, or usage scenarios.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero description coverage for parameters, and the description adds no meaning about any of the 10 parameters. Parameters like 'prompt', 'aspectRatio', 'durationSeconds' are not explained, so the agent cannot infer their purpose from the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool generates a video from an image, using a specific verb and resource. However, it does not differentiate from sibling tool 'generateVideoFromGeneratedImage', which has a similar purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when you have an image to convert to a video, but provides no explicit guidance on when to use this tool versus alternatives like 'generateVideoFromText' or 'generateVideoFromGeneratedImage'. No exclusions or context are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generateVideoFromTextC

Generate a video from a text prompt

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYes
aspectRatioNo16:9
autoDownloadNo
enhancePromptNo
negativePromptNo
numberOfVideosNo
durationSecondsNo
includeFullDataNo
personGenerationNodont_allow

TDQS

C2.6/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description bears full burden for behavioral transparency. It only states the basic function without disclosing any behavioral traits such as generation time, cost, success/failure handling, or return format. This is insufficient for responsible agent use.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise at one sentence, but given the tool's complexity (9 parameters), it is too brief. While front-loaded with the core action, it sacrifices necessary detail for brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Considering the tool's complexity (9 parameters, no output schema, no annotations), the description is critically incomplete. It lacks details on output, behavior, parameter effects, and usage context, making it inadequate for correct agent selection and invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, yet the description adds no value by explaining parameters like enhancePrompt, autoDownload, includeFullData, or personGeneration. Agents must infer meaning from names alone, risking misuse. At minimum, key parameters should be explained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Generate a video from a text prompt' clearly states the verb (Generate) and resource (video from text), distinguishing it from siblings like generateImage (image from text) and generateVideoFromImage (video from existing image). It's specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool vs alternatives. It doesn't mention scenarios, prerequisites, or exclusions, leaving the agent without context for appropriate usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

getImageB

Get a specific image by ID

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
includeFullDataNo

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description does not disclose behavioral traits beyond the basic operation. It lacks details about error handling, response format, or side effects. With no annotations, the description carries the full burden but provides minimal insight.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence. Every word is necessary and there is no wasted information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of an output schema, no parameter descriptions in the schema, and no annotations, the description is insufficient. It does not explain return values, error states, or the effect of optional parameters, leaving significant gaps for the agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description only hints at the 'id' parameter through the phrase 'by ID', but does not explain the 'includeFullData' parameter. Schema description coverage is 0%, so the description should compensate but does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get a specific image by ID' clearly states the verb (Get), resource (specific image), and method (by ID). It effectively distinguishes from sibling tools like generateImage, which are generative in nature.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. For example, it does not clarify that this tool is for retrieving existing images, while siblings handle generation or listing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

listGeneratedImagesB

List all generated images

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided; description does not disclose read-only nature, auth requirements, or any side effects. Simply states 'list all', offering no behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely concise at one sentence, no wasted words. Could be slightly improved by front-loading key info, but currently adequate.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Simple list tool with no params or output schema; however, lacks any mention of pagination or ordering, which may be necessary for completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Tool has zero parameters; schema coverage is 100%. Description adds no parameter info, but none is needed. Baseline 4 is appropriate for no-param tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states verb (list) and resource (generated images), distinguishing from siblings like getImage (single) and generateImage.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool vs alternatives; no mention of filtering, pagination, or context such as if the list is all images or scoped to a user.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

listGeneratedVideosB

List all generated videos

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, and the description merely repeats the tool name without disclosing any behavioral traits (e.g., pagination, scope of 'all', or read-only nature).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single, efficient sentence with no extraneous words. Appropriate for a simple list tool with no parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter list tool, the description is acceptable but lacks details like result format, sorting, or limits. Could be improved with a note on scope or output structure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters exist (schema is empty), so the description has no parameter semantics to add. Baseline 4 applies as the schema already covers 100% of parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'List all generated videos' clearly states the verb (List) and resource (generated videos), distinguishing it from siblings like generateImage, generateVideoFromText, and listGeneratedImages.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives such as getImage or generateVideoFromText. The description does not specify context, prerequisites, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 7 tool updatesv1.0.0
    • First observedgenerateImage
    • First observedgenerateVideoFromGeneratedImage
    • First observedgenerateVideoFromImage
    • First observedgenerateVideoFromText
    • First observedgetImage
    • First observedlistGeneratedImages
    • First observedlistGeneratedVideos

TDQS

B3/5.0

Scored across 7 tools

Disambiguation4/5

Most tools have distinct purposes, but generateVideoFromGeneratedImage and generateVideoFromImage overlap in functionality; one is described as 'one-step process' but the distinction is not immediately clear from names alone.

Naming Consistency4/5

Tool names follow a verb_noun pattern, but there is inconsistency: 'generateImage' uses a direct object while video tools use 'from' prepositional phrases. Also, 'getImage' differs from 'listGeneratedImages/listGeneratedVideos' in tense.

Tool Count5/5

With 7 tools, the set is well-scoped for an image/video generation server, covering creation and listing without being overly numerous.

Completeness3/5

The tool surface includes generation and listing but lacks retrieval of individual videos (no getVideo), update, and delete operations, leaving notable gaps in lifecycle management.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers