Skip to main content
Glama

media-mcp (Node.js)

中文 | English

npm version Node.js >=18 License: MIT

MCP 프로토콜 기반의 비디오 향상 및 이미지 세분화 서비스로, MCP Client-Server와 백엔드 HTTP Server 간의 상호작용을 수행합니다.

기능

다음과 같은 MCP 도구를 제공합니다:

  • create_task - 비디오 향상 작업 생성 (URL 또는 로컬 파일 업로드 지원)

  • get_task_status - 작업 상태 조회

  • enhance_video_sync - 동기식 비디오 향상 (완료될 때까지 차단)

  • sam3_predict - SAM3 이미지 세분화 (로컬 경로, URL 또는 Base64 이미지 지원)

Related MCP server: Grok Imagine Video MCP Server

사전 요구 사항

  • Node.js >= 18 (확인: node --version)

  • API Key (인증용, 서비스 제공업체에 문의하여 획득)

간편 설치 (권장)

사용 중인 AI 에이전트에 확정된 MCP 구성 경로가 있는 경우, 아래 문장을 복사하여 AI에게 보내십시오:

帮我安装 npm 包 @avclabs.ai/media-mcp 作为 MCP server。我的 API Key 是:sk-xxxxxxxx。

AI가 자동으로 다음 작업을 완료합니다:

  1. 사용 중인 MCP 클라이언트 감지

  2. 구성 파일 경로 찾기

  3. 올바른 구성 작성

  4. 클라이언트 재시작 안내

수동 설치

설치가 필요 없으며, MCP 클라이언트 구성에서 npx를 사용하여 직접 실행하십시오.

1. Claude Code (CLI)

Claude Code에서 실행:

/mcp

출력에서 **"User MCPs"**에 해당하는 구성 파일 경로를 확인한 후 해당 파일을 편집하십시오.

일반적인 경로 (/mcp를 사용할 수 없는 경우):

  • Windows: %USERPROFILE%\.claude.json

  • macOS: ~/.claude.json

  • Linux: ~/.claude.json

  • 구버전/대체: ~/.claude/mcp.json

다음 내용을 붙여넣으십시오 (your-api-key를 실제 API Key로 교체):

{
  "mcpServers": {
    "video-enhancement": {
      "command": "npx",
      "args": ["-y", "@avclabs.ai/media-mcp@latest"],
      "env": {
        "API_KEY": "your-api-key"
      }
    }
  }
}

저장 후 /mcp를 실행하여 로드 성공 여부를 확인하십시오.

2. Cursor

설정 > Tools & MCPs > Add New MCP Server로 이동:

  • Name: video-enhancement

  • Type: command

  • Command:

env HTTP_API_KEY=your-api-key npx -y @avclabs.ai/media-mcp@latest

또는 ~/.cursor/mcp.json 편집:

{
  "mcpServers": {
    "video-enhancement": {
      "command": "npx",
      "args": ["-y", "@avclabs.ai/media-mcp@latest"],
      "env": {
        "API_KEY": "your-api-key"
      }
    }
  }
}

설치 확인

클라이언트를 재시작한 후 도구가 성공적으로 로드되었는지 확인하십시오:

  1. 또는 AI에게 직접 질문: "사용 가능한 도구가 무엇인가요?"

  2. 다음이 보여야 합니다: create_task, get_task_status, enhance_video_sync, sam3_predict

구성 항목

변수명

필수

기본값

설명

API_KEY

예

-

API 인증 키 (비디오 향상 및 SAM3 공용)

HTTP_API_BASE_URL

아니요

https://mcp.avc.ai/enhance

비디오 향상 서비스 인터페이스 주소

SAM3_API_BASE_URL

아니요

https://mcp.avc.ai/sam

SAM3 서비스 인터페이스 주소

SAM3_POLL_INTERVAL

아니요

2000

폴링 간격 (밀리초)

SAM3_POLL_MAX_ATTEMPTS

아니요

60

최대 폴링 횟수

사용자 지정 서비스 주소

{
  "env": {
    "HTTP_API_BASE_URL": "https://your-endpoint.com",
    "API_KEY": "your-api-key",
    "SAM3_API_BASE_URL": "http://localhost:8001"
  }
}

또는 명령줄 인수를 통해:

npx -y @avclabs.ai/media-mcp@latest --base-url https://your-endpoint.com --api-key your-api-key --sam3-base-url http://localhost:8001

사용 예시

구성이 완료되면 AI에게 자연어로 다음과 같이 말하십시오:

"이 비디오를 1080p로 향상시켜 줘: https://example.com/video.mp4"

"내 바탕화면의 video.mp4를 2k 화질로 높여줘"

AI가 자동으로 해당 도구를 호출하여 작업을 완료합니다.

"이 이미지를 분석해서 안에 있는 모든 물체를 찾아줘: C:\Users\xxx\photo.png"

"SAM3로 이 이미지를 세분화해 줘, 프롬프트는 'find all cars'"

제공되는 도구

create_task

비디오 향상 작업 생성 (비동기).

매개변수

유형

필수

기본값

설명

video_source

string

예

-

비디오 URL 또는 로컬 파일 경로 (URL은 공개적으로 액세스 가능해야 하며, 로그인이나 서명이 필요한 링크는 지원하지 않음)

type

string

아니요

url

url 또는 local

resolution

string

아니요

720p

480p, 540p, 720p, 1080p, 2k

반환값:

{
  "success": true,
  "task_id": "xxx",
  "status": "wait"
}

get_task_status

작업 상태 조회.

매개변수

유형

필수

task_id

string

예

반환값:

{
  "success": true,
  "task_id": "xxx",
  "status": "completed",
  "progress": 100,
  "video_url": "https://..."
}

enhance_video_sync

동기식 비디오 향상 (완료될 때까지 차단).

매개변수

유형

필수

기본값

설명

video_source

string

예

-

비디오 URL 또는 로컬 파일 경로 (URL은 공개적으로 액세스 가능해야 함)

type

string

아니요

url

url 또는 local

resolution

string

아니요

720p

목표 해상도

poll_interval

number

아니요

5

폴링 간격 (초)

timeout

number

아니요

600

타임아웃 시간 (초)

sam3_predict

SAM3 세분화 API를 사용하여 이미지를 분석하고 추론 결과(masks, boxes, scores)를 생성합니다.

매개변수:

이미지 입력 (세 가지 중 하나 필수):

  • imagePath (string): 로컬 이미지의 절대 경로. 일반적인 이미지 형식(PNG, JPG, JPEG 등) 지원.

    • 예: "C:\\Users\\xxx\\photo.png", "/home/user/images/cat.jpg"

  • imageUrl (string): 공개적으로 액세스 가능한 이미지 URL.

    • 예: "https://example.com/photo.jpg"

  • imageBase64 (string): Base64 인코딩된 이미지 데이터.

    • 예: "iVBORw0KGgoAAAANSUhEUgAA..."

기타 매개변수:

  • prompt (string, required): 이미지에서 세분화할 대상 물체를 지정하는 영어 텍스트 프롬프트. 예: "person", "car", "a cat sitting on a sofa". SAM3 모델은 영어 프롬프트만 허용하므로 영어 설명을 권장합니다.

반환:

추론 완료 후 JSON 문자열을 직접 반환합니다. 이 JSON에는 다음 세 가지 필드가 포함됩니다:

  • masks: 2차원 배열. 각 요소는 입력 이미지와 크기가 같은 이진 마스크(0 또는 1)입니다.

  • boxes: 2차원 배열. 각 요소는 [x1, y1, x2, y2] 형식의 경계 상자 좌표입니다.

  • scores: 1차원 배열. 각 요소는 해당 감지 결과의 신뢰도 점수(0~1)입니다.

결과 JSON 내용 예시:

{
  "masks": [
    [[0, 0, 1, ...], [0, 1, 1, ...], ...],
    [[0, 0, 0, ...], [0, 0, 1, ...], ...]
  ],
  "boxes": [
    [120, 80, 300, 450],
    [400, 200, 600, 500]
  ],
  "scores": [0.95, 0.87]
}

자주 묻는 질문

첨부 파일을 드래그한 후 파일을 찾을 수 없다고 나오나요?

stdio MCP의 알려진 제한 사항입니다. Agent 인터페이스를 통해 파일을 드래그하거나 업로드할 때 파일 경로가 MCP Server로 전달되지 않을 수 있습니다.

해결 방법:

  1. 경로를 함께 제공 (권장): 이미지를 드래그한 후 텍스트에 로컬 절대 경로를 추가하십시오.

  2. 자동 인코딩 대기: Claude가 자동으로 이미지를 base64로 인코딩할 수 있습니다.

  3. 경로 질문에 답변: Claude가 경로를 물어보면 로컬 절대 경로를 직접 입력하십시오.

세 가지 입력 방식에 우선순위가 있나요?

엄격한 우선순위는 없습니다. Claude가 대화 맥락에 따라 가장 적절한 방식을 선택합니다.

지원되는 이미지 형식은 무엇인가요?

PNG, JPG, JPEG, BMP, WebP 등을 지원합니다. PNG 또는 JPG를 권장합니다.

URL 이미지 다운로드 실패 시 어떻게 하나요?

URL이 공개적으로 액세스 가능한지 확인하십시오. 로그인이 필요한 서비스의 경우 로컬로 다운로드한 후 imagePath를 사용하십시오.

파일 업로드 설명

type이 "local"일 때, MCP Server는 다음을 수행합니다:

  1. 로컬 파일 읽기

  2. 사전 서명된 URL을 통해 TOS 객체 스토리지로 직접 업로드

  3. 최대 파일 크기: 100MB

문제 해결

"command not found: npx"

Node.js >= 18을 설치하십시오: https://nodejs.org/

"오류: --api-key를 제공하거나 API_KEY를 설정해야 합니다"

API Key가 누락되었습니다. 구성의 env.API_KEY를 확인하십시오.

MCP Server가 클라이언트에서 빨간색/오류로 표시됨

로그 확인:

  • Claude Desktop macOS: ~/Library/Logs/Claude/mcp*.log

  • Claude Desktop Windows: %APPDATA%\Claude\logs\mcp*.log

  • Cursor: Output 패널 > MCP

전역 설치 (선택 사항)

매번 npx를 사용하고 싶지 않은 경우:

npm install -g @avclabs.ai/media-mcp

그런 다음 구성에서 "command": "media-mcp"와 "args": ["--api-key", "your-api-key"]를 함께 사용하십시오.

License

MIT License - 자세한 내용은 LICENSE 파일을 참조하십시오.

Available Tools

4 tools
create_taskB

创建视频增强任务(异步)

支持两种上传方式:

  1. URL 上传:提供视频 URL

  2. 本地上传:提供本地文件路径,MCP Server 自动上传到 TOS 对象存储

参数说明:

  • video_source: 视频 URL 或本地文件路径

  • type: "url" 或 "local"

  • resolution: 目标分辨率

ParametersJSON Schema
NameRequiredDescriptionDefault
video_sourceYes视频URL地址或本地文件路径(URL必须公网可访问,不支持需要登录或签名的链接)
typeNo上传类型:url=网络视频,local=本地文件url
resolutionNo目标分辨率,默认720p720p

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so the description must carry full burden. It notes async behavior and TOS upload but omits side effects, permissions, failure modes, or rate limits. The description only partially discloses behavioral traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with three bullet points and front-loaded purpose. Every sentence earns its place, but structure could be slightly improved with clearer differentiation from siblings.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers input parameters well but lacks output schema explanation (e.g., task ID or status). With no annotations and multiple siblings, more context on post-creation steps would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema has 100% description coverage, so baseline is 3. The description groups parameters and explains the two upload modes, but does not add new information beyond the schema's existing parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool creates an async video enhancement task and distinguishes between two upload methods (URL and local). It uses specific verbs and resources, and is not a tautology.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use each upload type (URL vs local) but does not explicitly guide when to use this async tool over its sync sibling (enhance_video_sync) or other tools like get_task_status.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

enhance_video_syncA

同步增强视频(阻塞等待完成)

支持两种上传方式:

  1. URL 上传:提供视频 URL

  2. 本地上传:提供本地文件路径,MCP Server 自动上传到 TOS 对象存储

参数说明:

  • video_source: 视频 URL 或本地文件路径

  • type: "url" 或 "local"

  • resolution: 目标分辨率

  • poll_interval: 轮询间隔(秒)

  • timeout: 超时时间(秒)

ParametersJSON Schema
NameRequiredDescriptionDefault
video_sourceYes视频URL地址或本地文件路径(URL必须公网可访问,不支持需要登录或签名的链接)
typeNo上传类型:url=网络视频,local=本地文件url
resolutionNo目标分辨率,默认720p720p
poll_intervalNo轮询间隔(秒),默认5
timeoutNo超时时间(秒),默认600

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description bears full responsibility for behavioral disclosure. It explicitly states 'blocking wait for completion', explains the automatic upload of local files to TOS storage, and mentions polling parameters, giving good transparency. It does not mention side effects, but given the nature, none are expected.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with bullet points and clear categorization of upload methods and parameters. Every sentence serves a purpose, and there is no redundancy or unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the blocking nature, upload methods, and all parameters thoroughly. It lacks an explicit description of the return value, but given the synchronous nature, it likely returns the enhanced video. Overall, it is fairly complete for a tool without an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds context beyond the schema, such as the automatic upload process for local files and that URLs must be publicly accessible. This extra information enhances understanding of parameter usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: to enhance video synchronously, with a blocking wait. It details two upload methods (URL and local), distinguishing it from sibling tools that handle different operations like task creation or status checking.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the two upload methods and the blocking nature, implicitly indicating when to use the tool. However, it does not explicitly contrast with siblings like create_task (likely async) or provide when-not-to-use guidance, making usage guidelines less explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_task_statusA

查询视频增强任务状态

ParametersJSON Schema
NameRequiredDescriptionDefault
task_idYes任务ID

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, and the description only states the purpose without disclosing any behavioral traits such as polling requirements, rate limits, or expected response behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence with no wasted words; efficient and to the point.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple status query with one parameter, the description is mostly complete but could benefit from mentioning possible return statuses or output format since no output schema is provided.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the description adds no additional meaning beyond what is already in the input schema, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'query' and the resource 'video enhancement task status', distinguishing from sibling tools 'create_task' and 'enhance_video_sync' which have different actions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives, but the context of sibling tools implies it is for checking status after creation or enhancement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sam3_predictA

Analyze an image using the SAM3 segmentation API to generate inference results (masks, boxes, scores). The image can be provided in one of three ways:

  1. imagePath: Absolute path of a local image file (e.g. C:\Users\xxx\photo.png). Use this when the user provides a local file path.

  2. imageUrl: Publicly accessible URL of the image (e.g. https://example.com/photo.jpg). Use this when the user provides a web link.

  3. imageBase64: Base64-encoded image data. Use this when the user uploads or drags-and-drops an image as an attachment and no local path is available. In this case, encode the image content as base64 and pass it via this parameter. If the user mentions an uploaded image but does not provide a path, URL, or base64 data, ask the user for the local absolute path. Prompt must be in English. If the user provides Chinese or other non-English text, translate it to English before calling this tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
imagePathNoAbsolute path of a local image file (e.g. C:\\Users\\xxx\\photo.png)
imageUrlNoPublicly accessible URL of the image to process
imageBase64NoBase64-encoded image data. Use this when the image is provided as an attachment without a local path
promptYesText prompt for mask generation. Must be in English. If the user provides Chinese or other non-English text, translate it to English before calling this tool

TDQS

A4.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It discloses that it calls an external API and generates masks, boxes, scores. However, it lacks details on potential side effects, authentication, error handling, or rate limits. It does not contradict annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with bullet points, front-loading the main purpose. Every sentence serves a purpose, explaining input methods and prompt requirements without redundancy. It is concise yet comprehensive.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers what the tool does (segmentation analysis), how to provide input (three methods), and what outputs are generated (masks, boxes, scores). Even without an output schema, it gives sufficient information for an agent to use the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema coverage is 100%, the description adds significant value by explaining usage contexts for each image parameter and specifying that the prompt must be in English, requiring translation if needed. This goes beyond schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'Analyze an image using the SAM3 segmentation API to generate inference results (masks, boxes, scores).' This specifies the verb (analyze), resource (image via SAM3 API), and output, effectively distinguishing it from siblings like create_task.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use each image input method (imagePath, imageUrl, imageBase64) and includes instructions for handling non-English prompts. However, it does not explicitly mention when not to use this tool or compare it to alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedcreate_task
    • First observedenhance_video_sync
    • First observedget_task_status
    • First observedsam3_predict

TDQS

B3.4/5.0

Scored across 4 tools

Disambiguation2/5

The first three tools are about video enhancement tasks with overlapping functionality (create_task and enhance_video_sync both appear to initiate enhancement), and the fourth tool (sam3_predict) is for image segmentation, a completely different domain. The descriptions are not clear enough to distinguish which tool to use for a given task, causing confusion.

Naming Consistency3/5

Tool names partially follow a verb_noun pattern (create_task, get_task_status), but 'enhance_video_sync' is awkward and 'sam3_predict' mixes model name with verb, introducing inconsistency.

Tool Count4/5

With 4 tools, the count is reasonable for a focused server, but the server actually combines two unrelated capabilities (video enhancement and image segmentation), making the scope unclear but the number itself is not extreme.

Completeness2/5

For video enhancement, there are create, sync enhance, and status query, but missing cancel, list, or delete operations. For image segmentation, only a single predict tool exists. The surface is incomplete for both domains.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers