Skip to main content
Glama

media-mcp (Node.js)

Deployment update (2026-09-06): the SAM3 source and release scripts are now maintained in sam3-http-server, rooted at /opt/ai/sam3. API schema candidate head is 20260906_0007. SAM3 production bootstrap and nine-tool E2E remain pending. See the four-project rollout.

English | 中文

npm version Node.js >=18 License: MIT

A video enhancement, image enhancement/colorization/denoising, and image segmentation service based on the MCP protocol, acting as an MCP Client-Server to interact with backend HTTP Servers.

Release status (2026-09-06): npm latest is 0.2.1; that published artifact contains 5 tools (3 video + 2 SAM3) and its runtime metadata incorrectly reports 0.3.0. This repository is the unreleased 0.3.0 candidate, which fixes version consistency and adds 4 image tools plus an optional IMAGE_API_BASE_URL override. Image and video share the same public HTTP server and /enhance base URL; the candidate backend keeps their prepare concurrency, Redis queues, workers, and processing concurrency separate. The image routes now exist in candidate code but are not deployed or real-AI/TOS verified, so the image tools below are not a production npm capability until all 9 tools pass smoke tests and 0.3.0 is published.

Features

Provides the following MCP Tools:

Video Enhancement

  • create_task - Create a video enhancement task (supports URL or local file upload)

  • get_task_status - Query task status

  • enhance_video_sync - Synchronously enhance video (blocking wait, truncated at ~50s by default)

Image Enhancement (unreleased 0.3.0 candidate)

  • enhance_image_sync - Enhance image quality and optimize faces (supports URL or local file upload)

  • colorize_image_sync - Colorize black-and-white photos (supports URL or local file upload)

  • denoise_image_sync - Remove noise from images (supports URL or local file upload)

  • get_image_task_status - Query image task status (for polling after sync timeout)

Image Segmentation (SAM3)

  • sam3_predict - SAM3 image segmentation (supports local path, URL, or Base64 image)

  • get_sam3_task_status - Query SAM3 task status (for polling after sync timeout)

Related MCP server: Grok Imagine Video MCP Server

Prerequisites

  • Node.js >= 18 (check: node --version)

  • API Key (required for authentication)

If your AI Agent has a known MCP config path, just copy the line below and send it to your AI:

Install the npm package @avclabs.ai/media-mcp as an MCP server. My API Key is: sk-xxxxxxxx.

The AI will automatically:

  1. Detect your MCP client

  2. Find the config file path

  3. Write the correct configuration

  4. Prompt you to restart the client

Manual Install

No installation needed. Use npx directly in your MCP client config.

1. Claude Code (CLI)

Run in Claude Code:

/mcp

Check the output for the "User MCPs" section to find the config file path, then edit that file.

Common paths (if /mcp is unavailable):

  • Windows: %USERPROFILE%\.claude.json

  • macOS: ~/.claude.json

  • Linux: ~/.claude.json

  • Legacy/Alternative: ~/.claude/mcp.json

Paste this (replace your-api-key):

{
  "mcpServers": {
    "video-enhancement": {
      "command": "npx",
      "args": ["-y", "@avclabs.ai/media-mcp@latest"],
      "env": {
        "API_KEY": "your-api-key"
      }
    }
  }
}

Save and run /mcp to verify it's loaded.

2. Cursor

Go to Settings > Tools & MCPs > Add New MCP Server:

  • Name: video-enhancement

  • Type: command

  • Command:

    env API_KEY=your-api-key npx -y @avclabs.ai/media-mcp@latest

Or edit ~/.cursor/mcp.json:

{
  "mcpServers": {
    "video-enhancement": {
      "command": "npx",
      "args": ["-y", "@avclabs.ai/media-mcp@latest"],
      "env": {
        "API_KEY": "your-api-key"
      }
    }
  }
}

Verify Installation

After restarting your client, check if the tools are available:

  1. Or ask: "What tools do you have available?"

  2. With the current npm latest=0.2.1, you should see create_task, get_task_status, enhance_video_sync, sam3_predict, and get_sam3_task_status.

  3. After 0.3.0 is published, you should additionally see enhance_image_sync, colorize_image_sync, denoise_image_sync, and get_image_task_status.

Configuration Options

Variable

Required

Default

Description

API_KEY

Yes

-

API authentication key (shared by video, image, and SAM3 services)

HTTP_API_BASE_URL

No

https://mcp.avc.ai/enhance

Shared video and image HTTP service endpoint

IMAGE_API_BASE_URL

No

Same as HTTP_API_BASE_URL

Optional image endpoint override; omit for the shared production server

SAM3_API_BASE_URL

No

https://mcp.avc.ai/sam

SAM3 service endpoint

SAM3_POLL_INTERVAL

No

2000

SAM3 polling interval (milliseconds)

SAM3_POLL_MAX_ATTEMPTS

No

25

SAM3 maximum polling attempts

IMAGE_API_BASE_URL is implemented by the unreleased 0.3.0 candidate as an optional override. Production uses the shared /enhance service, so it should normally be omitted; the candidate then resolves image and video calls to the same base URL.

Custom Endpoint (0.3.0 candidate)

{
  "env": {
    "HTTP_API_BASE_URL": "https://your-media-endpoint.com",
    "API_KEY": "your-api-key",
    "SAM3_API_BASE_URL": "https://your-sam3-endpoint.com"
  }
}

Or via CLI args:

npx -y @avclabs.ai/media-mcp@0.3.0 --base-url https://your-media-endpoint.com --api-key your-api-key --sam3-base-url https://your-sam3-endpoint.com

Run this command only after 0.3.0 is published. The optional --image-base-url remains available only for deployments that intentionally split the public image endpoint; npm 0.2.1 does not support it.

This project provides both synchronous and asynchronous modes.

Because MCP Agents typically enforce a ~60-second timeout per tool call, tasks with longer processing times (video enhancement) are strongly recommended to use asynchronous mode:

Video Enhancement:

  1. Call create_task to create a task → immediately get task_id

  2. Wait a few seconds, then call get_task_status to query the status

  3. If status is processing, continue waiting and repeat step 2

  4. If status is completed, the task is done and the result contains video_url

  5. If status is failed, the task failed and the result contains error_message

Synchronous Mode (Simple Scenarios)

Video Enhancement:

  • Call enhance_video_sync → the server polls internally

  • Defaults to a maximum wait of 50 seconds

  • If completed within 50 seconds, returns the result directly

  • If not completed within 50 seconds, returns task_id and instructions for the Agent to switch to get_task_status

Image Segmentation (SAM3):

  • Call sam3_predict → the server polls internally

  • Defaults to a maximum wait of 50 seconds (25 attempts × 2-second polling interval)

  • If completed within 50 seconds, returns the segmentation result directly

  • If not completed within 50 seconds, returns a truncation notice indicating the task is still processing

Usage Examples

Once configured, ask your AI agent naturally:

"Enhance this video to 1080p: https://example.com/video.mp4"

"Improve the quality of /Users/me/Desktop/video.mp4 to 2k"

"Enhance this image: https://example.com/photo.jpg"

"Colorize this black-and-white photo: /Users/me/Desktop/old_photo.png"

"Remove noise from this image: C:\Users\xxx\noisy.jpg"

"Analyze this image and find all objects: C:\Users\xxx\photo.png"

"Use SAM3 to segment this image, prompt: 'find all cars'"

The agent will automatically choose sync or async tools based on task complexity.

Provided Tools

Video Enhancement

create_task

Create an asynchronous video enhancement task.

Recommended for most use cases. Ideal for longer videos (over 10 seconds) to avoid timeouts and blocking the connection.

Parameter

Type

Required

Default

Description

video_source

string

Yes

-

Video URL or local file path (URL must be publicly accessible, links requiring login or signatures are not supported)

type

string

No

url

url or local

resolution

string

No

720p

480p, 540p, 720p, 1080p, 2k

Returns:

{
  "success": true,
  "task_id": "xxx",
  "status": "processing"
}

get_task_status

Query video enhancement task status.

The returned status field can be: processing, completed, or failed. If status is processing, you need to wait a few seconds and call this tool again.

Parameter

Type

Required

task_id

string

Yes

Returns:

{
  "success": true,
  "task_id": "xxx",
  "status": "completed",
  "progress": 100,
  "video_url": "https://...",
  "message": "Task is still processing, please check again later"
}

The message field only appears when status is processing, prompting the Agent to continue waiting.

enhance_video_sync

Synchronously enhance video (blocks until completion).

Best for short videos (estimated processing time < 1 minute). If the task is not completed within 50 seconds, the tool returns early with a task_id, and you need to use get_task_status to continue querying.

Parameter

Type

Required

Default

Description

video_source

string

Yes

-

Video URL or local file path

type

string

No

url

url or local

resolution

string

No

720p

Target resolution

poll_interval

number

No

5

Poll interval (seconds)

timeout

number

No

50

Sync wait timeout (seconds), returns early when exceeded

Truncated return example (not completed within 50s):

{
  "success": true,
  "status": "processing",
  "task_id": "xxx",
  "message": "Task is still processing (waited 50 seconds). Please use get_task_status to continue polling.",
  "note": "The synchronous wait for this long-running task has been truncated. Switch to get_task_status polling."
}

Image Enhancement

Three image processing tools are provided, each targeting a specific use case:

Tool

Function

Use Case

enhance_image_sync

Image quality enhancement & face optimization

Blurry, low-resolution, or degraded photos

colorize_image_sync

Black-and-white photo colorization

Restoring old B&W photos with realistic colors

denoise_image_sync

Image noise removal

Noisy/grainy photos taken in low light

All three tools share the same parameters and behavior pattern. They are synchronous — the tool blocks until the image is processed or the timeout is reached.

Supported image formats: PNG, JPG, JPEG, BMP, WebP, etc.

Two upload methods:

  1. URL upload: provide a publicly accessible image URL (type: "url")

  2. Local upload: provide a local file path, the MCP Server auto-uploads to TOS object storage (type: "local", max file size: 100MB)

enhance_image_sync

Synchronously enhance an image to improve quality and optimize faces.

The tool internally creates a task and polls for the result. If processing completes within the timeout (default 50s), the result is returned directly. If not, the tool returns early with a task_id — use get_image_task_status to continue polling.

Parameter

Type

Required

Default

Description

image_source

string

Yes

-

Image URL or local file path (URL must be publicly accessible, links requiring login or signatures are not supported)

type

string

No

url

url or local

scale

number

No

2

Enhancement scale multiplier (e.g. 2 for 2x, 4 for 4x upscaling)

poll_interval

number

No

5

Poll interval in seconds

timeout

number

No

50

Sync wait timeout in seconds, returns early when exceeded

Normal completion return:

{
  "success": true,
  "task_id": "xxx",
  "status": "completed",
  "progress": 100,
  "image_url": "https://..."
}

Truncated return (not completed within 50s):

{
  "success": true,
  "status": "processing",
  "task_id": "xxx",
  "message": "Task is still processing (waited 50 seconds). Please use get_image_task_status to continue polling.",
  "note": "The synchronous wait for this long-running task has been truncated. Switch to get_image_task_status polling."
}

colorize_image_sync

Synchronously colorize a black-and-white photo with AI.

Best for old black-and-white photos. The AI will add realistic colors to the image. Supports the same parameters and return format as enhance_image_sync.

Parameter

Type

Required

Default

Description

image_source

string

Yes

-

Image URL or local file path (URL must be publicly accessible, links requiring login or signatures are not supported)

type

string

No

url

url or local

poll_interval

number

No

5

Poll interval in seconds

timeout

number

No

50

Sync wait timeout in seconds, returns early when exceeded

Returns: Same format as enhance_image_sync.

denoise_image_sync

Synchronously remove noise from an image.

Best for grainy/noisy photos taken in low-light conditions or with high ISO settings. Supports the same parameters and return format as enhance_image_sync.

Parameter

Type

Required

Default

Description

image_source

string

Yes

-

Image URL or local file path (URL must be publicly accessible, links requiring login or signatures are not supported)

type

string

No

url

url or local

poll_interval

number

No

5

Poll interval in seconds

timeout

number

No

50

Sync wait timeout in seconds, returns early when exceeded

Returns: Same format as enhance_image_sync.

get_image_task_status

Query image processing task status. Used to poll for results when a sync tool times out.

The returned status field can be: processing, completed, or failed. If status is processing, wait a few seconds and call this tool again.

Parameter

Type

Required

task_id

string

Yes

Returns:

{
  "success": true,
  "task_id": "xxx",
  "status": "completed",
  "progress": 100,
  "image_url": "https://...",
  "message": "Task is still processing, please check again later"
}

The message field only appears when status is processing, prompting the Agent to continue waiting.

  1. For most images: Call enhance_image_sync / colorize_image_sync / denoise_image_sync directly — the tool handles everything and returns the result

  2. If truncated: The tool returns a task_id, then use get_image_task_status to poll until status becomes completed or failed

  3. If failed: Check the error_message field for details

Image Segmentation (SAM3)

sam3_predict

Analyze an image using the SAM3 segmentation API to generate inference results (masks, boxes, scores).

Parameters:

Image input (choose one, must provide exactly one):

  • imagePath (string): Absolute path of a local image file. Supports common formats (PNG, JPG, JPEG).

    • Example: "C:\\Users\\xxx\\photo.png", "/home/user/images/cat.jpg"

    • Use when: The user explicitly provides a local file path

  • imageUrl (string): Publicly accessible URL of the image.

    • Example: "https://example.com/photo.jpg"

    • Use when: The image is already online and the user provides a link

    • Note: The URL must be publicly accessible. Links requiring login or signatures are not supported

  • imageBase64 (string): Base64-encoded image data.

    • Example: "iVBORw0KGgoAAAANSUhEUgAA..."

    • Use when: The user drags or uploads an image attachment, and the Agent encodes it as base64

    • Note: Large images will produce very large base64 strings, which may slow transmission

Other parameters:

  • prompt (string, required): English text prompt specifying the target object to segment. Since the SAM3 model only accepts English prompts, provide an English description. If the user provides Chinese or other non-English text, the Agent will automatically translate it before calling the tool.

Normal completion return:

After inference completes, returns a JSON string containing three fields:

  • masks: 2D array. Each element is a binary mask (values 0 or 1) with the same dimensions as the input image, marking the pixel-level location of detected objects. The i-th mask corresponds to the i-th detected object instance.

  • boxes: 2D array. Each element is a bounding box in [x1, y1, x2, y2] format, representing the rectangular region of the detected object. x1, y1 are the top-left coordinates; x2, y2 are the bottom-right coordinates.

    Coordinate system: The top-left corner of the image is the origin (0, 0). The x-axis increases to the right, and the y-axis increases downward, in pixels. For example, [120, 80, 300, 450] means the region starts 120px from the left edge and 80px from the top edge, extending to 300px from the left and 450px from the top. Width = x2 - x1 = 180px, Height = y2 - y1 = 370px.

  • scores: 1D array. Each element is a confidence score for the corresponding detection result, ranging from 0 to 1. Higher scores indicate greater model confidence.

Example result JSON:

{
  "masks": [
    [[0, 0, 1, ...], [0, 1, 1, ...], ...],
    [[0, 0, 0, ...], [0, 0, 1, ...], ...]
  ],
  "boxes": [
    [120, 80, 300, 450],
    [400, 200, 600, 500]
  ],
  "scores": [0.95, 0.87]
}

Truncated return example (not completed within 50s):

{
  "success": true,
  "status": "processing",
  "task_id": "xxx",
  "message": "Task is still processing (waited about 50 seconds). Please retry later or record this task_id for manual follow-up.",
  "note": "The synchronous wait for this long-running task has been truncated."
}

get_sam3_task_status

Query SAM3 segmentation task status. Used to poll for results when sam3_predict times out.

The returned status field can be: processing, completed, or failed. If status is processing, wait a few seconds and call this tool again.

Parameter

Type

Required

task_id

string

Yes

Completed return:

{
  "success": true,
  "task_id": "xxx",
  "status": "completed",
  "result_url": "https://..."
}

Processing return:

{
  "success": true,
  "task_id": "xxx",
  "status": "processing",
  "message": "Task is still processing, please check again later."
}

Failed return:

{
  "success": false,
  "task_id": "xxx",
  "status": "failed",
  "error": "Task failed"
}

FAQ

Agent reports timeout when calling tools?

This is the primary issue this project addresses. MCP Agents (such as Claude, Cursor) typically enforce a ~60-second timeout per tool call. If task processing exceeds this limit, the Agent will error and disconnect.

Solutions:

  1. Prefer asynchronous tools: For video enhancement and other time-consuming tasks, always use create_task + get_task_status. These tools return instantly on each call and will not trigger timeouts.

  2. Sync tool truncation mechanism: enhance_video_sync has an internal 50-second truncation limit. If the task is not completed within 50 seconds, the tool proactively returns a task_id and instructs the Agent to use get_task_status to follow up.

  3. SAM3 truncation mechanism: sam3_predict defaults to 25 polling attempts (~50 seconds). If the task is not completed, it returns a truncation notice indicating the task is still processing.

  4. Adjust SAM3 polling parameters (advanced): If you are confident that SAM3 tasks are usually fast (e.g., under 10 seconds), you can increase polling attempts via environment variable:

    SAM3_POLL_MAX_ATTEMPTS=60

    But ensure the total wait time does not exceed your Agent's timeout limit.

Drag-and-drop attachment says file not found?

This is a known limitation of stdio MCP. When dragging or uploading an attachment through the Agent interface, the file path is usually not automatically passed to the MCP Server.

Solutions:

  1. Provide the path simultaneously (recommended): After dragging the image, add the local absolute path in your message:

    "Please analyze this image D:\\photos\\cat.jpg and find the cat"

  2. Wait for auto-encoding: Claude may automatically encode the image as base64. If successful, no extra action is needed.

  3. Reply to path inquiry: If Claude asks for the image path, simply reply with the local absolute path.

Is there a priority among the three input methods?

There is no strict priority. Claude will automatically choose the most appropriate method based on conversation context:

  • You provided a local path → uses imagePath

  • You provided a web link → uses imageUrl

  • You dragged an attachment without a path → tries imageBase64

What image formats are supported?

Common formats: PNG, JPG, JPEG, BMP, WebP, etc. PNG or JPG is recommended.

What if URL image download fails?

Ensure the URL is publicly accessible, requiring no login, cookies, or signatures. If the image is on a service requiring authentication (e.g., private S3 Bucket, login-required image host), download it locally first and use imagePath.

What if the base64 image is too large?

If the image is very large (e.g., 4K resolution), the base64-encoded data will be very large and may slow transmission. Suggestions:

  1. Use imagePath instead

  2. Or compress the image before encoding

File Upload Notes

When type is "local":

  1. File is read locally by the MCP Server

  2. Uploaded directly to TOS object storage via pre-signed URL

  3. Max file size: 100MB

Troubleshooting

"command not found: npx"

Install Node.js >= 18: https://nodejs.org/

"Error: --api-key argument or API_KEY environment variable is required"

Your API Key is missing. Double-check the env.API_KEY in your config.

MCP Server shows red/error in client

Check logs:

  • Claude Desktop macOS: ~/Library/Logs/Claude/mcp*.log

  • Claude Desktop Windows: %APPDATA%\Claude\logs\mcp*.log

  • Cursor: Output panel > MCP

"TOS upload failed"

Usually a signature mismatch. Ensure your IMAGE_API_BASE_URL (or its HTTP_API_BASE_URL fallback) and API_KEY are correct and active.

Global Install (Alternative)

If you prefer not using npx every time:

npm install -g @avclabs.ai/media-mcp

Then use "command": "media-mcp" with "args": ["--api-key", "your-api-key"] in your config.

Development and Release

This package runs locally in the MCP client and is released through npm; it is not a remote Node daemon. The sibling media-mcp-api-http-server is live as the shared /enhance video/image/account owner, although its current production release only implements video routes; candidate code adds the image routes, durable prepare/dispatch outboxes, task-level credit idempotency, and separate image/video queues. The Portal is live and SAM3 remains an external production dependency. Real image AI/TOS E2E, production backend rollout, and SAM3 JSON health are still pending, so 0.3.0 must not be published yet. Before publishing, run:

npm ci
npm run release:verify

See the release guide for version synchronization, publish order, smoke tests, and the relationship with the portal/backend deployment.

License

MIT License - See LICENSE file for details

Available Tools

4 tools
create_taskB

创建视频增强任务(异步)

支持两种上传方式:

  1. URL 上传:提供视频 URL

  2. 本地上传:提供本地文件路径,MCP Server 自动上传到 TOS 对象存储

参数说明:

  • video_source: 视频 URL 或本地文件路径

  • type: "url" 或 "local"

  • resolution: 目标分辨率

ParametersJSON Schema
NameRequiredDescriptionDefault
video_sourceYes视频URL地址或本地文件路径(URL必须公网可访问,不支持需要登录或签名的链接)
typeNo上传类型:url=网络视频,local=本地文件url
resolutionNo目标分辨率,默认720p720p

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so the description must carry full burden. It notes async behavior and TOS upload but omits side effects, permissions, failure modes, or rate limits. The description only partially discloses behavioral traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with three bullet points and front-loaded purpose. Every sentence earns its place, but structure could be slightly improved with clearer differentiation from siblings.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers input parameters well but lacks output schema explanation (e.g., task ID or status). With no annotations and multiple siblings, more context on post-creation steps would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema has 100% description coverage, so baseline is 3. The description groups parameters and explains the two upload modes, but does not add new information beyond the schema's existing parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool creates an async video enhancement task and distinguishes between two upload methods (URL and local). It uses specific verbs and resources, and is not a tautology.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use each upload type (URL vs local) but does not explicitly guide when to use this async tool over its sync sibling (enhance_video_sync) or other tools like get_task_status.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

enhance_video_syncA

同步增强视频(阻塞等待完成)

支持两种上传方式:

  1. URL 上传:提供视频 URL

  2. 本地上传:提供本地文件路径,MCP Server 自动上传到 TOS 对象存储

参数说明:

  • video_source: 视频 URL 或本地文件路径

  • type: "url" 或 "local"

  • resolution: 目标分辨率

  • poll_interval: 轮询间隔(秒)

  • timeout: 超时时间(秒)

ParametersJSON Schema
NameRequiredDescriptionDefault
video_sourceYes视频URL地址或本地文件路径(URL必须公网可访问,不支持需要登录或签名的链接)
typeNo上传类型:url=网络视频,local=本地文件url
resolutionNo目标分辨率,默认720p720p
poll_intervalNo轮询间隔(秒),默认5
timeoutNo超时时间(秒),默认600

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description bears full responsibility for behavioral disclosure. It explicitly states 'blocking wait for completion', explains the automatic upload of local files to TOS storage, and mentions polling parameters, giving good transparency. It does not mention side effects, but given the nature, none are expected.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with bullet points and clear categorization of upload methods and parameters. Every sentence serves a purpose, and there is no redundancy or unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the blocking nature, upload methods, and all parameters thoroughly. It lacks an explicit description of the return value, but given the synchronous nature, it likely returns the enhanced video. Overall, it is fairly complete for a tool without an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds context beyond the schema, such as the automatic upload process for local files and that URLs must be publicly accessible. This extra information enhances understanding of parameter usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: to enhance video synchronously, with a blocking wait. It details two upload methods (URL and local), distinguishing it from sibling tools that handle different operations like task creation or status checking.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the two upload methods and the blocking nature, implicitly indicating when to use the tool. However, it does not explicitly contrast with siblings like create_task (likely async) or provide when-not-to-use guidance, making usage guidelines less explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_task_statusA

查询视频增强任务状态

ParametersJSON Schema
NameRequiredDescriptionDefault
task_idYes任务ID

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, and the description only states the purpose without disclosing any behavioral traits such as polling requirements, rate limits, or expected response behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence with no wasted words; efficient and to the point.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple status query with one parameter, the description is mostly complete but could benefit from mentioning possible return statuses or output format since no output schema is provided.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the description adds no additional meaning beyond what is already in the input schema, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'query' and the resource 'video enhancement task status', distinguishing from sibling tools 'create_task' and 'enhance_video_sync' which have different actions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives, but the context of sibling tools implies it is for checking status after creation or enhancement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sam3_predictA

Analyze an image using the SAM3 segmentation API to generate inference results (masks, boxes, scores). The image can be provided in one of three ways:

  1. imagePath: Absolute path of a local image file (e.g. C:\Users\xxx\photo.png). Use this when the user provides a local file path.

  2. imageUrl: Publicly accessible URL of the image (e.g. https://example.com/photo.jpg). Use this when the user provides a web link.

  3. imageBase64: Base64-encoded image data. Use this when the user uploads or drags-and-drops an image as an attachment and no local path is available. In this case, encode the image content as base64 and pass it via this parameter. If the user mentions an uploaded image but does not provide a path, URL, or base64 data, ask the user for the local absolute path. Prompt must be in English. If the user provides Chinese or other non-English text, translate it to English before calling this tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
imagePathNoAbsolute path of a local image file (e.g. C:\\Users\\xxx\\photo.png)
imageUrlNoPublicly accessible URL of the image to process
imageBase64NoBase64-encoded image data. Use this when the image is provided as an attachment without a local path
promptYesText prompt for mask generation. Must be in English. If the user provides Chinese or other non-English text, translate it to English before calling this tool

TDQS

A4.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It discloses that it calls an external API and generates masks, boxes, scores. However, it lacks details on potential side effects, authentication, error handling, or rate limits. It does not contradict annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with bullet points, front-loading the main purpose. Every sentence serves a purpose, explaining input methods and prompt requirements without redundancy. It is concise yet comprehensive.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers what the tool does (segmentation analysis), how to provide input (three methods), and what outputs are generated (masks, boxes, scores). Even without an output schema, it gives sufficient information for an agent to use the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema coverage is 100%, the description adds significant value by explaining usage contexts for each image parameter and specifying that the prompt must be in English, requiring translation if needed. This goes beyond schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'Analyze an image using the SAM3 segmentation API to generate inference results (masks, boxes, scores).' This specifies the verb (analyze), resource (image via SAM3 API), and output, effectively distinguishing it from siblings like create_task.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use each image input method (imagePath, imageUrl, imageBase64) and includes instructions for handling non-English prompts. However, it does not explicitly mention when not to use this tool or compare it to alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedcreate_task
    • First observedenhance_video_sync
    • First observedget_task_status
    • First observedsam3_predict

TDQS

B3.4/5.0

Scored across 4 tools

Disambiguation2/5

The first three tools are about video enhancement tasks with overlapping functionality (create_task and enhance_video_sync both appear to initiate enhancement), and the fourth tool (sam3_predict) is for image segmentation, a completely different domain. The descriptions are not clear enough to distinguish which tool to use for a given task, causing confusion.

Naming Consistency3/5

Tool names partially follow a verb_noun pattern (create_task, get_task_status), but 'enhance_video_sync' is awkward and 'sam3_predict' mixes model name with verb, introducing inconsistency.

Tool Count4/5

With 4 tools, the count is reasonable for a focused server, but the server actually combines two unrelated capabilities (video enhancement and image segmentation), making the scope unclear but the number itself is not extreme.

Completeness2/5

For video enhancement, there are create, sync enhance, and status query, but missing cancel, list, or delete operations. For image segmentation, only a single predict tool exists. The surface is incomplete for both domains.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers