Skip to main content
Glama

media-mcp (Node.js)

Chinesisch | Englisch

npm version Node.js >=18 License: MIT

Ein auf dem MCP-Protokoll basierender Dienst für Videoverbesserung und Bildsegmentierung, der als MCP Client-Server mit einem Backend-HTTP-Server interagiert.

Funktionen

Stellt die folgenden MCP-Tools bereit:

  • create_task - Erstellt eine Videoverbesserungsaufgabe (unterstützt URL oder lokalen Datei-Upload)

  • get_task_status - Fragt den Aufgabenstatus ab

  • enhance_video_sync - Synchrones Verbessern von Videos (blockiert bis zum Abschluss)

  • sam3_predict - SAM3-Bildsegmentierung (unterstützt lokale Pfade, URLs oder Base64-Bilder)

Related MCP server: Grok Imagine Video MCP Server

Voraussetzungen

  • Node.js >= 18 (Überprüfung: node --version)

  • API-Key (zur Authentifizierung, bitte beim Dienstanbieter anfordern)

Einfache Installation (Empfohlen)

Wenn dein KI-Agent einen definierten MCP-Konfigurationspfad hat, kopiere einfach den folgenden Satz und sende ihn an die KI:

帮我安装 npm 包 @avclabs.ai/media-mcp 作为 MCP server。我的 API Key 是:sk-xxxxxxxx。

Die KI erledigt automatisch:

  1. Erkennung deines verwendeten MCP-Clients

  2. Auffinden des Konfigurationspfads

  3. Schreiben der korrekten Konfiguration

  4. Aufforderung zum Neustart des Clients

Manuelle Installation

Keine Installation erforderlich, einfach direkt in der MCP-Client-Konfiguration mit npx ausführen.

1. Claude Code (CLI)

Führe in Claude Code Folgendes aus:

/mcp

Suche in der Ausgabe den Konfigurationspfad unter "User MCPs" und bearbeite diese Datei.

Typische Pfade (falls /mcp nicht verfügbar ist):

  • Windows: %USERPROFILE%\.claude.json

  • macOS: ~/.claude.json

  • Linux: ~/.claude.json

  • Alt/Alternativ: ~/.claude/mcp.json

Füge den folgenden Inhalt ein (ersetze your-api-key durch deinen tatsächlichen API-Key):

{
  "mcpServers": {
    "video-enhancement": {
      "command": "npx",
      "args": ["-y", "@avclabs.ai/media-mcp@latest"],
      "env": {
        "API_KEY": "your-api-key"
      }
    }
  }
}

Nach dem Speichern /mcp ausführen, um die erfolgreiche Ladung zu überprüfen.

2. Cursor

Gehe zu Settings > Tools & MCPs > Add New MCP Server:

  • Name: video-enhancement

  • Type: command

  • Command:

    env HTTP_API_KEY=your-api-key npx -y @avclabs.ai/media-mcp@latest

Oder bearbeite ~/.cursor/mcp.json:

{
  "mcpServers": {
    "video-enhancement": {
      "command": "npx",
      "args": ["-y", "@avclabs.ai/media-mcp@latest"],
      "env": {
        "API_KEY": "your-api-key"
      }
    }
  }
}

Installation überprüfen

Starte den Client neu und bestätige, ob die Tools erfolgreich geladen wurden:

  1. Oder frage die KI direkt: "Welche Tools stehen dir zur Verfügung?"

  2. Du solltest sehen: create_task, get_task_status, enhance_video_sync, sam3_predict

Konfigurationsoptionen

Variablenname

Erforderlich

Standardwert

Beschreibung

API_KEY

Ja

-

API-Authentifizierungsschlüssel (wird für Videoverbesserung und SAM3 verwendet)

HTTP_API_BASE_URL

Nein

https://mcp.avc.ai/enhance

Basis-URL der Videoverbesserungs-API

SAM3_API_BASE_URL

Nein

https://mcp.avc.ai/sam

Basis-URL des SAM3-Dienstes

SAM3_POLL_INTERVAL

Nein

2000

Abfrageintervall (in Millisekunden)

SAM3_POLL_MAX_ATTEMPTS

Nein

60

Maximale Anzahl der Abfragen

Benutzerdefinierte Service-URLs

{
  "env": {
    "HTTP_API_BASE_URL": "https://your-endpoint.com",
    "API_KEY": "your-api-key",
    "SAM3_API_BASE_URL": "http://localhost:8001"
  }
}

Oder über Befehlszeilenargumente:

npx -y @avclabs.ai/media-mcp@latest --base-url https://your-endpoint.com --api-key your-api-key --sam3-base-url http://localhost:8001

Anwendungsbeispiele

Nach der Konfiguration kannst du die KI in natürlicher Sprache anweisen:

"Hilf mir, dieses Video auf 1080p zu verbessern: https://example.com/video.mp4"

"Verbessere das Video video.mp4 auf meinem Desktop auf 2k-Qualität"

Die KI ruft automatisch das entsprechende Tool auf, um die Aufgabe zu erledigen.

"Analysiere dieses Bild und finde alle Objekte darin: C:\Users\xxx\photo.png"

"Segmentiere dieses Bild mit SAM3, der Prompt ist 'find all cars'"

Verfügbare Tools

create_task

Erstellt eine Videoverbesserungsaufgabe (asynchron).

Parameter

Typ

Erforderlich

Standardwert

Beschreibung

video_source

string

Ja

-

Video-URL oder lokaler Dateipfad (URL muss öffentlich zugänglich sein, Links, die Login oder Signaturen erfordern, werden nicht unterstützt)

type

string

Nein

url

url oder local

resolution

string

Nein

720p

480p, 540p, 720p, 1080p, 2k

Rückgabewert:

{
  "success": true,
  "task_id": "xxx",
  "status": "wait"
}

get_task_status

Fragt den Aufgabenstatus ab.

Parameter

Typ

Erforderlich

task_id

string

Ja

Rückgabewert:

{
  "success": true,
  "task_id": "xxx",
  "status": "completed",
  "progress": 100,
  "video_url": "https://..."
}

enhance_video_sync

Synchrones Verbessern von Videos (blockiert bis zum Abschluss).

Parameter

Typ

Erforderlich

Standardwert

Beschreibung

video_source

string

Ja

-

Video-URL oder lokaler Dateipfad (URL muss öffentlich zugänglich sein)

type

string

Nein

url

url oder local

resolution

string

Nein

720p

Zielauflösung

poll_interval

number

Nein

5

Abfrageintervall (Sekunden)

timeout

number

Nein

600

Timeout (Sekunden)

sam3_predict

Analysiert ein Bild mit der SAM3-Segmentierungs-API und generiert Inferenz-Ergebnisse (Masken, Boxen, Scores).

Parameter:

Bildeingabe (eine von drei Optionen muss bereitgestellt werden):

  • imagePath (string): Absoluter Pfad zum lokalen Bild. Unterstützt gängige Formate (z. B. PNG, JPG, JPEG).

    • Beispiel: "C:\\Users\\xxx\\photo.png", "/home/user/images/cat.jpg"

    • Anwendungsfall: Der Benutzer hat explizit einen lokalen Dateipfad angegeben.

  • imageUrl (string): Öffentlich zugängliche Bild-URL.

    • Beispiel: "https://example.com/photo.jpg"

    • Anwendungsfall: Das Bild ist online und der Benutzer hat einen Link bereitgestellt.

    • Hinweis: Die URL muss öffentlich zugänglich sein; Links, die Login oder Signaturen erfordern, werden nicht unterstützt.

  • imageBase64 (string): Base64-kodierte Bilddaten.

    • Beispiel: "iVBORw0KGgoAAAANSUhEUgAA..."

    • Anwendungsfall: Der Benutzer hat ein Bild angehängt, das der Agent in Base64 kodiert hat.

    • Hinweis: Base64-Daten großer Bilder können sehr umfangreich sein, was die Übertragungszeit verlängern kann.

Weitere Parameter:

  • prompt (string, erforderlich): Englischer Text-Prompt zur Angabe der zu segmentierenden Zielobjekte im Bild. Zum Beispiel "person", "car", "a cat sitting on a sofa". Da das SAM3-Modell nur englische Prompts akzeptiert, wird empfohlen, englische Beschreibungen zu verwenden. Wenn der Benutzer chinesischen oder anderen nicht-englischen Text bereitstellt, übersetzt der Agent diesen automatisch ins Englische.

Rückgabe:

Nach Abschluss der Inferenz wird direkt ein JSON-String zurückgegeben. Dieses JSON enthält die folgenden drei Felder:

  • masks: Zweidimensionales Array. Jedes Element ist eine binäre Maske (Werte 0 oder 1) in der Größe des Eingabebildes, die die pixelgenaue Position des erkannten Objekts markiert. Die i-te Maske entspricht der i-ten erkannten Objektinstanz.

  • boxes: Zweidimensionales Array. Jedes Element ist ein Begrenzungsrahmen im Format [x1, y1, x2, y2], der den rechteckigen Bereich des erkannten Objekts darstellt. x1, y1 sind die Koordinaten der oberen linken Ecke, x2, y2 die der unteren rechten Ecke.

    Koordinatensystem: Ursprung (0, 0) in der oberen linken Ecke des Bildes, x-Achse wächst nach rechts, y-Achse wächst nach unten, Einheit in Pixeln. Zum Beispiel bedeutet [120, 80, 300, 450], dass der Objektbereich 120px vom linken Rand und 80px vom oberen Rand beginnt und bei 300px vom linken Rand und 450px vom oberen Rand endet. Die Breite beträgt x2 - x1 = 180px, die Höhe y2 - y1 = 370px.

  • scores: Eindimensionales Array. Jedes Element ist der Konfidenzwert des entsprechenden Erkennungsergebnisses im Bereich von 0 bis 1. Ein höherer Wert bedeutet eine höhere Sicherheit des Modells.

Beispiel für den JSON-Inhalt des Ergebnisses:

{
  "masks": [
    [[0, 0, 1, ...], [0, 1, 1, ...], ...],
    [[0, 0, 0, ...], [0, 0, 1, ...], ...]
  ],
  "boxes": [
    [120, 80, 300, 450],
    [400, 200, 600, 500]
  ],
  "scores": [0.95, 0.87]
}

Häufig gestellte Fragen

Fehlermeldung "Datei nicht gefunden" nach dem Anhängen per Drag & Drop?

Dies ist eine bekannte Einschränkung von stdio MCP. Wenn Dateien über die Agent-Oberfläche per Drag & Drop hochgeladen werden, wird der Dateipfad oft nicht automatisch an den MCP-Server übermittelt.

Lösung:

  1. Pfad zusätzlich angeben (Empfohlen): Nach dem Hochladen des Bildes den lokalen absoluten Pfad im Text angeben:

    "Bitte verarbeite dieses Bild D:\photos\cat.jpg und finde die Katze darin"

  2. Auf automatische Kodierung warten: Claude kodiert das Bild möglicherweise automatisch als Base64. Wenn dies gelingt, ist keine weitere Aktion erforderlich.

  3. Auf Pfad-Anfrage antworten: Wenn Claude nach dem Bildpfad fragt, antworte direkt mit dem lokalen absoluten Pfad.

Gibt es eine Priorität bei den drei Eingabemethoden?

Es gibt keine strikte Priorität. Claude wählt automatisch die am besten geeignete Methode basierend auf dem Kontext:

  • Du gibst einen lokalen Pfad an → imagePath wird verwendet

  • Du gibst einen Weblink an → imageUrl wird verwendet

  • Du lädst einen Anhang hoch ohne Pfad → imageBase64 wird versucht

Welche Bildformate werden unterstützt?

Unterstützt werden gängige Formate: PNG, JPG, JPEG, BMP, WebP usw. Es wird empfohlen, PNG oder JPG zu bevorzugen.

Was tun, wenn der Download des URL-Bildes fehlschlägt?

Stelle sicher, dass die URL öffentlich zugänglich ist und keinen Login, Cookies oder Signaturen erfordert. Wenn sich das Bild auf einem Dienst mit Authentifizierung befindet (z. B. privater S3-Bucket), lade es bitte zuerst lokal herunter und verwende imagePath.

Was tun, wenn das Base64-Bild zu groß ist?

Wenn das Bild sehr groß ist (z. B. 4K-Auflösung), werden die Base64-Daten sehr umfangreich, was die Übertragung verlangsamen kann. Empfehlung:

  1. Verwende stattdessen imagePath

  2. Oder komprimiere das Bild vor der Kodierung

Hinweise zum Datei-Upload

Wenn type auf "local" gesetzt ist, führt der MCP-Server folgende Schritte aus:

  1. Liest die lokale Datei

  2. Lädt sie über eine vorab signierte URL direkt in den TOS-Objektspeicher hoch

  3. Maximale Dateigröße: 100 MB

Fehlerbehebung

"command not found: npx"

Installiere Node.js >= 18: https://nodejs.org/

"Fehler: --api-key muss bereitgestellt oder API_KEY gesetzt werden"

Der API-Key fehlt, bitte überprüfe env.API_KEY in der Konfiguration.

MCP-Server wird im Client rot/fehlerhaft angezeigt

Logdateien prüfen:

  • Claude Desktop macOS: ~/Library/Logs/Claude/mcp*.log

  • Claude Desktop Windows: %APPDATA%\Claude\logs\mcp*.log

  • Cursor: Output-Panel > MCP

"TOS-Upload fehlgeschlagen"

Dies liegt normalerweise an einer nicht übereinstimmenden Signatur. Bitte stelle sicher, dass HTTP_API_BASE_URL und HTTP_API_KEY korrekt und gültig sind.

Globale Installation (Optional)

Wenn du nicht jedes Mal npx verwenden möchtest:

npm install -g @avclabs.ai/media-mcp

Verwende dann in der Konfiguration "command": "media-mcp" zusammen mit "args": ["--api-key", "your-api-key"].

Lizenz

MIT-Lizenz - siehe LICENSE-Datei für Details

Available Tools

4 tools
create_taskB

创建视频增强任务(异步)

支持两种上传方式:

  1. URL 上传:提供视频 URL

  2. 本地上传:提供本地文件路径,MCP Server 自动上传到 TOS 对象存储

参数说明:

  • video_source: 视频 URL 或本地文件路径

  • type: "url" 或 "local"

  • resolution: 目标分辨率

ParametersJSON Schema
NameRequiredDescriptionDefault
video_sourceYes视频URL地址或本地文件路径(URL必须公网可访问,不支持需要登录或签名的链接)
typeNo上传类型:url=网络视频,local=本地文件url
resolutionNo目标分辨率,默认720p720p

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so the description must carry full burden. It notes async behavior and TOS upload but omits side effects, permissions, failure modes, or rate limits. The description only partially discloses behavioral traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with three bullet points and front-loaded purpose. Every sentence earns its place, but structure could be slightly improved with clearer differentiation from siblings.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers input parameters well but lacks output schema explanation (e.g., task ID or status). With no annotations and multiple siblings, more context on post-creation steps would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema has 100% description coverage, so baseline is 3. The description groups parameters and explains the two upload modes, but does not add new information beyond the schema's existing parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool creates an async video enhancement task and distinguishes between two upload methods (URL and local). It uses specific verbs and resources, and is not a tautology.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use each upload type (URL vs local) but does not explicitly guide when to use this async tool over its sync sibling (enhance_video_sync) or other tools like get_task_status.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

enhance_video_syncA

同步增强视频(阻塞等待完成)

支持两种上传方式:

  1. URL 上传:提供视频 URL

  2. 本地上传:提供本地文件路径,MCP Server 自动上传到 TOS 对象存储

参数说明:

  • video_source: 视频 URL 或本地文件路径

  • type: "url" 或 "local"

  • resolution: 目标分辨率

  • poll_interval: 轮询间隔(秒)

  • timeout: 超时时间(秒)

ParametersJSON Schema
NameRequiredDescriptionDefault
video_sourceYes视频URL地址或本地文件路径(URL必须公网可访问,不支持需要登录或签名的链接)
typeNo上传类型:url=网络视频,local=本地文件url
resolutionNo目标分辨率,默认720p720p
poll_intervalNo轮询间隔(秒),默认5
timeoutNo超时时间(秒),默认600

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description bears full responsibility for behavioral disclosure. It explicitly states 'blocking wait for completion', explains the automatic upload of local files to TOS storage, and mentions polling parameters, giving good transparency. It does not mention side effects, but given the nature, none are expected.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with bullet points and clear categorization of upload methods and parameters. Every sentence serves a purpose, and there is no redundancy or unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the blocking nature, upload methods, and all parameters thoroughly. It lacks an explicit description of the return value, but given the synchronous nature, it likely returns the enhanced video. Overall, it is fairly complete for a tool without an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds context beyond the schema, such as the automatic upload process for local files and that URLs must be publicly accessible. This extra information enhances understanding of parameter usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: to enhance video synchronously, with a blocking wait. It details two upload methods (URL and local), distinguishing it from sibling tools that handle different operations like task creation or status checking.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the two upload methods and the blocking nature, implicitly indicating when to use the tool. However, it does not explicitly contrast with siblings like create_task (likely async) or provide when-not-to-use guidance, making usage guidelines less explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_task_statusA

查询视频增强任务状态

ParametersJSON Schema
NameRequiredDescriptionDefault
task_idYes任务ID

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, and the description only states the purpose without disclosing any behavioral traits such as polling requirements, rate limits, or expected response behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence with no wasted words; efficient and to the point.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple status query with one parameter, the description is mostly complete but could benefit from mentioning possible return statuses or output format since no output schema is provided.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the description adds no additional meaning beyond what is already in the input schema, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'query' and the resource 'video enhancement task status', distinguishing from sibling tools 'create_task' and 'enhance_video_sync' which have different actions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives, but the context of sibling tools implies it is for checking status after creation or enhancement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sam3_predictA

Analyze an image using the SAM3 segmentation API to generate inference results (masks, boxes, scores). The image can be provided in one of three ways:

  1. imagePath: Absolute path of a local image file (e.g. C:\Users\xxx\photo.png). Use this when the user provides a local file path.

  2. imageUrl: Publicly accessible URL of the image (e.g. https://example.com/photo.jpg). Use this when the user provides a web link.

  3. imageBase64: Base64-encoded image data. Use this when the user uploads or drags-and-drops an image as an attachment and no local path is available. In this case, encode the image content as base64 and pass it via this parameter. If the user mentions an uploaded image but does not provide a path, URL, or base64 data, ask the user for the local absolute path. Prompt must be in English. If the user provides Chinese or other non-English text, translate it to English before calling this tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
imagePathNoAbsolute path of a local image file (e.g. C:\\Users\\xxx\\photo.png)
imageUrlNoPublicly accessible URL of the image to process
imageBase64NoBase64-encoded image data. Use this when the image is provided as an attachment without a local path
promptYesText prompt for mask generation. Must be in English. If the user provides Chinese or other non-English text, translate it to English before calling this tool

TDQS

A4.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It discloses that it calls an external API and generates masks, boxes, scores. However, it lacks details on potential side effects, authentication, error handling, or rate limits. It does not contradict annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with bullet points, front-loading the main purpose. Every sentence serves a purpose, explaining input methods and prompt requirements without redundancy. It is concise yet comprehensive.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers what the tool does (segmentation analysis), how to provide input (three methods), and what outputs are generated (masks, boxes, scores). Even without an output schema, it gives sufficient information for an agent to use the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema coverage is 100%, the description adds significant value by explaining usage contexts for each image parameter and specifying that the prompt must be in English, requiring translation if needed. This goes beyond schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'Analyze an image using the SAM3 segmentation API to generate inference results (masks, boxes, scores).' This specifies the verb (analyze), resource (image via SAM3 API), and output, effectively distinguishing it from siblings like create_task.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use each image input method (imagePath, imageUrl, imageBase64) and includes instructions for handling non-English prompts. However, it does not explicitly mention when not to use this tool or compare it to alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedcreate_task
    • First observedenhance_video_sync
    • First observedget_task_status
    • First observedsam3_predict

TDQS

B3.4/5.0

Scored across 4 tools

Disambiguation2/5

The first three tools are about video enhancement tasks with overlapping functionality (create_task and enhance_video_sync both appear to initiate enhancement), and the fourth tool (sam3_predict) is for image segmentation, a completely different domain. The descriptions are not clear enough to distinguish which tool to use for a given task, causing confusion.

Naming Consistency3/5

Tool names partially follow a verb_noun pattern (create_task, get_task_status), but 'enhance_video_sync' is awkward and 'sam3_predict' mixes model name with verb, introducing inconsistency.

Tool Count4/5

With 4 tools, the count is reasonable for a focused server, but the server actually combines two unrelated capabilities (video enhancement and image segmentation), making the scope unclear but the number itself is not extreme.

Completeness2/5

For video enhancement, there are create, sync enhance, and status query, but missing cancel, list, or delete operations. For image segmentation, only a single predict tool exists. The surface is incomplete for both domains.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers