Skip to main content
Glama

xCOMET MCP Server

npm version CI MCP License: MIT

日本語版 README はこちら

⚠️ This is an unofficial community project, not affiliated with Unbabel.

Translation quality evaluation MCP Server powered by xCOMET (eXplainable COMET).

🎯 Overview

xCOMET MCP Server provides AI agents with the ability to evaluate machine translation quality. It integrates with the xCOMET model from Unbabel to provide:

  • Quality Scoring: Scores between 0-1 indicating translation quality

  • Error Detection: Identifies error spans with severity levels (minor/major/critical)

  • Batch Processing: Evaluate multiple translation pairs efficiently (optimized single model load)

  • GPU Support: Optional GPU acceleration for faster inference

graph LR
    A[AI Agent] --> B[Node.js MCP Server]
    B -- stdio JSON-RPC --> C[Python Worker]
    C --> D[xCOMET Model<br/>Persistent in Memory]
    D --> C
    C --> B
    B --> A

    style D fill:#9f9

Related MCP server: nativ-mcp

🔧 Prerequisites

Python Environment

  • Python 3.9 - 3.12 recommended (3.13+ is not yet supported by xCOMET dependencies)

xCOMET requires Python with several packages. We recommend using a virtual environment:

# If using uv (recommended - auto-downloads the correct Python version)
uv venv ~/.xcomet-venv --python 3.12
source ~/.xcomet-venv/bin/activate
uv pip install "unbabel-comet>=2.2.0"

# Or using standard venv (requires Python 3.9-3.12 already installed)
python3 -m venv ~/.xcomet-venv
source ~/.xcomet-venv/bin/activate  # Windows: ~/.xcomet-venv\Scripts\activate
pip install "unbabel-comet>=2.2.0"

Note (v0.5.0+): The Python worker now talks to Node.js over stdin/stdout (line-delimited JSON-RPC). FastAPI, uvicorn, and pydantic are no longer required — only unbabel-comet is.

Note: When using with Claude Desktop or other MCP hosts, set XCOMET_PYTHON_PATH to point to the venv Python (see Configuration).

Model Download

Important: XCOMET-XL and XCOMET-XXL are gated models on HuggingFace. You must:

  1. Create a HuggingFace account

  2. Visit Unbabel/XCOMET-XL and request access

  3. Login via CLI:

    source ~/.xcomet-venv/bin/activate
    huggingface-cli login

Unbabel/wmt22-comet-da does not require authentication (but requires reference translations).

After authentication, download the model (~14GB for XL, ~42GB for XXL):

source ~/.xcomet-venv/bin/activate
python -c "from comet import download_model; download_model('Unbabel/XCOMET-XL')"

Node.js

  • Node.js >= 22.0.0 (matches engines.node in package.json; CI runs on 22 and 24)

  • npm or yarn

📦 Installation

Note: If you just want to use xCOMET MCP Server, you do not need to clone this repository. Install the Python environment and model (see Prerequisites), then use npx (see Usage). The section below is for contributors and local development only.

Local Development

For contributors and local development:

# Clone the repository
git clone https://github.com/shuji-bonji/xcomet-mcp-server.git
cd xcomet-mcp-server

# Set up Python virtual environment and install dependencies
uv venv .venv --python 3.12    # or: python3 -m venv .venv
source .venv/bin/activate
pip install -r python/requirements.txt

# Install Node.js dependencies and build
npm install
npm run build

🚀 Usage

With Claude Desktop (npx)

Add to your Claude Desktop configuration (claude_desktop_config.json):

{
  "mcpServers": {
    "xcomet": {
      "command": "npx",
      "args": ["-y", "xcomet-mcp-server"],
      "env": {
        "XCOMET_PYTHON_PATH": "~/.xcomet-venv/bin/python3"
      }
    }
  }
}

Tip: If you installed Python packages system-wide or use pyenv, XCOMET_PYTHON_PATH may be omitted (auto-detection will find it). See Python Path Auto-Detection for details.

With Claude Code

claude mcp add xcomet --env XCOMET_PYTHON_PATH=~/.xcomet-venv/bin/python3 -- npx -y xcomet-mcp-server

Global Installation

If you prefer installing globally:

npm install -g xcomet-mcp-server

Then configure:

{
  "mcpServers": {
    "xcomet": {
      "command": "xcomet-mcp-server",
      "env": {
        "XCOMET_PYTHON_PATH": "~/.xcomet-venv/bin/python3"
      }
    }
  }
}

Local Development Build

If you cloned and built the repository locally (see Installation):

{
  "mcpServers": {
    "xcomet": {
      "command": "node",
      "args": ["/path/to/xcomet-mcp-server/dist/index.js"],
      "env": {
        "XCOMET_PYTHON_PATH": "~/.xcomet-venv/bin/python3"
      }
    }
  }
}

🛠️ Available Tools

xcomet_evaluate

Evaluate translation quality for a single source-translation pair.

Parameters:

Name

Type

Required

Description

source

string

Original source text

translation

string

Translated text to evaluate

reference

string

Reference translation

source_lang

string

Source language code (ISO 639-1)

target_lang

string

Target language code (ISO 639-1)

response_format

"json" | "markdown"

Output format (default: "json")

use_gpu

boolean

Use GPU for inference (default: false)

Example:

{
  "source": "The quick brown fox jumps over the lazy dog.",
  "translation": "素早い茶色のキツネが怠惰な犬を飛び越える。",
  "source_lang": "en",
  "target_lang": "ja",
  "use_gpu": true
}

Response:

{
  "score": 0.847,
  "errors": [],
  "summary": "Good quality (score: 0.847) with 0 error(s) detected."
}

xcomet_detect_errors

Focus on detecting and categorizing translation errors.

Parameters:

Name

Type

Required

Description

source

string

Original source text

translation

string

Translated text to analyze

reference

string

Reference translation

min_severity

"minor" | "major" | "critical"

Minimum severity (default: "minor")

response_format

"json" | "markdown"

Output format

use_gpu

boolean

Use GPU for inference (default: false)

xcomet_batch_evaluate

Evaluate multiple translation pairs in a single request.

Performance Note: With the persistent server architecture (v0.3.0+), the model stays loaded in memory. Batch evaluation processes all pairs efficiently without reloading the model.

Parameters:

Name

Type

Required

Description

pairs

array

Array of {source, translation, reference?} (max 500)

source_lang

string

Source language code

target_lang

string

Target language code

response_format

"json" | "markdown"

Output format

use_gpu

boolean

Use GPU for inference (default: false)

batch_size

number

Batch size 1-64 (default: 8). Larger = faster but uses more memory

Example:

{
  "pairs": [
    {"source": "Hello", "translation": "こんにちは"},
    {"source": "Goodbye", "translation": "さようなら"}
  ],
  "use_gpu": true,
  "batch_size": 16
}

🔗 Integration with Other MCP Servers

xCOMET MCP Server is designed to work alongside other MCP servers for complete translation workflows:

sequenceDiagram
    participant Agent as AI Agent
    participant DeepL as DeepL MCP Server
    participant xCOMET as xCOMET MCP Server
    
    Agent->>DeepL: Translate text
    DeepL-->>Agent: Translation result
    Agent->>xCOMET: Evaluate quality
    xCOMET-->>Agent: Score + Errors
    Agent->>Agent: Decide: Accept or retry?
  1. Translate using DeepL MCP Server (official)

  2. Evaluate using xCOMET MCP Server

  3. Iterate if quality is below threshold

Example: DeepL + xCOMET Integration

Configure both servers in Claude Desktop:

{
  "mcpServers": {
    "deepl": {
      "command": "npx",
      "args": ["-y", "@anthropic/deepl-mcp-server"],
      "env": {
        "DEEPL_API_KEY": "your-api-key"
      }
    },
    "xcomet": {
      "command": "npx",
      "args": ["-y", "xcomet-mcp-server"],
      "env": {
        "XCOMET_PYTHON_PATH": "~/.xcomet-venv/bin/python3"
      }
    }
  }
}

Then ask Claude:

"Translate this text to Japanese using DeepL, then evaluate the translation quality with xCOMET. If the score is below 0.8, suggest improvements."

⚙️ Configuration

Environment Variables

Variable

Default

Description

XCOMET_MODEL

Unbabel/XCOMET-XL

xCOMET model to use

XCOMET_PYTHON_PATH

(auto-detect)

Python executable path (see below)

XCOMET_PRELOAD

false

Pre-load model at startup (v0.3.1+)

XCOMET_DEBUG

false

Enable verbose debug logging (v0.3.1+)

XCOMET_NUM_WORKERS

1

DataLoader workers for model.predict() (v0.6.0+). Increase to better utilize idle CPU cores when running large batches, especially on GPU. Invalid values silently fall back to 1.

Model Selection

Choose the model based on your quality/performance needs:

Model

Parameters

Size

Memory

Reference

HF Auth

Quality

Use Case

Unbabel/XCOMET-XL

3.5B

~14GB

~8-10GB

Optional

✅ Required

⭐⭐⭐⭐

Recommended for most use cases

Unbabel/XCOMET-XXL

10.7B

~42GB

~20GB

Optional

✅ Required

⭐⭐⭐⭐⭐

Highest quality, requires more resources

Unbabel/wmt22-comet-da

580M

~2GB

~3GB

Required

Not required

⭐⭐⭐

Lightweight, faster loading

Important: XCOMET-XL and XCOMET-XXL are gated models on HuggingFace. Each model requires separate access approval. See Model Download for authentication setup.

Important: wmt22-comet-da requires a reference translation for evaluation. XCOMET models support referenceless evaluation.

Tip: If you experience memory issues or slow model loading, try Unbabel/wmt22-comet-da for faster performance with slightly lower accuracy (but remember to provide reference translations).

To use a different model, set the XCOMET_MODEL environment variable:

{
  "mcpServers": {
    "xcomet": {
      "command": "npx",
      "args": ["-y", "xcomet-mcp-server"],
      "env": {
        "XCOMET_MODEL": "Unbabel/XCOMET-XXL"
      }
    }
  }
}

Python Path Auto-Detection

The server automatically detects a Python environment with unbabel-comet installed:

  1. XCOMET_PYTHON_PATH environment variable (if set)

  2. pyenv versions (~/.pyenv/versions/*/bin/python3) - checks for comet module

  3. Homebrew Python (/opt/homebrew/bin/python3, /usr/local/bin/python3)

  4. Fallback: python3 command

This ensures the server works correctly even when the MCP host (e.g., Claude Desktop) uses a different Python than your terminal.

Example: Explicit Python path configuration

{
  "mcpServers": {
    "xcomet": {
      "command": "npx",
      "args": ["-y", "xcomet-mcp-server"],
      "env": {
        "XCOMET_PYTHON_PATH": "/Users/you/.pyenv/versions/3.11.0/bin/python3"
      }
    }
  }
}

⚡ Performance

Persistent Worker Architecture (v0.3.0+, stdio since v0.5.0)

The server uses a persistent Python worker process that keeps the xCOMET model loaded in memory. The Node.js MCP server talks to the worker over stdin/stdout using a line-delimited JSON-RPC protocol — no local HTTP listener, no port binding, no FastAPI.

Request

Time

Notes

First request

~25-90s

Model loading (varies by model size)

Subsequent requests

~500ms

Model already loaded

This provides a 177x speedup for consecutive evaluations compared to reloading the model each time.

Eager Loading (v0.3.1+)

Enable XCOMET_PRELOAD=true to pre-load the model at server startup:

{
  "mcpServers": {
    "xcomet": {
      "command": "npx",
      "args": ["-y", "xcomet-mcp-server"],
      "env": {
        "XCOMET_PRELOAD": "true"
      }
    }
  }
}

With preload enabled, all requests are fast (~500ms), including the first one.

graph LR
    A[MCP Request] --> B[Node.js Server]
    B -- stdio JSON-RPC --> C[Python Worker]
    C --> D[xCOMET Model<br/>in Memory]
    D --> C
    C --> B
    B --> A

    style D fill:#9f9

Batch Processing Optimization

The xcomet_batch_evaluate tool processes all pairs with a single model load:

Pairs

Estimated Time

10

~30-40 sec

50

~1-1.5 min

100

~2 min

GPU vs CPU Performance

Mode

100 Pairs (Estimated)

CPU (batch_size=8)

~2 min

GPU (batch_size=16)

~20-30 sec

Note: GPU requires CUDA-compatible hardware and PyTorch with CUDA support. If GPU is not available, set use_gpu: false (default).

Best Practices

1. Let the persistent server do its job

With v0.3.0+, the model stays in memory. Multiple xcomet_evaluate calls are now efficient:

✅ Fast: First call loads model, subsequent calls reuse it
   xcomet_evaluate(pair1)  # ~90s (model loads)
   xcomet_evaluate(pair2)  # ~500ms (model cached)
   xcomet_evaluate(pair3)  # ~500ms (model cached)

2. For many pairs, use batch evaluation

✅ Even faster: Batch all pairs in one call
   xcomet_batch_evaluate(allPairs)  # Optimal throughput

3. Memory considerations

  • XCOMET-XL requires ~8-10GB RAM

  • For large batches (500 pairs), ensure sufficient memory

  • If memory is limited, split into smaller batches (100-200 pairs)

Auto-Restart (v0.3.1+)

The server automatically recovers from failures:

  • Monitors health every 30 seconds

  • Restarts after 3 consecutive health check failures

  • Up to 3 restart attempts before giving up

📊 Quality Score Interpretation

Score Range

Quality

Recommendation

0.9 - 1.0

Excellent

Ready for use

0.7 - 0.9

Good

Minor review recommended

0.5 - 0.7

Fair

Post-editing needed

0.0 - 0.5

Poor

Re-translation recommended

🔍 Troubleshooting

Common Issues

"No module named 'comet'"

Cause: Python environment without unbabel-comet installed.

Solution:

# Check which Python is being used
python3 -c "import sys; print(sys.executable)"

# If using a virtual environment, make sure it's activated
source .venv/bin/activate
pip install -r python/requirements.txt

# For MCP hosts (e.g., Claude Desktop), specify the venv Python path
export XCOMET_PYTHON_PATH=~/.xcomet-venv/bin/python3

Model download fails or times out

Cause: Large model files (~14GB for XL) require stable internet connection. XCOMET models also require HuggingFace authentication (see Model Download).

Solution:

# Login to HuggingFace (required for XCOMET-XL/XXL)
huggingface-cli login

# Pre-download the model manually
python -c "from comet import download_model; download_model('Unbabel/XCOMET-XL')"

GPU not detected

Cause: PyTorch not installed with CUDA support.

Solution:

# Check CUDA availability
python -c "import torch; print(torch.cuda.is_available())"

# If False, reinstall PyTorch with CUDA
pip install torch --index-url https://download.pytorch.org/whl/cu118

Slow performance on Mac (MPS)

Cause: Mac MPS (Metal Performance Shaders) has compatibility issues with some operations.

Solution: The server automatically uses num_workers=1 for Mac MPS compatibility. For best performance on Mac, use CPU mode (use_gpu: false).

High memory usage or crashes

Cause: XCOMET-XL requires ~8-10GB RAM.

Solutions:

  1. Use the persistent server (v0.3.0+): Model loads once and stays in memory, avoiding repeated memory spikes

  2. Use a lighter model: Set XCOMET_MODEL=Unbabel/wmt22-comet-da for lower memory usage (~3GB)

  3. Reduce batch size: For large batches, process in smaller chunks (100-200 pairs)

  4. Close other applications: Free up RAM before running large evaluations

# Check available memory
free -h  # Linux
vm_stat | head -5  # macOS

VS Code or IDE crashes during evaluation

Cause: High memory usage from the xCOMET model (~8-10GB for XL).

Solution:

  • With v0.3.0+, the model loads once and stays in memory (no repeated loading)

  • If memory is still an issue, use a lighter model: XCOMET_MODEL=Unbabel/wmt22-comet-da

  • Close other memory-intensive applications before evaluation

Getting Help

If you encounter issues:

  1. Check the GitHub Issues

  2. Enable debug logging (check Claude Desktop's Developer Mode logs, or set XCOMET_DEBUG=true)

  3. Open a new issue with:

    • Your OS and Python version

    • The error message

    • Your configuration (without sensitive data)

🧪 Development

# Install dependencies
npm install

# Build TypeScript
npm run build

# Watch mode
npm run dev

# Run tests
npm test

# Test with MCP Inspector
npm run inspect

📋 Changelog

See CHANGELOG.md for version history and updates.

📝 License

MIT License - see LICENSE for details.

🙏 Acknowledgments

📚 References

Available Tools

3 tools
xcomet_batch_evaluateBatch Evaluate TranslationsA
Read-onlyIdempotent

Evaluate multiple translation pairs in a batch.

This tool processes multiple source-translation pairs and provides aggregate statistics along with individual results.

Args:

  • pairs (array): Array of translation pairs, each with:

    • source (string): Original source text

    • translation (string): Translated text

    • reference (string, optional): Reference translation

  • source_lang (string, optional): Source language code

  • target_lang (string, optional): Target language code

  • response_format ('json' | 'markdown'): Output format (default: 'json')

  • use_gpu (boolean, optional): Use GPU for inference if available (default: false)

  • batch_size (number, optional): Inference batch size, 1-64 (default: 8). Larger = faster but uses more memory.

Returns: { "average_score": number, "total_pairs": number, "results": [ { "index": number, "score": number, "error_count": number, "has_critical_errors": boolean } ], "summary": string }

Examples:

  • Evaluate entire translated document

  • Compare MT system quality across test set

  • Identify segments needing attention

ParametersJSON Schema
NameRequiredDescriptionDefault
pairsYesArray of translation pairs to evaluate
source_langNoSource language code
target_langNoTarget language code
response_formatNoOutput formatjson
use_gpuNoUse GPU for inference (faster if available). Default: false (CPU only)
batch_sizeNoBatch size for GPU processing (1-64). Larger = faster but uses more memory. Default: 8

Output Schema

ParametersJSON Schema
NameRequiredDescription
average_scoreYesAverage quality score across all pairs
total_pairsYesTotal number of evaluated pairs
resultsYesIndividual results for each pair
summaryYesOverall quality summary

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds value beyond annotations by detailing GPU usage, batch size limits, and memory implications; no contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured with sections, front-loaded purpose, but somewhat verbose; could trim repetitive parameter details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Fully covers all aspects given 6 params, output schema, and annotations; no gaps in usage or behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema already covers parameters well; description adds extra context like batch size range and examples, going beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the tool evaluates multiple translation pairs in a batch, distinguishing it from single-pair evaluate and error detection siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides examples and context but lacks explicit comparison to alternatives for when to choose batch vs single evaluate or detect-errors.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

xcomet_detect_errorsDetect Translation ErrorsA
Read-onlyIdempotent

Detect and categorize errors in a translation.

This tool focuses on error detection, providing detailed information about translation errors with their severity levels and positions.

Args:

  • source (string): Original source text

  • translation (string): Translated text to analyze

  • reference (string, optional): Reference translation

  • min_severity ('minor' | 'major' | 'critical'): Minimum severity to report (default: 'minor')

  • response_format ('json' | 'markdown'): Output format (default: 'json')

  • use_gpu (boolean, optional): Use GPU for inference if available (default: false)

Returns: { "total_errors": number, "errors_by_severity": { "minor": number, "major": number, "critical": number }, "errors": [ { "text": string, "start": number, "end": number, "severity": "minor" | "major" | "critical", "suggestion": string | null } ] }

Examples:

  • Find critical errors before publication

  • Identify areas needing post-editing

  • Quality gate for MT output

ParametersJSON Schema
NameRequiredDescriptionDefault
sourceYesOriginal source text
translationYesTranslated text to analyze
referenceNoOptional reference translation
min_severityNoMinimum severity level to report (minor, major, critical)minor
response_formatNoOutput formatjson
use_gpuNoUse GPU for inference (faster if available). Default: false (CPU only)

Output Schema

ParametersJSON Schema
NameRequiredDescription
total_errorsYesTotal number of errors detected
errors_by_severityYesError count by severity
errorsYesDetailed error list

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=true, which the description does not contradict. The description adds valuable behavioral context such as the return structure (errors with severity and positions) and GPU inference capability, going beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with sections (purpose, args, returns, examples) and front-loaded with the main action. It is appropriately sized for the complexity of the tool, with no wasted sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 6 parameters, 2 required, enums, and no output schema in the input schema, the description provides complete context: behavior, parameters, output structure, and examples. Everything an agent needs to select and invoke the tool correctly is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, so baseline is 3. The description adds meaning by listing parameters with defaults and enum values, and more importantly, provides a detailed output schema in the Returns section, which is not present in the input schema. This helps the agent understand what the tool returns.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool detects and categorizes translation errors, with a specific verb and resource. It distinguishes itself from sibling tools by focusing on error detection rather than batch evaluation or overall quality scoring.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage examples (e.g., 'Find critical errors before publication', 'Quality gate for MT output'), indicating when to use the tool. However, it does not explicitly state when not to use it or compare directly with sibling tools for alternative scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

xcomet_evaluateEvaluate Translation QualityA
Read-onlyIdempotent

Evaluate the quality of a translation using xCOMET model.

This tool analyzes a source text and its translation, providing:

  • A quality score between 0 and 1 (higher is better)

  • Detected error spans with severity levels (minor/major/critical)

  • A human-readable quality summary

Args:

  • source (string): Original source text to translate from

  • translation (string): Translated text to evaluate

  • reference (string, optional): Reference translation for comparison

  • source_lang (string, optional): Source language code (ISO 639-1)

  • target_lang (string, optional): Target language code (ISO 639-1)

  • response_format ('json' | 'markdown'): Output format (default: 'json')

  • use_gpu (boolean, optional): Use GPU for inference if available (default: false)

Returns: For JSON format: { "score": number, // Quality score 0-1 "errors": [ // Detected errors { "text": string, "start": number, "end": number, "severity": "minor" | "major" | "critical" } ], "summary": string // Human-readable summary }

Examples:

  • Evaluate EN→JA translation quality

  • Check if MT output needs post-editing

  • Compare translation against reference

ParametersJSON Schema
NameRequiredDescriptionDefault
sourceYesOriginal source text
translationYesTranslated text to evaluate
referenceNoOptional reference translation for comparison
source_langNoSource language code (ISO 639-1, e.g., 'en', 'ja')
target_langNoTarget language code (ISO 639-1, e.g., 'en', 'ja')
response_formatNoOutput format: 'json' for structured data or 'markdown' for human-readablejson
use_gpuNoUse GPU for inference (faster if available). Default: false (CPU only)

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYesQuality score between 0 and 1
errorsYesDetected error spans
summaryYesHuman-readable quality summary

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, idempotentHint=true. The description adds that it uses xCOMET model, returns quality scores, error spans (with severity), and a summary. No contradictions; behavior is fully disclosed with no hidden side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with a lead sentence, bullet points for outputs, labeled Args list, Returns section with JSON format, and usage examples. It is slightly verbose but each sentence serves a purpose, making it efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 7 parameters (2 required), 100% schema coverage, presence of output schema, and sibling tools, the description is complete: describes all parameters, return structure, usage examples, and model name. No gaps for agent decision-making.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description adds value by providing more context for each parameter (e.g., 'source: Original source text to translate from') and showing the return format with example output. This goes beyond schema terseness.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool evaluates translation quality using xCOMET model. It lists specific outputs (score, errors, summary) and examples. It distinguishes from siblings (xcomet_batch_evaluate for batch, xcomet_detect_errors for error-only) by focusing on single pair evaluation with full output.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides usage examples (evaluate EN→JA, check post-editing need, compare against reference) which imply when to use. It does not explicitly state when not to use or mention alternatives, but sibling tools are known. Still, clear context for single evaluation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

A4.5/5.0
Disambiguation5/5

Each tool has a clearly distinct purpose: batch evaluation, error detection, and single pair evaluation. There is no overlap; an agent can easily select the appropriate tool based on whether it needs aggregate stats, detailed error spans, or a single quality score.

Naming Consistency5/5

All tool names follow the consistent pattern 'xcomet_<verb>' in snake_case, making it easy to predict tool functionality from the name. No mixing of conventions or irregular naming.

Tool Count4/5

With 3 tools, the server is on the lower end of typical scope but still covers the core evaluation workflow. It is not unreasonably sparse; adding a few more tools (e.g., model info, comparison) would improve completeness.

Completeness4/5

The tools cover single evaluation, batch evaluation, and detailed error detection, which together address the primary use cases for translation quality assessment. Minor gaps exist, such as no tool for listing models or configuring settings, but the core functionality is present.

Maintenance

ActivityMaintained
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables translation and text rephrasing using DeepL's API through a cloud-deployed MCP server. Provides translation between multiple languages, text rephrasing, and language detection capabilities with professional deployment features.
    MIT
  • A
    license
    C
    quality
    C
    maintenance
    An MCP server for processing XLIFF and TMX translation files, enabling parsing, validation, and manipulation of translation units in localization workflows.
    9
    2
    MIT
  • A
    license
    B
    quality
    A
    maintenance
    Spanish dialect localization MCP server and CLI. It translates and QA-checks content across 25 regional variants with register control, structure preservation, and adversarial quality gates.
    16
    4
    Apache 2.0

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/shuji-bonji/xcomet-mcp-server'

If you have feedback or need assistance with the MCP directory API, please join our Discord server