xcomet-mcp-server
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| PORT | No | HTTP server port (when TRANSPORT=http) | 3000 |
| TRANSPORT | No | Transport mode: stdio or http | stdio |
| XCOMET_DEBUG | No | Enable verbose debug logging (v0.3.1+) | false |
| XCOMET_MODEL | No | xCOMET model to use (e.g., Unbabel/XCOMET-XL, Unbabel/XCOMET-XXL, or Unbabel/wmt22-comet-da) | Unbabel/XCOMET-XL |
| XCOMET_PRELOAD | No | Pre-load model at startup (v0.3.1+). Enabling this makes all requests fast (~500ms), including the first one. | false |
| XCOMET_PYTHON_PATH | No | Python executable path. If not set, the server automatically detects a Python environment with unbabel-comet installed. |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| xcomet_evaluateA | Evaluate the quality of a translation using xCOMET model. This tool analyzes a source text and its translation, providing:
Args:
Returns: For JSON format: { "score": number, // Quality score 0-1 "errors": [ // Detected errors { "text": string, "start": number, "end": number, "severity": "minor" | "major" | "critical" } ], "summary": string // Human-readable summary } Examples:
|
| xcomet_detect_errorsA | Detect and categorize errors in a translation. This tool focuses on error detection, providing detailed information about translation errors with their severity levels and positions. Args:
Returns: { "total_errors": number, "errors_by_severity": { "minor": number, "major": number, "critical": number }, "errors": [ { "text": string, "start": number, "end": number, "severity": "minor" | "major" | "critical", "suggestion": string | null } ] } Examples:
|
| xcomet_batch_evaluateA | Evaluate multiple translation pairs in a batch. This tool processes multiple source-translation pairs and provides aggregate statistics along with individual results. Args:
Returns: { "average_score": number, "total_pairs": number, "results": [ { "index": number, "score": number, "error_count": number, "has_critical_errors": boolean } ], "summary": string } Examples:
|
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 3 tools
Each tool has a clearly distinct purpose: batch evaluation, error detection, and single pair evaluation. There is no overlap; an agent can easily select the appropriate tool based on whether it needs aggregate stats, detailed error spans, or a single quality score.
All tool names follow the consistent pattern 'xcomet_<verb>' in snake_case, making it easy to predict tool functionality from the name. No mixing of conventions or irregular naming.
With 3 tools, the server is on the lower end of typical scope but still covers the core evaluation workflow. It is not unreasonably sparse; adding a few more tools (e.g., model info, comparison) would improve completeness.
The tools cover single evaluation, batch evaluation, and detailed error detection, which together address the primary use cases for translation quality assessment. Minor gaps exist, such as no tool for listing models or configuring settings, but the core functionality is present.