mcp-tournament
Related Servers
Alternatives to mcp-tournament
No user-submitted related servers found.
Related Servers
- AlicenseNot gradedqualityCmaintenanceEnables benchmarking of local LLM models (performance and quality) and sharing results to a public leaderboard via MCP tools.16 npm8Apache 2.0
- AlicenseNot gradedqualityAmaintenanceEnables provider-agnostic LLM discovery, capability testing, stability screening, and AIPerf benchmarking through MCP, with persistent SQLite history.MIT
- FlicenseNot gradedqualityBmaintenanceEnables running deterministic multi-model prompt regression and golden dataset benchmark suites, exposing structured results and telemetry through the Model Context Protocol for integration with MCP-compliant clients.8-

Patronus MCP Serverofficial
AlicenseNot gradedqualityDmaintenanceEnables running LLM evaluations, experiments, and custom evaluators through a standardized MCP interface.16Apache 2.0- AlicenseCqualityCmaintenanceEnables running, listing, inspecting, and comparing standardized model benchmarks through MCP.426 npmISC
- AlicenseNot gradedqualityAmaintenanceEnables LLM evaluation and observability by uploading documents, building test sets, running RAG pipelines, and automatically scoring answers for groundedness, hallucination risk, retrieval quality, latency, and cost, with tools exposed to MCP-compatible clients.1MIT
TDQS
Scored across 3 tools
Each tool has a distinct purpose: quick test runs a minimal scenario, leaderboard reads cached scores, and evaluate runs a full evaluation. No overlap or ambiguity.
Naming is inconsistent: 'quick_test' uses underscore and adjective+noun, 'leaderboard' is a single noun without underscore, and 'evaluate' is a bare verb. No consistent pattern in structure or part of speech.
With 3 tools, the server is at the low end of the typical 3-15 range but still reasonable for a focused evaluation service. The scope is narrow enough that each tool earns its place.
The tools cover the core workflow: quick test, full evaluation, and reading results. Minor gaps exist, such as no tool for configuring judges or scenarios, but these are likely predefined or managed externally.