multivon-mcp
OfficialRelated Servers
Alternatives to multivon-mcp
No user-submitted related servers found.
Related Servers
- AlicenseNot gradedqualityCmaintenanceMCP server that enables AI coding agents to communicate, share state, and coordinate work in real time via MCP tools or REST API.473 npm8MIT
- AlicenseNot gradedqualityBmaintenanceA local MCP server that provides tools for AI coding agents to execute, evaluate, and diff regression tests for LLM agents.3 npmMIT
- AlicenseNot gradedqualityBmaintenanceAn MCP server that provides AI coding agents with AST-accurate, context-budget-aware codebase querying, safety gates, and team policy integration via structured tools and a local plugin layer.104 npm4MIT
- AlicenseNot gradedqualityDmaintenanceAn MCP server that gives AI coding agents structured access to a project's architecture, rules, modules, and technical decisions.MIT
- FlicenseBqualityCmaintenanceMCP server that gives AI coding assistants persistent memory, structural code graph analysis, and safe multi-agent coordination, enabling them to answer architectural questions, track decisions across sessions, and coordinate safely in multi-agent workflows.394-
- AlicenseAqualityAmaintenanceMCP server that enables a coordinator AI agent to spawn, control, and supervise local coding agents with interactive gating for high-risk operations.108 npmMIT
TDQS
Scored across 23 tools
Most eval tools target distinct metrics, but several are easy to conflate: eval_faithfulness/eval_hallucination are inverse measures of the same construct, eval_vqa_faithfulness/eval_document_grounding both do vision-grounded checking, and eval_relevance/eval_answer_accuracy both score response quality. The detailed descriptions help, but the sheer number of similar score/pass/reason evaluators still creates real selection ambiguity.
All tools use snake_case and nearly all share the eval_ prefix followed by a metric noun (eval_toxicity, eval_context_precision), with a few verb-style exceptions (eval_discover, eval_ingest_trace, eval_generate_cases). The two pdfhell_* tools form a coherent sub-namespace rather than a violation, so overall naming is consistent but not perfectly uniform.
With 23 tools, this is on the heavy end for an MCP server and an agent must navigate a large surface. The breadth is defensible for a full LLM evaluation platform covering text, vision, RAG, safety, and PDF benchmarks, but some tools are closely related and could plausibly be consolidated.
The set covers a full eval lifecycle: case generation, trace ingestion, diverse text/vision/RAG/safety evaluators, report comparison, acceptance policies, and audit packaging. Minor gaps remain—such as no dedicated summarization or code-quality evaluator and no explicit tool for assembling arbitrary eval results into a saved report—but generic G-Eval and custom-rubric tools close most holes.