A
licenseNot graded
qualityC
maintenanceAn MCP server that allows agents to test and compare LLM prompts across OpenAI and Anthropic models, supporting single tests, side-by-side comparisons, and multi-turn conversations.
MIT
これは、MCP を使用して vLLM をインタラクティブにベンチマークする方法の概念実証です。
私たちはベンチマークに新しいわけではありません。私たちのブログをお読みください。
これは、MCP の可能性を探る試みにすぎません。
リポジトリをクローンする
MCP サーバーに追加します:
{
"mcpServers": {
"mcp-vllm": {
"command": "uv",
"args": [
"run",
"/Path/TO/mcp-vllm-benchmarking-tool/server.py"
]
}
}
}次に、たとえば次のようにプロンプトできます。
Do a vllm benchmark for this endpoint: http://10.0.101.39:8888
benchmark the following model: deepseek-ai/DeepSeek-R1-Distill-Llama-8B
run the benchmark 3 times with each 32 num prompts, then compare the results, but ignore the first iteration as that is just a warmup.Related MCP server: LLM API Benchmark MCP Server
vllm のランダムな出力により、無効な JSON が検出された可能性があります。まだ詳しく調べていません。
This server cannot be deployed
MCP server providing access to the Scorecard API to evaluate and optimize LLM systems.
MCP server for building and testing AI agents with multi-model experimentation and insights.
Hosted MCP server for LLM cost estimation, model comparison, and budget-aware routing.
A paid remote MCP for AI SDK benchmark dashboard, built to return verdicts, receipts, usage logs, an