phoenix-mcp-eval
phoenix-mcp-eval
MCP server for Arize Phoenix — LLM tracing, evaluation, and dataset management via AI agents.
What is this?
phoenix-mcp-eval is an MCP (Model Context Protocol) server that exposes Arize Phoenix's LLM observability capabilities to AI agents. It enables AI-driven analysis of traces, evaluation of LLM outputs, and management of evaluation datasets — directly from an MCP-compatible agent.
Built for platform engineers and ML teams running LLM pipelines on AI Foundry, LangChain, or LlamaIndex who need automated quality assurance and tracing.
Available Tools
Tool | Description |
| List all Phoenix tracing projects |
| Retrieve LLM traces for a project with filters |
| Get individual spans with input/output/latency data |
| List evaluation datasets in Phoenix |
| Fetch dataset examples for review or comparison |
| List evaluation runs and their scores |
| Get aggregated evaluation metrics (precision, recall, etc.) |
| Run structured queries over trace data |
Quick Start
Prerequisites
Python 3.11+
Arize Phoenix instance (self-hosted or cloud)
Phoenix API key or local server URL
Installation
git clone [https://github.com/akkireddy-challa/phoenix-mcp-eval.git](https://github.com/akkireddy-challa/phoenix-mcp-eval.git)
cd phoenix-mcp-eval
pip install -r requirements.txtConfiguration
export PHOENIX_HOST=http://localhost:6006
export PHOENIX_API_KEY=<your-api-key> # if using cloudRun
python server.pyMCP Client Config (Claude Desktop)
{
"mcpServers": {
"phoenix": {
"command": "python",
"args": ["/path/to/phoenix-mcp-eval/server.py"],
"env": {
"PHOENIX_HOST": "http://localhost:6006"
}
}
}
}Security Model
Connects to Phoenix via API key or local network only
All operations are read-only by default (trace/eval retrieval)
No model weights, prompts, or PII are transmitted outside Phoenix
API key stored in environment variables, never in code
Designed for internal network use within a Kubernetes cluster
Use Cases at Telia
This pattern is used to allow AI agents to:
Automatically review LLM trace quality after AI Foundry deployments
Surface failing evaluation metrics to on-call engineers without manual Phoenix access
Compare evaluation datasets across model versions
Trigger re-evaluation jobs based on trace anomaly detection
Roadmap
run_evaluation— trigger evaluation jobs programmaticallycreate_dataset— export traces to evaluation datasetsget_prompt_templates— retrieve versioned prompts from PhoenixIntegration with Azure AI Foundry deployment events
GitHub Actions workflow for CI validation
Related Projects
Repo | Purpose |
Kubernetes cluster diagnostics via MCP | |
Azure resource management via MCP | |
Grafana dashboards and alerts via MCP |
License
MIT License. See LICENSE for details.
Built by Akkireddy Challa — Platform Engineer at Telia, Stockholm.