phoenix-mcp-eval
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@phoenix-mcp-evallist evaluation scores for my chatbot project"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
phoenix-mcp-eval
MCP server for Arize Phoenix — LLM tracing, evaluation, and dataset management via AI agents.
What is this?
phoenix-mcp-eval is an MCP (Model Context Protocol) server that exposes Arize Phoenix's LLM observability capabilities to AI agents. It enables AI-driven analysis of traces, evaluation of LLM outputs, and management of evaluation datasets — directly from an MCP-compatible agent.
Built for platform engineers and ML teams running LLM pipelines on AI Foundry, LangChain, or LlamaIndex who need automated quality assurance and tracing.
Related MCP server: mcp-llm-eval
Available Tools
Tool | Description |
| List all Phoenix tracing projects |
| Retrieve LLM traces for a project with filters |
| Get individual spans with input/output/latency data |
| List evaluation datasets in Phoenix |
| Fetch dataset examples for review or comparison |
| List evaluation runs and their scores |
| Get aggregated evaluation metrics (precision, recall, etc.) |
| Run structured queries over trace data |
Quick Start
Prerequisites
Python 3.11+
Arize Phoenix instance (self-hosted or cloud)
Phoenix API key or local server URL
Installation
git clone https://github.com/akkireddy-challa/phoenix-mcp-eval
cd phoenix-mcp-eval
pip install -r requirements.txtConfiguration
export PHOENIX_HOST=http://localhost:6006
export PHOENIX_API_KEY=<your-api-key> # if using cloudRun
python server.pyMCP Client Config (Claude Desktop)
{
"mcpServers": {
"phoenix": {
"command": "python",
"args": ["/path/to/phoenix-mcp-eval/server.py"],
"env": {
"PHOENIX_HOST": "http://localhost:6006"
}
}
}
}Security Model
Connects to Phoenix via API key or local network only
All operations are read-only by default (trace/eval retrieval)
No model weights, prompts, or PII are transmitted outside Phoenix
API key stored in environment variables, never in code
Designed for internal network use within a Kubernetes cluster
Use Cases at Telia
This pattern is used to allow AI agents to:
Automatically review LLM trace quality after AI Foundry deployments
Surface failing evaluation metrics to on-call engineers without manual Phoenix access
Compare evaluation datasets across model versions
Trigger re-evaluation jobs based on trace anomaly detection
Roadmap
run_evaluation— trigger evaluation jobs programmaticallycreate_dataset— export traces to evaluation datasetsget_prompt_templates— retrieve versioned prompts from PhoenixIntegration with Azure AI Foundry deployment events
[ x GitHub Actions workflow for CI validation
Related Projects
Repo | Purpose |
Kubernetes cluster diagnostics via MCP | |
Azure resource management via MCP | |
Grafana dashboards and alerts via MCP |
License
MIT License. See LICENSE for details.
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
AlicenseBqualityAmaintenancePhoenix MCP Server is an implementation of the Model Context Protocol for the Arize Phoenix platform. It provides a unified interface to Phoenix's capabilites. You can use Phoenix MCP Server for: Prompts Management: Create, list, update, and iterate on prompts Datasets: Explore datasets, and synte272,83111,118Apache 2.0- AlicenseAqualityCmaintenanceA local MCP server that packages LLM evaluation gates as reusable CI/CD primitives, enabling AI agents to run datasets against models, score responses, and enforce quality thresholds.10MIT
- AlicenseAqualityDmaintenanceMCP server that gives AI agents access to your application's OpenTelemetry traces for querying, analysis, and debugging.5262MIT
- AlicenseNot gradedqualityAmaintenanceMCP server that enables AI agents to run a deterministic orchestration loop with decomposition, subagent execution, and review feedback across multiple LLM backends.54MIT
Related MCP Connectors
MCP server providing access to the Scorecard API to evaluate and optimize LLM systems.
MCP server for AI agents to plan, verify, and deploy Cloudflare-native apps.
An MCP server for Arcjet - the runtime security platform that ships with your AI code.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/akkireddy-challa/phoenix-mcp-eval'
If you have feedback or need assistance with the MCP directory API, please join our Discord server