Enables LLM evaluation and observability by uploading documents, building test sets, running RAG pipelines, and automatically scoring answers for groundedness, hallucination risk, retrieval quality, latency, and cost, with tools exposed to MCP-compatible clients.
Enables AI agents to access observability and evaluation data, including run history, span traces, LLM-as-judge evaluation results, and regression reports.