AgentLens MCP Server
README.md
<div align="center">

# π AgentLens
**Open-source observability and evaluation platform for AI agents.** Trace every step of your agent, score its behavior with LLM-as-judge evaluations, catch prompt regressions in CI, and give your coding agents (Claude Code, Cursor, etc.) native access to run history via MCP.
[](https://github.com/Nexus-universe-space/agentlens/actions/workflows/ci.yml)
[](https://www.python.org/)
[](https://github.com/Nexus-universe-space/agentlens/blob/main/LICENSE)
[](https://docs.astral.sh/ruff/)
**Zero-config by default.** One SQLite file, one `lens serve` command, done. Scale to Postgres when you need to.
</div>
---
## Why AgentLens?
Everyone is shipping agents now. Almost nobody knows what their agents actually do in production: which tools they call, what the LLM sees, how much it costs, and β crucially β whether yesterday's prompt change made the agent *worse*.
AgentLens brings the discipline of classical observability (traces, spans, cost accounting) and classical QA (regression suites, semantic scoring) to the agent world, in a single lightweight package with no infrastructure requirements.
| Pain point | AgentLens answer |
|---|---|
| "Which tool call blew up my agent's latency?" | Auto-nested async span traces via `@agent`, `@tool`, `@llm` decorators |
| "Is this prompt actually better than v2.3.1?" | Prompt regression suites with keyword + semantic checks, CI-gated |
| "How do I know the agent did a good job?" | LLM-as-judge evaluations with configurable rubrics |
| "What did that run cost me?" | Per-span token counting (tiktoken) and cost estimation |
| "I want my coding agent to look at run history" | Built-in MCP server, pluggable into Claude Code / Cursor |
## Quick start
```bash
pip install agentlens
# 1. Start the server + dashboard
lens serve # β http://localhost:3368
# 2. Run the traced demo agent
lens demo
# 3. Run evaluations and regression suites
lens regression init my_suite.yml
lens regression run my_suite.yml
```
### Instrument your agent (30 seconds)
```python
from agentlens.sdk import agent, tool, llm
from agentlens.evals import evaluate
from agentlens.regression import load_suite, run_suite
@llm(model="gpt-4o-mini")
async def answer(question: str) -> str:
# your LLM call here β tokens and cost are tracked automatically
...
@tool("search_index")
async def search_index(query: str) -> list[str]:
...
@agent("rag_agent")
async def rag_agent(question: str) -> str:
ctx = await search_index(question)
return await answer(f"Context: {ctx}\nQuestion: {question}")
# Span tree, parent-child nesting and cost β all automatic.
result = await rag_agent("How do I reset my password?")
```
## Features
### π§ Tracing SDK
Drop-in decorators (`@agent`, `@tool`, `@llm`, `@retriever`) that build a full span tree with automatic parent-child nesting through `contextvars`. Works for both async and synchronous functions. Each span records input/output, model, token counts, estimated cost (using real per-model pricing for OpenAI and Anthropic models, with an extensible registry for your own models) and status. Spans are flushed to the AgentLens API through a pluggable callback, so you can persist them, export them to OTLP, or mock them in tests.
### π§ββοΈ LLM-as-judge evaluations
Score agent outputs against built-in criteria (`CORRECTNESS`, `RELEVANCE`, `COHERENCE`, `SAFETY`, `HALLUCINATION`, `COMPLETENESS`) or define your own rubric. The judge is any OpenAI-compatible endpoint β OpenAI, Anthropic gateways, Ollama, LiteLLM β configured with two environment variables. Every case x criterion pair produces a numeric score, a human-readable reason, and a pass/fail verdict, aggregated into an evaluation report.
### π‘οΈ Prompt regression suites
YAML-defined test suites, the way prompt engineers have been wishing for:
```yaml
name: support-agent
agent: support_agent
model: gpt-4o-mini
version: "1.2.0"
cases:
- name: refund question
input: "When will I receive my refund?"
expected_contains: ["refund", "days"]
expected_not_contains: ["deny"]
min_score: 0.8
tags: ["billing"]
```
`lens regression run suite.yml` executes every case, applies keyword checks plus optional semantic scoring, and exits non-zero on failure β perfect for CI. Compare versions side by side in the dashboard.
### π° Cost tracking
Token counting via tiktoken (exact for OpenAI encoders, character-ratio fallback for unknown models), per-model pricing registry, and aggregated cost summaries across runs. Know exactly what your agent fleet spends.
### π€ Native MCP server
AgentLens ships as a proper MCP server (`mcp server`), exposing four tools:
| Tool | Purpose |
|---|---|
| `lens_search_runs` | Find recent agent runs, filter by project |
| `lens_run_summary` | Full summary + span tree of a run |
| `lens_eval_latest` | Latest LLM-as-judge evaluation results |
| `lens_regression_status` | Latest regression report |
Wire it into your favorite coding agent and let it debug your production agents for you:
```jsonc
// ~/.config/claude/claude_desktop_config.json
{
"mcpServers": {
"agentlens": {
"command": "lens",
"args": ["mcp-server"]
}
}
}
```
### π REST API + dashboard
A FastAPI application (`/v1/spans`, `/v1/runs`, `/v1/evals`, `/v1/regressions`, `/v1/summary`) with OpenAPI docs at `/docs`, plus a dark-themed embedded dashboard with run tables, interactive span trees, evaluation results, and regression history. SQLite zero-config by default; flip to Postgres with `LENS_STORAGE_BACKEND=postgres`.
## Architecture
```
ββββββββββββββ decorators ββββββββββββββββ flush ββββββββββββββββββββ
β Your code β ββββββββββββββΊ β Tracing SDK β ββββββββΊ β AgentLens API β
ββββββββββββββ ββββββββββββββββ β (FastAPI) β
ββββββββββββββ run ββββββββββββββββββββ β β
β YAML β ββββββββββΊ β Regression runnerβ βββ β βββββββββββββββ β
β suites β β (+ keyword/sem.) β β β β Dashboard / β β
ββββββββββββββ ββββββββββββββββββββ β β β MCP server β β
ββββββββββββββ evaluate ββββββββββββββββββββ β β βββββββββββββββ β
β Test casesβ ββββββββββΊ β LLM-as-judge β βββΌβββββββΊββββββββββ¬ββββββββββ
ββββββββββββββ ββββββββββββββββββββ β β
β ββββββββΌβββββββ
ββββββββββΊβ SQLite / β
β Postgres β
βββββββββββββββ
```
## CLI reference
| Command | Description |
|---|---|
| `lens serve` | Start API + dashboard (default port 3368) |
| `lens mcp-server` | Run the MCP server over stdio |
| `lens regression run <file>` | Run a regression suite (exits non-zero on failure) |
| `lens regression init <file>` | Scaffold a regression suite |
| `lens demo` | Run a traced demo agent, print the span tree |
Configuration happens through `LENS_` environment variables or a `.env` file (`LENS_PORT`, `LENS_STORAGE_BACKEND`, `LENS_DATABASE_URL`, `LENS_JUDGE_MODEL`, `LENS_JUDGE_BASE_URL`, `LENS_JUDGE_API_KEY`, ...). See `agentlens/core/settings.py` for the full list.
## Docker
```bash
docker compose up --build
# β http://localhost:3368
```
## Development
```bash
git clone https://github.com/Nexus-universe-space/agentlens.git
cd agentlens
pip install -e ".[dev]"
ruff check agentlens tests examples # lint
ruff format agentlens tests examples # format
pytest tests/unit tests/integration # 41 tests
mypy agentlens # strict-ish typing
```
CI runs lint, tests (Python 3.10β3.13 with coverage), and typing on every push and PR.
## Project layout
```
agentlens/
βββ agentlens/
β βββ api/ FastAPI REST API (spans, runs, evals, regressions, summary)
β βββ cli/ Click CLI (serve, mcp-server, regression, demo)
β βββ core/ Settings, Pydantic models, token/cost engine
β βββ evals/ LLM-as-judge evaluation engine
β βββ mcp/ MCP server with typed tools
β βββ regression/ YAML suite loader + regression runner
β βββ sdk/ Tracing decorators (agent / tool / llm / retriever)
β βββ storage/ Async SQLAlchemy store (SQLite + Postgres)
β βββ web/ Embedded React dashboard (zero build step)
βββ tests/ 41 tests β unit + integration (SQLite, API, MCP)
βββ examples/ Traced RAG agent + sample regression suite
βββ Dockerfile
βββ docker-compose.yml
βββ .github/workflows/ci.yml
```
## Roadmap
- [ ] OTLP trace exporter (Honeycomb, Grafana Tempo, Jaeger)
- [ ] Multi-agent session grouping and conversation views
- [ ] Real-time WebSocket updates in the dashboard
- [ ] Evaluation presets per domain (coding, RAG, customer support)
- [ ] Human-in-the-loop annotation of eval cases
- [ ] Prompt version diffing with blame attribution
## Contributing
Contributions are very welcome! Please read [CONTRIBUTING.md](CONTRIBUTING.md) for the workflow, and review our [Code of Conduct](CODE_OF_CONDUCT.md).
## License
MIT β see [LICENSE](LICENSE).