Skip to main content
Glama
TerraCo89

es-error-lens

by TerraCo89
README.md
# es-error-lens

![CI](https://github.com/TerraCo89/es-error-lens/actions/workflows/ci.yml/badge.svg)
![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)
![Python 3.10+](https://img.shields.io/badge/python-3.10%2B-blue.svg)
![Offline tests](https://img.shields.io/badge/tests-offline%2C%20no%20cluster%20needed-brightgreen.svg)

**An MCP server that gives LLM agents eyes on your Elasticsearch logs.** Point
it at a cluster holding [ECS](https://www.elastic.co/guide/en/ecs/current/index.html)-format
logs and any MCP client (Claude Desktop, Claude Code, or your own agent) can
search errors, detect recurring patterns, analyze error-rate trends, and pull
full trace context — the queries an engineer runs by hand during triage,
exposed as tools.

| Tool | What it answers |
|---|---|
| `search_errors` | "What's failing right now?" — filtered by window, service, level |
| `get_error_patterns` | "Is this systemic or a one-off?" — message aggregation with occurrence counts |
| `analyze_error_trend` | "Is it getting worse?" — time-series histogram, peak detection, trend verdict |
| `get_error_context` | "What led up to this?" — every log sharing the error's trace ID, in order |
| `compare_errors` | "Are these two failures related?" — attribute + fuzzy-message comparison |
| `health_check` | "Can I even reach the cluster?" |

Elasticsearch is spoken to over its plain REST API via `httpx` — no
Elasticsearch client dependency, and the HTTP layer is transport-injectable,
so the entire server runs offline against the bundled `FakeElasticsearch`
for tests and the demo.

## Demo (offline, no cluster, no API keys)

```bash
git clone https://github.com/TerraCo89/es-error-lens && cd es-error-lens
python -m venv .venv && source .venv/bin/activate   # .venv\Scripts\activate on Windows
pip install -e ".[dev]"

pytest              # 22 tests, all offline, < 1 second
es-error-lens-demo  # synthetic incident walk-through
```

The demo mounts `FakeElasticsearch` with a deterministic 60-document ECS
corpus simulating a small incident — a steady payment-provider failure, an
escalating search-timeout problem, and background noise — then triages it:

```text
[get_error_patterns] 2 recurring patterns in 57 errors:
   42x  [TimeoutError] Search query timed out after 30s  (search-api)
   13x  [UpstreamHTTPError] Payment provider returned 502 Bad Gateway  (checkout-api)
  error types: TimeoutError=42, UpstreamHTTPError=13, TemplateError=1, ConnectionResetError=1

[analyze_error_trend] search-api: 42 errors, trend=INCREASING
  peak: 8 errors at 2026-07-10T06:00:00Z

[get_error_context] trace-checkout-7f3a -> [UpstreamHTTPError] Payment provider returned 502 Bad Gateway
  3 related logs on the same trace:
    info  POST /checkout started
    info  Cart validated: 3 items
    warn  Payment provider latency 4100ms exceeds SLO

Demo result: OK
```

## Using it against a real cluster

```bash
export ELASTICSEARCH_URL=http://localhost:9200   # default
export ES_INDEX_PATTERN=logs-*                   # default

es-error-lens                                    # stdio, for Claude Desktop / Claude Code
es-error-lens --transport streamable-http --port 8080   # HTTP clients
```

Claude Desktop / Claude Code config:

```json
{
  "mcpServers": {
    "es-error-lens": {
      "command": "es-error-lens",
      "env": { "ELASTICSEARCH_URL": "http://localhost:9200" }
    }
  }
}
```

**Expected document shape** — standard ECS fields, the ones Filebeat and the
ECS logging libraries emit by default: `@timestamp`, `log.level`, `message`,
`service.name`, `error.type`, `error.stack_trace`, `trace.id`, `host.name`.
Anything missing degrades gracefully (fields come back `null`).

## Testing without a cluster

`FakeElasticsearch` (also exported) implements just enough of the `_search`
API for this server: `bool` queries with `term`/`range` clauses, sorting,
`terms` aggregations with `top_hits` sub-aggregations, and `date_histogram`
with fixed intervals. Mount it in your own tests:

```python
from es_error_lens import FakeElasticsearch, set_transport, search_errors

set_transport(FakeElasticsearch(my_ecs_docs).transport())
result = await search_errors(time_range="1h", service="checkout-api")
```

The test suite runs entirely through it — deterministic, offline, fast — and
a hygiene test enforces that all bundled demo data stays synthetic (hosts
under `.example`, no real domains or addresses).

## Provenance

Extracted from a personal observability platform where it fronts the
Elasticsearch instance that aggregates structured logs from a fleet of
side-project services, letting coding agents triage production errors
during development sessions. All demo and test data in this repository is
synthetic.

## License

MIT © 2026 Kris Cernjavic

TDQS

A3.8/5.0

Scored across 6 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: trend analysis, comparison, context retrieval, pattern detection, health check, and search. No overlapping functionality.

Naming Consistency4/5

Five tools follow verb_noun pattern (e.g., analyze_error_trend, search_errors); health_check is noun_verb, a minor deviation. Otherwise consistent snake_case.

Tool Count5/5

Six tools is well-scoped for an error analysis server. Each tool contributes a distinct capability without bloat.

Completeness5/5

Covers error search, trend analysis, pattern detection, context retrieval, comparison, and health check. No obvious gaps for the observation/analysis domain.

Maintenance

ActivityStale
ResponsivenessNo issues