observe-mcp
Summary: An MCP server for investigating production incidents by querying OpenObserve logs (plus optional read-only database correlation).
StreamList— list available OpenObserve streams with document counts, stored size, and newest timestamp; the starting point when you don't know what data exists.StreamSchema— get field names and types for a specific stream before writing SQL against it.SearchSQL— run SQL against a stream over a time window (defaults to last 24h), with flexiblestart/endformats (-24h,now, ISO, epoch s/ms/µs), row limits, and long-field truncation; supports time-bucketed histograms viahistogram(_timestamp, '1 hour').Answer "what actually happened in production?" — e.g. trace error spikes, correlate deploys with failures, and find root causes in log messages.
Query a database read-only (only when
O2_DB_URLis configured — theDbSchemaandDbQuerytools are otherwise hidden entirely) to correlate logs with the data they describe.Filter out known noise — excluded IPs/hostnames (monitors, CI, QA runners) are baked into the tools so counts aren't skewed.
Notable caveats surfaced by the tools: many log shippers emit several rows per request (naive count(*) overstates traffic — count only rows carrying a status/level field), and crawlers inflate unique-visitor counts (classify on user-agent).
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@observe-mcpwhy did signups drop this morning?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
observe-mcp
An MCP server that lets Claude answer "what actually happened in production?" — by reading your OpenObserve logs and, optionally, correlating them against your database through a strictly read-only connection.
You: Why did signups drop this morning?
Claude: [SearchSQL] → 14 signup requests returned 500 between 09:30 and 10:40
[SearchSQL] → all of them log "INTERNAL_SECRET is not set"
[DbQuery] → 9 users created today, 0 with a subscription
A deploy at 10:42 fixed it. Three verified users bounced off in between.Setup is a guided wizard that validates every answer before saving it, including proving that your "read-only" database credential genuinely cannot write.
Why this exists
OpenObserve ships its own MCP server, but on open-source builds it answers:
{"error":"MCP server is only available in enterprise edition"}The ordinary search API is available on every edition. This package wraps that, and
keeps the same tool names as the enterprise server (StreamList, StreamSchema,
SearchSQL) so prompts and habits transfer if you later license it. It then adds the
half the enterprise server does not have: a read-only SQL tool for correlating logs
against the data they describe.
The wizard will tell you if your instance does expose the official server, so you can use the first-party one instead.
Related MCP server: OpenSearch Logs MCP Server
Install
Requires Node 20 or newer. Works on Linux, macOS and Windows.
git clone <this repo> observe-mcp
cd observe-mcp
npm install
npm run setupThe wizard asks for:
OpenObserve URL | checked for reachability, and reports your build version |
Organization id | from your OpenObserve URL after |
Login email + password or token | verified by listing your streams before anything is saved |
Database URL (optional) | verified to be read-only — see below |
Traffic to ignore (optional) | monitors, QA runners, CI, your own crawlers — so they stop skewing counts |
Notes (optional) | anything about your data that would otherwise cost the assistant a few wasted queries |
Then it shows you everything for review, registers the server with your client, smoke-tests it by speaking MCP to the real process, and prints what to try first.
Nothing is written until you confirm, and secrets are never echoed to the terminal or stored anywhere in this repo.
Changing your mind
Type back (or <) at any question to return to the previous one; answers you have
already given come back as the defaults. Before anything is saved you get a review
screen, where Change one of the answers above re-runs just that step:
Review
1. OpenObserve instance https://o2.example.com
2. Organization and credentials default as me@example.com
3. Database for correlation configured, verified read-only
4. Traffic to ignore 10.0.0.5 (ci.example.com)
5. Notes for the assistant (none)
→ 1. Save and register — nothing has been written yet
2. Change one of the answers above
3. Cancel — discard everythingRe-running npm run setup later picks up your current configuration as the defaults, so
it doubles as an edit command.
Traffic to ignore
Monitors, QA runners and CI inflate request counts and unique-visitor counts, and the inflation is worst exactly when you are trying to work out whether something is wrong.
You can give hostnames as well as addresses — hostnames are what you actually know your own machines by — and they are resolved for you at setup time:
Excluded:
• 203.0.113.10 o2.example.com — the host your OpenObserve instance runs on
Keep these excluded? (Y/n)
Exclude anything else? (y/N) y
Addresses or hostnames: qa.example.com, 10.0.0.5
✓ 198.51.100.7 (qa.example.com)
✓ 10.0.0.5The package ships no built-in list — one deployment's monitor is another's real user. The single suggestion is derived from the instance you are configuring: the box running your observability stack is very often the box running your scheduled jobs too. An entry that fails to resolve is reported rather than silently dropped, because a typo in an exclusion list is invisible later — the counts are simply wrong.
Scripted / unattended setup
The wizard reads piped input, so it can be driven from a file or in CI:
printf '%s\n' "https://o2.example.com" "default" "me@example.com" "$O2_TOKEN" \
"y" "$READONLY_DB_URL" "" "" "4" | npm run setupOr skip it entirely and set the environment variables yourself (see Configuration).
The read-only guarantee
A role called "read only" is not necessarily read-only.
Managed Postgres providers often auto-grant a privileged group to roles created through
their web console. A role can be named Read_Only_role, be created expressly for
read-only access, and still hold INSERT, UPDATE, DELETE and BYPASSRLS — this is
not hypothetical, it is why the verification step exists.
So setup proves it instead of trusting the name. It reports the role's attributes and
group memberships, then deliberately switches off the session's read-only default
— that is a settable parameter, not a privilege, and any client can turn it off — and
attempts a write under BEGIN READ WRITE:
role: observer database: appdb
superuser=no createdb=no createrole=no bypassrls=no
inherits from: nothing
default_transaction_read_only=on read replica=false
tables visible: 27
✓ UPDATE refused permission denied for table …
✓ DELETE refused permission denied for table …
✓ TRUNCATE refused permission denied for table …
✓ CREATE TABLE refused permission denied for schema public
✓ CREATE ROLE refused permission denied to create role
✓ GRANT self INSERT refused no privileges were granted (expected)
✓ database credential is read-onlyEvery probe runs inside a transaction that is always rolled back. A probe that fails for any reason other than a privilege denial is reported as inconclusive rather than counted as evidence — a write that fails because the SQL was invalid proves nothing.
The GRANT probe is checked by re-reading has_table_privilege, not by whether the
statement threw: an unentitled GRANT in Postgres returns success and only emits
WARNING: no privileges were granted, so a naive check reports a no-op as an escalation.
If verification fails, the wizard shows you the SQL to create a proper role and offers to retry.
Creating a genuinely read-only role
Run this as the database owner, in a SQL client rather than your provider's "add role" button — roles created in SQL get no automatic group membership:
CREATE ROLE observer LOGIN PASSWORD '…';
GRANT CONNECT ON DATABASE yourdb TO observer;
GRANT USAGE ON SCHEMA public TO observer;
GRANT SELECT ON ALL TABLES IN SCHEMA public TO observer;
ALTER DEFAULT PRIVILEGES IN SCHEMA public GRANT SELECT ON TABLES TO observer;
ALTER ROLE observer SET default_transaction_read_only = on;ALTER DEFAULT PRIVILEGES only covers tables created by the role that runs it, so run it
as whoever owns your application's tables.
Stronger still, if your provider offers it: point O2_DB_URL at a read replica
endpoint. Those reject writes at the compute layer, so no grant mistake can matter. The
verifier reports read replica=true when it detects one.
Tools
Tool | What it does |
| Streams with document counts, size, newest timestamp. Start here. |
| Field names and types for one stream. |
| SQL over a stream, with a time window. |
| Tables and columns the connection can actually see. |
| A single read-only |
DbSchema and DbQuery are hidden entirely unless a database is configured.
Time windows accept whatever you would naturally type: -24h, -90m, now,
2026-10-01, an ISO timestamp, or epoch seconds/ms/µs. They default to the last 24 hours.
How DbQuery refuses writes
Four layers, because the one that matters most is the one you control:
a single statement — no semicolons;
it must begin with
SELECTorWITH;a keyword scan with string literals and comments stripped first — Postgres allows data-modifying CTEs like
WITH x AS (DELETE … RETURNING …)that sail past step 2, while a legitimateWHERE msg LIKE '%delete%'must still work;execution inside a
READ ONLYtransaction with a statement timeout.
None of this replaces a read-only role. It is the belt for that braces.
Configuration
Set by the wizard, or exported yourself:
Variable | Required | Meaning |
| yes | Base URL, e.g. |
| yes | Organization id |
| yes | Login email |
| yes | Password or API token |
| no | Read-only Postgres URL; omit to disable the database tools |
| no | Comma-separated IPs the assistant should filter out |
| no | Deployment notes injected into the |
| no | Statement timeout, default |
| no | Path to a |
OPENOBSERVE_* names are accepted as aliases, so an existing .env works unchanged.
The real environment always wins over a file, so a secret passed at registration time is
never shadowed by a stale copy.
Where your secrets end up
The wizard offers four routes, which differ only in that:
Claude Code, this project / all projects —
claude mcp add --scope local|user. Secrets go in your own Claude config, outside the project. Preferred.Write
.mcp.jsonhere — committable, because it is written with${VAR}placeholders rather than values. You supply the variables in the environment.Just print the config — for Claude Desktop, Cursor, VS Code and friends, with the usual file locations listed.
Codex CLI
MCP is an open protocol, so this is not Claude-only. Codex takes it directly:
codex mcp add observe \
--env O2_URL=https://o2.example.com --env O2_ORG=default \
--env O2_USER=you@example.com --env O2_TOKEN=… --env O2_DB_URL=… \
-- node /path/to/observe-mcp/src/server.mjsVerified against codex-cli 0.159.2: all five tools are discovered and callable, and
DbQuery still refuses a write. One Codex quirk — codex exec runs with
approval: never, so MCP calls are blocked there unless you pass
--dangerously-bypass-approvals-and-sandbox. Interactive codex prompts for approval
normally.
This repo never stores credentials. .gitignore covers .env and .mcp.json anyway.
Checking and troubleshooting
npm run doctor # re-run every check against the current configuration; changes nothing
npm test # protocol and SQL-gate tests; no live services neededdoctor verifies reachability, credentials, the read-only guarantee, and that the server
starts and lists its tools.
Symptom | Cause |
| Fixed in 1.0.1 — update, or re-run setup to rewrite the config |
| Wrong |
| Wrong |
| Expected on OSS builds; it is why this package exists |
| Run |
Tools missing in the client | Restart the client; it reads MCP config at startup |
| No |
Counting traffic correctly
Many log shippers emit several rows per request — one per output line — so a naive
count(*) overstates traffic, sometimes by more than 2×. Check the schema for a status
or level field and count only rows that carry one. The SearchSQL description tells the
model this, but it is worth knowing yourself when you check its work.
Unique-visitor counts have the mirror-image problem: crawlers inflate distinct-IP counts badly. Classify on the user-agent field before calling them users.
Licence
MIT.
Available Tools
3 toolsSearchSQLA
Run SQL against an OpenObserve stream and return the matching rows. The stream name is the FROM target. start and end accept an ISO timestamp, a plain date, epoch seconds/ms/µs, a relative offset like "-24h" or "-90m", or "now". They default to the last 24 hours. Bucket by time with histogram(_timestamp, '1 hour'). _timestamp is microseconds since the epoch. Call StreamList first if you do not know what exists, and StreamSchema before querying a stream whose fields you have not seen — field names differ per stream and guessing wastes a round trip. Beware that many log shippers emit SEVERAL rows per request (one per output line), so a naive count(*) overstates traffic. Check the schema for a status or level field and count only rows that carry one.
| Name | Required | Description | Default |
|---|---|---|---|
| end | No | Window end. Default now. | |
| sql | Yes | e.g. SELECT level, count(*) AS n FROM my_stream GROUP BY level ORDER BY n DESC | |
| size | No | Maximum rows to return. Default 50, maximum 1000. | |
| start | No | Window start. Default -24h. | |
| max_field_chars | No | Truncate long string fields to this length. Default 400. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full load and does well: it discloses the 24h default window, that `_timestamp` is microseconds, histogram bucketing, and the non-obvious pitfall that log shippers emit multiple rows per request so count(*) overstates traffic. It stops short of stating whether non-SELECT statements are permitted or what permissions/limits apply, which matters for a free-form SQL tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, followed by timestamp formats, routing guidance, and a query pitfall — each sentence earns its place. No filler or restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description should ideally hint at the returned row shape, which it only gestures at ('matching rows'). For a free-form SQL tool with 100% schema coverage on inputs, everything an agent needs to invoke it correctly is present, but return-shape detail is thin.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning the schema lacks: it enumerates the accepted formats for `start`/`end` (ISO, plain date, epoch s/ms/µs, relative offsets, 'now') and confirms the 24h default. That is genuine value beyond the terse schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Run SQL against an OpenObserve stream and return the matching rows') and clarifies that the stream name acts as the FROM target. It also names the two sibling tools it relates to, so an agent can place it without opening another definition.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit routing rules: call StreamList if you don't know what exists, and StreamSchema before querying an unseen stream, with the reason (field names differ and guessing wastes a round trip). This tells the agent both when to use this tool and when to call its siblings first.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
StreamListA
List OpenObserve streams with document counts, stored size and the newest document timestamp. Start here when you do not know what data is available.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the burden. It discloses the payload shape (counts, size, freshness timestamp) and the read-only nature implied by "List," which is meaningful context. It does not mention auth needs, result limits, or pagination, leaving some behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, both load-bearing: the first defines the payload, the second the usage trigger. The most decision-relevant content (what is returned) is front-loaded with zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must characterize returns, and it does list the three key fields. It omits ordering, volume/limit behavior, and how results relate to StreamSchema output, which are minor gaps for a zero-parameter discovery tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is no parameter semantics to document and the baseline is 4. The description correctly does not invent parameters, and the schema closes the object with additionalProperties=false.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb+resource ("List OpenObserve streams") and enumerates the returned fields (document counts, stored size, newest document timestamp). Sibling StreamSchema and SearchSQL are not named, so the agent must infer the distinction from the "start here" framing rather than an explicit contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Start here when you do not know what data is available" gives a concrete entry-point condition for selecting this tool over exploration via SearchSQL. It stops short of naming alternatives or stating exclusions, so it is clear context without full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
StreamSchemaA
Field names and types for one OpenObserve stream. Call this before writing SQL against a stream you have not queried before.
| Name | Required | Description | Default |
|---|---|---|---|
| stream | Yes | Stream name, as StreamList reports it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses what the tool returns (field names and types) and the recommended call ordering, which is useful, but says nothing about permissions, error behavior when a stream doesn't exist, or cost.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero waste, with the resource identified first and the usage rule trailing it. Nothing needs trimming.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter read tool with no output schema, the description covers the returned content and the call ordering. It stops short of documenting failure modes or whether results are cached, which are minor gaps here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and there is a single required parameter already documented as 'Stream name, as StreamList reports it.' The description adds only the implicit notion of 'one stream' and no format or naming detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource and output: field names and types for one OpenObserve stream. It clearly differs from StreamList (enumerating streams) and SearchSQL (executing queries), though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit conditional guidance: call this before writing SQL against a stream you have not queried before. That is a usable when-to-use rule, but it doesn't name the alternative tools or state when this call is unnecessary.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v1.0.1- First observed
SearchSQL - First observed
StreamList - First observed
StreamSchema
TDQS
Scored across 3 tools
Each tool has a clearly distinct purpose: listing streams, retrieving a schema, and executing SQL queries. The descriptions explicitly guide the correct order of use, eliminating ambiguity.
All names use PascalCase without separators, but the structural pattern is inconsistent: StreamList and StreamSchema are noun-first, while SearchSQL is verb-first. This mixed verb/noun style reduces predictability, though names remain readable.
Three tools cover the essential workflow of discovering streams, inspecting schemas, and querying data. This minimal, focused set avoids redundancy and each tool earns its place.
The core query lifecycle is well covered: list streams, get schema, run SQL. Minor gaps exist, such as no tool for cross-stream queries or saved query management, but agents can work around these with the provided SQL interface.
Maintenance
Related MCP Connectors
Read-only access to Auralogs production logs: search logs, inspect errors, review AI analyses.
- SuperlogOAuthsh.superlog
Open-source agent that observes and fixes your application. Query logs, traces, metrics, incidents.
Ingest and search LogsLoom logs from coding agents.
Query 40 databases from Claude, ChatGPT, or Cursor — on any device. Read-only, encrypted, audited.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables querying log data stored in SQLite databases through the Model Context Protocol, allowing natural language interactions with log analysis.-
- FlicenseAqualityDmaintenanceEnables querying OpenTelemetry logs stored in OpenSearch across development and production environments. Provides tools for searching logs by various criteria including free-text Lucene queries, trace IDs, service names, error levels, and specific fields.8-
- AlicenseAqualityCmaintenanceEnables read-only querying of Grafana Loki logs through tools like label discovery, LogQL queries, and recent log retrieval, without requiring knowledge of the Loki API.4MIT
- AlicenseAqualityCmaintenanceEnables AI agents to query OriginChain's telemetry agent for log patterns, semantic incident analysis, distributed traces, service graphs, and read-only SQL queries over telemetry data.104 npmMIT