Skip to main content
Glama

observe-mcp

An MCP server that lets Claude answer "what actually happened in production?" — by reading your OpenObserve logs and, optionally, correlating them against your database through a strictly read-only connection.

You:    Why did signups drop this morning?
Claude: [SearchSQL]  → 14 signup requests returned 500 between 09:30 and 10:40
        [SearchSQL]  → all of them log "INTERNAL_SECRET is not set"
        [DbQuery]    → 9 users created today, 0 with a subscription
        A deploy at 10:42 fixed it. Three verified users bounced off in between.

Setup is a guided wizard that validates every answer before saving it, including proving that your "read-only" database credential genuinely cannot write.


Why this exists

OpenObserve ships its own MCP server, but on open-source builds it answers:

{"error":"MCP server is only available in enterprise edition"}

The ordinary search API is available on every edition. This package wraps that, and keeps the same tool names as the enterprise server (StreamList, StreamSchema, SearchSQL) so prompts and habits transfer if you later license it. It then adds the half the enterprise server does not have: a read-only SQL tool for correlating logs against the data they describe.

The wizard will tell you if your instance does expose the official server, so you can use the first-party one instead.


Related MCP server: OpenSearch Logs MCP Server

Install

Requires Node 20 or newer. Works on Linux, macOS and Windows.

git clone <this repo> observe-mcp
cd observe-mcp
npm install
npm run setup

The wizard asks for:

OpenObserve URL

checked for reachability, and reports your build version

Organization id

from your OpenObserve URL after /web/, or Settings → Organizations

Login email + password or token

verified by listing your streams before anything is saved

Database URL (optional)

verified to be read-only — see below

Traffic to ignore (optional)

monitors, QA runners, CI, your own crawlers — so they stop skewing counts

Notes (optional)

anything about your data that would otherwise cost the assistant a few wasted queries

Then it shows you everything for review, registers the server with your client, smoke-tests it by speaking MCP to the real process, and prints what to try first.

Nothing is written until you confirm, and secrets are never echoed to the terminal or stored anywhere in this repo.

Changing your mind

Type back (or <) at any question to return to the previous one; answers you have already given come back as the defaults. Before anything is saved you get a review screen, where Change one of the answers above re-runs just that step:

Review
  1. OpenObserve instance           https://o2.example.com
  2. Organization and credentials   default as me@example.com
  3. Database for correlation       configured, verified read-only
  4. Traffic to ignore              10.0.0.5 (ci.example.com)
  5. Notes for the assistant        (none)

  → 1. Save and register — nothing has been written yet
    2. Change one of the answers above
    3. Cancel — discard everything

Re-running npm run setup later picks up your current configuration as the defaults, so it doubles as an edit command.

Traffic to ignore

Monitors, QA runners and CI inflate request counts and unique-visitor counts, and the inflation is worst exactly when you are trying to work out whether something is wrong.

You can give hostnames as well as addresses — hostnames are what you actually know your own machines by — and they are resolved for you at setup time:

Excluded:
  • 203.0.113.10 o2.example.com — the host your OpenObserve instance runs on

Keep these excluded? (Y/n)
Exclude anything else? (y/N) y
Addresses or hostnames: qa.example.com, 10.0.0.5
  ✓ 198.51.100.7 (qa.example.com)
  ✓ 10.0.0.5

The package ships no built-in list — one deployment's monitor is another's real user. The single suggestion is derived from the instance you are configuring: the box running your observability stack is very often the box running your scheduled jobs too. An entry that fails to resolve is reported rather than silently dropped, because a typo in an exclusion list is invisible later — the counts are simply wrong.

Scripted / unattended setup

The wizard reads piped input, so it can be driven from a file or in CI:

printf '%s\n' "https://o2.example.com" "default" "me@example.com" "$O2_TOKEN" \
              "y" "$READONLY_DB_URL" "" "" "4" | npm run setup

Or skip it entirely and set the environment variables yourself (see Configuration).


The read-only guarantee

A role called "read only" is not necessarily read-only.

Managed Postgres providers often auto-grant a privileged group to roles created through their web console. A role can be named Read_Only_role, be created expressly for read-only access, and still hold INSERT, UPDATE, DELETE and BYPASSRLS — this is not hypothetical, it is why the verification step exists.

So setup proves it instead of trusting the name. It reports the role's attributes and group memberships, then deliberately switches off the session's read-only default — that is a settable parameter, not a privilege, and any client can turn it off — and attempts a write under BEGIN READ WRITE:

  role: observer   database: appdb
  superuser=no  createdb=no  createrole=no  bypassrls=no
  inherits from: nothing
  default_transaction_read_only=on  read replica=false
  tables visible: 27

  ✓ UPDATE             refused  permission denied for table …
  ✓ DELETE             refused  permission denied for table …
  ✓ TRUNCATE           refused  permission denied for table …
  ✓ CREATE TABLE       refused  permission denied for schema public
  ✓ CREATE ROLE        refused  permission denied to create role
  ✓ GRANT self INSERT  refused  no privileges were granted (expected)

✓ database credential is read-only

Every probe runs inside a transaction that is always rolled back. A probe that fails for any reason other than a privilege denial is reported as inconclusive rather than counted as evidence — a write that fails because the SQL was invalid proves nothing.

The GRANT probe is checked by re-reading has_table_privilege, not by whether the statement threw: an unentitled GRANT in Postgres returns success and only emits WARNING: no privileges were granted, so a naive check reports a no-op as an escalation.

If verification fails, the wizard shows you the SQL to create a proper role and offers to retry.

Creating a genuinely read-only role

Run this as the database owner, in a SQL client rather than your provider's "add role" button — roles created in SQL get no automatic group membership:

CREATE ROLE observer LOGIN PASSWORD '…';
GRANT CONNECT ON DATABASE yourdb TO observer;
GRANT USAGE ON SCHEMA public TO observer;
GRANT SELECT ON ALL TABLES IN SCHEMA public TO observer;
ALTER DEFAULT PRIVILEGES IN SCHEMA public GRANT SELECT ON TABLES TO observer;
ALTER ROLE observer SET default_transaction_read_only = on;

ALTER DEFAULT PRIVILEGES only covers tables created by the role that runs it, so run it as whoever owns your application's tables.

Stronger still, if your provider offers it: point O2_DB_URL at a read replica endpoint. Those reject writes at the compute layer, so no grant mistake can matter. The verifier reports read replica=true when it detects one.


Tools

Tool

What it does

StreamList

Streams with document counts, size, newest timestamp. Start here.

StreamSchema

Field names and types for one stream.

SearchSQL

SQL over a stream, with a time window.

DbSchema

Tables and columns the connection can actually see.

DbQuery

A single read-only SELECT / WITH.

DbSchema and DbQuery are hidden entirely unless a database is configured.

Time windows accept whatever you would naturally type: -24h, -90m, now, 2026-10-01, an ISO timestamp, or epoch seconds/ms/µs. They default to the last 24 hours.

How DbQuery refuses writes

Four layers, because the one that matters most is the one you control:

  1. a single statement — no semicolons;

  2. it must begin with SELECT or WITH;

  3. a keyword scan with string literals and comments stripped first — Postgres allows data-modifying CTEs like WITH x AS (DELETE … RETURNING …) that sail past step 2, while a legitimate WHERE msg LIKE '%delete%' must still work;

  4. execution inside a READ ONLY transaction with a statement timeout.

None of this replaces a read-only role. It is the belt for that braces.


Configuration

Set by the wizard, or exported yourself:

Variable

Required

Meaning

O2_URL

yes

Base URL, e.g. https://o2.example.com

O2_ORG

yes

Organization id

O2_USER

yes

Login email

O2_TOKEN

yes

Password or API token

O2_DB_URL

no

Read-only Postgres URL; omit to disable the database tools

O2_EXCLUDE_IPS

no

Comma-separated IPs the assistant should filter out

O2_EXTRA_NOTES

no

Deployment notes injected into the SearchSQL description

O2_DB_TIMEOUT_MS

no

Statement timeout, default 15000

O2_ENV_FILE

no

Path to a KEY=value file to read as a fallback (local development)

OPENOBSERVE_* names are accepted as aliases, so an existing .env works unchanged. The real environment always wins over a file, so a secret passed at registration time is never shadowed by a stale copy.

Where your secrets end up

The wizard offers four routes, which differ only in that:

  • Claude Code, this project / all projects — claude mcp add --scope local|user. Secrets go in your own Claude config, outside the project. Preferred.

  • Write .mcp.json here — committable, because it is written with ${VAR} placeholders rather than values. You supply the variables in the environment.

  • Just print the config — for Claude Desktop, Cursor, VS Code and friends, with the usual file locations listed.

Codex CLI

MCP is an open protocol, so this is not Claude-only. Codex takes it directly:

codex mcp add observe \
  --env O2_URL=https://o2.example.com --env O2_ORG=default \
  --env O2_USER=you@example.com --env O2_TOKEN=… --env O2_DB_URL=… \
  -- node /path/to/observe-mcp/src/server.mjs

Verified against codex-cli 0.159.2: all five tools are discovered and callable, and DbQuery still refuses a write. One Codex quirk — codex exec runs with approval: never, so MCP calls are blocked there unless you pass --dangerously-bypass-approvals-and-sandbox. Interactive codex prompts for approval normally.

This repo never stores credentials. .gitignore covers .env and .mcp.json anyway.


Checking and troubleshooting

npm run doctor   # re-run every check against the current configuration; changes nothing
npm test         # protocol and SQL-gate tests; no live services needed

doctor verifies reachability, credentials, the read-only guarantee, and that the server starts and lists its tools.

Symptom

Cause

Cannot find module 'C:\\C:\\…'

Fixed in 1.0.1 — update, or re-run setup to rewrite the config

HTTP 401 / 403

Wrong O2_USER / O2_TOKEN, or the credential belongs to another org

HTTP 404 on search

Wrong O2_ORG, or the stream does not exist — run StreamList

MCP server is only available in enterprise edition

Expected on OSS builds; it is why this package exists

Cannot load the "postgres" driver

Run npm install in this directory

Tools missing in the client

Restart the client; it reads MCP config at startup

Db* tools missing

No O2_DB_URL configured — re-run setup

Counting traffic correctly

Many log shippers emit several rows per request — one per output line — so a naive count(*) overstates traffic, sometimes by more than 2×. Check the schema for a status or level field and count only rows that carry one. The SearchSQL description tells the model this, but it is worth knowing yourself when you check its work.

Unique-visitor counts have the mirror-image problem: crawlers inflate distinct-IP counts badly. Classify on the user-agent field before calling them users.


Licence

MIT.

Available Tools

3 tools
SearchSQLA

Run SQL against an OpenObserve stream and return the matching rows. The stream name is the FROM target. start and end accept an ISO timestamp, a plain date, epoch seconds/ms/µs, a relative offset like "-24h" or "-90m", or "now". They default to the last 24 hours. Bucket by time with histogram(_timestamp, '1 hour'). _timestamp is microseconds since the epoch. Call StreamList first if you do not know what exists, and StreamSchema before querying a stream whose fields you have not seen — field names differ per stream and guessing wastes a round trip. Beware that many log shippers emit SEVERAL rows per request (one per output line), so a naive count(*) overstates traffic. Check the schema for a status or level field and count only rows that carry one.

ParametersJSON Schema
NameRequiredDescriptionDefault
endNoWindow end. Default now.
sqlYese.g. SELECT level, count(*) AS n FROM my_stream GROUP BY level ORDER BY n DESC
sizeNoMaximum rows to return. Default 50, maximum 1000.
startNoWindow start. Default -24h.
max_field_charsNoTruncate long string fields to this length. Default 400.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full load and does well: it discloses the 24h default window, that `_timestamp` is microseconds, histogram bucketing, and the non-obvious pitfall that log shippers emit multiple rows per request so count(*) overstates traffic. It stops short of stating whether non-SELECT statements are permitted or what permissions/limits apply, which matters for a free-form SQL tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, followed by timestamp formats, routing guidance, and a query pitfall — each sentence earns its place. No filler or restatement of the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description should ideally hint at the returned row shape, which it only gestures at ('matching rows'). For a free-form SQL tool with 100% schema coverage on inputs, everything an agent needs to invoke it correctly is present, but return-shape detail is thin.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning the schema lacks: it enumerates the accepted formats for `start`/`end` (ISO, plain date, epoch s/ms/µs, relative offsets, 'now') and confirms the 24h default. That is genuine value beyond the terse schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Run SQL against an OpenObserve stream and return the matching rows') and clarifies that the stream name acts as the FROM target. It also names the two sibling tools it relates to, so an agent can place it without opening another definition.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit routing rules: call StreamList if you don't know what exists, and StreamSchema before querying an unseen stream, with the reason (field names differ and guessing wastes a round trip). This tells the agent both when to use this tool and when to call its siblings first.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

StreamListA

List OpenObserve streams with document counts, stored size and the newest document timestamp. Start here when you do not know what data is available.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the burden. It discloses the payload shape (counts, size, freshness timestamp) and the read-only nature implied by "List," which is meaningful context. It does not mention auth needs, result limits, or pagination, leaving some behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, both load-bearing: the first defines the payload, the second the usage trigger. The most decision-relevant content (what is returned) is front-loaded with zero filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description must characterize returns, and it does list the three key fields. It omits ordering, volume/limit behavior, and how results relate to StreamSchema output, which are minor gaps for a zero-parameter discovery tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no parameter semantics to document and the baseline is 4. The description correctly does not invent parameters, and the schema closes the object with additionalProperties=false.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb+resource ("List OpenObserve streams") and enumerates the returned fields (document counts, stored size, newest document timestamp). Sibling StreamSchema and SearchSQL are not named, so the agent must infer the distinction from the "start here" framing rather than an explicit contrast.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Start here when you do not know what data is available" gives a concrete entry-point condition for selecting this tool over exploration via SearchSQL. It stops short of naming alternatives or stating exclusions, so it is clear context without full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

StreamSchemaA

Field names and types for one OpenObserve stream. Call this before writing SQL against a stream you have not queried before.

ParametersJSON Schema
NameRequiredDescriptionDefault
streamYesStream name, as StreamList reports it.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It discloses what the tool returns (field names and types) and the recommended call ordering, which is useful, but says nothing about permissions, error behavior when a stream doesn't exist, or cost.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero waste, with the resource identified first and the usage rule trailing it. Nothing needs trimming.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter read tool with no output schema, the description covers the returned content and the call ordering. It stops short of documenting failure modes or whether results are cached, which are minor gaps here.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and there is a single required parameter already documented as 'Stream name, as StreamList reports it.' The description adds only the implicit notion of 'one stream' and no format or naming detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific resource and output: field names and types for one OpenObserve stream. It clearly differs from StreamList (enumerating streams) and SearchSQL (executing queries), though it never names those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit conditional guidance: call this before writing SQL against a stream you have not queried before. That is a usable when-to-use rule, but it doesn't name the alternative tools or state when this call is unnecessary.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv1.0.1
    • First observedSearchSQL
    • First observedStreamList
    • First observedStreamSchema

TDQS

A4.1/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: listing streams, retrieving a schema, and executing SQL queries. The descriptions explicitly guide the correct order of use, eliminating ambiguity.

Naming Consistency3/5

All names use PascalCase without separators, but the structural pattern is inconsistent: StreamList and StreamSchema are noun-first, while SearchSQL is verb-first. This mixed verb/noun style reduces predictability, though names remain readable.

Tool Count5/5

Three tools cover the essential workflow of discovering streams, inspecting schemas, and querying data. This minimal, focused set avoids redundancy and each tool earns its place.

Completeness4/5

The core query lifecycle is well covered: list streams, get schema, run SQL. Minor gaps exist, such as no tool for cross-stream queries or saved query management, but agents can work around these with the provided SQL interface.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables querying log data stored in SQLite databases through the Model Context Protocol, allowing natural language interactions with log analysis.
    -
  • F
    license
    A
    quality
    D
    maintenance
    Enables querying OpenTelemetry logs stored in OpenSearch across development and production environments. Provides tools for searching logs by various criteria including free-text Lucene queries, trace IDs, service names, error levels, and specific fields.
    8
    -
  • A
    license
    A
    quality
    C
    maintenance
    Enables read-only querying of Grafana Loki logs through tools like label discovery, LogQL queries, and recent log retrieval, without requiring knowledge of the Loki API.
    4
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Enables AI agents to query OriginChain's telemetry agent for log patterns, semantic incident analysis, distributed traces, service graphs, and read-only SQL queries over telemetry data.
    10
    4 npm
    MIT