graylog-mcp
Read-only integration with a Graylog (4.x–7.x) server for log investigation and incident analysis. Provides tools to search and count logs, follow a correlation/trace id across services, get messages and surrounding context, group errors with exact counts, build log histograms and top-value rankings, compare two time periods, list streams and fields, run configured query presets, and perform root cause analysis that ranks likely origins by comparing services' errors, traffic and latency against a baseline, pinning first-error timestamps, detecting deploys/restarts and inferring a service call graph from traces. Sensitive data (credentials, tokens, emails, card numbers, etc.) is masked before logs are returned.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@graylog-mcpfind the root cause of the checkout 500s in the last hour"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
graylog-mcp
A read-only Model Context Protocol server for Graylog 4.x to 7.x. It lets an LLM investigate incidents on its own: search logs, follow one request across services, group errors with exact counts, and find when a problem started. Sensitive data is masked before any log line leaves the machine.
Everything site-specific (trace field names, redaction rules, timezone, application packages, error query) lives in configuration, not in code.
How it differs from other Graylog MCP servers
Typical community server | graylog-mcp | |
Graylog versions | one range only (4.x/5.0 or 5.2+) | detects the version, works on 4.x through 7.x |
Statistics | raw messages, or counting over a sample | exact counts and groupings computed by Graylog |
Safety | no masking | masks by field name, regex and Luhn; per-country packs |
LLM context | dumps raw log lines | folds stack traces, groups repeated lines, hard size cap |
Writes to Graylog | some create saved searches | never writes anything |
Related MCP server: quickwit-mcp
Quick start
One command (needs uv), run in the folder holding your .graylog-mcp.toml or
anywhere for a fresh start:
uvx --from git+https://github.com/ntbang0901/graylog-mcp graylog-mcp startIt installs graylog-mcp, runs it in the background as one process for every Claude session, starts it again when you log in, connects Claude Code in every project and opens the admin page. The page's Setup checklist shows what is left (usually: add your Graylog URL and a token) with a button for each step.
Afterwards:
| open the admin page again |
| is it running, which version, is Claude Code connected |
| install the latest version and restart (or "Check for updates" on the page) |
| stop it and no longer start it at login |
Prefer the terminal? graylog-mcp init is a guided setup that writes .graylog-mcp.toml and registers the
server per repository (one process per session). See Setup helpers.
Or by hand:
Two environment variables are enough:
export GRAYLOG_URL=https://graylog.example.com
export GRAYLOG_TOKEN=<access token of a read-only user>
uvx graylog-mcp --check # connects, prints the detected version and APIs, exitsClaude Code:
claude mcp add graylog --env GRAYLOG_URL=https://graylog.example.com --env GRAYLOG_TOKEN=... -- uvx graylog-mcpClaude Desktop / any MCP client (mcpServers JSON):
{
"mcpServers": {
"graylog": {
"command": "uvx",
"args": ["graylog-mcp"],
"env": {
"GRAYLOG_URL": "https://graylog.example.com",
"GRAYLOG_TOKEN": "...",
"GRAYLOG_TIMEZONE": "Asia/Ho_Chi_Minh",
"GRAYLOG_REDACTION_PACKS": "vn"
}
}
}
}The Graylog user
Create a dedicated user with the Reader role and read access to the streams the model may see,
then generate an access token for it. The server only sends GET requests, plus POST to the four
search endpoints that execute or check a query without saving it (/views/search/sync, /search/messages,
/search/aggregate, /search/validate). Any other method or path is refused inside the client before a request is built.
Tools
All tools are annotated readOnlyHint and accept an optional instance.
Scan
scan— "is anything wrong?" in one call. Runs every scan rule concurrently (crashes, resource exhaustion, error spikes, new error types, HTTP 5xx, connectivity, database, auth failures, plus your own) with exact counts against a baseline, and returns only what fired, most severe first, with its query, top groups and a sample. Ad hoc rules (checks) cover a specific request in the same call. See Scan rules.list_scan_rules— the rules with their query, condition, severity and tags.
Root cause analysis
root_cause— "what broke first, and why?" in one call. Compares every service's errors, traffic and latency with a baseline window, pins each service's first error to the millisecond, detects deploys in the logs, infers the call graph from traces, and returns a ranked verdict with a timeline and evidence.detect_changes— deploys and restarts found in the logs themselves: a new value of a version field (app_version,build,commit, ...), a host rollout (new sources replacing old ones), start/stop lines. No CI/CD integration needed.service_map— which service calls which, inferred from sampled traces with no configuration: edges with traffic, error rate and p50/p95 latency.
Search
search_logs— Lucene query, relative (15m,2h) or absolute time (instance timezone or ISO 8601), streams by name or id, field selection, sort, paging; repeated lines grouped by default.count_logs— exact number of matching messages.get_message— one message with all fields, byindex/idref.
Investigate
trace_request— follows a correlation/request/trace id through the configuredtrace_fieldson all streams (falls back to full text); returns a cross-service timeline, per-service steps with durations and the first error.context_around— messages within ±N seconds of a message, limited to the same source, the same streams or everything.error_summary— groups errors by exception, logger, source or any field: exact count, first and last seen, one sample message per group.log_histogram— counts over time with automatic interval, peak bucket and the onset of a spike.top_values— top N values of a field with exact counts.compare_periods— two periods (e.g. before/after a deploy): which error groups are new, grew, disappeared or shrank, normalised per hour.
Discover
list_streams,list_fields— so the model writes queries with real names.list_presets,run_preset— named queries you define in the config.list_instances— instances with detected version and the API in use.
The server also sends the model instructions about Lucene syntax and a suggested investigation flow.
Scan rules
scan answers "is anything wrong?" or "scan X" with one call. Each rule is a Lucene query plus a condition, and
fires when:
Condition | Fires when | Use it for |
| more than N matches in the window ( | things that should never happen: crashes, OOM, data corruption |
| the rate is X times the baseline's (or the baseline had none), with at least M matches, and the rise is significant | things that always happen a little: timeouts, 5xx, auth failures |
| a value of the field (exception, logger, ...) appears that the baseline never saw, at least | new kinds of errors after a deploy |
What keeps false alarms out:
Share of traffic (
per_traffic = true): the rule compares matches / traffic instead of matches per hour, so errors that triple because a sale tripled the traffic stay quiet, and a rise from 1% to 10% of a busy hour fires. Traffic is every message in the scan's scope, ortraffic_query(e.g.path:/checkoutfor checkout errors). The built-in error, 5xx, connectivity and database rules use it.Seasonal baseline (
baseline_shift = "1d"or"7d",baseline_periods = 3): the window is compared with the same window one day (week) earlier, three times; the period with the median rate is the reference, so the morning peak is compared with yesterday's morning peak, and one bad day among the three changes nothing. Periods without any data (older than the index retention) are dropped and the finding says so (baseline_note). By default the baseline is the period right before the window, as long as the window, orbaseline = "24h".Significance: a growth fires only if it is unlikely to be chance, by an exact conditional binomial test of the window's count against the reference period's (the standard test comparing two Poisson rates). 3 errors then 7 (x2.3) stays quiet with a
notesaying so; 400 then 600 (x1.5, withgrowth = 1.4) fires. Each growth result carries itsconfidence; the rule'sconfidence(default 0.99) is the bar.
Every rule costs two exact counts, plus one per extra baseline period and for traffic; identical counts are made
once per scan (rules share the error query, the traffic) and run concurrently (limits.scan_concurrency). Groups
and one sample are fetched only for rules that fired or need them to decide, and a sample shown by one finding is
not repeated by the next. The result has findings (fired, most severe first), quiet (checked and normal, with
counts and trend) and skipped (could not be checked, e.g. a field these logs do not have).
Built-in rules: crash and resource_exhaustion (critical, any match); error_spike (share of errors x2),
new_error_types (exception unseen in 24 h), http_5xx (share x2, only where a status field exists),
connectivity (timeouts, refused/reset connections, DNS; share x3) and database (deadlocks, lock timeouts,
too many connections; share x3), all high; auth_failures (rate x3, at least 20), medium.
[scan]
disable = ["auth_failures"] # built-in rules to turn off
exclude = 'logger_name:HealthCheck OR "GET /health"' # noise dropped from every rule
[scan.rules.connectivity] # a built-in: only the keys you set change
min_count = 30
[scan.rules.card_declined]
description = "Bank declines"
query = 'service:payment AND message:"declined"'
severity = "high" # critical | high | medium | low
growth = 2.0
min_count = 20
per_traffic = true # compare declines / payment traffic, not declines per hour
traffic_query = "service:payment"
baseline_shift = "7d" # against the same hour of the last 3 weeks (payday, weekends)
group_by = "bank_code" # top groups shown with the finding
tags = ["payment"]
[scan.rules.ledger_mismatch]
query = '"ledger mismatch"' # no condition: any match fires
severity = "critical"
instances = ["payment/prod"] # instances, groups or environments; everywhere when omitted
[scan.rules.consumer_lag]
query = "consumer_lag:>10000"
requires = ["consumer_lag"] # skipped (and said so) where the field does not existOther keys: errors_only = true ANDs the instance's error_query, exclude drops noise for one rule, baseline
sets the rule's own baseline length, baseline_periods the number of shifted periods, confidence the
significance bar. A rule's own baseline wins over the call's. The config is validated at startup like the rest (unknown keys, a rule that
would fire on all traffic, new_groups without group_by).
Selecting what to run: scan(rules=["payment"]) takes rule names or tags, min_severity="high" leaves out
lower rules before any request, and checks=[{"name": "declined", "query": "message:declined", "growth": 2}]
adds rules written for the request at hand (same keys as above). A preset with tool = "scan" saves a scan
the team runs often, and the scan MCP prompt (/mcp__graylog__scan in Claude Code) runs scan, verifies each
finding and reports them as a table.
Writing rules that are fast and accurate:
Match on fields when you have them (
http_status:[500 TO 599],exception_class:...); otherwise quoted phrases ("Connection refused"). Avoid leading wildcards and regex: they scan every term in the index.Pick the condition from how often the event normally happens: never, then
threshold = 0; always a little, thengrowth(the significance test handles noise;min_countonly sets the smallest count worth a look).Anything that grows with traffic gets
per_traffic = true; traffic with a daily or weekly curve getsbaseline_shift.Drop known noise with
exclude(per rule or under[scan]) rather than raising thresholds.Use a longer baseline (
baseline = "24h") for rules on spiky or low-volume traffic, andrequiresfor rules that depend on a field only some systems log.Give each rule a severity that matches who should be woken up, and tags that match how people ask ("payment", "security", "dependencies").
Root cause analysis
Example verdict, from the scenario in the integration suite (a bad deploy of payment), identical on
Graylog 4.3, 5.0, 5.2, 6.1 and 7.0:
Most likely origin: payment (confidence high). errors began at 08:42:49.756, 182.5s after payment deployed 1.3.9 -> 1.4.0 (hosts pay-1, pay-2 -> pay-3, pay-4). Also affected: gateway errors (+50ms), bank-adapter traffic drop.
How it gets there:
Signals per service. Two pivots (errors, and traffic with average latency) split by service over the incident window plus a baseline window right before it.
Onsets. For each service, the first interval that is clearly abnormal against its own baseline (median and MAD, so one noisy minute in the baseline does not hide anything) and stays abnormal: error rise, traffic drop or spike, latency rise. Error onsets are then refined to the exact first error message.
Changes. Version fields, host rollouts and start/stop lines (see
detect_changes) in the window and the 24 hours before it. A service that merely went silent is reported as a traffic drop, not a change.Call graph. Traces sampled per service (and from failing requests) give caller -> callee edges.
Ranking. Earliest onset (exact timestamps break ties inside the first interval), a change shortly before the onset, callers failing after it, and a dependency failing before it (which points further down) all move the score. The result lists the reasons for each candidate, a merged timeline, the first error message with its
ref, and the next calls to verify the hypothesis.
It is a ranked hypothesis with its evidence, not a certainty. Field names come from config:
service_fields, trace_fields, version_fields, latency_fields, change_query.
Version detection
At startup the server reads GET /api/system (falling back to GET /api/ when the token may not read
system info), picks the APIs below and caches the result; list_instances shows the choice. An API that
answers 404/405/501 is dropped and the next one is used.
Version | Messages | Aggregations |
4.x – 5.1 |
|
|
5.2 – 6.x | universal → views → Scripting API |
|
7.x | same as 5.2+ | same as 5.2+ |
Pivot requests use "field": "x" on 4.x and "fields": [...] from 5.0. Histograms always use a views
pivot with a time grouping. Query syntax errors and zero-result queries are checked with
POST /search/validate so the model gets the actual problem ("incomplete query", "unknown field: sevrity")
instead of a bare "all shards failed". Graylog 7 ships its own MCP endpoint; this server is still useful
there for masking and compact output.
Every tool, each API path forced on its own (universal, views, Scripting API), the read-only guarantee and the no-leak check pass against real containers of:
Graylog | Search backend | Result |
4.3.15 | OpenSearch 1.3 | pass |
5.0.13 | OpenSearch 2.15 | pass |
5.2.12 | OpenSearch 2.15 | pass |
6.1.16 | OpenSearch 2.15 | pass |
7.0.13 | OpenSearch 2.15 | pass |
Behaviour observed on real servers and handled: 5.x+ pivots return an (Empty Value) bucket for documents
without the grouped field (reported separately as without_field); the Scripting API returns "-" for
missing fields and no message index; invalid queries come back as HTTP 500 from universal search on
4.x-5.x; universal search is still present in 7.0.
Output shaping
Masking (before anything else):
always on: sensitive field names (
password,token,secret,authorization,cookie, ...), emails,Bearer/Basiccredentials, JWTs,key=value/"key": "value"secrets, credentials in URLs, PEM private keys, card numbers that pass the Luhn check;packs enabled in config:
vn(phone numbers, 12-digit CCCD, optional 9-digit CMND),us(SSN),eu(IBAN with checksum),uk(NINO),in(Aadhaar, PAN);your own regexes, plus an allow-list and excluded field names to avoid false positives. Grouping keys (
top_values,error_summary) are masked too.
Stack traces (Java/Kotlin, Python, .NET, Go, Node): exception lines,
Caused byblocks and frames fromapp_packagesare kept; the rest becomes… N frames. Withoutapp_packages, the first N frames are kept (the last N for Python).Repeated lines: numbers, UUIDs, hex, IPs and timestamps are normalised and identical templates grouped with
count,first,lastand one sample.Size: per-value and per-call character caps, with
truncatedandnext_offset.Time: shown in the configured timezone with its offset. Each line carries
ref: index/idfor follow-up calls.Errors: clear messages for 401, 403, query syntax errors (with position), timeouts, unknown streams (with suggestions) and unsupported versions.
Configuration
Environment variables:
Variable | Purpose |
| single instance with token auth |
| basic auth instead of a token |
| TLS and proxy for the env-only instance |
| display timezone and zone for naive times (default UTC) |
| e.g. |
| e.g. |
| path of a TOML config file |
| bearer token required by the HTTP transport |
A TOML file (--config or GRAYLOG_MCP_CONFIG, else .graylog-mcp.toml in the current directory or a
parent up to the repository root, else ~/.config/graylog-mcp/config.toml) covers the
rest: several instances, token or basic auth, TLS verification and CA bundle, proxy, timezone,
trace_fields, error_query (because level is a syslog number or a string depending on how logs are
shipped), app_packages, redaction packs and custom patterns, presets and limits. Secrets are never
written in the file; it names the environment variable that holds them (token_env, password_env).
The file is validated at startup and any mistake (unknown key, bad regex, unknown timezone, missing
secret) stops the server with a precise message.
See examples/config.toml for every option.
One repository, several environments
Commit a .graylog-mcp.toml at the root of your application repository with one instance per environment.
The server finds it from the current directory or any parent up to the repository root, so everyone working
in the repo gets the same setup. Tokens stay out of the file: each instance names its environment variable.
# .graylog-mcp.toml
default_instance = "staging" # what tools use when no environment is named
timezone = "Asia/Ho_Chi_Minh"
[redaction]
packs = ["vn"]
[investigation]
trace_fields = ["traceId", "X-Request-ID"]
[instances.dev]
url = "https://graylog-dev.example.com"
token_env = "GRAYLOG_DEV_TOKEN"
description = "Development cluster, noisy, debug logs on"
[instances.staging]
url = "https://graylog-staging.example.com"
token_env = "GRAYLOG_STAGING_TOKEN"
description = "Staging, deployed on every merge to main"
[instances.prod]
url = "https://graylog.example.com"
token_env = "GRAYLOG_PROD_TOKEN"
description = "Production"
error_query = "level:<=3 AND NOT logger_name:healthcheck" # any key can differ per environmentRegister the server once for the project (Claude Code reads .mcp.json at the repo root and expands
${VAR} from each developer's environment):
{
"mcpServers": {
"graylog": {
"command": "uvx",
"args": ["--from", "git+https://github.com/ntbang0901/graylog-mcp", "graylog-mcp"],
"env": {
"GRAYLOG_DEV_TOKEN": "${GRAYLOG_DEV_TOKEN}",
"GRAYLOG_STAGING_TOKEN": "${GRAYLOG_STAGING_TOKEN}",
"GRAYLOG_PROD_TOKEN": "${GRAYLOG_PROD_TOKEN}"
}
}
}
}Then ask in plain words: "why is checkout failing on staging?", "compare errors on prod before and
after 14:00". Every tool takes instance, list_instances shows each environment with its description
and detected Graylog version (environments may run different versions), and results always name the
instance they come from. A developer who has no token for an environment simply gets a clear error for that
one; the others keep working.
Clients that do not start the server inside the repository (Claude Desktop) need the path explicitly:
GRAYLOG_MCP_CONFIG=/path/to/repo/.graylog-mcp.toml.
Setup helpers and admin UI
Command | What it does |
| Guided setup: environments, connection test, field detection, config file, client registration. |
| Asks for each missing token/password, tests it, and saves it for this user in |
| Checks every environment (token set, reachable, TLS, version, readable streams, data, configured fields exist, error query matches, redaction) and prints a fix for each problem. Exit code 1 on failure, |
| Suggests |
| Registers the server in |
| Local admin web UI (below). |
graylog-mcp ui opens a page on 127.0.0.1 with:
Overview: environment cards and the doctor checks with fixes;
Environments: add, edit, delete, set default; type the token or password once, test the connection, and save (the secret goes to your per-user secrets file, the variable name is chosen for you);
Field mapping: run detection on an environment, review the evidence, apply the selected settings;
Redaction: toggle country packs, add custom patterns and allow-list entries, and see live which rules mask your sample text;
Playground: run any tool against any environment and see the exact output, its size and an approximate token count;
Connect clients: ready-to-copy snippets and one-click install for each client;
Config file: edit the TOML with validation and an automatic backup.
The UI listens on loopback only, checks the Host header, and every API call needs the random token in the
link printed at startup. It writes the config file and client configs when you ask. A token or password
entered in the form is saved, when you tick "Save it on this machine", to the same per-user secrets file as
graylog-mcp login, never to the config file.
*_env settings hold the name of a variable (GRAYLOG_PROD_TOKEN), never the secret: a value that is not
a valid variable name is refused, and never echoed back, since it is most likely a secret typed into the wrong
field.
Groups and your own environments
Large organisations often run one Graylog per system and per environment: ERP, CXP, PAYMENT... each with its own dev/uat/prod (or sandbox, dr, prod-eu...). Environment names are yours; nothing assumes dev/staging/prod, and each group can have a different set.
# graylog-org.toml: one file for the whole company, e.g. in a shared platform repository
timezone = "Asia/Ho_Chi_Minh"
default_environment = "uat" # used when a question names only the system
[environments.prod] # declare environments once; every group inherits these settings
description = "Production"
error_query = "level:<=2"
ca_bundle = "/etc/ssl/corp-ca.pem"
[environments.uat]
description = "User acceptance"
[groups.erp]
description = "ERP"
trace_fields = ["correlationId"] # anything set on a group applies to all its environments
[groups.erp.environments.uat]
url = "https://graylog-erp-uat.corp"
token_env = "GRAYLOG_ERP_UAT_TOKEN"
[groups.erp.environments.prod]
url = "https://graylog-erp.corp"
token_env = "GRAYLOG_ERP_PROD_TOKEN"
[groups.payment]
description = "Payment platform"
default_environment = "sandbox"
[groups.payment.environments.sandbox]
url = "https://graylog-pay-sbx.corp"
token_env = "GRAYLOG_PAYMENT_SANDBOX_TOKEN"
[groups.payment.environments.prod]
url = "https://graylog-pay.corp"
token_env = "GRAYLOG_PAYMENT_PROD_TOKEN"Settings are layered: [environments.<env>] < [groups.<group>] < [groups.<group>.environments.<env>].
Each pair becomes an instance named <group>/<environment>.
A service repository then only says which group it belongs to:
# payment-api/.graylog-mcp.toml
include = "../platform/graylog-org.toml" # relative to this file; a list is allowed
default_group = "payment"
only_groups = "payment" # load only this group here (a list, or "*" for every group)How the model (and you) pick an instance, in every tool's instance argument:
You say | Instance |
|
|
| the group's |
|
|
nothing |
|
list_instances returns the groups with their environments and descriptions, graylog-mcp init asks for
groups first, and the admin UI shows environments by group (instances from an included file are marked and
edited in that file).
Repositories of a group
List the repositories each group serves, as local folders (paths relative to the config file, ~ allowed) or
git remotes (git@gitlab.corp:f88/payment-api.git, f88/payment-api or just payment-api):
[groups.payment]
repos = ["~/code/payment-api", "../payment-worker", "f88/payment-gateway"]When the server runs inside one of them (a folder or any subfolder, or a clone whose origin matches), it
loads only that group: "errors on prod" means that system's production, and the other groups' Graylog
servers are not reachable from that repository. Asking for one returns an error that says why and how to
enable it. list_instances shows the repository and the scope; graylog-mcp doctor shows which repository
was recognised.
To reach more groups from a repository, set only_groups = ["payment", "erp"] (or "*" for all) in its
.graylog-mcp.toml, or GRAYLOG_MCP_GROUPS=payment,erp in the MCP client's environment. only_groups also
works without repositories, e.g. in a personal config. The admin UI always shows every group.
Manage them in the admin UI (Environments > Groups > Repositories: add a local folder, see its git remote and
whether it is set up, "Set up" writes .graylog-mcp.toml with an include of the shared file and
only_groups, so the scope holds on every machine whatever its folder layout, and registers
Claude Code there) or from the command line:
graylog-mcp repo --config ../platform/graylog-org.toml add payment # current folder; also sets it up
graylog-mcp repo list
graylog-mcp repo remove payment ~/code/payment-apiFocus: search this repository's service by default
Inside a repository, "errors in the last hour" means that service's errors. search_logs, count_logs,
error_summary, log_histogram, top_values, compare_periods and detect_changes add a filter on the
service field and say so in their result (focus). Other services and streams are searched only when asked for:
streams=["*"] (everything), explicit streams, or a query that names a service field
(application:"other-service"). trace_request, service_map and root_cause always span every service.
Each repository sets its focus in its own .graylog-mcp.toml:
[focus]
service = "cobra-mdm-service" # value of the service field; a list for several; false = no service filter
streams = ["MDM"] # optional: streams searched by default
# field = "application" # optional: the service field (default: the first of service_fields present)Without [focus], the service is guessed from the repository name (folder or git remote) by matching it against
the values of the service fields (service, application, ...) over the last 24 hours: cobra-mdm finds
cobra-mdm-service. No clear match, no filter; list_instances shows what was picked. Set it from the admin UI
(Groups & repositories > Focus), with graylog-mcp repo focus cobra-mdm-service --streams MDM inside the
repository, or for one MCP client with GRAYLOG_MCP_SERVICE=<name> (off disables it).
Settings by scope
The admin UI's Settings page edits field names, queries and connection options for a chosen scope: global, one environment (for every group), one group, or a single instance. Each field shows the value currently in effect; leaving it empty inherits it.
One process for every session
With stdio, the client starts a server process for every session: ten Claude Code sessions are ten
Python processes, each with its own connections and version detection. graylog-mcp start runs one server
for all of them instead, and still answers each session with the configuration of the repository it works
in (its .graylog-mcp.toml, group, focus), as if it had been started there.
What start does, so you know what is on your machine:
A stable copy:
uv tool install(in~/.local/share/uv/tools), not uvx's temporary cache.updatereinstalls it from the same source; the version shown everywhere includes the git commit.The server:
graylog-mcp serve --shared --admin --port 8000 --config <file>on 127.0.0.1, started at login by a LaunchAgent (macOS), a systemd user unit (Linux) or the Startup folder (Windows); elsewhere it runs until you log out. Log:~/Library/Logs/graylog-mcp.logor~/.config/graylog-mcp/server.log.The admin page at
http://127.0.0.1:8000/admin/, with the access token kept in~/.config/graylog-mcp/admin-token(owner only);graylog-mcp uiopens it.Claude Code, registered once for every project (user scope):
claude mcp add --transport http graylog --scope user http://127.0.0.1:8000/mcp --header 'X-Graylog-MCP-Repo: ${PWD:-}'. Claude Code fills in the folder it runs in; the server walks up to the repository. A project whose committed.mcp.jsonstill starts graylog-mcp itself gets a private override (local scope) instead: the file stays as it is for your teammates.
How the server picks the configuration of a session:
The repository's own
.graylog-mcp.toml; without one, the--configfile (startuses the one found where you ran it, else~/.config/graylog-mcp/config.toml), with the group matched from itsrepos.Edits (admin page,
login,repo focus) apply within seconds, without a restart.One Graylog client per instance, shared by every repository using it: one connection pool, one version detection, one stream and field cache.
Tokens come from the server's environment or
graylog-mcp login(the admin page saves them there too).Other clients: Cursor and VS Code send
${workspaceFolder}(graylog-mcp install cursor --shared), any client can add?repo=<folder>to the URL. Claude Desktop already runs one server for all its chats.
Shared HTTP server and Docker
stdio is the default. For one server shared by a team, use streamable HTTP with its own bearer token:
GRAYLOG_MCP_HTTP_TOKEN=$(openssl rand -hex 32) graylog-mcp --transport streamable-http --host 0.0.0.0 --port 8000The server refuses to listen on a non-loopback address without a token (unless --allow-no-auth is
given, e.g. behind an authenticating proxy). /healthz is open; the MCP endpoint is /mcp.
docker build -t graylog-mcp .
docker run -p 8000:8000 \
-e GRAYLOG_URL=https://graylog.example.com -e GRAYLOG_TOKEN=... \
-e GRAYLOG_MCP_HTTP_TOKEN=... graylog-mcpTesting
uv sync
uv run pytest # unit + contract tests (no network), coverage >= 85%
uv run ruff check && uv run ruff format --check && uv run mypyUnit tests: masking (including values that must not be masked), stack traces in five languages, line grouping, time parsing, pivot building and parsing, config validation, HTTP auth.
Contract tests: every tool against an in-memory Graylog (
tests/fake_graylog.py, served throughhttpx.MockTransport) that reproduces the response shapes of 4.3, 5.0, 5.2, 6.1 and 7.0, with planted sensitive values that must never reach the output, an assertion that no request could change state, and a size budget per call.Integration tests:
tests/integration/docker-compose.ymlstarts Graylog 4.3, 5.0, 5.2, 6.1 and 7.0, each with MongoDB and OpenSearch;seed.pyships the same GELF dataset, then every tool runs against it:tests/integration/run.sh v50 # or v43 v52 v61 v70CI runs the whole matrix weekly, on demand, and on pull requests that change the backends.
CI on every push: ruff, mypy, tests on Linux (Python 3.11-3.13), Windows and macOS, wheel build with metadata check and a smoke test in a clean environment, and a Docker build with an HTTP smoke test. CodeQL scans the code and workflows; Dependabot keeps dependencies and actions current. See CONTRIBUTING.md and SECURITY.md.
Project layout
src/graylog_mcp/
config.py environment + TOML, validated at startup (fail fast)
client.py async httpx client: auth, TLS, proxy, read-only guard, HTTP error mapping
backends/ version detection and API selection: universal.py, views.py, scripting.py
timerange.py time parsing, display, interval selection
redact.py core rules, country packs, Luhn / IBAN checks
shaping.py field selection, truncation, stack traces, line grouping, output budget
tools.py tool implementations
rca.py root cause analysis: onsets, change detection, service map, ranking
server.py MCP tool declarations and model instructionsLicense
MIT
Available Tools
17 toolscompare_periodsCompare periodsARead-onlyIdempotent
Compare two periods (default: last window vs the one before; or around split_at). Lists groups that are new, increased, gone or decreased, normalised per hour. Defaults to errors only.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Groups to return | |
| query | No | Extra Lucene filter | |
| window | No | Length of each period when using split_at or the default | 1h |
| streams | No | Stream titles or ids to search in; all streams when omitted | |
| group_by | No | 'exception', 'logger', 'source' or any field | exception |
| instance | No | Graylog instance (environment) from list_instances, e.g. 'staging' or 'prod'; the default instance when omitted | |
| split_at | No | Point in time (e.g. a deploy); compares [split-window, split] with [split, split+window] | |
| current_to | No | Explicit current period end; default now | |
| baseline_to | No | Explicit baseline end | |
| errors_only | No | AND the configured error query (default true) | |
| current_from | No | Explicit current period start | |
| baseline_from | No | Explicit baseline start |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so safety is covered; the description adds real value by disclosing that results are normalised per hour, bucketed into new/increased/gone/decreased groups, and filtered to errors by default.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences that front-load the core operation, then the split_at variant, then the output categories and default filter. No filler, though the parenthetical could be slightly cleaner.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter, no-output-schema tool, the description covers the default time basis, the split_at alternative, the grouping output, normalisation, and the errors-only default. It leaves period-length details to the schema, which is reasonable given 100% coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all 12 parameters including split_at, window, and errors_only. The description restates the default window comparison and split_at semantics, which reinforces but does not meaningfully extend the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (compare) and resource (two periods of grouped log data), and clarifies the default comparison basis and the split_at variant. It does not name the closest sibling (detect_changes or error_summary), so an agent must still infer when to prefer this over those.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains the default behavior (last window vs the one before, or around split_at) and the errors-only default, which implies typical usage, but it never says when to choose this tool over detect_changes, error_summary, or log_histogram.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
context_aroundContext aroundCRead-onlyIdempotent
Messages logged within ±N seconds of a given message.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | 'index/message_id' of the anchor message | |
| limit | No | Max messages (split before/after) | |
| query | No | Extra Lucene filter inside the window | |
| scope | No | 'source' (same host/app, default), 'stream' (same streams) or 'all' | source |
| seconds | No | Window before and after the message | |
| instance | No | Graylog instance (environment) from list_instances, e.g. 'staging' or 'prod'; the default instance when omitted |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so the safety profile is covered structurally. The description adds nothing beyond that: no note on how the window is measured, what happens at log boundaries, result ordering, or the effect of the scope/query narrowing. For an open-world tool it under-discloses.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with zero filler and the core scoping concept ('±N seconds') front-loaded. Its brevity is arguably under-specification rather than good economy, but there is no wasted text to trim.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With six parameters, no output schema and an open-world read tool, the description leaves the agent without any guidance on how scope/query/instance interact or how to narrow the window sensibly. The schema covers individual params, but the description contributes no operational context for a moderately complex retrieval tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so scope, query, limit, instance and seconds are already documented in the schema, and the baseline is 3. The description's '±N seconds' does add the useful symmetry detail (seconds applies both before and after the anchor), but nothing else is elaborated beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The fragment 'Messages logged within ±N seconds of a given message' does convey a specific resource and a temporal-scope operation, so an agent can infer it fetches messages near an anchor. However, it is a bare noun phrase with no verb and no differentiation from siblings like get_message or search_logs, which also retrieve messages. Purpose is identifiable but not sharp.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance at all — nothing says this is for pulling surrounding traffic around a known message versus using get_message for one message or search_logs for a query. The '±N seconds' framing only weakly implies a 'when you need surrounding context' scenario. No alternatives or exclusions are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
count_logsCount logsARead-onlyIdempotent
Exact number of messages matching a query. Cheap; use it to size a problem.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Lucene query, '*' for everything | * |
| range | No | Relative range ending now (or at to_time): '15m', '2h', '1d', '1h30m' | |
| streams | No | Stream titles or ids to search in; all streams when omitted | |
| to_time | No | Absolute end, same formats as from_time; default now | |
| instance | No | Graylog instance (environment) from list_instances, e.g. 'staging' or 'prod'; the default instance when omitted | |
| from_time | No | Absolute start: ISO 8601 or 'YYYY-MM-DD HH:MM' in the instance timezone; overrides range |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description's only added behavioral claim is 'Cheap', which is a genuine cost signal not present in the structured fields, but it says nothing about count caps, accuracy limits, or whether the count is bounded.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero filler, and the returned value is front-loaded ahead of the cost hint. Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameter-rich but conceptually simple read tool with a full-coverage schema, a complete annotation safety profile and a stated return value, the description is adequate. It omits only what to do with the count or how it compares to the histogram, which is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% across all six parameters, so query syntax, range formats, streams, instance and the from_time-overrides-range rule are all documented in the schema. The description adds no parameter detail, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: returns the exact number of messages matching a query. The word 'Exact' usefully contrasts with the histogram sibling, but no sibling is named, so an agent still has to infer the boundary with log_histogram or search_logs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'use it to size a problem' implies a usage pattern (cheap pre-flight sizing before heavier queries) but gives no explicit when-not guidance and never names the alternatives (search_logs, log_histogram) that an agent should pick when it needs the messages themselves rather than a tally.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
detect_changesDetect changesBRead-onlyIdempotent
Deploys and restarts found in the logs themselves: new values of version fields (app_version, build, commit...), host rollouts (new sources replacing old ones), and start/stop lines. No CI/CD integration needed.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Optional Lucene filter for scope | |
| range | No | Relative range ending now (or at to_time): '15m', '2h', '1d', '1h30m' | 6h |
| streams | No | Stream titles or ids to search in; all streams when omitted | |
| to_time | No | Absolute end, same formats as from_time; default now | |
| instance | No | Graylog instance (environment) from list_instances, e.g. 'staging' or 'prod'; the default instance when omitted | |
| from_time | No | Absolute start: ISO 8601 or 'YYYY-MM-DD HH:MM' in the instance timezone; overrides range |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint false, and openWorldHint, covering the safety profile. The description adds that detection is based on log patterns themselves (no CI/CD), which is useful behavioral context but does not cover return format or performance.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that efficiently conveys the detection targets and the key differentiator (log-based, no CI/CD). Every phrase earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only detection tool with full parameter documentation and no output schema, the description sufficiently explains what is detected and how. It could mention the shape of the output (e.g., list of change events), but the purpose and mechanism are clear.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six parameters are documented in the schema. The description adds no parameter-specific syntax or filtering guidance, leaving the baseline score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific detection task: finding deploys and restarts in logs via version field changes, host rollouts, and start/stop lines. It is clear what the tool does, though it does not explicitly differentiate itself from siblings like compare_periods or error_summary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance or alternatives are provided. The only contextual note is 'No CI/CD integration needed,' which is a feature, not usage guidance; an agent must infer that this tool is for log-based change detection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
error_summaryError summaryBRead-onlyIdempotent
Group errors (configured error query) by exception, logger, source or any field, with exact counts, first/last seen and a sample message per group.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Number of groups | |
| query | No | Extra Lucene filter, ANDed with the error query | |
| range | No | Relative range ending now (or at to_time): '15m', '2h', '1d', '1h30m' | 1h |
| samples | No | Attach one sample message per group | |
| streams | No | Stream titles or ids to search in; all streams when omitted | |
| to_time | No | Absolute end, same formats as from_time; default now | |
| group_by | No | 'exception', 'logger', 'source' (mapped via config) or any field name | exception |
| instance | No | Graylog instance (environment) from list_instances, e.g. 'staging' or 'prod'; the default instance when omitted | |
| from_time | No | Absolute start: ISO 8601 or 'YYYY-MM-DD HH:MM' in the instance timezone; overrides range |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, non-destructive and openWorld, so the safety profile is covered. The description adds that it uses a preconfigured error query and returns exact counts plus first/last seen and a sample per group, but says nothing about limits on result size, cost, or timeouts.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the operation and immediately states the grouping dimensions and the returned metrics. Every clause carries information; slightly long parenthetical nesting keeps it from being maximally crisp.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 9 optional parameters, no output schema and read-only annotations, the description covers the purpose and the shape of results adequately, but omits routing guidance relative to the many sibling log-analysis tools. For a tool with this many knobs it is the minimum viable description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter including group_by, range, query and instance is already documented in the schema. The description only echoes the group_by options already listed there, so no meaning is added beyond the structured fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (group) applied to a specific resource (errors) and enumerates the grouping dimensions and returned aggregates. An agent can tell it apart from most siblings, though it does not explicitly distinguish itself from the similarly grouping-oriented top_values.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance and no named alternative. It implies the tool always operates on a 'configured error query', which is useful context, but it never says when an agent should prefer this over count_logs, top_values or search_logs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_messageGet messageBRead-onlyIdempotent
One message with all its fields (redacted), by ref.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | 'index/message_id' from a previous result's 'ref', or a message id | |
| fields | No | Only these fields; all fields when omitted | |
| instance | No | Graylog instance (environment) from list_instances, e.g. 'staging' or 'prod'; the default instance when omitted | |
| compact_stacktrace | No | Fold framework stack frames (default true); false for the raw trace |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint, so the safety profile is covered. The description adds one genuinely useful behavioral fact — that returned fields are redacted — but says nothing about pagination, error behavior, or what happens with an invalid ref.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with no filler, and the key retrieval scope and the redaction caveat are front-loaded. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read tool with full schema coverage and clear annotations, the definition is nearly sufficient — it states the resource, the key, and that fields are redacted. With no output schema, a little more on the returned payload shape would help, but the gap is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents ref format, field filtering, instance selection, and stacktrace folding; the description's 'by ref' merely echoes the schema. Baseline 3 applies when structured data carries the parameter burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb-action (get) and resource (one message) plus the lookup key ('by ref'), so the agent knows exactly what it retrieves. It does not, however, distinguish itself from siblings like context_around or search_logs, which also surface message content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no exclusions, and no mention of alternatives such as context_around for surrounding logs or search_logs for finding messages. The agent must infer usage purely from the name and ref semantics.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_fieldsList fieldsARead-onlyIdempotent
Field names (and types where available) present in the indices, plus the configured trace fields and error query. Use before writing queries on unfamiliar logs.
| Name | Required | Description | Default |
|---|---|---|---|
| contains | No | Only fields whose name contains this text | |
| instance | No | Graylog instance (environment) from list_instances, e.g. 'staging' or 'prod'; the default instance when omitted | |
| include_internal | No | Include Graylog internal gl2_* fields |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds useful content-level detail (that output includes configured trace fields and the error query), but says nothing about result size, truncation, or cost on large indices.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with what is returned before the guidance sentence. No filler. The parenthetical and the clause about trace/error fields make it slightly denser than ideal, but nothing is wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must carry return-value meaning; it does so by enumerating what the listing contains. For a simple read-only introspection tool with a fully documented schema and annotations, this is nearly complete, missing only result-size or pagination behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (contains, instance, include_internal) are already fully documented in the schema. The description adds no parameter-level meaning beyond that, which is the baseline 3 case when structured data does the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: listing field names and types. It also scopes the content to fields 'present in the indices' plus trace fields and the error query, which distinguishes it from siblings like list_instances or list_streams. It stops short of naming an alternative tool, so not a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear usage condition: 'Use before writing queries on unfamiliar logs.' That tells the agent when this tool is the right first step. There is no explicit when-not or named alternative, so it lands at 4 rather than 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_instancesList instancesBRead-onlyIdempotent
Configured Graylog instances with detected version, the API used for messages and aggregations, and active redaction rules.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, fully covering the tool's behavior. The description adds only output content (version, API, redaction rules) and no behavioral context such as ordering, scope, or side effects. No behavioral traits beyond the annotations are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence with no wasted words and the resource is front-loaded. However, the fragmentary noun-phrase structure does not clearly state the action, which slightly reduces clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters and no output schema, the description usefully enumerates the fields returned by the listing. It stops short of explicitly saying the tool lists all configured instances, but the title and name fill that gap sufficiently for a simple read-only tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so the baseline is 4. The description does not need to explain parameter semantics, and the empty schema is self-evident.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the specific resource (Graylog instances) but never states the action. It reads as an output summary rather than a purpose statement; the agent must infer 'list' from the tool name and title. The resource is distinct from siblings like list_streams, but the verb is missing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of alternatives. The description only describes the returned data, leaving the agent to infer the tool's role entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_presetsList presetsCRead-onlyIdempotent
Named queries defined in the server configuration.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is fully covered by structured data. The description adds only that presets originate in server configuration (i.e., they are static/predefined), with no return shape, auth needs, or refresh behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single short fragment with no wasted words, but it is under-specified rather than concise — no leading verb and no statement of what is returned.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters, no output schema and complete annotations, the description is minimally adequate, but it never states what listing yields (names only vs. full query definitions) and never links to run_preset, which is the main thing an agent would need to know.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate. Baseline 4 applies; no parameter meaning is missing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is a noun phrase, not a verb+resource statement; the listing action is only implied by the tool name. It does define what a preset is ('named queries defined in the server configuration'), which is slightly more than restating the title, but it gives no scope or differentiation from siblings like run_preset or list_fields.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no mention of run_preset as the natural follow-up, and no exclusions. An agent must infer that this is a discovery step before executing a preset.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_streamsList streamsBRead-onlyIdempotent
Streams (id, title, description) the token can read.
| Name | Required | Description | Default |
|---|---|---|---|
| instance | No | Graylog instance (environment) from list_instances, e.g. 'staging' or 'prod'; the default instance when omitted | |
| include_disabled | No | Also list paused streams |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint, covering the safety profile. The description adds one piece of behavioral context — that results are limited to what the token can read — but says nothing about return format or pagination, so it clears the lowered bar only modestly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Extremely short with no wasted words; the field list is front-loaded. The sentence-fragment style is terse to the point of under-specification, which keeps it just below a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple zero-required-parameter list tool with full schema coverage and annotations covering safety, the description is adequate but thin — it omits any mention of pagination or result limits and gives no usage routing, leaving gaps an agent must infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (instance, include_disabled) are already fully documented in the schema. The description adds no parameter-level detail, making the baseline 3 appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb and resource ('Streams ... the token can read') and enumerates the returned fields (id, title, description). It is specific enough to distinguish from most siblings, but does not explicitly contrast with list_instances, list_fields, or list_presets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use guidance, no prerequisites, and names no alternatives. The only implicit context is the 'token can read' scoping, which hints at permissions but does not route the agent among sibling list tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
log_histogramLog histogramARead-onlyIdempotent
Message counts over time (exact). Reports the peak bucket, first/last non-empty bucket and the 'onset' of a spike, to find when a problem started.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Lucene query, '*' for everything | * |
| range | No | Relative range ending now (or at to_time): '15m', '2h', '1d', '1h30m' | 1h |
| streams | No | Stream titles or ids to search in; all streams when omitted | |
| to_time | No | Absolute end, same formats as from_time; default now | |
| instance | No | Graylog instance (environment) from list_instances, e.g. 'staging' or 'prod'; the default instance when omitted | |
| interval | No | Bucket size such as '1m', '5m', '1h'; chosen automatically when omitted | |
| from_time | No | Absolute start: ISO 8601 or 'YYYY-MM-DD HH:MM' in the instance timezone; overrides range |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, openWorld, so the safety profile is covered. The description adds genuinely useful behavioral detail by naming exactly what is reported (peak bucket, first/last non-empty bucket, spike onset) and flagging 'exact' counting, which the annotations do not.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the core 'what', then the purpose clause. Little waste, though the parenthetical '(exact)' is slightly awkward placement.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given annotations carry the safety profile and the schema fully documents all 7 params, the description's job is to convey output and intent, which it does. No output schema exists, but the description summarizes the returned figures, making it reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every one of the 7 parameters is already documented in the schema. The description adds nothing parameter-specific beyond the 'exact' precision note, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Message counts over time (exact)'. Clear what it produces, but it does not distinguish itself from close siblings like count_logs or search_logs by name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'to find when a problem started' implies a usage context, which is better than nothing. However, it gives no explicit when-not guidance or named alternatives against siblings such as count_logs or compare_periods.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
root_causeRoot causeARead-onlyIdempotent
Find which service broke first and why. Compares every service's errors, traffic and latency with a baseline, pins the first error of each to the millisecond, detects deploys/restarts/host rollouts from the logs, infers the call graph from traces, and returns a ranked verdict with a timeline and evidence. Start here for 'what is causing this incident?'.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Optional Lucene filter for scope, e.g. 'env:prod' | |
| range | No | Window to analyse (the incident), e.g. '1h', '30m' | 1h |
| streams | No | Stream titles or ids to search in; all streams when omitted | |
| to_time | No | Absolute end, same formats as from_time; default now | |
| baseline | No | Length of the normal period right before the window; same as the window by default | |
| instance | No | Graylog instance (environment) from list_instances, e.g. 'staging' or 'prod'; the default instance when omitted | |
| from_time | No | Absolute start: ISO 8601 or 'YYYY-MM-DD HH:MM' in the instance timezone; overrides range |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare the safety profile (readOnly, idempotent, non-destructive, openWorld); the description goes well beyond that by disclosing the comparison-against-baseline method, millisecond-level error pinning, deploy/restart/host-rollout detection, call-graph inference from traces, and the ranked-verdict-with-timeline output shape. That is substantive behavioral context an agent cannot get from structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose is front-loaded in the first clause, the middle enumerates only capabilities that actually differentiate the tool, and the final short sentence gives the usage cue. No sentence is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must carry return-value semantics, and it does ('returns a ranked verdict with a timeline and evidence'). All 7 optional parameters are schema-documented, and the description supplies the analytical context an agent needs to call this correctly for incident triage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline for parameter understanding is already met by the schema itself. The description alludes to the baseline-comparison concept, which lightly reinforces the 'baseline' parameter, but adds no syntax, format, or defaulting detail beyond what the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific, high-signal statement of the job ('Find which service broke first and why') and then enumerates the concrete analytical steps it performs. An agent can distinguish it from error_summary, trace_request, and service_map because it explicitly fuses error, traffic, latency, deploy, and trace analysis into one ranked verdict.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The closing line 'Start here for "what is causing this incident?"' gives a clear routing condition and makes this the default entry point for incident triage. It does not name sibling alternatives (e.g. error_summary or trace_request) or state when NOT to use it, so it falls short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_presetRun presetCRead-onlyIdempotent
Run a named preset query, optionally overriding its arguments.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Preset name from list_presets | |
| instance | No | Graylog instance (environment) from list_instances, e.g. 'staging' or 'prod'; the default instance when omitted | |
| overrides | No | Arguments overriding the preset, e.g. {"range": "6h"} |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true and destructiveHint=false, covering the safety profile. The description adds essentially no behavioral context beyond that – it says nothing about what a preset executes, whether it hits a remote instance, or what the result contains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the core action first and the optional capability second; nothing is wasted. It is terse but not at the cost of clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 3-parameter, non-destructive, read-only tool with full schema coverage and no output schema, the description is minimally adequate. It leaves the agent without context on what a preset is or what running one returns, but the safety profile is carried by annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with name, instance and overrides each documented in the schema itself (including the 'staging'/'prod' example and the override shape). The description's phrase "optionally overriding its arguments" adds no syntax or format detail beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ("Run a named preset query") and notes the override capability, which is enough to distinguish it from the sibling list_presets. It stops short of explicitly naming alternatives such as search_logs, so it is clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance; the agent must infer that a preset must already exist (via list_presets) and that this is the execution step. No alternatives or preconditions are named, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_logsSearch logsARead-onlyIdempotent
Search log messages with a Lucene query. Returns compact, redacted lines with a 'ref' (index/id) usable by get_message and context_around. Default range: last 15 minutes.
| Name | Required | Description | Default |
|---|---|---|---|
| sort | No | 'timestamp:desc' (default), 'asc', or '<field>:asc|desc' | timestamp:desc |
| limit | No | Messages to fetch (capped by config) | |
| query | No | Lucene query, '*' for everything | * |
| range | No | Relative range ending now (or at to_time): '15m', '2h', '1d', '1h30m' | |
| fields | No | Fields to return; default timestamp/source/level/message, ['*'] for all | |
| offset | No | Skip this many messages (use next_offset) | |
| streams | No | Stream titles or ids to search in; all streams when omitted | |
| to_time | No | Absolute end, same formats as from_time; default now | |
| instance | No | Graylog instance (environment) from list_instances, e.g. 'staging' or 'prod'; the default instance when omitted | |
| from_time | No | Absolute start: ISO 8601 or 'YYYY-MM-DD HH:MM' in the instance timezone; overrides range | |
| dedup_lines | No | Group lines that only differ in numbers/ids/timestamps (default true) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the bar is lower. The description adds useful behavior: results are 'compact, redacted lines' with a 'ref', and it discloses the default range of last 15 minutes, which is not in the schema. It doesn't mention rate limits or the dedup default, but this is solid context beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the verb and query language, then output format, then default behavior. No filler and every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers purpose, output shape, downstream usage, and default range for a read-only, idempotent search tool with a fully documented schema and no output schema. It does not explain pagination (offset/next_offset) or the default field set, but those are in the schema and the core call contract is clear.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 11 parameters thoroughly. The description only adds the default range behavior and the 'ref' output concept; it doesn't explain query syntax or field filtering beyond what the schema provides. Baseline 3 is correct when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Search) and resource (log messages) plus the query dialect (Lucene). The return type is described with concrete fields ('ref', index/id), giving it more precision than siblings like count_logs or log_histogram.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the downstream tools that consume the 'ref' (get_message, context_around) and states the default range, which helps the agent know what this tool is for. It does not explicitly contrast with count_logs or log_histogram, so it's a clear context but not full alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
service_mapService mapARead-onlyIdempotent
Which service calls which, inferred from sampled traces (no configuration): edges with traffic, error rate and p50/p95 latency, plus entry points.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Optional Lucene filter for scope | |
| range | No | Relative range ending now (or at to_time): '15m', '2h', '1d', '1h30m' | 1h |
| sample | No | Number of traces to sample | |
| streams | No | Stream titles or ids to search in; all streams when omitted | |
| to_time | No | Absolute end, same formats as from_time; default now | |
| instance | No | Graylog instance (environment) from list_instances, e.g. 'staging' or 'prod'; the default instance when omitted | |
| from_time | No | Absolute start: ISO 8601 or 'YYYY-MM-DD HH:MM' in the instance timezone; overrides range |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and non-destructive behavior, so the safety profile is covered. The description adds genuinely useful context beyond that: results are statistically inferred from sampling rather than from configuration, which tells the agent output is approximate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the core meaning ('which service calls which') and then enumerates outputs. No wasted words, nothing under- or over-specified.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description compensates by naming the returned content (edges with traffic, error rate, p50/p95 latency, entry points). Combined with full schema coverage and rich annotations, an agent has enough to call it correctly; only minor gaps like result size limits remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (query, range, sample, streams, time bounds, instance) is already documented in the schema. The description alludes to sampling but adds no syntax or format detail beyond what the schema provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (service dependency map) and how it's derived ('inferred from sampled traces'), plus the output shape (edges with traffic, error rate, p50/p95 latency, entry points). It clearly reads as a topology/dependency tool, distinct from trace_request or root_cause, though it never names siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical '(no configuration)' implies the tool is the right choice when no instrumentation/config exists, which is useful implied context. However, there is no explicit when-to-use vs when-not guidance and no reference to alternatives such as trace_request or root_cause.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
top_valuesTop valuesBRead-onlyIdempotent
Top N values of a field with exact counts and percentages.
| Name | Required | Description | Default |
|---|---|---|---|
| field | Yes | Field to group by, e.g. source, http_status, user_agent | |
| limit | No | Number of values | |
| query | No | Lucene query, '*' for everything | * |
| range | No | Relative range ending now (or at to_time): '15m', '2h', '1d', '1h30m' | 1h |
| streams | No | Stream titles or ids to search in; all streams when omitted | |
| to_time | No | Absolute end, same formats as from_time; default now | |
| instance | No | Graylog instance (environment) from list_instances, e.g. 'staging' or 'prod'; the default instance when omitted | |
| from_time | No | Absolute start: ISO 8601 or 'YYYY-MM-DD HH:MM' in the instance timezone; overrides range |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=true, so the safety profile is fully covered. The description adds only that output includes exact counts and percentages, which is modest extra context and says nothing about aggregation cost, limits, or missing-field behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short, front-loaded sentence with zero padding. It is efficient, though arguably too sparse to count as optimally structured for an 8-parameter aggregation tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an aggregation tool with 8 parameters but no output schema, the description at least signals the shape of the result (counts and percentages). However it omits any guidance on limiting large result sets or on the query/range/time relationships, leaving the schema to carry the whole load.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and all 8 parameters carry clear descriptions with examples, defaults, and bounds, so the schema does the heavy lifting. The description adds no syntax, interaction, or precedence detail (e.g. from_time overriding range) beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific operation (top N values of a field) plus the return semantics (exact counts and percentages), so an agent can tell it apart from count_logs or log_histogram. It does not, however, name any sibling or scope the tool to a data source beyond the schema's implicit fields.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use or when-not-to-use guidance and no alternatives. It never hints at how this differs from count_logs, log_histogram, or list_fields, so the agent must infer selection from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
trace_requestTrace requestARead-onlyIdempotent
Follow one request across services: searches the configured trace fields (falls back to full text) on all streams and returns a chronological timeline, per-service steps with durations, and the first error.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max messages in the timeline | |
| range | No | Relative range ending now (or at to_time): '15m', '2h', '1d', '1h30m' | 24h |
| streams | No | Stream titles or ids to search in; all streams when omitted | |
| to_time | No | Absolute end, same formats as from_time; default now | |
| instance | No | Graylog instance (environment) from list_instances, e.g. 'staging' or 'prod'; the default instance when omitted | |
| trace_id | Yes | Correlation / request / trace id to follow | |
| from_time | No | Absolute start: ISO 8601 or 'YYYY-MM-DD HH:MM' in the instance timezone; overrides range |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish the safe/idempotent/read-only profile, and the description adds real behavioral context beyond that: the trace-field search with full-text fallback, and the shape of the result (chronological timeline, per-service steps with durations, first error). With no output schema present, disclosing the return structure is especially valuable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with a colon-delimited enumeration of the returned artifacts. Dense but every clause carries information; minor cost is that the enumeration makes it feel packed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter, single-required-param tool with no output schema, the description covers core behavior and the return shape well. Gaps are modest: no pagination/truncation note, no statement on what happens when the trace id matches nothing, and no instance/time-range interaction guidance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 7 parameters are self-documented with defaults and formats. The description adds no parameter-level meaning (e.g. how limit interacts with the timeline, or that streams are optional), so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: follows one request across services using trace fields, and specifies scope (all streams). It's clearly a trace-correlation tool rather than a generic log search, though it never names a sibling like search_logs or context_around to sharpen the contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: you need a correlation/trace id and want a cross-service timeline. There is no explicit when-to-use vs when-not guidance, no mention of alternatives (e.g. search_logs for free-text exploration), and no prerequisites or instance-selection advice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
17 tool updates
v0.1.0- First observed
compare_periods - First observed
context_around - First observed
count_logs - First observed
detect_changes - First observed
error_summary - First observed
get_message - First observed
list_fields - First observed
list_instances - First observed
list_presets - First observed
list_streams - First observed
log_histogram - First observed
root_cause - First observed
run_preset - First observed
search_logs - First observed
service_map - First observed
top_values - First observed
trace_request
TDQS
Scored across 17 tools
Each tool targets a distinct query or aggregation pattern (search, count, get, histogram, top values, trace, service map). The main overlap is that root_cause conceptually aggregates error_summary, compare_periods, detect_changes and service_map, which could make an agent unsure whether to use the composite or the individual tools, but descriptions clarify root_cause as the 'start here' entry point.
All names are snake_case with a consistent verb_noun or noun form (search_logs, count_logs, get_message, compare_periods). Some are noun phrases (error_summary, log_histogram, service_map) and others are verb-led, but the style is uniform and readable throughout.
17 tools is on the heavier side but justified for a rich observability domain covering search, aggregation, tracing, topology and infrastructure discovery. No tool appears redundant or purely decorative.
Covers the full investigation lifecycle: search/count/get/context, time-series and top-value aggregation, error grouping, period comparison, tracing, service topology, deploy detection, and discovery of streams/fields/instances/presets. Minor gaps exist (no alert or user/config management), but they are outside the core log-investigation purpose.
Maintenance
Related MCP Connectors
Enable secure connectivity between Sentry issues and debugging data, and LLM clients, using a Model Context Protocol (MCP) server.
The Grafbase MCP server sits in front of a GraphQL API and exposes an MCP protocol-compliant interface that allows AI agents and LLMs to explore and query GraphQL APIs using natural language. It provides tools to search schemas, introspect types and fields, and execute GraphQL queries while minimizing context bloat by returning only relevant schema subsets, with built-in support for authentication, authorization, and configurable access control.
The Cortex MCP server provides read-only access to real-time engineering context from the Cortex developer portal, allowing AI coding assistants to answer natural language questions about your organization's catalog (microservices, libraries, domains, teams, infrastructure), scorecards (engineering standards and best practices), initiatives (goals and deadlines), and Engineering Intelligence metrics. It includes tools for querying documentation, tracking personal entities, and accessing AI-assisted insights across the entire Cortex ecosystem.
Model Context Protocol server for the Apideck Unified API. Connect any MCP-compatible agent framework to 100+ accounting systems, HRIS platforms, file storage providers, and more through one integration. More information https://www.apideck.com/mcp-server
Related MCP Servers
- AlicenseDqualityBmaintenanceA Model Context Protocol server that provides AI agents with controlled read access to Datalust Seq instances for log analysis and monitoring. It enables agents to search events, execute data queries, and retrieve information about signals, dashboards, and alerts.1003MIT
- AlicenseNot gradedqualityFmaintenanceA read-only MCP server that exposes Quickwit log search and aggregations to LLM clients, enabling natural language log investigation.Apache 2.0
- AlicenseAqualityDmaintenanceAn MCP server that gives AI assistants direct access to your Graylog logs -- search, aggregate, analyze, and cluster log data through natural language.2319 npmMIT
- AlicenseBqualityBmaintenanceA portable, read-only Model Context Protocol server for turning observability data into bounded evidence that AI agents can inspect safely.7Apache 2.0