Skip to main content
Glama
crunchtools

mcp-syslog

by crunchtools

mcp-syslog-crunchtools

MCP server for the logs collected by crunchtools/syslog.

Built for RT #1460, to close a specific gap: Hermes gets paged by Nagios and can restart a service, but it cannot read the service's logs — so every remediation is a blind restart. This turns that into an informed one.

Capabilities

Tool

What it answers

syslog_sources_tool

What can I query?

syslog_search_tool

Show me ERRs from this service in the last 15 minutes

syslog_grep_tool

Where does this string appear across the whole fleet?

syslog_tail_tool

What are this service's most recent lines?

syslog_context_tool

What was everything saying around 03:14?

syslog_stats_tool

Which service is loudest, and which is actually unhealthy?

Related MCP server: Log Analyzer MCP Server

The triage loop

nagios_current_problems_tool          → what is broken
syslog_search_tool(source=…,          → why it broke
                   severity="ERR",
                   since="15m")
syslog_context_tool(timestamp=…)      → what else was happening at that moment
nagios_schedule_check_tool            → confirm the fix

syslog_context_tool without a source is the one that earns its keep. It spans every source at once, which is how "the app died" gets connected to "the database container OOMed four seconds earlier".

Design notes

Every result is bounded, and says when it is. Logs are unbounded and this output lands in a model's context window. Each tool caps its results, caps how many lines it will scan, and annotates the answer when either limit is hit:

[!] Stopped after the 2,000,000-line scan limit, so this result is INCOMPLETE
    and an empty or short result does not mean nothing happened.

That annotation is load-bearing. A caller that cannot distinguish "no errors occurred" from "I stopped looking" will draw the wrong conclusion from an empty result — and this server exists to inform remediation decisions.

Time filtering is cheap. The collector puts the date in the filename, so a ten-minute query opens one file rather than reading ninety days of history.

Source names are untrusted. They come from a model and are used to build a filesystem path. Each is resolved and then confirmed to still be inside the log root, which catches traversal, absolute paths, and symlinks pointing out of the tree — see tests/test_security.py.

Severity is "at least this severe". severity="ERR" returns ERR, CRIT, ALERT and EMERG. An unrecognised severity is kept rather than dropped, on the grounds that hiding a line you do not understand is worse than showing it — but it is not counted as an error in syslog_stats_tool, or a healthy service would report a 57% error rate.

Severity is not badness. Podman records anything a container writes to stderr at priority err, and plenty of services log routine INFO there. One chatty httpx-based service in this fleet sits around 65% "ERR" while being entirely healthy:

PRIORITY=3 | 2026-08-23 15:55:24 INFO  httpx: HTTP Request: GET https://... "200 OK"

The collector is reporting the journal faithfully; the journal is reporting the file descriptor. Read the message, and prefer a change in error rate to its absolute value. This caveat is in the server's MCP instructions too, so an agent querying it is told the same thing.

Two line formats are parsed. The collector emitted five fields before 2026-08-23 and six after, and the old lines stay in retention for 90 days. Which layout a line uses is decided by where a real severity sits, not by counting fields.

Log format

The collector writes six space-delimited fields:

2026-08-23T15:41:52+00:00 crunchtools.com crunchtools.com httpd ERR AH00169: caught SIGTERM
└─ timestamp ───────────┘ └─ host ──────┘ └─ source ────┘ └prog┘ └sev┘ └─ message ────────┘

source is the log stream — normally a container name. program is the process inside it, which matters for systemd containers where httpd, php-fpm and mariadb all file under one service name.

Configuration

Variable

Default

Purpose

SYSLOG_LOG_ROOT

/logs

Collector log root, mounted read-only

SYSLOG_MAX_RESULTS

200

Cap on entries returned per call

SYSLOG_SCAN_LIMIT

2000000

Cap on lines examined per call

No credentials — the server reads files off a read-only bind mount.

Running

podman run -d --name mcp-syslog \
  --network crunchtools \
  -p 127.0.0.1:8027:8027 \
  -v /path/to/syslog/data/logs:/logs:ro \
  quay.io/crunchtools/mcp-syslog:latest \
  --transport streamable-http --host 0.0.0.0 --port 8027

Mount :ro. This server never needs to write, and a read-only mount means a bug here cannot destroy the forensic record it exists to protect.

Development

uv sync
uv run ruff check src tests
uv run mypy src
uv run pytest -v

Available Tools

6 tools
syslog_context_toolSyslog Context ToolA

Return log entries surrounding a specific moment.

Use this after an alert names a time. Omitting source spans the whole fleet, which is how a failure gets correlated with whatever else was happening.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum entries to return.
sourceNoRestrict to one source. Omit to span all sources.
severityNoOptional minimum severity.
timestampYesThe moment of interest — ISO-8601, or relative like '30m' ago.
after_secondsNoHow far forward from the timestamp to include.
before_secondsNoHow far back from the timestamp to include.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden. It does add real behavioral context — omitting 'source' spans the whole fleet, which is how correlation happens — but says nothing about ordering, whether the ±60s window is a hard bound, result caps, or that this is a pure read operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero filler, and the purpose is front-loaded before the usage cue. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists and parameter coverage is complete, so return values and inputs need no explanation here. The description covers purpose and the key triggering context; the main remaining gap is the lack of any routing guidance relative to the five sibling syslog tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter including the relative-timestamp format ('30m') and the before/after window defaults is already documented in the schema. The description only reinforces the 'source' omission behavior, adding no new parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Return log entries surrounding a specific moment'), which is a distinct capability from siblings like syslog_search_tool, syslog_grep_tool, and syslog_tail_tool. However, it never names or contrasts with those siblings, so the agent must infer the boundary itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use this after an alert names a time' gives a concrete triggering condition, which is genuinely useful routing guidance. It stops short of stating when-not to use it or naming alternatives such as syslog_search_tool for non-time-anchored queries.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

syslog_grep_toolSyslog Grep ToolC

Regex search across all sources, or within one named source.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum entries to return.
sinceNoHow far back to search — relative ('15m', '2h', '3d') or ISO-8601.24h
sourceNoRestrict to one source. Omit to search the whole fleet.
patternYesCase-insensitive regular expression matched against the message.
severityNoOptional minimum severity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It implies a read-only regex search but discloses nothing about auth requirements, rate limits, result ordering, or the case-insensitive matching behavior, all of which the agent must infer or read from the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single, front-loaded sentence with no wasted words. The brevity is appropriate for the core action, though for a five-parameter tool with no annotations it borders on under-specification.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema covers return values and the input schema fully documents parameters, so those needs are met. The remaining gap is behavioral and sibling-routing context, which the one-sentence description does not supply for an unannotated search tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description reinforces the source parameter by mentioning 'one named source', but adds no syntax, format, or default-value detail beyond what the schema already documents.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Regex search') and resource ('sources' / 'named source'), which is clear enough. However, it does not distinguish this tool from the sibling syslog_search_tool, leaving an agent to guess which search tool applies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'across all sources, or within one named source' implies a scope choice, but there is no explicit when-to-use, when-not-to-use, or named alternative. With a sibling search tool present, the lack of routing guidance is a clear gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

syslog_search_toolSyslog Search ToolC

Search collected logs by source, time window, severity and pattern.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum entries to return.
sinceNoStart of the window — relative ('15m', '2h', '3d') or ISO-8601.1h
untilNoEnd of the window. Omit for "up to now".
sourceNoContainer or service name. Omit to search every source.
patternNoOptional case-insensitive regular expression matched against the message.
programNoRestrict to one program within the source (e.g. 'httpd' inside a web container).
severityNoMinimum severity: EMERG, ALERT, CRIT, ERR, WARNING, NOTICE, INFO, DEBUG. 'ERR' returns ERR and anything more severe.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations at all, the description carries the full burden of behavioral disclosure, yet it says nothing about read-only nature, result limits, ordering, or truncation behavior. It is implied to be a read/search operation but nothing confirms the safety profile or output behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with the key filter dimensions front-loaded and no filler. It is appropriately short for a broad search tool, though it is arguably too sparse to earn a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values needn't be explained, and the schema fully documents all seven parameters. The remaining gap is routing guidance among the many sibling log tools, which the description does not address.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both the description and schema document the parameters; the schema in fact adds far more (relative/ISO-8601 time formats, severity ordering, regex matching, program scoping). The description only names the filter dimensions already covered by the schema, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (search) and resource (collected logs) plus the filterable dimensions, so the purpose is unambiguous. However, it offers no differentiation from the many similar siblings (syslog_grep_tool, syslog_tail_tool, syslog_context_tool), leaving the agent to guess which search variant applies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance and no mention of alternatives, despite a crowded sibling set of overlapping log-query tools. The agent gets no signal about when this broad search is preferable to grep, tail, or context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

syslog_sources_toolSyslog Sources ToolA

List log sources available to query, with size and last-write time.

ParametersJSON Schema
NameRequiredDescriptionDefault
patternNoOptional case-insensitive substring to filter source names.
include_internalNoInclude the collector's own '_collector' statistics.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It implies a read-only enumeration and discloses returned metadata, but says nothing about permissions, cost, pagination, or whether hidden sources are omitted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the action and the returned fields are stated immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be re-explained, and the two optional params are schema-documented. However, for a discovery tool in a six-tool family, the absence of any routing guidance to sibling tools leaves a real gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (pattern, include_internal) are already documented in the schema. The description adds no filtering or formatting detail beyond that baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (log sources), plus the metadata returned (size, last-write time), which distinguishes it from the search/grep/tail/stats siblings by resource type. It stops short of explicitly contrasting itself with any named sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"available to query" implies this is a discovery step before using syslog_search_tool or syslog_grep_tool, but no alternative is named and no when-not-to-use condition is given. Usage must be inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

syslog_stats_toolSyslog Stats ToolA

Summarise log volume and error rate per source.

ParametersJSON Schema
NameRequiredDescriptionDefault
topNoHow many sources to list, ranked by volume.
sinceNoWindow to summarise — relative ('1h', '24h') or ISO-8601.1h
sourceNoRestrict to one source. Omit to cover every source.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose the core behavior — aggregated volume and error-rate metrics keyed by source — which tells the agent this is a read/aggregate operation, but it says nothing about cost, permissions, or whether it scans the whole window, and there is no note that it does not return raw log lines.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler or repetition. It is efficient, though for a tool in a six-member family it is arguably thinner than ideal, which keeps it out of 5 territory.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the input schema is fully described at 100% coverage. The description supplies the aggregation semantics the structured fields cannot express, which is sufficient for an agent to call this low-complexity stats tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so 'top', 'since', and 'source' are already fully documented with defaults and formats. The description's 'per source' phrase loosely connects to the 'source' and 'top' parameters but adds no syntax, ordering, or edge-case meaning beyond the schema; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (summarise) and resource (log volume and error rate), plus the grouping key (per source), so an agent can tell this is an aggregation tool rather than a content-retrieval one. It stops short of naming or contrasting any of the five syslog siblings (search, grep, tail, context, sources), so differentiation is inferred rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the aggregation framing suggests using this for metrics overviews and the siblings for log content. There is no explicit when-to-use, when-not-to-use, or named alternative, so the agent must infer routing from the verb alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

syslog_tail_toolSyslog Tail ToolC

Return the most recent entries for one source.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoNumber of entries to return.
sourceYesContainer or service name.
severityNoOptional minimum severity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it discloses almost nothing: it does not say whether it reads live/streams, the ordering of results, how 'minimum severity' is interpreted, or that limit defaults to 50. Only the read-like nature and recency scope are conveyed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the resource and scope front-loaded and zero filler. It is efficient, though its brevity borders on under-specification rather than optimal conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the schema covers all parameters. However, for a no-annotation tool, key behavioral facts (live tail vs snapshot, ordering, severity semantics) remain undisclosed, leaving real gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (limit, source, severity) are already documented in the schema. The description adds no syntax or semantic detail beyond the schema, which is the baseline case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Return the most recent entries') with a clear scope ('for one source'), which distinguishes it from search/grep/stats siblings by implication. It does not explicitly name an alternative, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of when not to use it, and no reference to siblings like syslog_search_tool or syslog_context_tool. Only the phrase 'most recent' hints at its purpose, so usage must be inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 6 tool updatesv0.1.0
    • First observedsyslog_context_tool
    • First observedsyslog_grep_tool
    • First observedsyslog_search_tool
    • First observedsyslog_sources_tool
    • First observedsyslog_stats_tool
    • First observedsyslog_tail_tool

TDQS

A3.5/5.0

Scored across 6 tools

Disambiguation4/5

Most tools target distinct operations: listing sources, tailing, stats, and context retrieval are clearly separated. However, syslog_search_tool (pattern search) and syslog_grep_tool (regex search across/within sources) overlap substantially and an agent could easily pick the wrong one for a pattern-matching query.

Naming Consistency5/5

All six tools use the identical syslog_<action>_tool convention, making the pattern predictable and easy to scan. The redundant '_tool' suffix is uniform, so it does not hinder consistency.

Tool Count5/5

Six tools is well-scoped for a log-query server, each covering a distinct access pattern (discovery, search, grep, tail, aggregate, context). No filler or redundancy beyond the search/grep overlap.

Completeness4/5

The surface covers the core log-investigation lifecycle: discovering sources, filtering, tailing, aggregating, and contextual surrounding entries. Minor gaps exist, such as no explicit pagination/offset control or config/retention management, but agents can work around these.

Maintenance

ActivityActive
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables querying and analyzing logs from multiple remote Unix hosts via the Log Collector API, with tools for search, error detection, and summary generation.
    -
  • F
    license
    Not graded
    quality
    B
    maintenance
    Provides telemetry tools for retrieving recent logs and system metrics to support root-cause analysis of infrastructure incidents. Enables autonomous incident triage with grounded verification and human-in-the-loop remediation.
    1
    -
  • F
    license
    A
    quality
    C
    maintenance
    Enables debugging of distributed transactions by continuously ingesting Docker container logs, indexing them by trace/request ID, and exposing MCP tools to search, tail, and correlate logs across services.
    7
    -