Skip to main content
Glama

search_messages

Search Kafka messages using a JavaScript predicate on value, key, headers, or metadata. Returns matching records with offset, partition, and timestamp.

Instructions

Search message key, value, headers or metadata with a JavaScript predicate. The script returns true for a match and receives value (parsed JSON, a decoded Avro/Protobuf/JSON Schema record, or text), key, headers, partition, offset and timestamp. Omit it to match all messages. Schema-encoded values are searched by field exactly like JSON.

Every partition is read together, one chunk deep at a time, so a limited newest-first search returns the newest matches in the topic rather than the newest in whichever partition was read first. Kafka orders records only within a partition, so matches are merged by timestamp; producers set timestamps unless the topic uses LogAppendTime.

Kafka has no server-side search, so scans are bounded. Check complete, stopped_reason and scanned_ranges before treating no matches as conclusive. Use count_only or output_file for large result sets. A script that runs past timeout_seconds is interrupted and counted in script_errors; scripts are not memory-sandboxed, so keep predicates simple.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
topicYesTopic to search. Matched exactly and case-sensitively.
scriptNoOptional JavaScript that decides whether a message matches. Return true to keep it. In scope: value (the parsed document for JSON, and the decoded record for Avro, Protobuf, JSON Schema or a configured format; the raw text otherwise), key (string, decoded document when the key has a schema, or null), headers (object of header name to string), partition, offset and timestamp (a Date). Examples: return key === 'order-123'; return value.eventType === 'NEW' && value.payload.amount >= 500; return value.payload.cancelledAt === null. Omit to match every message.
directionNoOptional scan direction: newest_first (default) or oldest_first. Decides which matches are kept when max_matches cuts the search short. Every partition is read together, so newest_first means newest in the topic, ordered by timestamp, rather than newest in one partition.
to_offsetNoOptional exclusive offset to stop scanning at, applied to every searched partition.
count_onlyNoOptional. When true, scan the whole range and return only how many messages matched, with a per-partition breakdown and no message bodies. Use this first when a query may match a great many messages, then ask the user how they want them before fetching any.
partitionsNoOptional partitions to restrict the search to. Defaults to every partition. Do not guess a partition from a message key: producers may set the partition explicitly, so the key does not determine it.
from_offsetNoOptional inclusive offset to start scanning from, applied to every searched partition.
max_matchesNoOptional maximum number of matches to return. Defaults to 10.
output_fileNoOptional file name to write every match to, as one JSON message per line. Use this instead of returning thousands of messages. A name only, not a path: the server chooses the directory. The response reports the path, the number written and a short preview.
parallelismNoOptional number of concurrent readers, from 1 to 16. Defaults to 1. It splits a single-partition topic's offsets between readers, which makes a full scan of one large partition faster. A multi-partition topic is already read across its partitions together, so this does not apply there. Worth using for count_only, output_file or a full scan of one partition.
to_timestampNoOptional exclusive end time (RFC3339). Resolved to the first offset at or after this time.
from_timestampNoOptional inclusive start time (RFC3339). Resolved to the first offset at or after this time.
max_value_bytesNoOptional maximum value bytes to return per match. Defaults to 512. Longer values are cut and flagged with truncated=true.
timeout_secondsNoOptional wall-clock limit for the scan in seconds. Defaults to 30.
max_messages_scannedNoOptional maximum number of messages to read before giving up. Defaults to 10000.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
topicYes
matchesYes
completeYes
match_countYes
output_fileNo
script_errorsNo
scanned_rangesYes
stopped_reasonYes
scanned_messagesYes
written_messagesNo
matches_by_partitionNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.2.0

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses that Kafka has no server-side search so scans are bounded, that every partition is read one chunk deep and merged by timestamp, that scripts are interrupted past timeout_seconds and counted in script_errors, and that scripts are not memory-sandboxed. These are exactly the costly, non-obvious traits an agent needs before calling it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose, then layers scan semantics; nearly every sentence earns its place given 15 parameters. Minor duplication with the schema ("omit it to match all messages" appears in both script and description) and the timestamp-ordering paragraph is dense but justified.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be spelled out, yet the description still flags the completion fields (complete, stopped_reason, scanned_ranges) an agent must inspect. Combined with coverage of truncation, timeouts, parallelism and memory limits, nothing needed to invoke this tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds real meaning: it notes schema-encoded values are searched by field exactly like JSON, restates the script's in-scope bindings, and frames direction as deciding which matches survive max_matches. Most of the script detail still duplicates the schema, keeping it just above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific verb (search), the searchable surface (key, value, headers, metadata) and the exact mechanism (JavaScript predicate). This mechanism is distinctive from every sibling that merely reads messages (get_message, sample_messages), so an agent can route on it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear operational direction: omit the script to match all, use count_only or output_file for large result sets, check complete/stopped_reason/scanned_ranges before concluding no matches. It never explicitly contrasts with sibling readers such as sample_messages or get_message, so the when-not and alternatives are left implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.