search_messages
Search Kafka messages using a JavaScript predicate on value, key, headers, or metadata. Returns matching records with offset, partition, and timestamp.
Instructions
Search message key, value, headers or metadata with a JavaScript predicate. The script returns true for a match and receives value (parsed JSON, a decoded Avro/Protobuf/JSON Schema record, or text), key, headers, partition, offset and timestamp. Omit it to match all messages. Schema-encoded values are searched by field exactly like JSON.
Every partition is read together, one chunk deep at a time, so a limited newest-first search returns the newest matches in the topic rather than the newest in whichever partition was read first. Kafka orders records only within a partition, so matches are merged by timestamp; producers set timestamps unless the topic uses LogAppendTime.
Kafka has no server-side search, so scans are bounded. Check complete, stopped_reason and scanned_ranges before treating no matches as conclusive. Use count_only or output_file for large result sets. A script that runs past timeout_seconds is interrupted and counted in script_errors; scripts are not memory-sandboxed, so keep predicates simple.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| topic | Yes | Topic to search. Matched exactly and case-sensitively. | |
| script | No | Optional JavaScript that decides whether a message matches. Return true to keep it. In scope: value (the parsed document for JSON, and the decoded record for Avro, Protobuf, JSON Schema or a configured format; the raw text otherwise), key (string, decoded document when the key has a schema, or null), headers (object of header name to string), partition, offset and timestamp (a Date). Examples: return key === 'order-123'; return value.eventType === 'NEW' && value.payload.amount >= 500; return value.payload.cancelledAt === null. Omit to match every message. | |
| direction | No | Optional scan direction: newest_first (default) or oldest_first. Decides which matches are kept when max_matches cuts the search short. Every partition is read together, so newest_first means newest in the topic, ordered by timestamp, rather than newest in one partition. | |
| to_offset | No | Optional exclusive offset to stop scanning at, applied to every searched partition. | |
| count_only | No | Optional. When true, scan the whole range and return only how many messages matched, with a per-partition breakdown and no message bodies. Use this first when a query may match a great many messages, then ask the user how they want them before fetching any. | |
| partitions | No | Optional partitions to restrict the search to. Defaults to every partition. Do not guess a partition from a message key: producers may set the partition explicitly, so the key does not determine it. | |
| from_offset | No | Optional inclusive offset to start scanning from, applied to every searched partition. | |
| max_matches | No | Optional maximum number of matches to return. Defaults to 10. | |
| output_file | No | Optional file name to write every match to, as one JSON message per line. Use this instead of returning thousands of messages. A name only, not a path: the server chooses the directory. The response reports the path, the number written and a short preview. | |
| parallelism | No | Optional number of concurrent readers, from 1 to 16. Defaults to 1. It splits a single-partition topic's offsets between readers, which makes a full scan of one large partition faster. A multi-partition topic is already read across its partitions together, so this does not apply there. Worth using for count_only, output_file or a full scan of one partition. | |
| to_timestamp | No | Optional exclusive end time (RFC3339). Resolved to the first offset at or after this time. | |
| from_timestamp | No | Optional inclusive start time (RFC3339). Resolved to the first offset at or after this time. | |
| max_value_bytes | No | Optional maximum value bytes to return per match. Defaults to 512. Longer values are cut and flagged with truncated=true. | |
| timeout_seconds | No | Optional wall-clock limit for the scan in seconds. Defaults to 30. | |
| max_messages_scanned | No | Optional maximum number of messages to read before giving up. Defaults to 10000. |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| topic | Yes | ||
| matches | Yes | ||
| complete | Yes | ||
| match_count | Yes | ||
| output_file | No | ||
| script_errors | No | ||
| scanned_ranges | Yes | ||
| stopped_reason | Yes | ||
| scanned_messages | Yes | ||
| written_messages | No | ||
| matches_by_partition | No |