Kafka MCP
kafka-mcp
An MCP server that exposes Kafka debugging as tools an LLM can call. It speaks MCP over stdio by default, optionally serves streamable HTTP, and talks to Kafka with franz-go.
One process can connect to several Kafka clusters. In stdio mode, one endpoint is selected for the session. In HTTP mode, each endpoint has its own path, so a session is bound to one cluster by how it connects rather than by a parameter a caller could forget to send.
Install
Every release ships the same server four ways:
Binary:
kafka-mcp_<Os>_<arch>archives on the releases page. Putkafka-mcpon yourPATH.MCP Bundle:
kafka-mcp_<version>_<os>_<arch>.mcpbon the same page, for macOS on Apple Silicon, Linux (amd64, arm64) and Windows (amd64). Open it in a client that installs bundles, such as Claude Desktop. It asks for your configuration file and, when that file has several endpoints, which one to serve. Intel Macs use the binary or the image.Container image:
ghcr.io/denizgursoy/kafka-mcp:<tag>. It starts with--server; for stdio, mount the file and pass--server=false:docker run -i --rm \ --mount type=bind,src=$PWD/kafka-mcp.yaml,dst=/config/kafka-mcp.yaml,readonly \ -e CONFIG_FILE=/config/kafka-mcp.yaml \ ghcr.io/denizgursoy/kafka-mcp:latest --server=false --endpoint localMCP Registry: listed as
io.github.denizgursoy/kafka-mcpin the official registry, with the image and the bundles, so a client that reads the registry can install it from there.
How releases reach the registry and the other catalogs is in PUBLISHING.md.
Related MCP server: kafka-mcp
Configuration
kafka-mcp.{toml,yaml,yml,json} in the working directory, ~/.config/kafka-mcp/ or /etc configures the server. The http block is used only with --server.
http:
address: ":8090"
base_path: /kafka-mcp # optional; prefixes every endpoint and /healthz
output_dir: /var/tmp/kafka-mcp
clusters:
prod:
brokers:
- kafka-1:9093
- kafka-2:9093
security:
tls:
enabled: true
ca_file: /etc/kafka/ca.pem
# cert_file: /etc/kafka/client.pem # optional mTLS; requires key_file
# key_file: /run/secrets/client.key
sasl:
- scram:
enabled: true
algorithm: SCRAM-SHA-256 # or SCRAM-SHA-512
user: kafka-mcp-readonly
pass: "{env:KAFKA_PASSWORD}" # or password_file: /run/secrets/kafka
preprod:
brokers: kafka-preprod:9093
endpoints:
prod-read:
cluster: prod
path: /mcp
description: Production investigation and debugging
read_only: true
prod-write:
cluster: prod
path: /mcp/rw
description: Approved production changes
tools:
copy_message: false
preprod:
cluster: preprod
path: /mcp/preprodclusters owns Kafka connection details: brokers, TLS and SASL. endpoints
owns the policy and, in HTTP mode, the MCP route. Stdio selects an endpoint by
name and does not use its path. Several endpoints may reference one cluster, so
the example reuses one production connection at /kafka-mcp/mcp in read-only
mode and /kafka-mcp/mcp/rw in writable mode. Paths are exact: /mcp does not
capture /mcp/rw. description is optional and is reported by server_config
so a caller knows what the endpoint is intended for.
http.base_path prefixes endpoint paths and the liveness route. A leading or
trailing slash on an endpoint path is optional. Paths must be unique and may
not be /, /healthz, or contain a query or fragment. When path is omitted,
it defaults to /mcp/<endpoint-name>.
For compatibility, a file with no endpoints block still creates one endpoint
per cluster at /mcp/<cluster-name>. Existing cluster-level read_only and
tools values continue to apply as a lower bound during migration; an explicit
endpoint cannot widen them. New configurations should put both fields under
endpoints.
An endpoint may switch individual tools off, by name:
endpoints:
prod-read:
cluster: prod
path: /mcp
read_only: true
tools:
create_topic: false
commit_offset: falseA tool the map does not mention stays exposed, so the file states only what it
withholds rather than relisting every tool and silently losing whatever is added
later. Names are exact and lowercase, as listed by server_config. They are
checked at startup: a name that is not a tool stops the
server, because a typo would leave the tool it was meant to withhold exposed.
server_config cannot be switched off, since it is how a session learns which
cluster and policy it reached and which tools that endpoint has. The tools
map only narrows an endpoint: read_only: true still withholds the writing
tools regardless of what the map says.
A minimal local file:
clusters:
local:
brokers: localhost:19092
schema_registry:
urls: ["http://localhost:18081"]
endpoints:
local:
cluster: local
path: /mcp/localRun a custom configuration over stdio with
CONFIG_FILE=/path/to/config.yaml go run ./cmd/server. If it defines several
endpoints, select one with --endpoint <name>. Add --server to serve every
configured endpoint over HTTP instead; --endpoint is not used in HTTP mode.
Without CONFIG_FILE, chu discovers kafka-mcp.{toml,yaml,yml,json} first in
the working directory, then in the operating system's user config directory
(~/.config/kafka-mcp/ on Linux), and finally in /etc. It uses the first
matching file rather than merging files. Its standard loader order is defaults,
file, HTTP, then environment; environment overrides use the KAFKA_MCP_ prefix (for example,
KAFKA_MCP_HTTP_ADDRESS=:9000 or
KAFKA_MCP_HTTP_BASE_PATH=/kafka-mcp). Logging can be configured with
LOG_LEVEL and LOG_PRETTY.
At least one cluster with a broker is required. http.address defaults to
:8090, http.base_path defaults to the HTTP root, and output_dir defaults
to the system temp directory. Exports are confined to that directory:
output_file takes a file name, never a path.
brokers accepts either one address as a scalar or several addresses as a YAML
list. Each address is passed to Kafka as a separate seed broker.
Keep secrets out of the file with {env:VAR} or password_file. Unknown fields
are ignored by chu.
TLS is configured with twmb/tlscfg, using its TLS 1.2 minimum and recommended
cipher suites. TLS uses the system trust store;
ca_file adds a custom CA. For mTLS, supply both cert_file and key_file.
security.sasl is a preference-ordered list: each entry enables either scram
or plain, or uses oauth as shown below. For PLAIN, use plain: {enabled: true, user: alice, pass: "{env:KAFKA_PASSWORD}"}.
Both accept optional zid (authorization identity) and password_file instead
of pass; SCRAM also accepts is_token: true for delegation tokens. Algorithms
are case-insensitive. Disabled entries are ignored; repeated mechanisms and
entries enabling multiple mechanisms are rejected. Fallback negotiates a
broker-supported mechanism; it does not retry invalid credentials.
All Kafka connections, including message-reading sessions, use these settings.
Legacy per-cluster tls and sasl: {mechanism, user, password, password_file}
remain supported, but cannot be combined with security on the same cluster.
For OAUTHBEARER client credentials, use this cluster security block:
security:
tls:
enabled: true
sasl:
- oauth:
enabled: true
token_url: https://identity.example.com/realms/apps/protocol/openid-connect/token
client_id: kafka-mcp
client_secret: "{env:KAFKA_CLIENT_SECRET}"
scopes: [kafka]
timeout: 10s
# zid: optional-authorization-id
# extensions: {tenant: example}Alternatively, use oauth: {enabled: true, token: "{env:KAFKA_TOKEN}"} for a
static token. Configure exactly one mode. Client-credentials tokens are fetched
on authentication, cached per cluster across admin and reading connections,
and renewed near expiry using expires_in. Token requests use the current
authentication context and a timeout (default 10 seconds). Static tokens are
not renewed. The broker must support OAUTHBEARER and trust the identity provider.
Broker security.tls settings apply to Kafka connections; HTTPS token endpoints
use the HTTP client's system trust store.
A cluster that cannot be reached at startup is still served, and
list_clusters reports it as disconnected. One cluster being down must not
block debugging the others.
Message formats
Messages are decoded to JSON for every tool that shows or searches them, by the first rule that applies:
The topic has a format in
topic_formats.The bytes carry a Schema Registry header (Confluent wire format) and the cluster has a
schema_registry: Avro, Protobuf (with message index and referenced schemas) or JSON Schema.Otherwise JSON, then UTF-8 text, then base64.
clusters:
prod:
brokers: kafka-1:9093
schema_registry:
urls: ["https://schema-registry:8081"]
user: kafka-mcp # or bearer_token: "{env:SR_TOKEN}"
password: "{env:SR_PASSWORD}" # or password_file
# tls: {enabled: true, ca_file: /etc/sr/ca.pem}
topic_formats: # topics whose messages carry no schema id
"orders.*": # exact name or glob; longest match wins
value:
format: protobuf
proto_files: [shop/order.proto]
import_paths: [./protos] # or descriptor_set: ./orders.binpb
message_type: shop.Order
events:
value: {format: avro, schema_file: ./schemas/event.avsc}
metrics:
key: {format: text}
value: {format: msgpack}Formats are avro, protobuf, json, msgpack, text and binary. Schema
files and .proto sources are loaded at startup, so a missing file or type stops
the server. A message that names a schema but cannot be decoded is returned as
base64 with decode_error saying why, never silently. JSON output has sorted
keys; Avro longs beyond 2^53 keep their exact value; unset Protobuf fields are
shown with their defaults.
Each rendered message reports format, schema_id and message_type for the
value, and key_format / key_schema_id when the key is not plain text.
server_config reports the endpoint name, exact path, description, policy and
the cluster it serves, and never the password. Its authentication and
sasl_user describe the first configured option; sasl_options lists all
configured identities in preference order, not the mechanism negotiated by an
individual broker connection. It also reports schema_registry URLs and
topic_formats.
Browser clients (CORS)
A command-line client sends no Origin header and needs none of this. A client
running in a browser does: the browser discards the response unless the server
allows the origin, so the defaults already cover the MCP transport.
http:
address: ":8090"
cors:
allow_origins: ["https://mcp-client.example"]
allow_methods: [GET, POST, DELETE, OPTIONS]
allow_headers: [content-type, accept, authorization, cache-control, last-event-id, mcp-session-id, mcp-protocol-version]
expose_headers: [Mcp-Session-Id]
allow_private_network: true
max_age: 600http.cors is ada's CORS middleware configuration, read straight from the
file, so every option that middleware has is available here. Each key is
optional and keeps its own default, so setting allow_origins alone does not
drop the rest. The defaults are the values above with allow_origins: ["*"].
Three of them are load-bearing for the MCP transport. allow_methods needs
GET, POST and DELETE: requests are posted, the event stream is a GET,
and a client ends its session with DELETE. allow_headers needs
mcp-session-id and mcp-protocol-version, which the client sends from the
second request onwards, and a header missing there fails the whole preflight
rather than being dropped. Mcp-Session-Id must stay in expose_headers: the
session id arrives on the initialize response, and a page that cannot read
it cannot make a second call.
allow_private_network answers Chrome's Private Network Access preflight,
which a page on a public address must pass before it may reach a server on a
private or loopback address; it defaults to on, and is granted only on a
preflight the rest of the policy already allowed. Setting allow_credentials
together with a wildcard allow_origins is refused by the middleware at
startup unless unsafe_wildcard_origin_with_allow_credentials is also set,
which it should not be.
These endpoints have no authentication of their own. An allowed origin can
drive every tool with the server's Kafka credentials, from any page the
browser's user happens to visit. allow_origins is the only built-in HTTP
barrier, so narrow it to the pages that should have that power, protect writable
paths in a reverse proxy, and set read_only: true on endpoints that should not
write. A less obvious path such as /mcp/rw is not authentication.
Connecting a client
For the default stdio transport, configure the client to launch the binary. A configuration with one endpoint needs no arguments:
{
"mcp": {
"kafka-local": {
"type": "local",
"command": ["kafka-mcp"],
"environment": {"CONFIG_FILE": "/path/to/kafka-mcp.yaml"}
}
}
}When the file contains several endpoints, add the endpoint selection to the
command, for example "command": ["kafka-mcp", "--endpoint", "prod-read"].
The process refuses to start without it so it cannot silently connect a session
to the wrong cluster or permission policy.
To use HTTP, start kafka-mcp --server and register one remote client entry per
endpoint:
{
"mcp": {
"kafka-local": {
"type": "remote",
"url": "http://localhost:8090/mcp/local"
},
"kafka-prod-read": {
"type": "remote",
"url": "http://localhost:8090/mcp",
"enabled": false
}
}
}The key becomes the tool prefix, so these appear as kafka-local_list_topics
and kafka-prod-read_list_topics. Name entries after both the cluster and the
endpoint policy: the prefix is the clearest signal of what a call can hit and
change. Use server_config to confirm rather than trusting the client-side
name.
Each enabled cluster costs context: these tools are roughly 8k tokens of definitions. Enable only what you need, and put production behind an agent:
{
"tools": { "kafka-prod*": false },
"agent": {
"kafka-prod": {
"description": "Debugging against production Kafka.",
"tools": { "kafka-prod*": true }
}
}
}Permissions
Three layers, and only one of them is real security:
Layer | Protects against | Real security? |
| An LLM changing things on one ambiguous request | No — a guardrail |
endpoint | Accidental writes with the server's Kafka identity | No — anyone who can edit the config can turn it off |
Kafka ACLs on the SASL principal | An unauthorised person | Yes — the broker decides |
What a read-only endpoint exposes
read_only: true does more than refuse a write: the endpoint does not list the
tools whose only purpose is to write. add_partitions, alter_topic_config,
commit_offset, create_topic, delete_consumer_group, delete_records and
delete_topic are absent from tools/list on a read-only endpoint, so a client never sees a tool it could not have used, and their
preview cannot describe a change this endpoint would never apply.
A writable endpoint can withhold individual tools too, with its tools map.
That is the same mechanism seen from the client: the tool is
not registered, so it is absent from tools/list and from server_config.
copy_message and produce_message stay, because read_only protects the
cluster being written to and the destination is chosen per call. Copying a
message out of a read-only production cluster, or seeding a writable preprod
cluster from a protected session, is exactly what they are for. Both refuse
outright when the destination is the read-only cluster itself.
Hiding a tool decides what is advertised, not what is permitted: both tools
still refuse at the point of mutation, so a registration mistake cannot turn
into a write. server_config reports the tools the endpoint actually exposes,
which is how a session can tell the two cases apart.
Giving two people different permissions
Kafka enforces permissions against the SASL principal, so two people running
the same server with different credentials get different rights. Adding
partitions requires ALTER on the topic:
# ali may read but not reshape topics
rpk acl create --allow-principal User:ali \
--operation read,describe --topic orders
# deniz may also add partitions
rpk acl create --allow-principal User:deniz \
--operation read,describe,alter --topic ordersEach points CONFIG_FILE at their own file, differing only in security.sasl[0].scram.user
and the password. When ali calls add_partitions, the broker refuses:
not authorized to add partitions to "orders": the broker refused this request.
Adding partitions requires ALTER permission on the topic for the principal
this server connects asali cannot bypass that by editing config or rebuilding the binary, because the
decision is made by Kafka rather than by this server. On a cluster without
ACLs, read_only: true is the available protection.
Audit logging
Every tool call is logged. One slog record per call, on the server's own log
stream, so nothing extra has to be configured:
{"level":"INFO","msg":"tool call","tool":"produce_message","outcome":"ok",
"duration_ms":12,"endpoint":"prod-write","cluster":"prod","read_only":false,
"principal":"kafka-mcp-rw","session":"QY7MZM...","client":"claude-code",
"client_version":"1.0.0","confirm":true,"item_count":1,"targets":"orders"}The tools that change a cluster — add_partitions, alter_topic_config,
commit_offset, create_topic, copy_message, delete_consumer_group,
delete_records, delete_topic, produce_message — are logged at INFO.
Everything else is logged at DEBUG, because reads are constant and change
nothing, so recording them at the same level would bury the writes among them.
Raise the log level to see them.
Two fields carry most of the weight. confirm separates a real write from a
preview, and targets names the topic, group, partition and offset each item
pointed at, so a record says which topic was touched rather than only that some
topic was.
Message content is never logged. produce_message and copy_message carry
arbitrary payloads, and an audit log is usually readable by more people than the
data it describes, so keys, values and headers are left out. The audit code has
no field to unmarshal them into, so content cannot reach a log even by mistake.
None of the recorded identities is a person, and the server cannot make one up. What each actually means:
Field | What it is |
| The Kafka credential this server connects as, which is what ACLs are enforced against. Everyone reaching the same endpoint shares it |
| Which endpoint policy allowed the call, and so which cluster was touched |
| One MCP session, which groups a sequence of calls into one investigation |
| The program that connected, as it identified itself at |
| Present only when an inbound bearer token established it. This server installs no token verifier, so it is absent today |
| The HTTP request id, which joins a record to the access log |
For attribution to a person, give each person their own endpoint and SASL
credentials, as above: the principal in the record is then the answer to who
acted. A shared credential cannot be made to answer it.
Tools
Batch operations
describe_topic, sample_messages, get_message, get_schema,
consumer_lag, describe_consumer_group, open_transactions,
add_partitions, alter_topic_config, create_topic, commit_offset,
delete_consumer_group, delete_records, delete_topic, copy_message and
produce_message take their target only as a required items array. There
is no single-target form: one operation is an items array of length one.
{"items": [{"topic": "orders"}]}
{"items": [{"topic": "orders"}, {"topic": "payments"}]}Everything naming or shaping an operation lives on the item, so each field and
its description exist in exactly one place. What governs the whole call stays at
the top level, which in practice means confirm.
The message-heavy tools accept at most 20 items; lag, schema and administrative tools accept at most 100.
Every response is the same envelope. Results stay in input order, each entry
carrying index and either result or error, followed by succeeded,
failed, applied and atomic: false. An item's failure is data in the
response rather than an error for the call, so it never hides the items that
worked. Batch writes preview every item first, then apply the valid ones only
when the top-level confirm is true. They are not transactions: Kafka cannot
roll back a topic, partition, offset or produced message after a later item
fails. Duplicate write targets are refused before anything changes, except in
produce_message, where two identical items mean two messages rather than a
mistake.
list_clusters
Lists the clusters this server serves, with whether each is reachable and
whether it accepts writes. Takes no parameters. Available from every endpoint,
so a session can discover what copy_message, produce_message and
compare_clusters may target.
{"clusters": [
{"name": "prod", "connected": true, "read_only": true},
{"name": "preprod", "connected": true, "read_only": false}
], "count": 2}connected is checked when you call, not recorded at startup, so a cluster
that has since gone down is reported honestly. Only the name, reachability and
writability are reported: broker addresses and credentials are deliberately
not, because this tool is reachable from every endpoint.
compare_clusters
Compares the topics of other clusters against the one this endpoint serves, and reports what differs. Use it to find what preproduction has that production does not, or to check whether two environments still match.
Parameter | Type | Required | Meaning |
| object[] | yes | 1 to 100 clusters to compare against |
Item fields:
Field | Type | Required | Meaning |
| string | yes | The other cluster. Use |
| string | no | Case-insensitive substring a topic name must contain |
| bool | no | Default false: internal topics are excluded |
{"here": "preprod", "there": "prod",
"here_cluster": {"name": "preprod", "brokers": 1, "topics": 12},
"there_cluster": {"name": "prod", "brokers": 3, "topics": 11},
"only_here": [{"topic": "orders-v2", "partitions": 6, "replication_factor": 1,
"configs": {"retention.ms": "604800000"}}],
"only_there": [],
"differing": [{"topic": "orders", "differences": ["partitions"],
"here": {"partitions": 1}, "there": {"partitions": 12}}],
"in_both": 11}only_here and only_there name the direction, which is decided by the
endpoint you call: "here" is always the cluster this endpoint serves. Entries
carry the partition count, replication factor and explicitly-set configs of the
cluster that has the topic, so they can be passed straight to create_topic.
This tool creates and changes nothing. To create the missing topics, hand
the chosen entries to create_topic, which previews them against the broker
first and warns that a partition count can never be reduced.
differing is usually the more valuable half: a topic that exists on both sides
with a different partition count or retention.ms is the common reason a bug
reproduces in one environment and not the other. Only configs a topic sets for
itself are compared, because two clusters may carry different broker defaults
and comparing inherited values would report every topic as different.
A difference is not necessarily a mistake. A topic missing from production is often deliberate, so the report says what differs, never what is correct.
Topic listings come from the Kafka client's metadata cache, which is a few
seconds old, so a topic created moments earlier may still appear in
only_there. Repeat the comparison rather than creating it twice.
list_topics
Lists the topics on the cluster, sorted by name, each with its partition count, replication factor and the configs it sets for itself.
Parameter | Type | Required | Meaning |
| string | no | JavaScript predicate deciding whether a topic is listed |
| int | no | Limit for evaluating the script. Default 30 |
The predicate sees:
Variable | Type | Meaning |
| string | The topic name |
| number | Partition count |
| number | Replicas of the first partition |
| boolean | Kafka's own topics, such as |
| object | Values this topic sets for itself, e.g. |
return topic.indexOf('orders') >= 0
return partitions > 6
return configs['cleanup.policy'] === 'compact'
return replication_factor === 1 && !internal{"name": "list_topics", "arguments": {"script": "return partitions > 6"}}{"topics": [{"topic": "orders", "partitions": 12, "replication_factor": 3,
"configs": {"retention.ms": "604800000"}}],
"count": 1}Filtering is JavaScript only, as it is for search_messages: a name match is
return topic.indexOf('orders') >= 0. The predicate can also answer what a
substring never could — which topics have more than six partitions, only one
replica, or a compacted cleanup policy.
Only configs a topic sets for itself are reported. Inherited cluster defaults are excluded, because including them would make every topic look configured.
A topic whose predicate throws, or is cut short by the timeout, is counted in
script_errors rather than listed, so a broken filter is never mistaken for an
empty cluster.
describe_topic
Reports a topic's partitions, offset ranges, message count, time span and full configuration. Use it before searching to see how much data a search would read and how far back the topic can hold data at all.
Parameter | Type | Required | Meaning |
| object[] | yes | 1 to 20 topics to describe |
Item fields:
Field | Type | Required | Meaning |
| string | yes | Topic to describe |
{"results": [{"index": 0, "result": {
"topic": "orders", "partition_count": 1, "message_count": 3,
"partitions": [{"partition": 0, "start_offset": 0, "end_offset": 3, "message_count": 3}],
"configs": [{"key": "cleanup.policy", "value": "delete", "source": "DYNAMIC_TOPIC_CONFIG", "is_default": false},
{"key": "retention.ms", "value": "604800000", "source": "DEFAULT_CONFIG", "is_default": true}]}}],
"succeeded": 1, "failed": 0, "applied": 0, "atomic": false}configs lists every topic config as the string Kafka reports, where -1
means unlimited. is_default is true when the value is inherited rather than
set on the topic. Two entries decide whether a message can still exist at all:
retention.ms (how long messages are kept) and cleanup.policy (compact
keeps only the latest message per key).
sample_messages
Reads a small sample of the newest messages and reports what they look like: value formats, field paths with their types, key statistics, which value fields carry the message key, and the schemas the values were written with. Avro, Protobuf and JSON Schema values are decoded first, so their field paths are reported like JSON. Use it before searching to decide how to search, and before producing to find the schema to write against.
Parameter | Type | Required | Meaning |
| object[] | yes | 1 to 20 topics to sample |
Item fields:
Field | Type | Required | Meaning |
| string | yes | Topic to sample |
| int | no | Messages to read in total. Default 20 |
| int[] | no | Restrict to these partitions |
| int | no | Value bytes per message. Default 512 |
{"results": [{"index": 0, "result": {
"value_formats": {"json": 20, "text": 0, "binary": 0},
"json_fields": [{"path": "payload.amount", "types": ["number"], "present": 20, "example": "500"}],
"key_stats": {"present": 20, "absent": 0, "unique": 20, "all_unique": true},
"key_in_value": ["payload.orderId"],
"schemas": [{"format": "avro", "schema_id": 7, "message_type": "shop.Order", "count": 20}],
"sampled_ranges": [{"partition": 0, "start": 980, "end": 1000}]}}],
"succeeded": 1, "failed": 0, "applied": 0, "atomic": false}key_in_value naming a field means the key is that identifier, so searching
the key alone is the precise, cheap lookup. value_formats also counts avro,
protobuf, json_schema, msgpack, null and undecodable — values that
named a schema but could not be decoded, which is a configuration problem, not
binary data.
search_messages
Scans a bounded range of a topic, filtering messages with a JavaScript expression. Kafka has no server-side search, so this reads messages and filters them client-side; the result reports what was covered.
Parameter | Type | Required | Meaning |
| string | yes | Topic to search |
| string | no | JavaScript filter. Omit to match every message |
| int | no | Concurrent readers, 1–16. Default 1 |
| int[] | no | Restrict to these partitions |
| int | no | Offset window, end exclusive |
| string | no | RFC3339 time window |
| string | no |
|
| int | no | Stop after this many matches. Default 10 |
| int | no | Read at most this many. Default 10000 |
| int | no | Value bytes per match. Default 512 |
| int | no | Wall-clock limit. Default 30 |
| bool | no | Return counts only, no message bodies |
| string | no | Write every match to this file as JSONL |
The script
Return true to keep a message. In scope:
Name | Value |
| parsed JSON document; the decoded record for Avro, Protobuf, JSON Schema or a configured format; the raw text otherwise |
| string, the decoded document when the key has a schema, or |
| object of header name to string |
| numbers |
| a |
return key === 'order-123'
return value.eventType === 'NEW' && value.payload.amount >= 500
return value.payload.cancelledAt === null // present and null
return value.payload.cancelledAt === undefined // field absent
return headers['correlation-id'] === 'corr-999'
return /ORD-\d{4}/.test(value.payload.orderId)
return value.indexOf('ERROR') >= 0 // non-JSON topic: value is a stringSearching by key exactly is far more precise than searching the body: 123
also appears inside "amount": 1123, and those false positives can fill
max_matches and hide the message wanted.
A script that throws on a message is counted in script_errors and the scan
continues, so a broken script is distinguishable from a genuine absence of
matches. Scripts run in a sandbox with no filesystem, network or host access,
and are stopped if they exceed the search timeout. They are not bounded by
memory: something like 'x'.repeat(1e12) can exhaust the server process.
Scripts are trusted input; the blast radius is this server, not the cluster.
Scan order
Every partition is read together, one chunk deep at a time: the newest chunk of every partition, then the chunk behind it, and so on. A limited newest-first search therefore returns the newest matches in the topic, not the newest in whichever partition happened to be read first.
Kafka orders records within a partition and never across them, so matches are
merged and reported by timestamp, with partition and offset breaking ties.
Timestamps are set by the producer unless the topic uses LogAppendTime, so
they can be skewed; it is still the only thing comparable between partitions.
One scan reads every partition through a single connection, so a wide topic
costs no more connections than a narrow one. It does read more: a limited
search on a 12-partition topic examines the newest chunk of all twelve rather
than stopping inside the first. max_messages_scanned still bounds it, and may
become the stopped_reason on a wide topic sooner than on a narrow one.
Parallelism
parallelism splits a single-partition topic's offset range between that
many readers, which is what makes a full scan of one large partition fast. A
multi-partition topic is already read in parallel across its partitions, so the
setting does not apply there, and a range too small to divide is read by one
reader.
It pays off for count_only, output_file and full scans of a single
partition.
Result
{"topic": "orders", "match_count": 1,
"matches": [{"partition": 0, "offset": 17, "key": "order-42", "value": "...", "encoding": "utf8"}],
"scanned_messages": 120, "scanned_ranges": [{"partition": 0, "start": 0, "end": 120}],
"stopped_reason": "range_exhausted", "complete": true}complete is true only when the whole range was read. An empty match list
means "not there" only if complete is true; otherwise check stopped_reason
and narrow the search.
For large result sets, use count_only to learn how many matches exist, then
output_file to write them out instead of returning them.
get_message
Reads messages at exact offsets, plus optional neighbours.
Parameter | Type | Required | Meaning |
| object[] | yes | 1 to 20 addresses to read |
Item fields:
Field | Type | Required | Meaning |
| string | yes | Topic to read from |
| int | yes | Partition to read from |
| int | yes | Exact offset to read |
| int | no | Also return this many messages either side |
| int | no | Value bytes to return. Default 4096 |
{"name": "get_message", "arguments": {"items": [
{"topic": "orders", "partition": 0, "offset": 17, "context": 1}]}}Schema-encoded values are decoded to JSON, with format, schema_id and
message_type saying what they were. Values that could not be decoded and are
not valid UTF-8 are base64 encoded, with encoding set to base64;
decode_error explains a value that named a schema but could not be decoded.
{"partition": 0, "offset": 17, "key": "o-1",
"value": "{\"amount\":42,\"id\":\"o-1\"}", "format": "avro",
"schema_id": 7, "message_type": "shop.Order", "encoding": "utf8", "value_bytes": 12}get_schema
Reads schemas from the cluster's Schema Registry, by subject or by the
schema_id a message carries. Read the schema before producing to a
schema-encoded topic: it names every field and enum a value needs, which a
sampled message may not show.
Parameter | Type | Required | Meaning |
| object[] | yes | 1 to 100 schemas to look up |
Item fields, giving subject or id:
Field | Type | Required | Meaning |
| string | either | Subject, usually |
| int | no | Subject version. Defaults to the latest |
| int | either | Schema id, e.g. from |
{"results": [{"index": 0, "result": {
"schema_id": 7, "subject": "orders-value", "version": 3, "versions": [1, 2, 3],
"type": "avro", "schema": "{...}", "references": [],
"message_types": null, "used_by": null}}],
"succeeded": 1, "failed": 0, "applied": 0, "atomic": false}message_types lists Protobuf messages, which produce_message takes as
message_type. Looking up an id fills used_by with the subject versions
that use it. The call fails as a whole when the cluster has no
schema_registry.
list_consumer_groups
Lists consumer groups with their state, member count and the topics they consume.
Parameter | Type | Required | Meaning |
| string | no | Only groups consuming or committed to this topic |
| string[] | no | Filter by state, e.g. |
{"groups": [{"group": "payments", "state": "Stable", "members": 2, "topics": ["orders"]}], "count": 1}A group in state Empty can still report lag: committed offsets outlive the
consumers that made them. Kafka has no topic-to-group index, so filtering by
topic describes every group on the cluster.
describe_consumer_group
Describes consumer groups in detail: who the members are, which partitions each one owns, and where the group stands on every partition. Use it to turn "a partition is stuck" into "this pod on this host is stuck".
Parameter | Type | Required | Meaning |
| object[] | yes | 1 to 100 groups, each |
{"results": [{"index": 0, "result": {
"group": "payments", "state": "Stable", "protocol_type": "consumer", "assignor": "cooperative-sticky",
"coordinator": 1, "total_lag": 4200,
"members": [{"member_id": "payments-7-…", "client_id": "payments-7", "host": "/10.0.4.17",
"assignments": [{"topic": "orders", "partitions": [0, 1]}]}],
"partitions": [{"topic": "orders", "partition": 0, "has_commit": true, "committed_offset": 812,
"end_offset": 5012, "lag": 4200, "member_id": "payments-7-…",
"client_id": "payments-7", "host": "/10.0.4.17"}]}}],
"succeeded": 1, "failed": 0, "applied": 0, "atomic": false}Partitions are the union of what members own and what the group has committed,
so an Empty group still shows its positions. has_commit: false means the
group owns a partition it has never committed on, so where it starts is decided
by the consumer's auto.offset.reset, not by an offset.
open_transactions
Finds open transactions holding back read_committed consumers. A transactional
producer that hangs or dies mid-transaction leaves the partition's last stable
offset stuck, and every read_committed consumer stops there. In
consumer_lag that looks exactly like a poison message.
Parameter | Type | Required | Meaning |
| object[] | yes | 1 to 100 topics, each |
{"results": [{"index": 0, "result": {
"topic": "orders", "blocked": true,
"partitions": [{"partition": 0, "last_stable_offset": 812, "high_watermark": 5012,
"unreadable_messages": 4200,
"producers": [{"producer_id": 2004, "producer_epoch": 0, "transaction_start_offset": 812,
"transactional_id": "payments-writer-1", "state": "Ongoing",
"started_at": "2026-10-02T08:14:03Z", "open_for": "41m12s", "timeout_ms": 900000}]}]}}],
"succeeded": 1, "failed": 0, "applied": 0, "atomic": false}The fix is in the producer: restart or fence the application named by
transactional_id, or wait for timeout_ms, after which the broker aborts the
transaction. Moving the consumer's offset does not help.
consumer_lag
Measures how far behind a topic's consumers are, how fast messages are produced and consumed, and when the backlog will clear.
Parameter | Type | Required | Meaning |
| object[] | yes | 1 to 100 measurements to take |
Item fields:
Field | Type | Required | Meaning |
| string | yes | Topic to measure |
| string | no | Defaults to every group consuming the topic |
| int | no | Consume-rate sample window. Default 5. The call blocks |
| bool | no | Return immediately, without a rate or estimate |
Sampling windows run concurrently, so several measurements do not add their waits together.
{"results": [{"index": 0, "result": {
"topic": "orders", "total_lag": 4200,
"produce_rate": {"last_minute": {"messages": 3000, "per_second": 50, "per_minute": 3000, "per_hour": 180000}},
"groups": [{"group": "payments", "state": "Stable", "members": 2, "lag": 4200,
"consume_rate": {"per_second": 120, "sampled_seconds": 5},
"drain_per_second": 70, "eta_seconds": 60, "eta_human": "1m 0s",
"status": "draining"}]}}],
"succeeded": 1, "failed": 0, "applied": 0, "atomic": false}The two rates are measured differently, and the output says so:
produce_rate— historical fact, from message timestamps, over the last second, minute and hour.window_truncatedmarks a topic younger than the window.consume_rate— a sample: the committed offset is read, then read againsample_secondslater.sample_inconclusivemeans nothing moved.
The backlog drains at the consume rate minus the produce rate. status
says what the numbers mean: caught_up, draining (with an ETA), growing
(never clears, with growing_by_per_minute), stalled, no_active_consumers,
or not_measured. An ETA is only given when the lag is genuinely shrinking.
cluster_health
Checks the cluster this endpoint serves in one call: brokers, controller, and every partition that is not fully healthy.
Parameter | Type | Required | Meaning |
| string | no | Only topics whose name contains this. Case-insensitive |
| bool | no | Also check |
{"cluster_id": "…", "controller": 1, "healthy": false,
"brokers": [{"id": 1, "host": "kafka-1", "port": 9092, "rack": "eu-1a", "controller": true, "leaders": 61}],
"summary": {"topics": 40, "partitions": 182, "offline": 0, "under_replicated": 3, "under_min_isr": 1, "errored": 0},
"problems": [{"topic": "orders", "partition": 4, "issues": ["under_replicated", "under_min_isr"],
"leader": 1, "replicas": [1, 2, 3], "isr": [1], "min_insync_replicas": 2}],
"warnings": []}offline means the partition has no leader, so nothing can be read or written.
under_replicated means a replica is out of sync. under_min_isr is the
condition behind NOT_ENOUGH_REPLICAS: producers using acks=all fail until
the in-sync replicas recover. A broker leading no partitions while others lead
many is usually one that restarted and was never given leadership back.
Some Kafka-compatible brokers, Redpanda among them, do not report
min.insync.replicas. The topics affected are listed in min_isr_unknown, and
under_min_isr is not judged for them rather than guessed.
list_acls
Lists access control entries, for when a client fails with
TOPIC_AUTHORIZATION_FAILED or GROUP_AUTHORIZATION_FAILED.
Parameter | Type | Required | Meaning |
| string | no | Only this principal, e.g. |
| string | no |
|
| string | no | Every ACL applied to this name, including prefixed and |
{"acls": [{"principal": "User:payments", "host": "*", "resource_type": "topic", "resource_name": "orders",
"pattern_type": "literal", "operation": "read", "permission": "allow"}], "count": 1}A deny overrides every allow. A broker without an authorizer fails with
SECURITY_DISABLED, which means ACLs are not enforced at all.
server_config
Reports the effective configuration: endpoint name, exact path, description, cluster, brokers, authentication mechanism and principal, TLS, read-only state, export directory and the tools this endpoint exposes. Takes no parameters. The password is never reported.
tools is the list for this endpoint, not for the deployment: a read-only
endpoint omits add_partitions, alter_topic_config, commit_offset,
create_topic, delete_consumer_group, delete_records and delete_topic,
because it
does not register them, and any endpoint omits whatever its tools
configuration switches off.
Use it when a result is surprising: an empty topic list means something very different on a local broker than on production.
add_partitions
Adds partitions to a topic. Irreversible — Kafka cannot reduce a partition count. Not exposed on a read-only endpoint.
Parameter | Type | Required | Meaning |
| object[] | yes | 1 to 100 topics to change |
| bool | no | Default false: preview only, nothing changes |
Item fields:
Field | Type | Required | Meaning |
| string | yes | Topic to change |
| int | yes | Final total, not the number to add. Repeating a call is safe |
| bool | no | Required when messages are keyed |
| int | no | Messages inspected for keys. Default 20 |
Without confirm it reports what would happen: current and target counts,
whether messages are keyed, which consumer groups will rebalance, and warnings.
Adding partitions changes which partition a key hashes to, so existing keys
lose their ordering guarantee. A keyed topic therefore requires
acknowledge_key_ordering as well. Requesting fewer partitions than the topic
has is refused with an explanation rather than attempted.
alter_topic_config
Changes topic-level configuration: retention, cleanup policy, maximum message size and anything else Kafka allows per topic. Changes are incremental, so every key not named keeps its value. Not exposed on a read-only endpoint.
Parameter | Type | Required | Meaning |
| object[] | yes | 1 to 100 topics to change |
| bool | no | Default false: preview only, nothing changes |
Item fields:
Field | Type | Required | Meaning |
| string | yes | Topic to change |
| object | no | Keys to set, e.g. |
| string[] | no | Overrides to remove, so the cluster default applies again |
{"results": [{"index": 0, "result": {
"topic": "orders", "applied": false, "messages_past_retention": 18000,
"changes": [{"key": "retention.ms", "action": "set", "current": "604800000",
"current_source": "DYNAMIC_TOPIC_CONFIG", "requested": "86400000"}],
"warnings": ["18000 message(s) are already older than the new retention of 24h0m0s and become eligible for deletion as soon as it applies; Kafka cannot bring them back"]}}],
"succeeded": 1, "failed": 0, "applied": 0, "atomic": false}The preview asks the broker to validate the change, so an unknown key or an
invalid value is refused before confirm. Shortening retention.ms reports how
many messages are already past the new limit; changing cleanup.policy or
setting retention.bytes is warned about. After applying, every key is re-read
from the broker, which is the only way the inherited value of a deleted override
is known.
create_topic
Creates a topic. Refuses a topic that already exists rather than adjusting it. Not exposed on a read-only endpoint.
Parameter | Type | Required | Meaning |
| object[] | yes | 1 to 100 topics to create |
| bool | no | Default false: the broker validates the requests and creates nothing |
Item fields:
Field | Type | Required | Meaning |
| string | yes | Name of the topic to create |
| int | no | Omit for the broker default on Kafka 2.4+. Can grow later, never shrink |
| int | no | Omit for the broker default on Kafka 2.4+. Cannot exceed the broker count |
| map | no | Topic-level config, such as |
Without confirm the request is sent to the broker with ValidateOnly, so the
preview reports the cluster's own answer — an invalid name, an unknown config
key, a replication factor larger than the cluster — rather than a guess. With
confirm the topic is created and the resulting partition count and
replication factor are read back from the cluster, which is how an omitted
count is reported as the number the broker actually chose.
Using broker defaults requires Kafka 2.4 or newer, whose CreateTopics v4 API
introduced -1 as "use the broker default". On an older broker, pass both
counts explicitly.
A replication factor larger than the number of brokers is refused here, with the broker count in the message, because brokers differ on whether a validate-only request catches it.
Use add_partitions to change an existing topic's partition count; this tool
never modifies a topic it did not create.
delete_topic
Deletes topics. The most destructive tool here: deleting a topic destroys every message in it, Kafka has no undo, and any consumer group reading it breaks. Not exposed on a read-only endpoint.
Parameter | Type | Required | Meaning |
| object[] | yes | 1 to 100 topics to delete |
| bool | no | Default false: preview only, nothing is deleted |
Item fields:
Field | Type | Required | Meaning |
| string | yes | Topic to delete. It must exist |
| bool | no | Required when the topic still holds messages |
Without confirm it reports, per topic, how many messages would be destroyed,
how many partitions it had, and which consumer groups had committed offsets for
it:
{"results": [{"index": 0, "result": {
"topic": "orders-old", "partitions": 6, "message_count": 41207,
"consumer_groups": ["payments"], "deleted": false, "would_delete": true,
"warnings": ["41207 message(s) would be destroyed, and Kafka cannot restore them: the only recovery is a backup taken beforehand"]}}],
"succeeded": 1, "failed": 0, "applied": 0, "atomic": false}Two separate acknowledgements are required, because they answer different
questions. confirm says the caller meant to delete; acknowledge_data_loss
says they know what is inside. A topic holding messages is refused without both,
and the count is re-read at deletion time, so a topic that gained messages since
the preview is still caught.
Internal topics such as __consumer_offsets are refused outright at any level
of acknowledgement: they hold cluster state rather than a caller's data, and
deleting one breaks every consumer at once.
Duplicate topics in one batch are refused before anything is deleted. Deletion is not atomic — topics removed before a later item failed stay removed.
delete_records
Deletes the oldest messages of a partition without deleting the topic. The topic, its configuration and its consumer groups stay. Irreversible. Not exposed on a read-only endpoint.
Parameter | Type | Required | Meaning |
| object[] | yes | 1 to 100 partitions to truncate |
| bool | no | Default false: preview only, nothing is deleted |
Item fields:
Field | Type | Required | Meaning |
| string | yes | Topic to delete from |
| int | yes | Partition to delete from |
| int | yes | Everything below this offset goes; this one becomes the first |
| bool | no | Required to apply |
{"results": [{"index": 0, "result": {
"topic": "orders", "partition": 0, "start_offset": 0, "end_offset": 5012, "before_offset": 812,
"messages_deleted": 812, "would_delete": true, "deleted": false,
"affected_groups": [{"group": "replay-job", "committed_offset": 100, "unprocessed_lost": 712}]}}],
"succeeded": 1, "failed": 0, "applied": 0, "atomic": false}affected_groups lists every group committed below the cut, with how many
messages it would lose without ever processing them; such a group resumes from
the new start. Use the partition's end offset as before_offset to empty it.
delete_consumer_group
Deletes consumer groups and their committed offsets. Use it to clean up groups whose consumers were decommissioned: their commits keep reporting lag that nobody will ever drain. Not exposed on a read-only endpoint.
Parameter | Type | Required | Meaning |
| object[] | yes | 1 to 100 groups, each |
| bool | no | Default false: preview only, nothing is deleted |
The preview lists each group's committed offsets and lag. A group with active
members is refused, and re-checked at deletion time. A consumer that later
starts with the same group id begins from its auto.offset.reset, not from
where the group left off.
commit_offset
Moves consumer groups' committed offsets. Forward to skip messages, backward to replay them, or to a point in time to reprocess everything since. Irreversible in the sense that skipped messages are never processed. Not exposed on a read-only endpoint.
Parameter | Type | Required | Meaning |
| object[] | yes | 1 to 100 moves to make |
| bool | no | Default false: preview only, nothing changes |
Item fields — give exactly one of offset, timestamp or position:
Field | Type | Required | Meaning |
| string | yes | Topic whose offset is moving |
| string | yes | Consumer group to move |
| int | no | Required with |
| int | — | The offset the group reads next. To skip offset 42, commit 43 |
| string | — | RFC3339. Each partition moves to its first message at or after this time |
| string | — |
|
| bool | no | Proceed despite running consumers |
{"items": [{"topic": "orders", "group": "payments", "timestamp": "2026-10-01T09:00:00Z"}]}{"results": [{"index": 0, "result": {
"topic": "orders", "group": "payments", "state": "Empty", "members": 0,
"partitions": [{"partition": 0, "current_offset": 5012, "target_offset": 4100, "replayed_messages": 912},
{"partition": 1, "current_offset": 4990, "target_offset": 4021, "replayed_messages": 969}],
"replayed_messages": 1881, "applied": false}}],
"succeeded": 1, "failed": 0, "applied": 0, "atomic": false}A partition with no message at or after timestamp moves to its end, and its
entry carries a note saying so. A whole-topic item and a partition item for
the same group and topic in one batch are refused, because which one wins would
depend on order.
The group must have no active members. A running consumer keeps its position in memory and only reads the committed offset when it joins, so a commit made while it runs is overwritten by its next commit and the group does not move. Stop the consumers first.
copy_message
Copies messages to another topic, preserving key, value and headers. Takes each message's address, never its content, so it can only duplicate a message the cluster already holds.
Parameter | Type | Required | Meaning |
| object[] | yes | 1 to 20 copies to make |
| bool | no | Default false: preview only, nothing is written |
Item fields:
Field | Type | Required | Meaning |
| yes | Message to copy | |
| string | yes | Where to write it. Must already exist |
| string | no | Another cluster to write to. Defaults to this endpoint's own |
| bool | no | Re-register a schema id in the destination's registry |
| int | no | Preview value limit. The whole value is always copied |
Set destination_cluster to copy into another cluster this server serves,
which is how a production message is taken into a preproduction topic to be
debugged safely. Use list_clusters to see which names are valid.
Every copy carries provenance headers — kafka-mcp-copied-from-cluster,
-from-topic, -from-partition, -from-offset, -copied-at,
-copied-by-tool, -copied-by-principal — so a message in a dead letter or
preproduction topic can be traced back to its original. If the message already
carries one of those headers, the original is kept and the collision is
reported.
Key and value bytes are copied unchanged. A schema id is only meaningful in the
registry that issued it, so when the destination cluster uses a different
registry, or none, the response warns. With translate_schema the schema is
registered in the destination registry under <destination_topic>-value (or
-key), with its references, and the id in the copy is rewritten; the payload
is untouched. That registration is a write to the destination registry and
happens only with confirm.
For a copy within the endpoint's own cluster, its read_only policy protects
the destination. A read-only endpoint can still be the source of a
cross-cluster copy, because copying out changes nothing there. A different
destination cluster is writable when it has at least one writable endpoint;
list_clusters reports that effective state. The tool is refused entirely
when the destination is read-only, preview included, because writing is all it
does.
produce_message
Writes new messages to existing topics. Unlike copy_message, the caller
supplies the content, so this can put a message into a topic that no producer
ever sent.
Parameter | Type | Required | Meaning |
| object[] | yes | 1 to 20 messages to write |
| bool | no | Default false: preview only, nothing is written |
Item fields:
Field | Type | Required | Meaning |
| string | yes | Existing topic to write to |
| string | yes | The message body |
| string | no | Decides the partition when |
| object | no | Header name to value |
| int | no | Exact partition. Omit to let the key decide |
| string | no |
|
| object | no | Encode |
| object | no | Encode |
| string | no | Another cluster to write to. Defaults to this endpoint's own |
| int | no | Preview value limit. The whole value is always written |
Every message carries kafka-mcp-produced-at, kafka-mcp-produced-by-tool,
kafka-mcp-produced-by-principal and, when the client identifies itself,
kafka-mcp-produced-by-client, so a fabricated message stays distinguishable
from a genuine one. A header the caller supplies under one of those names is
kept as given and the collision is reported.
value_schema and key_schema take {subject, version, id, message_type},
all optional: {} means the latest version of <topic>-value (or -key) in
the destination cluster's registry, id pins an exact schema, and
message_type picks a Protobuf message. The value is given as JSON, validated
against the schema in the preview — a missing or misspelt field is refused by
name — and written framed with the schema id, as registry-aware consumers
expect. A topic with a format in topic_formats is encoded to it without being
asked. Neither can be combined with encoding: base64. The response reports
the schema used in value_encoding / key_encoding.
The topic must already exist: a missing one is refused rather than left to
auto-creation. Omit partition unless the exact partition is the point — the
key decides placement, and naming a partition puts a keyed message where its
key does not hash to, which breaks ordering for that key. The response warns
whenever an explicit partition is used.
read_only protects the cluster being written to, so a read-only endpoint may
still produce into a different, writable cluster, and is refused outright —
preview included — when writing to its own. A produced message cannot be
deleted; it stays until retention removes it.
Skills
skills/kafka-debugging/SKILL.md is the one skill an agent loads. It routes to
the scenario guides under skills/kafka-debugging/references/, rather than
holding every workflow itself, so a session reads only the one it needs:
find-message.md— locating a message from something the user knows about it.check-lag.md— measuring lag and throughput, and judging when a backlog will clear.scale-partitions.md— deciding whether more partitions will help, and adding them safely.skip-poison-message.md— unblocking a consumer stuck on a message it cannot process, preserving the message first.create-topic.md— creating a topic with a partition count and retention chosen on purpose, including as acopy_messagedestination.produce-message.md— writing a message: repairing and re-injecting one, reproducing a failure in another cluster, or seeding a topic.compare-clusters.md— finding what differs between two environments, and creating the topics one of them is missing.delete-topic.md— removing a topic and everything in it, after the user has seen what that destroys.replay-messages.md— reprocessing everything since a moment, after a bug fix ships.tune-topic-config.md— changing retention, cleanup policy or message size on a topic, knowing what the change does to data already there.cluster-health.md— finding offline and under-replicated partitions behindNOT_ENOUGH_REPLICASand unreadable topics.purge-messages.md— deleting old messages from a partition while keeping the topic, or removing abandoned consumer groups.authorization-error.md— working out which ACL is refusing a client.
The umbrella also resolves the overlap between them: "the consumer is behind"
opens three of these guides, and consumer_lag's status is what decides which
one is right.
Development
Command | Purpose |
| Start / stop Redpanda and Console |
| Build with goreleaser into |
| Run the HTTP server from source |
| All tests, including container tests |
| Tests that need no containers |
| Vet all packages |
The documentation site lives in _docs (Vite, pnpm). pnpm install && pnpm dev
there serves it locally; pushing changes under _docs/ to main publishes it to
GitHub Pages through .github/workflows/docs.yml.
Tests run against real containers started by internal/domain/testenv (a Redpanda
broker plus Console), so Docker must be available for the full suite.
Start the HTTP server
Requires Go 1.27 or later. Configuration is loaded with chu; set CONFIG_FILE
to select a YAML or JSON file. into manages the process lifecycle, ada serves
HTTP with context-driven shutdown, and logi initializes structured logging.
make env-up # local Redpanda + Console
make run # runs with --server and serves kafka-mcp.local.yamlkafka-mcp.local.yaml is committed and points at the compose broker, so a
clone works without writing any configuration. make env-up publishes the broker
on localhost:19092, the Schema Registry on localhost:18081 and the Redpanda
Console on http://localhost:8080.
Check it is up:
curl http://localhost:8090/healthzAvailable Tools
24 toolsadd_partitionsA
Increase the partition count of 1 to 100 topics in one call through items, and report each topic's current count, affected consumer groups, sampled key usage and warnings. Changing one topic is an items array of length one.
Kafka cannot remove partitions, and changing the count can break ordering for keyed messages. No change is made unless confirm is true; one confirm covers the whole batch, and changes are not atomic because Kafka cannot roll a successful item back. Results follow items order, each carrying index with result or error.
Keyed samples also require acknowledge_key_ordering on the item. Requires Kafka ALTER permission.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | The topics to change, 1 to 100 of them. Changing one topic is an array of length one. Duplicate topic names are refused before anything changes. | |
| confirm | No | Optional. When false or omitted, nothing is changed and the response describes what would happen for every item. Must be true to actually add partitions. One confirm covers the whole batch. |
Output Schema
| Name | Required | Description |
|---|---|---|
| atomic | Yes | |
| failed | Yes | |
| applied | Yes | |
| results | Yes | |
| succeeded | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses irreversibility (Kafka cannot remove partitions), the ordering risk for keyed messages, the confirm gate, batch-level non-atomicity with no rollback, item-ordered results with index, and the required Kafka ALTER permission. This is rich behavioral context an agent could not derive from structured fields alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded and sentences are dense with distinct facts (batch size, irreversibility, confirm gate, non-atomicity, result shape, permission). Minor redundancy in restating the single-topic-as-length-one case, but no filler sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although an output schema exists and return values need not be described, the definition still explains result ordering and the per-item index/result-or-error shape and covers confirmation, atomicity, permissions, and ordering risk. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds batch-level meaning the schema does not convey on its own: results follow items order and each carries index with result or error, one confirm covers the whole batch, and keyed samples additionally require acknowledge_key_ordering on the item. It exceeds the baseline with cross-parameter interaction guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource+scope: 'Increase the partition count of 1 to 100 topics in one call.' The batch framing and the 1-100 bound are concrete. It does not explicitly differentiate from the nearest sibling alter_topic_config, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear conditions for use (only increasing counts, ballled with confirm=true to take effect) and the constraint that Kafka cannot remove partitions. No explicit when-to-use-this-vs-alternative routing against siblings like alter_topic_config or create_topic, so not a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
alter_topic_configA
Change topic-level configuration of 1 to 100 topics in one call through items. Each item sets keys (set) and removes overrides so the cluster default applies again (delete). Every key not named keeps its value: changes are incremental, never a full replace. Changing one topic is an items array of length one.
No change is made unless confirm is true. The preview asks the broker to validate every item and returns, per key, the current value and its source (DYNAMIC_TOPIC_CONFIG means a deliberate topic setting; anything else is inherited) next to the requested value. Shortening retention.ms reports how many messages are already older than the new limit, because they become eligible for deletion straight away. Changing cleanup.policy or retention.bytes is warned about.
One confirm covers the whole batch, and applying is not atomic: topics changed before a later item failed stay changed. Results follow items order, each carrying index with result or error. Requires Kafka ALTER_CONFIGS permission.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | The topics to change, 1 to 100 of them. Changing one topic is an array of length one. Duplicate topic names are refused before anything changes. | |
| confirm | No | Optional. When false or omitted, nothing is changed: the broker validates every item and the response shows each key's current and requested value. Must be true to apply. One confirm covers the whole batch. |
Output Schema
| Name | Required | Description |
|---|---|---|
| atomic | Yes | |
| failed | Yes | |
| applied | Yes | |
| results | Yes | |
| succeeded | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden and discharges it: incremental rather than full-replace semantics, non-atomic partial application ('topics changed before a later item failed stay changed'), duplicate rejection, the DYNAMIC_TOPIC_CONFIG source meaning, and the ALTER_CONFIGS permission requirement. These are exactly the caveats an agent cannot infer from the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence and each subsequent sentence adds a distinct behavior (incremental, confirm gate, preview sources, warnings, non-atomicity, permission). It is dense prose rather than scannable bullets, which costs a point, but there is little wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a batch mutation tool with no annotations, the description covers everything an agent needs: permission, confirmation semantics, atomicity, per-item result shape, and which keys trigger warnings. An output schema also exists, so the description's brief nod to result structure is sufficient rather than required.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema: set/delete mutual exclusivity is stated structurally, 'one confirm covers the whole batch', and results map back to items order via index. The set/delete field-level details in the schema are not repeated, which is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource+scope: 'Change topic-level configuration of 1 to 100 topics in one call through items.' The qualifier 'topic-level' implicitly separates it from the cluster-level server_config sibling, and 'Every key not named keeps its value: changes are incremental' makes the modification model unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operating context: the confirm gate ('No change is made unless confirm is true'), the preview/validation flow, and explicit warnings for retention.ms, cleanup.policy, and retention.bytes. It never names an alternative tool (e.g. server_config vs topic config) or states when-not to use it, so it stops short of 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cluster_healthA
Check the health of the cluster this endpoint serves in one call: cluster id, controller, every broker (id, host, port, rack, and how many partitions it leads), a summary count, and every partition that is not fully healthy.
A problem partition lists its issues:
offline: no leader, so nothing can be produced to or consumed from it.
under_replicated: fewer in-sync replicas than replicas, so a broker is down or falling behind.
under_min_isr: fewer in-sync replicas than the topic's min.insync.replicas, so producers using acks=all fail with NOT_ENOUGH_REPLICAS.
error: the broker returned an error for the partition.
healthy is true only when there are no problems. Use search to limit the check to topics whose name contains a substring (case-insensitive); internal topics are skipped unless include_internal is true. min_isr_unknown lists topics whose min.insync.replicas the broker did not report, so under_min_isr was not judged for them.
| Name | Required | Description | Default |
|---|---|---|---|
| search | No | Optional case-insensitive substring that a topic name must contain to be checked. Omit to check every topic. Brokers and the controller are always reported. | |
| include_internal | No | Optional. When true, also check Kafka's internal topics such as __consumer_offsets. Defaults to false. An unhealthy internal topic breaks every consumer group, so include them when groups fail across the board. |
Output Schema
| Name | Required | Description |
|---|---|---|
| brokers | Yes | |
| healthy | Yes | |
| summary | Yes | |
| problems | Yes | |
| warnings | Yes | |
| cluster_id | Yes | |
| controller | Yes | |
| min_isr_unknown | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and largely succeeds: it discloses the return shape, the meaning of each problem type, the rule that internal topics are skipped unless include_internal is true, and the caveat that min_isr_unknown topics were not judged for under_min_isr. It omits operational details like required permissions or cost/latency, which is the only notable gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Content is front-loaded (purpose first, then the issue taxonomy, then parameter behavior) and the bulleted enumeration of offline/under_replicated/under_min_isr/error is efficient and skimmable. It is somewhat long, but nearly every sentence carries distinct diagnostic meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value structure is covered, yet the description still supplies the interpretive layer an agent needs: the healthy flag's condition and the semantics of each reported problem. Combined with the two fully documented parameters, nothing required to invoke or interpret the tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters. The description nevertheless adds semantic value: it clarifies that search matches a case-insensitive substring and that brokers and the controller are always reported regardless of the filter, and ties include_internal to the consequence of an unhealthy internal topic.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Check the health of the cluster') and enumerates exactly what one call returns: cluster id, controller, per-broker details, a summary count, and unhealthy partitions. It is clearly distinguishable from siblings like compare_clusters, list_clusters, or describe_topic, which serve different scopes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states when to use the tool ('check the health ... in one call') and gives concrete conditions for the optional filters, e.g. use search to limit to topics by substring and include_internal 'when groups fail across the board'. It does not explicitly route the agent away from compare_clusters or describe_topic, but the usage context is clear enough to act on.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
commit_offsetA
Move 1 to 100 consumer group positions in one call through items. Each item names a group and topic and exactly one target: offset (one partition, the next offset the group reads), timestamp (RFC3339; each partition moves to its first message at or after that time, which is how messages are replayed since a moment), or position (earliest or latest). With timestamp or position, omit partition to move every partition of the topic. Moving one offset is an items array of length one.
The response previews, per partition and in total, how many messages each move would skip or replay. No change is made unless confirm is true; one confirm covers the whole batch, and applying is not atomic because Kafka cannot roll back commits that succeeded before a later item failed. Results follow items order, each carrying index with result or error.
Active groups are refused unless allow_active_members is set on the item, because running consumers may overwrite the commit. Requires Kafka offset-commit permission.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | The moves to make, 1 to 100 of them. Moving one offset is an array of length one. Two items that would move the same group, topic and partition, including a whole-topic item and a partition item, are refused before anything changes. | |
| confirm | No | Optional. When false or omitted, nothing is changed and the response describes what would happen for every item. Must be true to actually move the offsets. One confirm covers the whole batch. |
Output Schema
| Name | Required | Description |
|---|---|---|
| atomic | Yes | |
| failed | Yes | |
| applied | Yes | |
| results | Yes | |
| succeeded | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: dry-run by default with preview counts of skipped/replayed messages, a single whole-batch confirm, explicit non-atomicity because Kafka cannot roll back already-successful commits, refusal of active groups, and the required offset-commit permission. These are exactly the behavioral risks an agent must know before calling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action in the first sentence, then layered behavior, then constraints. It runs four dense paragraphs and restates some schema content (length-one array, exactly-one-of), but nearly every sentence contributes operational meaning rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the description need not catalog return fields; it still usefully characterizes the response as a per-partition and total preview and as index-ordered result-or-error entries. Combined with the permission, confirmation, and atomicity disclosures, nothing needed to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so the per-field docs already define offset, timestamp, position, partition, and allow_active_members. The description adds value beyond that by tying the fields into item-level semantics: exactly one target per item, whole-topic vs partition behavior, batch ordering of results by index, and one confirm governing the entire array.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (move consumer group positions) and immediately scopes it to batch semantics via items. It is trivially distinguishable from siblings like describe_consumer_group, consumer_lag, and delete_consumer_group, which never mutate committed offsets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear guidance on which of the three mutually exclusive targets to choose and why (timestamp for replaying since a moment, position for earliest/latest, offset for a single partition), plus the omit-partition rule for whole-topic moves. It stops short of explicitly naming sibling tools or stating when not to use it, so it is strong context rather than full when/when-not routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compare_clustersA
Compare the topics of 1 to 100 other clusters against the one this endpoint serves, through items. Reports which topics only this cluster has, which only the other has, which exist on both but disagree, and how many brokers and topics each side has. Comparing one cluster is an items array of length one.
Use it to find what preproduction has that production does not, or to check whether two environments still match. Topics reported as only on the other cluster carry their partition count, replication factor and explicitly-set configs, so the report can be handed straight to create_topic.
This tool creates and changes nothing. To create the missing topics, pass them to create_topic, which previews them against the broker first.
Only configs a topic sets for itself are compared. Two clusters may carry different broker defaults, and comparing inherited values would report every topic as different. Internal topics are excluded unless include_internal is set.
Topic listings come from the client's metadata cache, so a topic created within the last few seconds may still be reported as missing. Repeat the comparison after a moment rather than creating it twice.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | The clusters to compare against, 1 to 100 of them. Comparing one cluster is an array of length one. Results follow this order and an unreachable cluster is reported against its own item. |
Output Schema
| Name | Required | Description |
|---|---|---|
| atomic | Yes | |
| failed | Yes | |
| applied | Yes | |
| results | Yes | |
| succeeded | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: it declares the tool is read-only, explains that only explicitly-set topic configs are compared (and why inherited defaults would falsely flag everything), notes internal topics are excluded by default, discloses metadata-cache staleness that can report a just-created topic as missing, and states how unreachable clusters are reported.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and output shape are front-loaded, and each paragraph carries distinct value (semantics, workflow, exclusions, caveat). It runs long and repeats the 'array of length one' framing that the schema already states, which keeps it short of a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, yet the description still characterizes the report well enough to know it is directly consumable and even speculates on the actionable payload (partition count, replication factor, explicit configs). Combined with the exclusions and staleness caveats, nothing needed to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the nested item fields (search, cluster, include_internal) are already documented in the schema. The description reinforces scope ('1 to 100', array of length one) but adds little syntax or format detail beyond what the schema already states, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — comparing topics across 1-100 clusters against the endpoint's own cluster — and enumerates exactly what the comparison produces (only-here, only-there, disagreeing, broker/topic counts). An agent can distinguish this from list_topics or describe_topic without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete scenarios ('find what preproduction has that production does not', 'check whether two environments still match') and explicitly routes the agent onward: missing topics should be handed to create_topic rather than re-created by hand. It also states this tool changes nothing, which tells the agent when this tool is not the right one.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
consumer_lagA
Measure consumer lag for 1 to 100 topics in one call through items: production and consumption rates, and whether each topic's backlog is caught up, draining, growing, stalled or has no active consumers. Returns an ETA only when lag is shrinking.
Results follow items order, each carrying index with result or error. Measuring one topic is an items array of length one.
Consumption rate is sampled for sample_seconds, so the call waits that long. Windows run concurrently, so several measurements do not add their waits together. Set skip_consume_rate for an immediate result without an ETA.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | The measurements to take, 1 to 100 of them. Measuring one topic is an array of length one. Sampling windows run concurrently, so several measurements do not add their wait times together. |
Output Schema
| Name | Required | Description |
|---|---|---|
| atomic | Yes | |
| failed | Yes | |
| applied | Yes | |
| results | Yes | |
| succeeded | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does so well: it discloses that the call blocks for sample_seconds, that sampling windows run concurrently so waits don't accumulate, why sampling is required (Kafka stores no commit history), that ETA appears only when lag is shrinking, and that per-item results/errors are indexed. It omits auth/permission requirements, which is the remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and output shape are front-loaded, followed by behavioral caveats in short paragraphs; nothing is filler. Minor redundancy appears where concurrency and the blocking sample wait are re-explained in both description and schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a rich input schema and an existing output schema, the description need not explain return values, and it covers the essential behavior: latency cost, concurrency, ETA conditionality, and per-item error handling. Only permissions and any rate limits are unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents topic, group, sample_seconds and skip_consume_rate; the description largely restates these. It adds the outer 1–100 sizing constraint and the single-topic-array idiom, but that is marginal beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Measure consumer lag'), names exactly what is returned (production and consumption rates, backlog state, ETA), and scopes it to 1–100 topics. An agent can distinguish this from describe_consumer_group or list_consumer_groups without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives operational conditions ('Set skip_consume_rate for an immediate result without an ETA') but never states when to prefer this over siblings such as describe_consumer_group or sample_messages. Usage is implied rather than compared against alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
copy_messageA
Copy 1 to 20 existing messages in one call through items, each identified by topic, partition and offset, to an existing topic on this or another configured cluster. Preserves key, value and headers and adds traceable provenance headers. Copying one message is an items array of length one.
No message is written unless confirm is true; one confirm covers the whole batch, and writing is not atomic because Kafka cannot retract a record produced before a later item failed. Results follow items order, each carrying index with result or error.
The destination must be writable and requires Kafka write permission; a read-only cluster may still be the source.
A key or value carrying a Schema Registry id is copied byte for byte. When the destination cluster uses a different registry the id may mean nothing there, so the response warns; set translate_schema to register the schema in the destination registry and rewrite the id.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | The copies to make, 1 to 20 of them. Copying one message is an array of length one. Two items naming the same source and destination are refused before anything is written. | |
| confirm | No | Optional. When false or omitted, nothing is written and the response shows the messages that would be copied. Must be true to actually write them. One confirm covers the whole batch. |
Output Schema
| Name | Required | Description |
|---|---|---|
| atomic | Yes | |
| failed | Yes | |
| applied | Yes | |
| results | Yes | |
| succeeded | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: it discloses the confirm gate ('No message is written unless confirm is true'), non-atomicity ('Kafka cannot retract a record produced before a later item failed'), permission requirements (destination must be writable), the read-only source allowance, and Schema Registry id translation warnings with the registry side effect of translate_schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and scope, then organized into confirm/atomicity, permissions, and schema-translation paragraphs. Dense with earning content, though slightly long for a two-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite an output schema making return values optional to explain, the description adds result ordering ('Results follow items order, each carrying index with result or error') and the duplicate-item refusal rule. An agent has everything needed to invoke this correctly, including safety and side-effect context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both top-level parameters and the nested item fields are already fully documented. The description reinforces confirm and items semantics ('one confirm covers the whole batch') and adds provenance-header behavior, but contributes no syntax or format detail the schema lacks, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Copy 1 to 20 existing messages'), names the addressing scheme (topic/partition/offset) and the destination constraint. It is clearly distinguishable from siblings like get_message, produce_message, or search_messages.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the operating context well: batches are capped at 20, a single message is an array of length one, confirm gates any write, and list_clusters is named for resolving destination_cluster names. It does not explicitly contrast with produce_message for creating new messages, so it stops short of full alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_topicA
Create 1 to 100 Kafka topics in one call through items, each with optional partition count, replication factor and topic-level configs. Omitted values use broker defaults. Existing topics are refused rather than modified. Creating one topic is an items array of length one.
No topic is created unless confirm is true; otherwise the broker only validates the requests. One confirm covers the whole batch, and creation is not atomic: topics created before a later item failed stay, because Kafka cannot roll them back. Results follow items order, each carrying index with result or error.
Partition counts cannot be reduced later. Requires Kafka CREATE permission.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | The topics to create, 1 to 100 of them. Creating one topic is an array of length one. Duplicate topic names are refused before anything is created. | |
| confirm | No | Optional. When false or omitted, nothing is created: every item is validated by the broker and the response describes what would happen. Must be true to actually create. One confirm covers the whole batch. |
Output Schema
| Name | Required | Description |
|---|---|---|
| atomic | Yes | |
| failed | Yes | |
| applied | Yes | |
| results | Yes | |
| succeeded | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so richly: confirm=false means validation-only, one confirm covers the whole batch, creation is non-atomic with partial results persisted, results follow items order with per-item index, and the call requires Kafka CREATE permission. These are exactly the traits an agent cannot infer from the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and count range, then organized into short paragraphs on defaults, the confirm gate, non-atomicity, and permissions. Slightly redundant in restating the existing-topic refusal and the confirm rule that the schema also states, but no sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema is present and the description still usefully notes that results mirror items order with index plus result-or-error. Combined with the permission requirement and partial-failure caveat, an agent has everything needed to invoke and interpret this batch mutation correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema by explaining the batch-level semantics of confirm (one flag governs all items) and that omitted values fall back to broker defaults. It does not add new per-item syntax detail beyond what the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and resource scoped to a batch ('Create 1 to 100 Kafka topics in one call through items'), and the scope immediately separates it from siblings like delete_topic, add_partitions and alter_topic_config. Adds the key constraint that existing topics are refused rather than modified.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the confirm-gate workflow and clarifies that a single topic is just an items array of length one, which tells the agent exactly how to invoke it. It notes partition counts can only grow, hinting at add_partitions as the later alternative, but never names a sibling tool or gives an explicit when-to-use-this-vs-that comparison.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_consumer_groupA
Delete 1 to 100 consumer groups in one call through items, removing each group's committed offsets. Use it to clean up groups whose consumers were decommissioned: their commits keep reporting lag that nobody will ever drain. Deleting one group is an items array of length one.
No group is deleted unless confirm is true. The preview lists each group's state, the topics it has committed on, every committed offset with its lag, and the total lag that would disappear with it. A group with active members is refused: it is not abandoned, and deleting it would reset where its consumers resume.
If a consumer later starts with the same group id, it begins wherever its auto.offset.reset points, not where the group left off. One confirm covers the whole batch, and deletion is not atomic: groups deleted before a later item failed stay deleted. Results follow items order, each carrying index with result or error. Requires Kafka DELETE permission on the group.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | The groups to delete, 1 to 100 of them. Deleting one group is an array of length one. Naming the same group twice is refused before anything is deleted. | |
| confirm | No | Optional. When false or omitted, nothing is deleted and the response shows each group's committed offsets and lag. Must be true to delete. One confirm covers the whole batch. |
Output Schema
| Name | Required | Description |
|---|---|---|
| atomic | Yes | |
| failed | Yes | |
| applied | Yes | |
| results | Yes | |
| succeeded | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: confirm gating, refusal on active members, non-atomic batch semantics ('groups deleted before a later item failed stay deleted'), per-item result/error ordering, required Kafka DELETE permission, and the offset-reset consequence for a re-created group id.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and scope, then proceeds through preview, refusal, and failure semantics. Length is justified by the destructive, unannotated nature of the tool, and no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists so return values needn't be fully described, yet the description still covers result ordering and error shape. Safety, permissions, batch limits, and idempotency-adjacent caveats are all present for a high-risk mutation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema: 'Deleting one group is an items array of length one,' 'One confirm covers the whole batch,' and the ordered per-item index/result/error shape.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Delete 1 to 100 consumer groups'), names the batch mechanism (items), and specifies the side effect ('removing each group's committed offsets'). This clearly separates it from siblings like delete_topic or describe_consumer_group.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives the explicit use case (groups whose consumers were decommissioned, still reporting undrainable lag), the refusal condition (active members), and the safe-preview alternative via confirm=false. An agent knows both when to call it and when not to.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_recordsA
Delete the oldest messages of 1 to 100 partitions in one call through items, without deleting the topic. Each item removes every message of one partition with an offset lower than before_offset; the message at before_offset becomes the first readable one. Use it to purge test data, or a run of bad messages at the head of a partition, while keeping the topic, its configuration and its consumer groups. Truncating one partition is an items array of length one.
No message is deleted unless confirm is true and the item sets acknowledge_data_loss. The preview reports the partition's start and end offsets, how many messages would be removed, and every consumer group whose committed offset is below the cut, with how many messages it would lose without ever processing them. Such a group resumes from the new start offset.
Kafka cannot restore deleted records. One confirm covers the whole batch, and deletion is not atomic: partitions truncated before a later item failed stay truncated. Results follow items order, each carrying index with result or error. Compacted topics are not supported by every broker. Requires Kafka DELETE permission on the topic.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | The partitions to truncate, 1 to 100 of them. Truncating one partition is an array of length one. Two items naming the same topic and partition are refused before anything is deleted. | |
| confirm | No | Optional. When false or omitted, nothing is deleted and the response shows how many messages each item would remove and which consumer groups have not read them yet. Must be true to delete. One confirm covers the whole batch. |
Output Schema
| Name | Required | Description |
|---|---|---|
| atomic | Yes | |
| failed | Yes | |
| applied | Yes | |
| results | Yes | |
| succeeded | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so thoroughly: the two-key safety gate (confirm plus per-item acknowledge_data_loss), the dry-run preview contents, irreversibility, the fact that one confirm covers the whole batch, non-atomic partial failure leaving earlier partitions truncated, and the Kafka DELETE permission requirement. These are exactly the traits an agent cannot infer from the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and semantics, and every paragraph carries real information. Minor redundancy remains, such as the length-one array note appearing in both description and schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive batch operation with an output schema present, the description covers the safety model, failure mode, permissions, and per-item result shape without needing to explain return values. An agent has everything required to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so items, topic, partition, before_offset, acknowledge_data_loss and confirm are already fully documented in the schema. The description restates much of this (batch-wide confirm, length-one array) rather than adding new syntax or constraints, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with precise scope: 'Delete the oldest messages of 1 to 100 partitions in one call ... without deleting the topic.' It immediately distinguishes itself from the sibling delete_topic by clarifying the topic itself survives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete use cases ('purge test data, or a run of bad messages at the head of a partition') and an implicit when-not ('while keeping the topic, its configuration and its consumer groups'). It never names delete_topic directly as the alternative, so routing relies on inference rather than an explicit pointer.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_topicA
Delete 1 to 100 Kafka topics in one call through items. Deleting one topic is an items array of length one.
This is the most destructive tool here. Deleting a topic destroys every message in it, and Kafka has no undo: recreating the topic does not bring the data back. Any consumer group reading it breaks.
Nothing is deleted unless confirm is true; otherwise the response reports, per topic, how many messages would be destroyed, how many partitions it had, and which consumer groups had committed offsets for it. A topic that still holds messages also requires acknowledge_data_loss on its item, so destroying data is never a single unconsidered call.
One confirm covers the whole batch, and deletion is not atomic: topics deleted before a later item failed stay deleted. Results follow items order, each carrying index with result or error. Internal topics are refused. Requires Kafka DELETE permission.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | The topics to delete, 1 to 100 of them. Deleting one topic is an array of length one. Naming the same topic twice is refused before anything is deleted. | |
| confirm | No | Optional. When false or omitted, nothing is deleted and the response describes what each deletion would destroy. Must be true to actually delete. One confirm covers the whole batch. |
Output Schema
| Name | Required | Description |
|---|---|---|
| atomic | Yes | |
| failed | Yes | |
| applied | Yes | |
| results | Yes | |
| succeeded | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so thoroughly: irreversibility ('no undo', recreating does not restore data), consumer-group breakage, confirm gating, per-topic preview output, non-atomic batch behavior, and DELETE permission requirement. This is exactly the behavioral disclosure a destructive tool needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and scope, then layers the destructive warning and the safety gates. Slightly verbose and some content duplicates the schema, but every sentence carries substantive information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive batch tool with no annotations and an output schema present, the description covers safety, permissions, failure modes, and batch semantics completely. Nothing an agent needs to call it safely is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already documents items, topic, acknowledge_data_loss, and confirm in detail, so the baseline is 3. The description largely restates schema semantics ('one confirm covers the whole batch', items ordering) rather than adding new meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Delete 1 to 100 Kafka topics in one call through items') and immediately scopes the batch semantics. An agent can distinguish it from create_topic, delete_consumer_group, and delete_records without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives strong usage context: preview-vs-delete via confirm, the acknowledge_data_loss gate, and refusal of internal topics. It does not name an alternative sibling tool for inspection (e.g. describe_topic), but the workflow guidance is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
describe_consumer_groupA
Describe 1 to 100 consumer groups in one call through items: state, assignor, coordinator broker, every running member (member_id, client_id, host and the partitions assigned to it), and for every partition the group owns or has committed on, the committed offset, end offset, lag, and the member, client id and host consuming it. Partitions and members are sorted.
Use it to find which consumer instance owns a stuck or lagging partition, so the operator knows which pod or host to inspect. has_commit false means the group owns the partition but never committed there, so its starting point is decided by the consumer's auto.offset.reset rather than by an offset. An Empty group has no members but keeps its commits.
Results follow items order, each carrying index with result or error.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | The groups to describe, 1 to 100 of them. Describing one group is an array of length one. Results follow this order and a group that does not exist is reported against its own item. |
Output Schema
| Name | Required | Description |
|---|---|---|
| atomic | Yes | |
| failed | Yes | |
| applied | Yes | |
| results | Yes | |
| succeeded | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does substantial work: it explains has_commit semantics, that an Empty group keeps its commits, that partitions/members are sorted, and that non-existent groups are reported per item. It omits permission/auth requirements and any cost concerns, keeping it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core capability before usage and edge-case detail, with no filler sentences. It is dense but each sentence carries information; slight length keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, yet the description still covers the parameter contract, per-item error reporting, and edge-case semantics. An agent has everything needed to call it correctly and interpret results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single parameter is fully documented in the schema, so the baseline is 3. The description restates the items-ordering behavior that the schema already specifies, adding no new syntactic or format guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Describe) and resource (consumer groups) and enumerates the exact fields returned (state, assignor, coordinator, members, partitions, offsets, lag), which goes well beyond the name. It does not explicitly distinguish itself from close siblings such as consumer_lag or list_consumer_groups, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a concrete use case: 'find which consumer instance owns a stuck or lagging partition, so the operator knows which pod or host to inspect.' This gives clear situational context, but it names no alternative tool or exclusion condition, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
describe_topicA
Describe 1 to 20 topics in one call through items: partitions, offset ranges, approximate message count, oldest/newest timestamps and complete effective configuration. Config entries identify whether values are inherited or topic-specific.
Results follow items order, each carrying index with result or error, so a topic that does not exist is reported against its own item rather than failing the call. Describing one topic is an items array of length one.
Message counts are offset spans and may overcount after retention or compaction.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | The topics to describe, 1 to 20 of them. Describing one topic is an array of length one. Results follow this order and a missing topic is reported against its own item. |
Output Schema
| Name | Required | Description |
|---|---|---|
| atomic | Yes | |
| failed | Yes | |
| applied | Yes | |
| results | Yes | |
| succeeded | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does so well: it discloses result ordering, per-item error isolation instead of call-level failure, config inheritance vs topic-specific values, and the important caveat that message counts can overcount after retention or compaction. It stops short of stating read-only semantics or performance/authorization behavior, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and its scope, then layered with result shape, error behavior, and the counting caveat. Every sentence carries distinct information and none is redundant filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values are already structured; the description usefully supplements it with ordering, per-item error semantics, and the overcount caveat. For a read-only batch-describe tool with one well-documented parameter, nothing needed for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single parameter is thoroughly documented there, so the baseline is 3. The description adds the 1-to-20 bound and ordering semantics, but those are already present in the schema's items description, so no meaningful value is added beyond the structured field.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Describe 1 to 20 topics') and then enumerates exactly what the description returns: partitions, offset ranges, message counts, timestamps, and effective configuration. This is clearly distinguishable from list_topics (enumeration) and describe_consumer_group (different resource), even without an explicit sibling callout.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains how to describe a single topic (items array of length one) and how per-item errors are reported, which is genuine usage context. However, it never states when to prefer this over list_topics or when describing is unnecessary, and gives no prerequisites. Usage is implied rather than directed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_messageA
Read 1 to 20 Kafka messages at exact addresses in one call through items, and optionally nearby messages for context. Returns key, value, headers, timestamp and original value size.
Values carrying a Schema Registry id (Avro, Protobuf, JSON Schema) and topics with a configured format are decoded to JSON; format, schema_id and message_type say what the bytes were. Anything else is returned as text, or base64 when it is binary, and decode_error explains a value that named a schema but could not be decoded.
Results follow items order, each carrying index with result or error, so an invalid partition or an offset beyond the partition end is reported against its own item rather than failing the call. Reading one message is an items array of length one.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | The message addresses to read, 1 to 20 of them. Reading one message is an array of length one. Results follow this order and an address that cannot be read is reported against its own item. |
Output Schema
| Name | Required | Description |
|---|---|---|
| atomic | Yes | |
| failed | Yes | |
| applied | Yes | |
| results | Yes | |
| succeeded | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does substantial work: it discloses result ordering, per-item error isolation ('an invalid partition or an offset beyond the partition end is reported against its own item rather than failing the call'), decoding rules, base64 fallback for binary, and decode_error semantics. It omits two relevant traits — that reading does not commit offsets (important given the commit_offset and consumer_lag siblings) and any auth/permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded and tightly written: the purpose sentence comes first, then behavior, then error semantics. The decoding paragraph is dense but the sentences each carry a distinct fact; the only mild waste is restating return fields (key, value, headers, timestamp) when an output schema already exists.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, nested-array read tool with an output schema, the description covers what to build, how results map back to requests, and how failures surface — enough to call it correctly. Its remaining gap is the absence of any statement about read-only/non-mutating behavior and permissions, which matters because siblings like commit_offset and delete_records mutate the same cluster.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3 and the schema already documents topic, partition, offset, context and max_value_bytes. The description adds real meaning beyond that: the items-level semantic that all addresses share one call, the 1-20 cardinality, and that 'reading one message is an items array of length one', which is exactly the construct an agent must build.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a precise verb and resource with scope: 'Read 1 to 20 Kafka messages at exact addresses in one call through items'. The 'exact addresses' framing plus the 1-20 bound functionally separates it from the sampling and searching siblings (sample_messages, search_messages) without either schema being opened.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly conveys the situation this tool serves — reading known offsets, optionally with nearby context via 'context' — and explains the single-message case explicitly. It never names an alternative (search_messages, sample_messages) or states when not to use it, so routing between siblings is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_schemaA
Read 1 to 100 schemas from this cluster's Schema Registry in one call through items, each by subject (latest version unless version is given) or by the schema id a message carries. Returns the schema text, its type (avro, protobuf or json_schema), id, version, every version of the subject, the schemas it references, and for an id the subject versions that use it.
Read the schema before producing to a schema-encoded topic: it names every field and enum a value needs, which a sampled message may not show. For Protobuf, message_types lists the names produce_message accepts as message_type. The subject for a topic's values is usually -value.
Results follow items order, each carrying index with result or error. Fails as a whole when the cluster has no schema_registry configured.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | The schemas to look up, 1 to 100 of them. Looking up one schema is an array of length one. Results follow this order and a subject or id that does not exist is reported against its own item. |
Output Schema
| Name | Required | Description |
|---|---|---|
| atomic | Yes | |
| failed | Yes | |
| applied | Yes | |
| results | Yes | |
| succeeded | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and largely does: it discloses the whole-call failure mode when no schema_registry is configured, per-item error reporting against each item, result ordering, and that every version of the subject plus referenced schemas are returned. Read-only behavior is only implied by 'Read,' not stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action in the first sentence, then layers usage and failure context. Dense but every sentence carries operational information; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, yet the description goes further by covering failure mode, per-item error semantics, and the pre-produce workflow step. Nothing needed to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds meaning beyond it: it explains the subject-vs-id choice, the default-to-latest behavior unless version is given, the exact-match/case-sensitive convention, and the protobuf message_type linkage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with explicit scope: 'Read 1 to 100 schemas from this cluster's Schema Registry in one call.' It names the two lookup modes (by subject/latest-version or by schema id), which no sibling tool offers, so an agent can distinguish it immediately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a real when-to-use rule: 'Read the schema before producing to a schema-encoded topic,' plus practical routing hints (the <topic>-value convention, message_types for Protobuf). It doesn't name an alternative or an explicit when-not-to-use case, but no sibling competes for this job.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_aclsA
List the access control entries (ACLs) on the cluster this endpoint serves. Each entry has principal, host, resource_type, resource_name, pattern_type (literal, prefixed), operation (read, write, describe, ...) and permission (allow or deny). Results are sorted.
Use it when a client fails with TOPIC_AUTHORIZATION_FAILED, GROUP_AUTHORIZATION_FAILED or similar: filter by the client's principal, or by the resource it was refused on. Filtering by resource_name returns every ACL the broker applies to that name, including prefixed and wildcard entries, so the answer covers what actually decides access. A deny overrides any allow.
All filters are optional and combine. Fails with SECURITY_DISABLED when the broker has no authorizer, which means ACLs are not enforced at all. Needs DESCRIBE permission on the cluster.
| Name | Required | Description | Default |
|---|---|---|---|
| principal | No | Optional principal to list ACLs for, including its type, such as User:payments. Matched exactly and case-sensitively. Omit for every principal. | |
| resource_name | No | Optional resource name, such as a topic or group. Returns every ACL the broker applies to that name: the exact name, prefixed ACLs whose prefix it starts with, and the * wildcard. Case-sensitive. Needs resource_type. | |
| resource_type | No | Optional resource type: topic, group, cluster, transactional_id or delegation_token. Case-insensitive. Omit for every type. Required when resource_name is given. |
Output Schema
| Name | Required | Description |
|---|---|---|
| acls | Yes | |
| count | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden and does so well: it discloses the DESCRIBE permission requirement on the cluster, the SECURITY_DISABLED failure mode meaning no authorizer is present, that results are sorted, and the semantic rule that a deny overrides any allow.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with what the tool does, then usage, then caveats. Mostly every sentence earns its place, though phrasing like 'so the answer covers what actually decides access' is slightly wordy for the value it adds.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need no explanation, and the description covers purpose, triggers, permission requirements, failure modes, and filter combination semantics. Nothing needed to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: it explains that resource_name matches exact, prefixed, and wildcard entries and requires resource_type, and that a deny overrides an allow. That goes beyond the field-level schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (access control entries/ACLs) plus the exact fields each entry contains (principal, host, resource_type, resource_name, pattern_type, operation, permission). No sibling tool covers ACLs, so the agent can route to it unambiguously.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger scenario (client fails with TOPIC_AUTHORIZATION_FAILED, GROUP_AUTHORIZATION_FAILED) and prescribes how to narrow: filter by the client's principal or by the refused resource. This is a concrete when-to-use with actionable filtering guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_clustersA
List the Kafka clusters this server serves, with whether each is reachable and whether it accepts writes. Connectivity is checked at call time. Use returned names as copy_message destination_cluster values. Broker and credential details are not exposed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| count | Yes | |
| clusters | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full burden and does well: it discloses that connectivity is evaluated live at call time (not cached), that write-acceptance is surfaced, and explicitly that broker and credential details are withheld. That is meaningful behavioral context beyond a bare listing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, each earning its place: purpose, live-check caveat, cross-tool usage, and an explicit non-disclosure boundary. The core purpose is front-loaded in the first clause.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, yet the description still names the key fields (reachable, accepts writes). For a no-arg read-only listing tool this is essentially complete; only the missing sibling routing keeps it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There are no argument semantics to add, and the description correctly avoids inventing any.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (list Kafka clusters the server serves) and enriches it with the exact per-cluster attributes returned: reachability and write acceptance. Clear enough to distinguish from active-inspection siblings like cluster_health or compare_clusters, though it never names them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a concrete downstream usage rule ('use returned names as copy_message destination_cluster values') and a timing note ('connectivity is checked at call time'). However, it never states when to call this vs. cluster_health or compare_clusters, so the when-to-use guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_consumer_groupsA
List the consumer groups on the cluster, with their state, member count and the topics they consume. Filter by exact topic or group state. Empty groups may still hold committed offsets and lag; they have no active consumers to drain it. Topic filtering may inspect every group because Kafka has no topic-to-group index.
| Name | Required | Description | Default |
|---|---|---|---|
| topic | No | Optional topic name. When given, only groups that consume or have committed offsets for this topic are returned. Matched exactly and case-sensitively. | |
| states | No | Optional group states to return, such as Stable, Empty, PreparingRebalance or Dead. Defaults to every state. |
Output Schema
| Name | Required | Description |
|---|---|---|
| count | Yes | |
| groups | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it delivers two non-obvious behaviors: empty groups can still retain committed offsets and lag with no active consumers to drain them, and topic filtering may scan every group because Kafka lacks a topic-to-group index. That performance caveat is genuinely useful. It stops short of 5 by omitting any auth, pagination, or cost/limit disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and returned fields, then the filtering and caveats. Four sentences with no filler; the final clause is slightly dense but each sentence carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be re-explained, and the description supplies the scope, filter semantics, and the empty-group/scan caveats. What is missing is routing guidance relative to the several sibling consumer-group tools, which an agent would still have to infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters (topic, states) are already fully documented, including exact and case-sensitive matching. The description's 'filter by exact topic or group state' restates the schema, and the only added semantics concern the cost of topic filtering rather than the parameters themselves; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List the consumer groups on the cluster') and enumerates the returned fields (state, member count, topics consumed). It distinguishes itself from describe_consumer_group and consumer_lag only implicitly, via the word 'list' and cluster scope, so it falls short of an explicit sibling contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description says which filters exist ('filter by exact topic or group state') but never says when to choose this tool over describe_consumer_group, consumer_lag, or delete_consumer_group, and gives no when-not guidance. Usage must be inferred from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_topicsA
List the topics on the endpoint's cluster, sorted by name, each with its partition count, replication factor and the configs it sets for itself.
Filter with an optional JavaScript predicate. A name match is return topic.indexOf('orders') >= 0, and the predicate can also read partitions, replication_factor, internal and configs — questions a substring filter cannot express, such as which topics have more than six partitions, only one replica, or a compacted cleanup policy.
Only configs a topic sets for itself are reported, because inherited cluster defaults would make every topic look configured. Topics whose predicate throws are counted in script_errors rather than listed, so a broken filter is never mistaken for an empty cluster.
| Name | Required | Description | Default |
|---|---|---|---|
| script | No | Optional JavaScript that decides whether a topic is listed. Return true to keep it. In scope: topic (the name), partitions (number), replication_factor (number), internal (true for Kafka's own topics such as __consumer_offsets) and configs (an object of the values this topic sets for itself, such as configs['retention.ms']; inherited cluster defaults are not included). Examples: return topic.indexOf('orders') >= 0; return partitions > 6; return configs['cleanup.policy'] === 'compact'; return replication_factor === 1 && !internal. Omit to list every topic. | |
| timeout_seconds | No | Optional wall-clock limit in seconds for evaluating the script. Defaults to 30. A topic whose evaluation is cut short is counted in script_errors rather than listed. |
Output Schema
| Name | Required | Description |
|---|---|---|
| count | Yes | |
| topics | Yes | |
| script_errors | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and mostly succeeds: it discloses sort order, that only self-set configs are reported (inherited defaults excluded), and that throwing or timed-out predicates land in script_errors rather than silently shrinking results. It stops short of stating key behavioral facts such as read-only safety, pagination, or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and return shape, then filter mechanics, then edge-case behavior. Mostly every sentence earns its place, though the 'only configs a topic sets for itself' point and the predicate examples echo what the input schema already states.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be re-explained, and the description covers everything else an agent needs: scope, ordering, filter semantics, error accounting, and the config-inheritance caveat. Nothing material is missing for a read-only list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3; the schema already documents in-scope variables, examples, and the timeout default. The description reinforces the semantics (self-set configs only, predicate failures counted as errors) but adds little the schema does not already say.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with scope ('List the topics on the endpoint's cluster') and enumerates the returned fields (partition count, replication factor, self-set configs), so an agent knows exactly what it gets. It does not, however, name or contrast itself against the closest sibling, describe_topic, leaving that distinction to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description is explicit about when the script predicate is warranted versus a simple substring match ('questions a substring filter cannot express'), which is useful. But it never says when to call list_topics rather than describe_topic, cluster_health, or list_clusters, and offers no when-not guidance at the tool level.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
open_transactionsA
Find open transactions holding back read_committed consumers on 1 to 100 topics in one call through items. A consumer with isolation.level=read_committed cannot read past the first message of a transaction that has not been committed or aborted, so a transactional producer that hangs or dies mid-transaction stalls every such consumer on that partition. In consumer_lag this looks exactly like a poison message: members present, nothing consumed.
For each topic, blocked says whether any partition's last stable offset trails its high watermark. Each such partition reports both offsets, how many messages read_committed consumers cannot see, and every producer with an open transaction there: producer id and epoch, the offset the transaction started at, and where the coordinator knows it, the transactional id (which names the application), state, when it started, how long it has been open and its timeout, after which the broker aborts it.
The fix is in the producer, not the consumer: restart or fence the producing application, or wait for the timeout. Skipping offsets does not help. Results follow items order, each carrying index with result or error. Needs DESCRIBE on the topic and on transactional ids.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | The topics to check, 1 to 100 of them. Checking one topic is an array of length one. Results follow this order and a missing topic is reported against its own item. |
Output Schema
| Name | Required | Description |
|---|---|---|
| atomic | Yes | |
| failed | Yes | |
| applied | Yes | |
| results | Yes | |
| succeeded | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does: it explains why the condition matters (consumers cannot read past the first message of an uncommitted transaction), details what each result field conveys (blocked flag, both offsets, hidden message count, producer id/epoch, transactional id, state, age, timeout), and states the required DESCRIBE permissions. This is far more than a restatement of the tool name.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and the differential-vs-consumer_lag point are front-loaded, and subsequent sentences each carry operational payload (output fields, remediation, permissions). It is on the longer side and some sentences (e.g. the producer-side fix) are advisory rather than selection-critical, but nothing is truly wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only diagnostic tool, the description covers the trigger scenario, the output interpretation, the required permissions, and result ordering. An output schema exists, yet the description still maps the meaningful fields an agent needs for diagnosis, leaving no material gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single items/topic parameter, and the schema already documents exact case-sensitive matching and the order-preservation behavior. The description reinforces the 1-to-100 scope and per-item error reporting but adds little syntax or semantics beyond what the schema already states, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Find open transactions') plus the exact condition of interest ('holding back read_committed consumers'). It further differentiates itself from the sibling tool consumer_lag by noting how the symptom presents there ('members present, nothing consumed'). An agent can pick this out without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides strong differential-diagnosis context: read_committed consumers stall on uncommitted transactions, and 'In consumer_lag this looks exactly like a poison message,' which routes the agent from a lag symptom to this diagnostic tool. It also advises on the remediation path (producer-side, not consumer). It stops short of an explicit 'use X when, use this when not Y' statement, so not quite a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
produce_messageA
Write 1 to 20 new messages in one call through items, to existing topics on this or another configured cluster. The caller supplies the key, value and headers, so unlike copy_message this can write content no producer ever sent. Writing one message is an items array of length one.
Every message carries provenance headers naming this tool, the time and the principal, so a fabricated message stays distinguishable from a genuine one.
Nothing is written unless confirm is true; one confirm covers the whole batch, and writing is not atomic because Kafka cannot retract a record produced before a later item failed. Results follow items order, each carrying index with result or error.
A produced message cannot be deleted: it stays until retention removes it, and any consumer reading the topic will process it. The destination must be writable and requires Kafka write permission; the endpoint you call may itself be read-only.
Omit partition unless the exact partition matters. The key decides placement, and naming a partition puts a keyed message where its key does not hash to, which breaks ordering for that key.
A topic whose consumers read Avro, Protobuf or JSON Schema needs value_schema: give the value as JSON and it is encoded to the destination registry's schema, refused with the offending field when it does not fit. A topic with a format configured in topic_formats is encoded to it without being asked. Check get_message or sample_messages first: a schema_id on the existing messages means the topic needs value_schema, and writing plain JSON there breaks its consumers.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | The messages to write, 1 to 20 of them. Writing one message is an array of length one. Two identical items mean two messages, which is allowed: appending the same payload twice is a real request. | |
| confirm | No | Optional. When false or omitted, nothing is written and the response shows the messages that would be produced. Must be true to actually write them. One confirm covers the whole batch. |
Output Schema
| Name | Required | Description |
|---|---|---|
| atomic | Yes | |
| failed | Yes | |
| applied | Yes | |
| results | Yes | |
| succeeded | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and discharges it: nothing is written unless confirm=true, one confirm covers the batch, the write is non-atomic (Kafka cannot retract), produced records cannot be deleted and will be consumed, destination must be writable and needs write permission, the endpoint may be read-only, and provenance headers make fabricated messages distinguishable. This is exactly the operational context an agent needs before a destructive-adjacent write.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Six paragraphs is long, but the length is justified by the tool's complexity and each paragraph carries distinct semantics (purpose, provenance, confirmation/atomicity, retention, partitioning, schema encoding). Purpose is front-loaded in the first sentence; a little tightening in the schema paragraph would make it a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, yet the description still tells the agent results follow items order with index and result/error — the one structural fact worth knowing. Combined with permission, schema-encoding, and confirmation guidance, nothing needed to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description goes beyond it by explaining the confirm gating semantics, the partition/key-hashing ordering trap, and the decision rule for when value_schema is required (schema_id present on existing messages). Some overlap with schema text remains, but the cross-parameter interplay and when-to-set reasoning add real value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb and resource ('Write 1 to 20 new messages ... to existing topics') and immediately bounds scope (batch size, this or another configured cluster). It explicitly distinguishes itself from the sibling copy_message, so an agent can route between them without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use and when-not-to guidance: use copy_message for existing content, 'Omit partition unless the exact partition matters', and 'Check get_message or sample_messages first' before writing to a schemaed topic. Alternatives and their selecting conditions are named, not implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sample_messagesA
Sample the newest messages of 1 to 20 topics in one call through items, and summarize value formats, field paths, key usage, the schemas values were written with, and sampled offset ranges. Use this to design a search_messages predicate, and to find the schema to produce against.
Avro, Protobuf and JSON Schema values carrying a Schema Registry id, and topics with a configured format, are decoded, so their field paths are reported like JSON. value_formats.undecodable counts values that named a schema but could not be decoded; decode_error on each message says why.
Results follow items order, each carrying index with result or error. Sampling one topic is an items array of length one. The sample describes recent data only, and keys must not be used to guess partitions.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | The topics to sample, 1 to 20 of them. Sampling one topic is an array of length one. Results follow this order and a topic that cannot be sampled is reported against its own item. |
Output Schema
| Name | Required | Description |
|---|---|---|
| atomic | Yes | |
| failed | Yes | |
| applied | Yes | |
| results | Yes | |
| succeeded | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries full burden and does so well: it discloses decoding of Avro/Protobuf/JSON Schema values with a Schema Registry id, the meaning of value_formats.undecodable and per-message decode_error, ordered results with per-item index/error, and the caveats that the sample covers recent data only and keys must not be used to guess partitions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core verb and scope, then decode behavior and result shape. It is fairly long but nearly every sentence adds operational value (decoding rules, error fields, order semantics); a little compression is possible in the decode paragraph.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value documentation is not required, yet the description still explains result ordering, per-item error reporting and undecodable counts. Combined with the fully covered input schema, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema already documents topic, partitions, sample_size and max_value_bytes with defaults, so the baseline is 3. The description adds the 1-to-20 bound and order-preservation semantics of items, but no extra syntax or default detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (sample) and resource (newest messages of 1-20 topics) plus the summarized dimensions (value formats, field paths, key usage, schemas, offset ranges). It also names the sibling it feeds into (search_messages), so an agent can distinguish it from search_messages, get_message and describe_topic without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent when to reach for it: 'Use this to design a search_messages predicate, and to find the schema to produce against.' Context is clear, but it gives no when-not or explicit exclusion against alternatives like get_message or search_messages for direct reads.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_messagesA
Search message key, value, headers or metadata with a JavaScript predicate. The script returns true for a match and receives value (parsed JSON, a decoded Avro/Protobuf/JSON Schema record, or text), key, headers, partition, offset and timestamp. Omit it to match all messages. Schema-encoded values are searched by field exactly like JSON.
Every partition is read together, one chunk deep at a time, so a limited newest-first search returns the newest matches in the topic rather than the newest in whichever partition was read first. Kafka orders records only within a partition, so matches are merged by timestamp; producers set timestamps unless the topic uses LogAppendTime.
Kafka has no server-side search, so scans are bounded. Check complete, stopped_reason and scanned_ranges before treating no matches as conclusive. Use count_only or output_file for large result sets. A script that runs past timeout_seconds is interrupted and counted in script_errors; scripts are not memory-sandboxed, so keep predicates simple.
| Name | Required | Description | Default |
|---|---|---|---|
| topic | Yes | Topic to search. Matched exactly and case-sensitively. | |
| script | No | Optional JavaScript that decides whether a message matches. Return true to keep it. In scope: value (the parsed document for JSON, and the decoded record for Avro, Protobuf, JSON Schema or a configured format; the raw text otherwise), key (string, decoded document when the key has a schema, or null), headers (object of header name to string), partition, offset and timestamp (a Date). Examples: return key === 'order-123'; return value.eventType === 'NEW' && value.payload.amount >= 500; return value.payload.cancelledAt === null. Omit to match every message. | |
| direction | No | Optional scan direction: newest_first (default) or oldest_first. Decides which matches are kept when max_matches cuts the search short. Every partition is read together, so newest_first means newest in the topic, ordered by timestamp, rather than newest in one partition. | |
| to_offset | No | Optional exclusive offset to stop scanning at, applied to every searched partition. | |
| count_only | No | Optional. When true, scan the whole range and return only how many messages matched, with a per-partition breakdown and no message bodies. Use this first when a query may match a great many messages, then ask the user how they want them before fetching any. | |
| partitions | No | Optional partitions to restrict the search to. Defaults to every partition. Do not guess a partition from a message key: producers may set the partition explicitly, so the key does not determine it. | |
| from_offset | No | Optional inclusive offset to start scanning from, applied to every searched partition. | |
| max_matches | No | Optional maximum number of matches to return. Defaults to 10. | |
| output_file | No | Optional file name to write every match to, as one JSON message per line. Use this instead of returning thousands of messages. A name only, not a path: the server chooses the directory. The response reports the path, the number written and a short preview. | |
| parallelism | No | Optional number of concurrent readers, from 1 to 16. Defaults to 1. It splits a single-partition topic's offsets between readers, which makes a full scan of one large partition faster. A multi-partition topic is already read across its partitions together, so this does not apply there. Worth using for count_only, output_file or a full scan of one partition. | |
| to_timestamp | No | Optional exclusive end time (RFC3339). Resolved to the first offset at or after this time. | |
| from_timestamp | No | Optional inclusive start time (RFC3339). Resolved to the first offset at or after this time. | |
| max_value_bytes | No | Optional maximum value bytes to return per match. Defaults to 512. Longer values are cut and flagged with truncated=true. | |
| timeout_seconds | No | Optional wall-clock limit for the scan in seconds. Defaults to 30. | |
| max_messages_scanned | No | Optional maximum number of messages to read before giving up. Defaults to 10000. |
Output Schema
| Name | Required | Description |
|---|---|---|
| topic | Yes | |
| matches | Yes | |
| complete | Yes | |
| match_count | Yes | |
| output_file | No | |
| script_errors | No | |
| scanned_ranges | Yes | |
| stopped_reason | Yes | |
| scanned_messages | Yes | |
| written_messages | No | |
| matches_by_partition | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses that Kafka has no server-side search so scans are bounded, that every partition is read one chunk deep and merged by timestamp, that scripts are interrupted past timeout_seconds and counted in script_errors, and that scripts are not memory-sandboxed. These are exactly the costly, non-obvious traits an agent needs before calling it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads purpose, then layers scan semantics; nearly every sentence earns its place given 15 parameters. Minor duplication with the schema ("omit it to match all messages" appears in both script and description) and the timestamp-ordering paragraph is dense but justified.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be spelled out, yet the description still flags the completion fields (complete, stopped_reason, scanned_ranges) an agent must inspect. Combined with coverage of truncation, timeouts, parallelism and memory limits, nothing needed to invoke this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds real meaning: it notes schema-encoded values are searched by field exactly like JSON, restates the script's in-scope bindings, and frames direction as deciding which matches survive max_matches. Most of the script detail still duplicates the schema, keeping it just above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific verb (search), the searchable surface (key, value, headers, metadata) and the exact mechanism (JavaScript predicate). This mechanism is distinctive from every sibling that merely reads messages (get_message, sample_messages), so an agent can route on it without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operational direction: omit the script to match all, use count_only or output_file for large result sets, check complete/stopped_reason/scanned_ranges before concluding no matches. It never explicitly contrasts with sibling readers such as sample_messages or get_message, so the when-not and alternatives are left implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
server_configA
Report this endpoint's name, path, purpose, cluster, brokers, authentication, TLS, read-only policy, Schema Registry, configured topic formats and exposed tools. Use it to confirm the target and permissions before acting; several endpoints may target one cluster with different policies. Passwords are never returned.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| tls | Yes | |
| note | Yes | |
| path | Yes | |
| tools | Yes | |
| brokers | Yes | |
| cluster | Yes | |
| endpoint | Yes | |
| read_only | Yes | |
| sasl_user | No | |
| output_dir | Yes | |
| config_file | No | |
| description | No | |
| http_address | No | |
| sasl_options | No | |
| topic_formats | No | |
| authentication | Yes | |
| schema_registry | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it does meaningful work: it discloses that passwords are never returned (secret redaction) and surfaces the read-only policy and authentication/TLS posture. It stops short of explicitly stating that the call has no side effects, which an agent must infer from 'Report'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the verb and resource, then enumerates the reported fields, then closes with the usage condition. The field list is long but every item is substantive; structure is sound with minimal waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the description need not explain return values, and it already covers purpose, the when-to-use condition, and the security behavior. For a zero-parameter introspection tool this is complete enough to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate. Baseline 4 applies; the description correctly adds no fabricated parameter guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Report) and a well-defined resource (this endpoint's configuration), then enumerates exactly what is reported: name, path, purpose, cluster, brokers, auth, TLS, read-only policy, Schema Registry, topic formats and exposed tools. That is far more specific than a tautology, though it never names a sibling to differentiate itself from, e.g., list_clusters or describe_topic.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for use — confirm the target and permissions before acting — and adds the non-obvious rationale that several endpoints may target one cluster with different policies. No explicit when-not or named alternative, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
24 tool updates
v0.2.0- First observed
add_partitions - First observed
alter_topic_config - First observed
cluster_health - First observed
commit_offset - First observed
compare_clusters - First observed
consumer_lag - First observed
copy_message - First observed
create_topic - First observed
delete_consumer_group - First observed
delete_records - First observed
delete_topic - First observed
describe_consumer_group - First observed
describe_topic - First observed
get_message - First observed
get_schema - First observed
list_acls - First observed
list_clusters - First observed
list_consumer_groups - First observed
list_topics - First observed
open_transactions - First observed
produce_message - First observed
sample_messages - First observed
search_messages - First observed
server_config
TDQS
Scored across 24 tools
Each tool targets a distinct Kafka resource and operation. List/describe pairs (e.g., list_topics vs. describe_topic, list_consumer_groups vs. describe_consumer_group) are clearly separated by scope, and read/write message tools (get_message, sample_messages, search_messages, produce_message, copy_message) have non-overlapping intents.
All names use snake_case with no camelCase or other style mixing. However, a few monitoring tools (cluster_health, consumer_lag, server_config) are noun phrases rather than the predominant verb_noun pattern, which is a minor deviation.
24 tools is above the typical 3-15 range, but the breadth reflects a genuinely complex Kafka administration domain and no tool appears redundant. The set is heavy yet each tool earns its place for a comprehensive surface.
Core lifecycle operations for topics, partitions, configs, messages, consumer groups, and clusters are well covered. Direct schema registration and ACL write operations are absent, and some areas (broker config, transaction control) are read-only, which are minor-to-moderate gaps for a full admin surface.
Maintenance
Related MCP Connectors
MCP server for secureFlows: token-free URL builders and integration-linting tools for AI agents.
An MCP server that provides an API to LLMs to manage their JumpCloud resources.
MCP server providing access to the Scorecard API to evaluate and optimize LLM systems.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceAn MCP server that enables interaction with Kafka clusters to manage topics, monitor consumer groups, and stream messages. It provides a comprehensive suite of tools for broker metadata inspection and local Kafka user management.MIT
- AlicenseBqualityBmaintenanceMCP server for Apache Kafka that allows LLM agents to inspect topics, consumer groups, and safely manage offsets (reset, rewind).1913Apache 2.0
- AlicenseNot gradedqualityAmaintenanceAn MCP server that enables AI assistants to safely interact with Apache Kafka clusters, providing tools for topic management, message operations, consumer groups, and cluster information.3MIT
- FlicenseCqualityDmaintenanceExposes Kafka administration operations as MCP tools, enabling AI agents to inspect Kafka clusters using natural language.1-