get_message_history
Get the event history for a message, showing each step in the delivery pipeline (enqueued, sent, delivered, etc.).
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | Filter by event type | |
| message_id | Yes | The message ID |
Get the event history for a message, showing each step in the delivery pipeline (enqueued, sent, delivered, etc.).
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | Filter by event type | |
| message_id | Yes | The message ID |
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description is consistent with that. It adds value beyond the annotations by describing the returned information as a sequence of delivery stages, which tells the agent what kind of lifecycle data to expect.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The entire description is one focused sentence that front-loads the core action and immediately clarifies the return content with examples. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only lookup with fully documented parameters and a safe annotation, the description is nearly complete. It gives enough examples to understand the output even without an output schema, though it does not specify ordering or pagination.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with message_id and type both documented in the schema. The tool description adds no additional parameter semantics beyond the schema, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Get the event history for a message.' It also adds concrete examples of the pipeline stages (enqueued, sent, delivered), which clearly distinguishes this from sibling tools like get_message or get_message_content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes it clear this is the tool for delivery-pipeline event history, giving enough context for an agent to select it over get_message or get_audit_event. It does not explicitly name alternatives or exclusions, but the use case is strongly implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Add one secure layer between your agents and this server.
Tools are mostly organized as distinct resource/action pairs, but several clusters are easy to confuse: list subscription tools (add_subscribers_to_list vs bulk_subscribe_to_list vs subscribe_user_to_list), message vs message-content vs message-history retrieval, and the many journey/journey-template list/get tools. Detailed descriptions rescue most selections, but the sheer number of near-identical verb/resource names creates real misselection risk.
Almost all tools follow a snake_case verb_noun pattern (create_, get_, list_, replace_, send_, publish_, archive_). Minor deviations keep it from a perfect score: courier_installation_guide is noun-first, and add_bulk_users sits awkwardly next to the bulk_add_* family, but the overall convention is predictable and readable.
144 tools is an extreme working-set size for an agent to hold and choose from, far beyond the reasonable 3–15 range. Even for a broad platform like Courier, this should be split into focused sub-servers (templates, journeys, users, lists, preferences, etc.) to remain usable.
The surface is remarkably comprehensive, covering sending, templates, journeys, automations, users, tenants, lists, preferences, providers, routing, brands, audiences, translations, digests, bulk jobs, and audit events. Notable gaps exist—automation template CRUD and digest schedule management are missing—but most workflows can still be completed with workarounds.