Skip to main content
Glama

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already supply the safety profile: readOnlyHint=false, destructiveHint=false, idempotentHint=true. The description adds upsert semantics ('Create or update') but no extra detail about whether existing translations are overwritten wholesale or how partial updates behave, so it adds some but not rich context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One concise sentence that front-loads the action and resource, plus a relevant API reference link. There is no filler or redundant repetition of schema content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple upsert operation with 100% parameter coverage, annotations for safety, and no output schema, the description is largely sufficient. The main missing piece is sibling differentiation, which is already captured in the usage_guidelines dimension.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema covers all parameters with descriptions, so the baseline is 3 under the rubric. The description's phrase 'for a specific locale' loosely maps to the locale parameter but adds no new meaning beyond the schema fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Create or update') and resource ('a translation for a specific locale'), making the core function immediately clear. It does not explicitly distinguish from sibling locale-related tools like put_notification_locale or put_journey_template_locale, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as get_translation, put_notification_locale, or put_journey_template_locale. The description explains what the tool does but not which scenario selects it over siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

B3.1/5.0
Disambiguation3/5

Tools are mostly organized as distinct resource/action pairs, but several clusters are easy to confuse: list subscription tools (add_subscribers_to_list vs bulk_subscribe_to_list vs subscribe_user_to_list), message vs message-content vs message-history retrieval, and the many journey/journey-template list/get tools. Detailed descriptions rescue most selections, but the sheer number of near-identical verb/resource names creates real misselection risk.

Naming Consistency4/5

Almost all tools follow a snake_case verb_noun pattern (create_, get_, list_, replace_, send_, publish_, archive_). Minor deviations keep it from a perfect score: courier_installation_guide is noun-first, and add_bulk_users sits awkwardly next to the bulk_add_* family, but the overall convention is predictable and readable.

Tool Count1/5

144 tools is an extreme working-set size for an agent to hold and choose from, far beyond the reasonable 3–15 range. Even for a broad platform like Courier, this should be split into focused sub-servers (templates, journeys, users, lists, preferences, etc.) to remain usable.

Completeness4/5

The surface is remarkably comprehensive, covering sending, templates, journeys, automations, users, tenants, lists, preferences, providers, routing, brands, audiences, translations, digests, bulk jobs, and audit events. Notable gaps exist—automation template CRUD and digest schedule management are missing—but most workflows can still be completed with workarounds.