Skip to main content
Glama

prompt_effectiveness

Analyze Agent template performance by aggregating success rates, durations, failure reasons, and lessons to identify prompts that need improvement.

Instructions

Return effectiveness statistics for Agent templates.

Frozen: still callable, no longer developed.

Aggregates activity records to compute success rate, average duration, and top failure reasons per template. Also counts failure_analysis lessons per template, matched through the failed task's assigned agent.

Use this to identify which Agent templates perform well and which need prompt improvement.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
template_nameNoOptional filter (e.g. "engineering-backend-architect"). Leave empty to return stats for all templates.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.9.0

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does a solid job: it discloses the lifecycle caveat ('Frozen: still callable, no longer developed'), which an agent could not infer otherwise, and explains the computation pipeline (aggregating activity records, counting failure_analysis lessons matched through the failed task's assigned agent). It never explicitly says the operation is read-only or mentions cost/latency, which keeps it short of 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose, then delivers the frozen-lifecycle caveat, the computation detail, and the intended use in short, scannable blocks. The sentence about failure_analysis lesson matching is dense but earns its place by explaining a non-obvious metric derivation; nothing reads as filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-shape explanation is unnecessary, and the description covers purpose, mechanics, and the frozen status. For a single-optional-param, read-only analytics tool this is nearly complete; only the absence of explicit read-only framing and alternative-tool routing leaves a small gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one optional parameter, and the schema already documents it at 100% coverage including the empty-string-means-all behavior. The description adds no syntax or format guidance beyond the schema, so the baseline of 3 is appropriate when structured fields do the lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Return effectiveness statistics for Agent templates') and even enumerates the computed metrics (success rate, average duration, top failure reasons), so the agent knows exactly what comes back. It does not, however, differentiate itself from nearby siblings such as failure_analysis, agent_activity_query, or agent_template_list by name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states the purpose-driven use case: 'Use this to identify which Agent templates perform well and which need prompt improvement.' That gives clear context for when to reach for it. There is no when-not guidance and no pointer to alternatives like failure_analysis for drilling into individual failures, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools