compare_models
Compare API response costs across multiple LLM models
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| models | No | Array of model IDs to compare | |
| content | Yes | API response content | |
| input_tokens | No | Optional input token count |
Compare API response costs across multiple LLM models
| Name | Required | Description | Default |
|---|---|---|---|
| models | No | Array of model IDs to compare | |
| content | Yes | API response content | |
| input_tokens | No | Optional input token count |
Changes observed during successful MCP inspections.
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden for behavioral disclosure. It only states the comparison purpose, omitting any mention of side effects, permissions, output format, or how the parameters interact. This is similar to the update_drive calibration case, which scored 2.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, front-loaded with the core action, and contains no redundant words. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotations, the description is too sparse. It fails to explain what the tool returns, whether models is actually required to compare, or how content and input_tokens factor into the cost comparison. This leaves critical gaps for the agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (all three parameters have descriptions), so the baseline is 3. The description adds no extra semantic meaning beyond the schema, and its implication that 'models' is needed ('across multiple LLM models') slightly conflicts with the schema marking models as optional.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'compare' and clearly identifies the resource ('API response costs across multiple LLM models'). This distinguishes it from sibling tools like 'estimate_cost' (likely singular cost) and 'analyze_response' (likely content analysis).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives, no exclusions, and no naming of alternative tools. It is a bare functional statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Add one secure layer between your agents and this server.