Skip to main content
Glama

compare_reports

Read-onlyIdempotent

Compare two MetaTrader tester reports to show absolute and percent changes per key, with optional guard thresholds to flag profit metrics below or drawdown/loss limits above.

Instructions

Diff two tester reports key by key (absolute and percent deltas), optionally checking guard thresholds.

With guards, a profit-style metric below its threshold or a drawdown/loss metric above it is a violation. Accepting only improvements across many iterations is selection bias; treat guards as a sanity gate, not proof of robustness. Replaces regression_check.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
guardsNoOptional percent thresholds per summary key, e.g. {"net_profit": -5, "profit_factor": -10, "max_drawdown": 25}; when given, `violations` and `ok` are added.
baselineYesAbsolute path to the baseline tester report (.htm).
candidateYesAbsolute path to the candidate tester report (.htm).

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed4 schema fields changedv0.5.0
    • addedInput schema / properties / baseline / description
      Added value: +"Absolute path to the baseline tester report (.htm)."
    • addedInput schema / properties / candidate / description
      Added value: +"Absolute path to the candidate tester report (.htm)."
    • addedInput schema / properties / guards
      Added value: +{
      +  "anyOf": [
      +    {
      +      "additionalProperties": true,
      +      "type": "object"
      +    },
      +    {
      +      "type": "null"
      +    }
      +  ],
      +  "default": null,
      +  "description": "Optional percent thresholds per summary key, e.g. {\"net_profit\": -5, \"profit_factor\": -10, \"max_drawdown\": 25}; when given, `violations` and `ok` are added.",
      +  "title": "Guards"
      +}
    • changedOutput schema / (root)
      Previous value: -nullNew value: +{
      +  "additionalProperties": true,
      +  "title": "compare_reportsDictOutput",
      +  "type": "object"
      +}
  2. First observedv0.4.1

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only and idempotent behavior. The description adds useful guard semantics and a caution about selection bias, but it does not disclose deeper behavioral details like output size, performance, or whether files are consulted beyond what the schema implies.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The main purpose is stated in one crisp sentence, with a focused second paragraph elaborating on guard semantics. The caveat about selection bias is relevant, though slightly tangential to simply invoking the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and annotations covering safety, the description provides enough context to know when to combine variables. It explains the optional guard behavior and its interpretation, covering the essential decision space.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers all parameters, but the description adds key meaning. Specifically, it explains that profit-style metrics below a threshold and drawdown/loss metrics above are violations, going beyond the schema's 'percent thresholds per summary key'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Diff two tester reports key by key (absolute and percent deltas)', which uses a specific verb and resource distinctly. It also explicitly says 'Replaces regression_check', differentiating it from a sibling tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states it is the replacement for regression_check, giving clear guidance on which sibling to use. It also frames when to use guards and how to interpret thresholds, though it does not enumerate contexts where other comparison tools should be preferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.