Gecko Scorecard
get_scorecardBenchGecko Gecko Scorecard: every model across BenchGecko's own profile tests with a letter grade per test and an overall Gecko Score (behavior, not intelligence).
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
get_scorecardBenchGecko Gecko Scorecard: every model across BenchGecko's own profile tests with a letter grade per test and an overall Gecko Score (behavior, not intelligence).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
Changes observed during successful MCP inspections.
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint=true, idempotentHint=true, openWorldHint=false, and destructiveHint=false, so the safety profile is covered. The description adds a genuinely useful semantic note that the score measures behavior rather than intelligence, but says nothing about result volume or how the 'every model' claim interacts with a limiting parameter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that identifies the tool and its return contents with no filler. The parenthetical '(behavior, not intelligence)' earns its place as disambiguation, though the sentence is dense enough that a second short clause about scope would have helped.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does describe the shape at a high level (per-test letter grades plus an overall Gecko Score). It stops short of covering the pagination/limit behavior implied by the schema's default of 15 against the claim of 'every model'.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single `limit` parameter, and the description never mentions it, so an agent gets no indication of whether it caps models, tests, or rows, nor of its default of 15 versus the maximum of 100. One parameter with no textual compensation falls below the baseline for documented params.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (Gecko Scorecard), its coverage (every model across BenchGecko's profile tests), and its contents (letter grade per test plus an overall Gecko Score). This implicitly distinguishes it from get_gecko_test (single test) and get_model, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: an agent can infer this is the aggregate-scorecard view, in contrast to get_gecko_test or compare_models. There is no explicit 'use this when' statement, no exclusion, and no mention of when a per-model or per-test tool would be preferable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Add one secure layer between your agents and this server.