Nonobench
Server Details
An open-source benchmark of how well LLMs solve nonogram puzzles, from 5x5 to 20x20. Query every result through a REST API or MCP server.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 11 tools
Most tools target distinct query types, but get_model_results, compare_models, and get_leaderboard all surface model accuracy metrics, and get_model_puzzles vs get_puzzle_results are transposed views of the same outcome data. Descriptions clarify the intended scope reasonably well (one model vs side-by-side vs ranked; by-model vs by-puzzle), so confusion is limited.
Every tool follows a clean verb_noun pattern: get_*/list_*/check_*/compare_* with snake_case throughout. The domain nouns (model, puzzle, leaderboard, families, providers, runs, solution) are singular/plural used consistently.
11 tools is well within the ideal 3-15 range for a benchmark-inspection server. Each tool maps to a distinct facet (leaderboards, per-model results, per-puzzle results, puzzle retrieval, solution checking, reference listings), so none feel redundant.
The surface covers the full read/inspect lifecycle: leaders, per-model and per-puzzle breakdowns, puzzle content, solution verification, and family/provider/run listings. Minor gaps like puzzle filtering/search or aggregate cost/latency trends are workable around.
Available Tools
11 toolscheck_solutionCheck solutionARead-onlyInspect
Check a grid against a puzzle's clues, using the same rule as the benchmark grader. Reports which rows and columns do not match.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Puzzle id from list_puzzles | |
| grid | Yes | Row-major string of 0 (empty) and 1 (filled), width × height characters |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | |
| correct | Yes | |
| rowViolations | Yes | |
| columnViolations | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds meaningful context beyond that: it uses 'the same rule as the benchmark grader' and reports specific mismatched rows and columns, which helps set expectations for the check behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the tool's purpose and followed by the reporting behavior. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity read-only check tool, the description covers purpose and key behavior, while the schema documents parameters and an output schema exists for return values. It is nearly complete, though usage context relative to sibling tools is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the two parameters (id and grid) are fully documented in the schema, including the grid's row-major 0/1 format. The description adds no parameter-level detail beyond the schema, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Check a grid against a puzzle's clues' and clarifies it uses the benchmark grader's rule. It does not explicitly distinguish itself from siblings such as get_puzzle_results, so it is clear but lacks sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description says what the tool does but gives no guidance on when to use it versus alternatives like get_puzzle_results or get_puzzle. There are no prerequisites, exclusions, or alternative tools named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compare_modelsCompare modelsARead-onlyInspect
Side-by-side core overall and per-size accuracy, cost, latency and token results for model or family names.
| Name | Required | Description | Default |
|---|---|---|---|
| models | Yes | 2 to 20 model variant ids or family names, e.g. claude-opus-5.5 or gpt-6-astra-xhigh |
Output Schema
| Name | Required | Description |
|---|---|---|
| models | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds which metric families are compared (accuracy, cost, latency, tokens; overall and per-size), which is useful, but since an output schema exists this is largely a preview of return contents rather than disclosure of behavior an agent couldn't infer.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no filler, front-loading the comparison framing and then the dimensions compared. Nothing could be removed without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, one-parameter comparison tool with a full output schema, the description covers what is compared closely enough for correct invocation. The remaining gap is routing guidance versus sibling tools such as get_model_results or get_leaderboard.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the single 'models' array is fully documented, including the 2–20 bound and id examples. The description's 'model or family names' merely restates the schema, adding no new format or syntax guidance; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific action (side-by-side comparison) plus the resource (models/families) and enumerates the compared dimensions: overall and per-size accuracy, cost, latency, and token results. An agent can distinguish it from single-model tools like get_model_results, though the description never explicitly names that sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The word 'side-by-side' and the plural 'names' imply the multi-model use case, and the schema's 2–20 min/max reinforces it, but there is no explicit 'use this when…' statement and no reference to get_model_results, get_leaderboard, or other alternatives for single-model or ranked lookups.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_leaderboardGet leaderboardBRead-onlyInspect
Models ranked by accuracy; defaults to all effort levels for compatibility.
| Name | Required | Description | Default |
|---|---|---|---|
| size | No | Grid size to filter on | |
| effort | No | best, all (default), or one effort level; empty means all | |
| family | No | Comma-separated family ids; empty means no filter | |
| version | No | Comma-separated benchmark versions: 1.0, 1.1, 1.2; empty means all | |
| provider | No | Comma-separated provider ids; empty means no filter | |
| reasoning | No | Only reasoning (true) or non-reasoning (false) variants | |
| min_correct | No | Minimum puzzles solved in the selected tier; default 0 includes unsolved variants | |
| open_weights | No | Only open-weight (true) or closed (false) models |
Output Schema
| Name | Required | Description |
|---|---|---|
| size | Yes | |
| models | Yes | |
| updatedAt | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds one genuine behavioral fact — that omitted effort defaults to all levels for compatibility — but says nothing about ranking metric mechanics or tie-breaking. Adequate but thin against an annotation-covered read tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight clause with the core ranking statement front-loaded and the default-behavior caveat trailing. No padding, though it is arguably under-specified rather than optimally concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a full output schema and 100% parameter coverage, the description is not required to explain returns or params. However, for an eight-filter leaderboard tool it omits any hint of sort order, scope semantics, or how 'accuracy' is measured, leaving an agent to infer ranking details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% across all eight parameters, including the effort default and empty-means-all conventions, so the schema carries the load. The description's effort-default remark marginally reinforces the schema rather than adding new meaning. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource and operation — models ranked by accuracy — so an agent knows it returns an ordered list of models. It does not distinguish itself from lookalike siblings such as get_model_results or compare_models, which also surface model performance data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use or when-not-to-use guidance and names no alternative among the ten siblings. The phrase about effort-level defaults is a behavioral note, not routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_model_puzzlesGet model puzzlesBRead-onlyInspect
Which puzzles one model solved, missed, timed out on, or has not run.
| Name | Required | Description | Default |
|---|---|---|---|
| model | Yes | Model variant id as listed on the leaderboard, e.g. claude-opus-5.5-high |
Output Schema
| Name | Required | Description |
|---|---|---|
| model | Yes | |
| solved | Yes | |
| puzzles | Yes | |
| attempted | Yes | |
| displayName | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this a safe read operation (readOnlyHint=true, openWorldHint=false), so the safety profile is covered. The description adds the useful behavioral detail that results are partitioned into solved/missed/timed-out/not-run categories, but says nothing about auth, scope, or cost; with an output schema present, disclosing return values is not required.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short sentence fragment with zero filler and the key content front-loaded. It is arguably too terse — a verbless clause rather than a call-to-action sentence — which slightly limits how self-explanatory it reads in a tool list.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and the only parameter fully documented in the schema, the description carries little remaining burden. The one gap is routing: it never clarifies how this differs from get_model_results, which matters given how many retrieval siblings exist.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'model' parameter is clearly documented there with an example id format. The description adds no syntax or format detail beyond the schema, which is the expected baseline when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (puzzles for one model) and enumerates the outcome buckets returned (solved, missed, timed out, not run), which is more specific than restating the title. It does not, however, distinguish this from the close sibling get_model_results, so an agent cannot tell the two apart from the description alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as get_model_results or get_puzzle_results. The agent must infer from the name alone which of the sibling retrieval tools to call.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_model_resultsGet model resultsBRead-onlyInspect
Accuracy, cost, latency and token use for one model, broken down by grid size.
| Name | Required | Description | Default |
|---|---|---|---|
| model | Yes | Model name as listed on the leaderboard, e.g. gpt-5.4-xhigh |
Output Schema
| Name | Required | Description |
|---|---|---|
| model | Yes | Model variant id, e.g. claude-opus-5.5-high |
| total | Yes | |
| bySize | Yes | |
| effort | Yes | |
| family | Yes | |
| correct | Yes | |
| accuracy | Yes | Percentage of puzzles solved |
| provider | Yes | |
| reasoning | Yes | |
| failedRuns | Yes | |
| displayName | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds the useful detail that results are segmented by grid size, but says nothing about aggregation, time window, or what happens for models without runs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; every word (metrics list, single-model scope, grid-size breakdown) carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be restated, and the parameter list is trivially small. The remaining gap is routing context: nothing tells the agent why to pick this over get_model_puzzles or compare_models.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single 'model' parameter, including a naming example ('gpt-5.4-xhigh'), so the schema does the heavy lifting. The description's phrase 'for one model' adds no format or constraint detail beyond that.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the specific resource (results for one model) and enumerates the metrics returned (accuracy, cost, latency, token use) with a scope qualifier ('broken down by grid size'). It does not explicitly distinguish itself from close siblings like get_model_puzzles or get_puzzle_results, so 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use guidance, no prerequisites, and names no alternatives. With siblings such as compare_models, get_leaderboard, and get_model_puzzles in the same family, the agent gets no signal on when this tool is the right choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_puzzleGet puzzleARead-onlyInspect
One puzzle, including the clue text models were prompted with. The reference solution is only included on request; some puzzles have several valid solutions.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Puzzle id from list_puzzles | |
| include_solution | No | Include a reference solution |
Output Schema
| Name | Required | Description |
|---|---|---|
| id | Yes | |
| url | Yes | |
| size | Yes | |
| index | Yes | |
| width | Yes | |
| height | Yes | |
| prompt | Yes | Clue text as models received it |
| rowClues | Yes | |
| columnClues | Yes | |
| referenceSolution | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds genuine behavioral context beyond that: the clue text is included, the reference solution is opt-in, and some puzzles have several valid solutions — a useful caveat that affects downstream comparison/checking. It stops short of auth or rate-limit details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero filler. The core result (the puzzle and its clue text) is front-loaded, followed by the solution-conditional caveat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value detail is unnecessary, and the description covers the non-obvious piece an agent needs — the optional solution and the existence of multiple valid solutions. It could add a note about the id source or sibling routing, but is largely complete for a two-param read.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3 and the schema already documents both parameters. The description nevertheless adds meaning: it explains that the solution is 'only included on request,' clarifying the conditional semantics of include_solution beyond its terse schema label.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('One puzzle') and clarifies it returns the clue text models were prompted with, which distinguishes it from the plural list_puzzles sibling. It does not, however, explicitly name any sibling or state the singular-by-id retrieval role to contrast with get_puzzle_results or get_model_puzzles.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied (fetch a single puzzle, pass an id from list_puzzles), and it notes the solution is only included on request, which guides use of include_solution. But there is no explicit when-to-use/when-not guidance and no named alternative for adjacent needs like results or model comparisons.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_puzzle_resultsGet puzzle resultsBRead-onlyInspect
Per-model outcomes for one puzzle. Answers are omitted unless requested.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Puzzle id from list_puzzles | |
| effort | No | best, all (default), or one effort level | |
| family | No | Comma-separated family ids; empty means no filter | |
| provider | No | Comma-separated provider ids; empty means no filter | |
| reasoning | No | Only reasoning (true) or non-reasoning (false) variants | |
| open_weights | No | Only open-weight (true) or closed (false) models | |
| include_answers | No | Include each model's answer grid |
Output Schema
| Name | Required | Description |
|---|---|---|
| runs | Yes | |
| size | Yes | |
| index | Yes | |
| solved | Yes | |
| attempts | Yes | |
| puzzleId | Yes | |
| updatedAt | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description usefully discloses a non-obvious default (answer grids are omitted unless requested), but says nothing about pagination, result volume, or behavior when filters match nothing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no waste, with the resource scope front-loaded and the output-default caveat trailing. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the description covers the resource scope plus the answers-omission behavior. For a 7-param filtered query tool the definition is adequate, though the absence of any usage routing leaves a small gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and there are 7 well-documented parameters, so the schema carries the parameter burden. The description adds nothing about parameter syntax, defaults, or filter interaction beyond what the schema already states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Per-model outcomes for one puzzle" gives a specific verb+resource with a clear scope (one puzzle, all models), which implicitly separates it from get_puzzle (single puzzle metadata) and get_model_results (one model across puzzles). It never names a sibling, so differentiation is inferable rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no exclusions, and no routing to alternatives such as get_puzzle or get_model_results. The scoping phrase hints at the context, but the agent must infer when this is the right call.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_familiesList familiesCRead-onlyInspect
Model families, available efforts and best variants.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| families | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, covering the safety profile. The description adds no behavioral details such as whether the list is complete, static, paginated, or otherwise constrained, and its content is largely restated by the output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The seven-word fragment is very concise and front-loads the returned content categories. It is not a complete sentence, but it wastes no words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-parameter list tool with an output schema and read-only annotations, the description supplies the core resource and return categories. It still omits usage context relative to sibling list/get tools, so it is only minimally complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there are no parameter semantics for the description to clarify. The baseline for zero parameters is 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (model families) and return content (available efforts and best variants), which helps distinguish it from siblings like list_providers. However, it is a noun phrase with no explicit verb or scope statement, so the tool's action is only implied.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this tool versus list_providers, list_puzzles, or get_leaderboard. Usage is only inferable from the resource name, with no alternatives or exclusions stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_providersList providersCRead-onlyInspect
Provider ids, names, families and variant counts.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| providers | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds no behavioral context beyond the annotations: no statement about ordering, cost, or whether results are cached or exhaustive. What it does say (field names) duplicates the existing output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single short fragment with no filler, but it is telegraphic rather than front-loaded prose — a nominal phrase with no verb, which forces the reader to reconstruct intent. Brief but under-specified.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters and an output schema present, the description needn't explain return values or arguments, so the scope of what's missing is narrow. The remaining gap is any routing guidance among the many sibling list_* and get_* tools, which goes unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric this is baseline 4. There is nothing parameter-related for the description to compensate for.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The fragment 'Provider ids, names, families and variant counts' tells the agent the resource is providers and enumerates the fields returned, so the purpose is inferable but never stated as an action. It also does not distinguish itself from the sibling list_families, which is a plausible alternate target. Vague-but-decodable rather than a clear verb+resource statement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no condition selecting this tool over list_families or check_solution, and no prerequisites. The agent must infer usage entirely from the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_puzzlesList puzzlesCRead-onlyInspect
The benchmark puzzles with their ids and row/column clues.
| Name | Required | Description | Default |
|---|---|---|---|
| size | No | Grid size to filter on |
Output Schema
| Name | Required | Description |
|---|---|---|
| puzzles | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds only that ids and row/column clues are included, which overlaps with the existing output schema, and says nothing about volume, pagination, or result size. Little behavioral value beyond the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no filler, front-loaded with the resource and its contents. It is efficient, though arguably too terse for a listing tool, bordering on under-specification rather than maximal information density.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema exists, so return values need not be spelled out, and annotations cover safety. Still, for a listing tool the description omits anything about result volume, ordering, or how it relates to get_puzzle. Adequate but with clear gaps given the crowded sibling set.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one optional parameter (size) with 100% schema description coverage and an enum of grid sizes, so the schema fully documents it. The description does not mention filtering by size at all, adding no meaning beyond the schema. Per the baseline rule for high schema coverage, 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (benchmark puzzles) and the fields returned (ids and row/column clues), so an agent knows this is a listing of puzzle definitions. However, it is a noun phrase with no verb and no differentiation from siblings like get_puzzle or get_model_puzzles, which also concern puzzles. Purpose is decipherable but not sharply scoped.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no indication of when to use this tool versus get_puzzle (single puzzle), get_model_puzzles, or get_puzzle_results. No prerequisites, no exclusions, no alternatives are mentioned. The agent must infer usage purely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_runsList runsBRead-onlyInspect
Individual benchmark runs (one model on one puzzle), optionally with the raw prompt and model output.
| Name | Required | Description | Default |
|---|---|---|---|
| size | No | Grid size to filter on | |
| limit | No | Default 100 | |
| model | No | Only runs of this model variant id | |
| offset | No | Runs to skip, for paging; default 0 | |
| puzzle_id | No | Only runs on this puzzle id from list_puzzles | |
| include_output | No | Include raw prompt and model output (large) |
Output Schema
| Name | Required | Description |
|---|---|---|
| runs | Yes | |
| limit | Yes | |
| total | Yes | |
| offset | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds only that output can optionally include raw prompt/model output, which largely restates the schema's include_output note rather than disclosing new behavior like pagination defaults or payload size limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no wasted words or filler. It is efficient, though it is arguably too thin to carry the tool's full scope.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and annotations covering the safety profile, return values and safety need not be explained. However, a six-parameter listing tool whose description omits any filtering or sibling-routing guidance is only minimally complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and every parameter (size, limit, model, offset, puzzle_id, include_output) is documented in the schema itself. The description adds no syntax, format, or interaction detail beyond what the schema provides, so it meets the baseline only.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the specific resource ('individual benchmark runs') and defines it precisely as 'one model on one puzzle', which is genuinely clarifying. It does not, however, distinguish this list tool from siblings like get_model_results or get_puzzle_results that may return overlapping data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no condition selecting this over the sibling result tools, and no mention of how filtering should be applied. The agent must infer usage entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
11 tool updates
- Changed
check_solution1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": {}, + "properties": { + "columnViolations": { + "items": {}, + "type": "array" + }, + "correct": { + "type": "boolean" + }, + "error": { + "type": "string" + }, + "rowViolations": { + "items": {}, + "type": "array" + } + }, + "required": [ + "correct", + "rowViolations", + "columnViolations" + ], + "type": "object" +}
- Changed
compare_models2 fields changed- added
Input schema / properties / models / descriptionAdded value: +"2 to 20 model variant ids or family names, e.g. claude-opus-5.5 or gpt-6-astra-xhigh" - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": false, + "properties": { + "models": { + "items": { + "additionalProperties": {}, + "properties": { + "accuracy": { + "description": "Percentage of puzzles solved", + "type": "number" + }, + "bySize": { + "items": { + "additionalProperties": {}, + "properties": { + "size": { + "type": "string" + } + }, + "required": [ + "size" + ], + "type": "object" + }, + "type": "array" + }, + "correct": { + "type": "number" + }, + "displayName": { + "type": "string" + }, + "effort": { + "type": [ + "string", + "null" + ] + }, + "failedRuns": { + "type": "number" + }, + "family": { + "type": "string" + }, + "model": { + "description": "Model variant id, e.g. claude-opus-5.5-high", + "type": "string" + }, + "provider": { + "type": [ + "string", + "null" + ] + }, + "reasoning": { + "type": "boolean" + }, + "total": { + "type": "number" + } + }, + "required": [ + "model", + "displayName", + "family", + "effort", + "provider", + "reasoning", + "accuracy", + "correct", + "total", + "failedRuns", + "bySize" + ], + "type": "object" + }, + "type": "array" + } + }, + "required": [ + "models" + ], + "type": "object" +}
- Changed
get_leaderboard3 fields changed- added
Input schema / properties / open_weights / descriptionAdded value: +"Only open-weight (true) or closed (false) models" - added
Input schema / properties / reasoning / descriptionAdded value: +"Only reasoning (true) or non-reasoning (false) variants" - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": false, + "properties": { + "models": { + "items": { + "additionalProperties": {}, + "properties": { + "accuracy": { + "description": "Percentage of puzzles solved", + "type": "number" + }, + "correct": { + "type": "number" + }, + "displayName": { + "type": "string" + }, + "effort": { + "type": [ + "string", + "null" + ] + }, + "family": { + "type": "string" + }, + "model": { + "description": "Model variant id, e.g. claude-opus-5.5-high", + "type": "string" + }, + "provider": { + "type": [ + "string", + "null" + ] + }, + "rank": { + "type": "number" + }, + "reasoning": { + "type": "boolean" + }, + "total": { + "type": "number" + }, + "totalCostUsd": { + "type": "number" + } + }, + "required": [ + "model", + "displayName", + "family", + "effort", + "provider", + "reasoning", + "accuracy", + "rank", + "correct", + "total", + "totalCostUsd" + ], + "type": "object" + }, + "type": "array" + }, + "size": { + "type": "string" + }, + "updatedAt": { + "type": "string" + } + }, + "required": [ + "updatedAt", + "size", + "models" + ], + "type": "object" +}
- Changed
get_model_puzzles2 fields changed- added
Input schema / properties / model / descriptionAdded value: +"Model variant id as listed on the leaderboard, e.g. claude-opus-5.5-high" - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": false, + "properties": { + "attempted": { + "type": "number" + }, + "displayName": { + "type": "string" + }, + "model": { + "type": "string" + }, + "puzzles": { + "items": { + "additionalProperties": {}, + "properties": { + "id": { + "type": "string" + }, + "index": { + "type": "number" + }, + "size": { + "type": "string" + }, + "state": { + "enum": [ + "solved", + "wrong", + "cut-off", + "not-run" + ], + "type": "string" + } + }, + "required": [ + "id", + "index", + "size", + "state" + ], + "type": "object" + }, + "type": "array" + }, + "solved": { + "type": "number" + } + }, + "required": [ + "model", + "displayName", + "solved", + "attempted", + "puzzles" + ], + "type": "object" +}
- Changed
get_model_results1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": {}, + "properties": { + "accuracy": { + "description": "Percentage of puzzles solved", + "type": "number" + }, + "bySize": { + "items": { + "additionalProperties": {}, + "properties": { + "size": { + "type": "string" + } + }, + "required": [ + "size" + ], + "type": "object" + }, + "type": "array" + }, + "correct": { + "type": "number" + }, + "displayName": { + "type": "string" + }, + "effort": { + "type": [ + "string", + "null" + ] + }, + "failedRuns": { + "type": "number" + }, + "family": { + "type": "string" + }, + "model": { + "description": "Model variant id, e.g. claude-opus-5.5-high", + "type": "string" + }, + "provider": { + "type": [ + "string", + "null" + ] + }, + "reasoning": { + "type": "boolean" + }, + "total": { + "type": "number" + } + }, + "required": [ + "model", + "displayName", + "family", + "effort", + "provider", + "reasoning", + "accuracy", + "correct", + "total", + "failedRuns", + "bySize" + ], + "type": "object" +}
- Changed
get_puzzle1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": {}, + "properties": { + "columnClues": { + "items": { + "items": { + "type": "number" + }, + "type": "array" + }, + "type": "array" + }, + "height": { + "type": "number" + }, + "id": { + "type": "string" + }, + "index": { + "type": "number" + }, + "prompt": { + "description": "Clue text as models received it", + "type": "string" + }, + "referenceSolution": { + "type": "string" + }, + "rowClues": { + "items": { + "items": { + "type": "number" + }, + "type": "array" + }, + "type": "array" + }, + "size": { + "type": "string" + }, + "url": { + "type": "string" + }, + "width": { + "type": "number" + } + }, + "required": [ + "id", + "index", + "size", + "width", + "height", + "rowClues", + "columnClues", + "url", + "prompt" + ], + "type": "object" +}
- Changed
get_puzzle_results8 fields changed- added
Input schema / properties / effort / descriptionAdded value: +"best, all (default), or one effort level" - added
Input schema / properties / family / descriptionAdded value: +"Comma-separated family ids; empty means no filter" - added
Input schema / properties / id / descriptionAdded value: +"Puzzle id from list_puzzles" - added
Input schema / properties / include_answers / descriptionAdded value: +"Include each model's answer grid" - added
Input schema / properties / open_weights / descriptionAdded value: +"Only open-weight (true) or closed (false) models" - added
Input schema / properties / provider / descriptionAdded value: +"Comma-separated provider ids; empty means no filter" - added
Input schema / properties / reasoning / descriptionAdded value: +"Only reasoning (true) or non-reasoning (false) variants" - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": false, + "properties": { + "attempts": { + "type": "number" + }, + "index": { + "type": "number" + }, + "puzzleId": { + "type": "string" + }, + "runs": { + "items": { + "additionalProperties": {}, + "properties": { + "answer": { + "type": [ + "string", + "null" + ] + }, + "correct": { + "type": "boolean" + }, + "displayName": { + "type": "string" + }, + "model": { + "type": "string" + }, + "status": { + "type": "string" + } + }, + "required": [ + "model", + "correct", + "status", + "displayName" + ], + "type": "object" + }, + "type": "array" + }, + "size": { + "type": "string" + }, + "solved": { + "type": "number" + }, + "updatedAt": { + "type": "string" + } + }, + "required": [ + "updatedAt", + "puzzleId", + "index", + "size", + "attempts", + "solved", + "runs" + ], + "type": "object" +}
- Changed
list_families1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": false, + "properties": { + "families": { + "items": { + "additionalProperties": {}, + "properties": { + "bestVariant": { + "type": "string" + }, + "displayName": { + "type": "string" + }, + "efforts": { + "items": { + "type": "string" + }, + "type": "array" + }, + "family": { + "type": "string" + }, + "provider": { + "type": [ + "string", + "null" + ] + } + }, + "required": [ + "family", + "displayName", + "provider", + "bestVariant", + "efforts" + ], + "type": "object" + }, + "type": "array" + } + }, + "required": [ + "families" + ], + "type": "object" +}
- Changed
list_providers1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": false, + "properties": { + "providers": { + "items": { + "additionalProperties": {}, + "properties": { + "families": { + "items": { + "type": "string" + }, + "type": "array" + }, + "id": { + "type": "string" + }, + "name": { + "type": "string" + }, + "variantCount": { + "type": "number" + } + }, + "required": [ + "id", + "name", + "variantCount", + "families" + ], + "type": "object" + }, + "type": "array" + } + }, + "required": [ + "providers" + ], + "type": "object" +}
- Changed
list_puzzles1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": false, + "properties": { + "puzzles": { + "items": { + "additionalProperties": {}, + "properties": { + "columnClues": { + "items": { + "items": { + "type": "number" + }, + "type": "array" + }, + "type": "array" + }, + "height": { + "type": "number" + }, + "id": { + "type": "string" + }, + "index": { + "type": "number" + }, + "rowClues": { + "items": { + "items": { + "type": "number" + }, + "type": "array" + }, + "type": "array" + }, + "size": { + "type": "string" + }, + "url": { + "type": "string" + }, + "width": { + "type": "number" + } + }, + "required": [ + "id", + "index", + "size", + "width", + "height", + "rowClues", + "columnClues", + "url" + ], + "type": "object" + }, + "type": "array" + } + }, + "required": [ + "puzzles" + ], + "type": "object" +}
- Changed
list_runs4 fields changed- added
Input schema / properties / model / descriptionAdded value: +"Only runs of this model variant id" - added
Input schema / properties / offset / descriptionAdded value: +"Runs to skip, for paging; default 0" - added
Input schema / properties / puzzle_id / descriptionAdded value: +"Only runs on this puzzle id from list_puzzles" - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": false, + "properties": { + "limit": { + "type": "number" + }, + "offset": { + "type": "number" + }, + "runs": { + "items": { + "additionalProperties": {}, + "properties": { + "correct": { + "type": "boolean" + }, + "model": { + "type": "string" + }, + "puzzleId": { + "type": "string" + }, + "size": { + "type": "string" + }, + "status": { + "type": "string" + } + }, + "required": [ + "model", + "puzzleId", + "size", + "correct", + "status" + ], + "type": "object" + }, + "type": "array" + }, + "total": { + "type": "number" + } + }, + "required": [ + "total", + "limit", + "offset", + "runs" + ], + "type": "object" +}
11 tool updates
- First observed
check_solution - First observed
compare_models - First observed
get_leaderboard - First observed
get_model_puzzles - First observed
get_model_results - First observed
get_puzzle - First observed
get_puzzle_results - First observed
list_families - First observed
list_providers - First observed
list_puzzles - First observed
list_runs
Related MCP Servers
- AlicenseAqualityCmaintenanceEnables brand visibility monitoring across major AI platforms like ChatGPT, Claude, Gemini, and Perplexity. It allows users to track visibility scores, analyze competitor data, and receive actionable insights to improve AI-generated brand recommendations.1622 npm1MIT
- AlicenseCqualityAmaintenanceCompetitor Monitor AI - MCP server providing AI-powered tools and automation by MEOK AI Labs119 npm37 PyPIMIT
- AlicenseNot gradedqualityBmaintenanceEnables tracking competitor websites, changelogs, blog feeds, and pricing pages with meaningful diffs, classification, and Markdown digests via MCP tools for listing, adding, removing competitors, running checks, and retrieving digests or changes.MIT

industrylens-mcpofficial
AlicenseNot gradedqualityBmaintenanceBrowse IndustryLens's published competitive-intelligence reports and head-to-head competitor comparisons from any AI agent — real, source-backed data.MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.