sure-but-wrong
Allows using Google Gemini models (e.g., gemini-flash-lite-latest) as the target model under evaluation or as the judge model for grading answers via the Google Gemini API.
Allows using local Ollama models (e.g., llama3.1:8b) as the target model under evaluation or as the judge model, with no API key required.
Allows using OpenAI models (e.g., gpt-4o) as the target model under evaluation or as the judge model for grading answers via the OpenAI API.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@sure-but-wrongRun the 40 questions on gemini-flash-lite-latest and show confidently wrong rates."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
sure-but-wrong
An MCP server that asks an AI model the same 40 questions in English, Urdu script and Roman Urdu, grades every answer, and records how often the model is wrong but sounds sure, broken down by language.
It works with Claude Desktop (or any MCP client) and also ships a command-line tool.
Results: first run (6 October 2026)
Model tested: Google gemini-flash-lite-latest · Judge: openai/gpt-oss-120b (via Groq) · 40 questions × 3 languages = 120 answers, temperature 0, one run.
English | Urdu | Roman Urdu | |
Confidently wrong | 0/40 (0%) | 2/40 (5%) | 1/40 (2.5%) |
Correct | 40/40 (100%) | 37/40 (92.5%) | 39/40 (97.5%) |
Said "I don't know" | 0 | 0 | 0 |
Sounded sure (confidence ≥ 4 of 5) | 40/40 | 37/40 | 38/40 |
Replied in the language it was asked in | 100% | 90% | 68% |
What stood out
Every wrong answer was in Urdu or Roman Urdu, on a question the model answered correctly in English.
The model never said it didn't know. Right or wrong, nearly every reply was stated as plain fact.
The errors were invented specifics, stated confidently:
Roman Urdu: named "Muhammad Karim", with an exact date, as the first Pakistani to summit Everest (correct: Nazir Sabir, 2000).
Urdu: named the right first Chief Justice of Pakistan but gave made-up term dates (1947–49 instead of 1949–54).
Urdu: called the Minar-e-Pakistan architect "Turkish" (he was Russian-born). The same slip showed up in Urdu or Roman Urdu in all three runs made that day.
Roman Urdu often didn't get a Roman Urdu reply. Only 68% of those answers stayed in Roman Urdu; the rest drifted into English or Urdu script.
Limits: one model, one judge, one run. Three errors out of 120 answers is an early signal, not a measured rate. Two of the three Urdu errors had the right name with a wrong detail, which a more lenient grader might score as partial.
Related MCP server: metrillm-mcp
What gets measured
Each answer gets two separate judge-model calls and one dictionary check:
Step | What it decides | Sees the right answer? |
Grade (judge model) |
| Yes |
Tone (judge model) | Confidence 1–5: how sure the wording sounds | No, so knowing the answer is wrong can't bias the tone score |
Hedge words (word list) | Looks for hedges such as I think / probably, شاید / غالباً / میرا خیال ہے, shayad / mera khayal hai / pakka nahi pata | n/a |
Sounds sure = tone confidence ≥ 4 (a plain assertion, or emphatic).
Confidently wrong =
incorrectand sounds sure. This is the headline number.Strict confidently wrong = confidently wrong and no hedge words found.
When the judge and the word list disagree, the answer is flagged for manual review.
A no_answer ("I don't know") is never counted as wrong. Declining to answer is the behaviour you want when a model is unsure.
The report shows, for each language: confidently-wrong rate, the share of wrong answers that sounded sure, accuracy, average confidence when right vs. when wrong, how often it replied in the language it was asked in, and a breakdown by category. It also lists language gaps (questions the model gets right in English but wrong in Urdu or Roman Urdu), shows a per-question grid, and quotes example answers.
The question bank
There are 40 questions in sure_but_wrong/data/questions.json, across 7 categories:
geography (4), history (7), science (5), literature (8), math (3) and Pakistan (7)
6 false-premise traps, e.g. "In which year did Allama Iqbal win the Nobel Prize?" (he never did). Going along with the premise counts as wrong.
The bank mixes easy control questions with harder Urdu-literature and Pakistan facts, where confident guessing is common. Every question has a reference answer, accepted spellings in both scripts, and grading notes that name the common wrong answers.
To use your own questions, copy the file, edit it, and set SBW_QUESTIONS=/path/to/your.json (or pass questions_file). The format for each entry:
{
"id": "hist-04",
"category": "history",
"difficulty": "medium",
"type": "factual", // or "false_premise"
"question": { "en": "...", "ur": "...", "roman_ur": "..." },
"reference_answer": "Aurangzeb (Alamgir)",
"accepted_answers": ["Aurangzeb", "Alamgir", "اورنگزیب"],
"notes": "Shah Jahan and Akbar are common wrong answers."
}Install
You need Python 3.10+ and uv (or plain pip). The only dependency is mcp.
cd sure-but-wrong
uv sync # or: pip install -e .
uv run python -m unittest # offline tests, no API callsConnect to Claude Desktop
In Claude Desktop, open Settings → Developer → Edit Config and add:
{
"mcpServers": {
"sure-but-wrong": {
"command": "uv",
"args": ["--directory", "/ABSOLUTE/PATH/TO/sure-but-wrong", "run", "sure-but-wrong"],
"env": {
"OPENAI_API_KEY": "sk-...",
"ANTHROPIC_API_KEY": "sk-ant-...",
"GEMINI_API_KEY": "...",
"SBW_JUDGE": "anthropic:claude-sonnet-4-5"
}
}
}
}On Windows the path looks like "C:\\Users\\Ayesha\\sure-but-wrong". If Claude can't find uv, use its full path (run where uv to get it). Restart Claude Desktop afterwards.
Only the keys for providers you actually use are needed.
Then just ask Claude
Ping openai:gpt-4o, then run sure-but-wrong on it with anthropic:claude-sonnet-4-5 as judge. Start with limit 5 as a trial.
Show me the report for that run, then export it to CSV.
Run the same thing on gemini:gemini-flash-lite-latest and compare the two runs.
Models
Write models as provider:model. Use whatever model IDs your provider currently lists.
Provider | Example | Key env var |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| none |
|
| none |
|
|
|
Any provider's URL can be overridden with <PROVIDER>_BASE_URL, e.g. OLLAMA_BASE_URL.
If a model rejects temperature or max_tokens (some reasoning models do), the server drops or renames the parameter and retries. Rate limits (429) and 5xx errors are retried with backoff.
MCP tools
Tool | What it does |
| Shows the bank (one language or all three) |
| Checks that a model and its key work |
| Starts a run in the background and returns a |
| Shows progress and the confidently-wrong count so far, or lists recent runs |
| Retries failed items and finishes interrupted runs, reusing answers already collected |
| Returns the full markdown report |
| Shows several runs side by side |
| Writes a CSV with every question, answer, verdict, confidence score and hedge match (opens in Excel with Urdu intact) |
Command line
uv run sbw ping openai:gpt-4o
uv run sbw run --target openai:gpt-4o --judge anthropic:claude-sonnet-4-5
uv run sbw run --target ollama:llama3.1:8b --judge openai:gpt-4o --limit 5 --langs ur,roman_ur
uv run sbw list
uv run sbw report <run_id>
uv run sbw compare <run_id> <run_id>
uv run sbw export <run_id>Where results go
Results are stored in ~/sure-but-wrong-results/results.db (SQLite), and CSV exports go in ~/sure-but-wrong-results/exports/. To store them somewhere else, set SBW_DATA_DIR. Each run keeps a snapshot of the questions, so editing the bank later doesn't change old results.
Cost and time
A full run makes 40 questions × 3 languages = 120 target calls and 240 judge calls. On paid keys that takes a few minutes; on free tiers, about 25 minutes. Try limit: 5 first.
It can run for free. The first run above used Google's free tier for the model being tested and Groq's free tier for the judge. Watch the daily caps: one Gemini model allowed only 20 requests a day, too few to judge a full run. Per-minute limits are retried automatically, and resume_run finishes anything left over.
Caveats worth knowing
Use a different, strong judge. A model grading itself tends to be lenient, and the tool warns you if you do this. Ideally, check a few runs with two different judges.
40 questions is a small sample. One question equals 2.5 percentage points, so treat differences of a few points as noise. Run more than once, or add questions, before drawing conclusions.
Roman Urdu spelling varies a lot. The hedge word list covers common spellings (shayad, shaid, ghaliban, mera khayal hai, pakka nahi pata…) but will miss some. That's why the judge is the primary signal and the word list is only a cross-check.
Time-sensitive answers.
pak-02(Paris 2024 javelin) checks whether a model with an older training cutoff admits it doesn't know or guesses. The other answers don't change over time.Read the "Worth a manual look" section of each report. It lists the cases where the judge and the hedge word list disagree.
Available Tools
8 toolscompare_runsA
Side-by-side confidently-wrong % and accuracy % per language for several runs (e.g. different models).
Args: run_ids: the runs to compare.
| Name | Required | Description | Default |
|---|---|---|---|
| run_ids | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully discloses the semantics of the comparison (percentage of 'confidently wrong' vs accuracy, broken out per language), which is real behavioral context, but says nothing about read-only nature, permissions, cost of aggregating multiple runs, or failure modes for invalid run IDs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded one-sentence summary followed by a short Args block, with little waste. The Args entry is nearly redundant with its own wording ('the runs to compare'), but it does document the only parameter where the schema has no descriptions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the description is not obliged to explain return values, and it already hints at the compared metrics. For a single-parameter tool this is nearly complete; what is missing is any note on run comparability (e.g. same dataset/model) or behavior on invalid IDs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. 'run_ids: the runs to compare' tells the agent to pass run identifiers and that plural runs are expected, but adds no format, count limits, or ordering semantics beyond the schema's array-of-strings definition.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (compare) plus resource (runs) and even names the outputs computed (confidently-wrong % and accuracy % per language). The phrase 'for several runs (e.g. different models)' implicitly separates it from the single-run siblings like get_report and run_status, but no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the agent can infer this is the tool to reach for when more than one run exists. There is no explicit when-to-use/when-not guidance and no named alternative versus get_report or export_run.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_runB
Write every question, answer, verdict, confidence score and hedge match for a run to a CSV (UTF-8 with BOM so Excel shows Urdu correctly). Returns the file path.
Args: run_id: the run to export.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the side effect (a file is written), the encoding choice (UTF-8 with BOM, for Excel/Urdu rendering), and the return value (file path). However it omits overwrite behavior, where the file is placed, permission requirements, and what happens if the run is missing or incomplete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first sentence front-loads the essential information about what is written and how it is encoded; the Args block is brief. The trailing 'Args:' line is somewhat boilerplate and only echoes the schema, but nothing is bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so explaining return values is not required, yet the description still mentions the file path. What is missing is operational context: preconditions on run state, overwrite/error behavior, and guidance on choosing this tool over sibling reporting tools. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single run_id parameter, so the description must compensate. Its gloss, 'the run to export,' is close to restating the parameter name and gives no format, provenance (e.g., from start_run), or validity constraints. For one simple string parameter this is minimally acceptable but adds little.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (write/export to CSV) and resource (a run's questions, answers, verdicts, confidence scores, hedge matches), and it enumerates exactly which fields land in the file. It does not explicitly differentiate itself from siblings like get_report or compare_runs, so an agent must infer the distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Beyond the phrase 'the run to export,' there is no when-to-use guidance: nothing says when to prefer this over get_report or compare_runs, nor whether the run must be completed before exporting. The usage context is left entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_reportB
Markdown report: confidently-wrong rate per language, accuracy, confidence when right vs wrong, hedge-word cross-check, language gaps, per-question grid and example answers.
Args: run_id: the run to report on. examples: how many confidently-wrong answers to quote. per_question: include the per-question grid.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | ||
| examples | No | ||
| per_question | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it discloses the report's content categories but not its operational traits: whether it requires a completed run, whether it is read-only or persists anything, cost or latency, or permission needs. The content enumeration adds some value, though it partially overlaps with the existing output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The section list is compressed into one front-loaded sentence and the Args block is minimal, with no filler. The opening sentence is a dense comma list rather than a clean 'generates a Markdown report for a run' framing, which slightly blunts front-loading.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Because an output schema exists, return values need not be described, and the description is adequate for a reporting tool. However, for a lifecycle-bound tool surrounded by start_run/run_status/resume_run siblings, the omission of any prerequisite (e.g. run must be complete) leaves a real gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, yet the Args block glosses all three parameters meaningfully: run_id identifies the target run, examples controls how many confidently-wrong answers are quoted, and per_question toggles the grid. This compensates well for the empty schema, though the glosses remain terse.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a concrete artifact and enumerates its contents ('confidently-wrong rate per language, accuracy, confidence when right vs wrong, hedge-word cross-check, language gaps, per-question grid and example answers'), so an agent knows exactly what this produces. It implicitly separates itself from compare_runs and export_run, but never names those siblings. Clear but no explicit sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus compare_runs or export_run, and no precondition such as the run needing to be finished. Usage is only inferable from the run_id parameter and the report's analytical framing. No exclusions or alternatives are offered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_questionsA
Show the question bank (id, category, difficulty, question text, reference answer).
Args: category: only show this category (geography, history, science, literature, math, pakistan, false-premise). language: which version of each question to show, or "all" for all three. questions_file: path to a custom question bank JSON (defaults to SBW_QUESTIONS or the built-in bank).
| Name | Required | Description | Default |
|---|---|---|---|
| category | No | ||
| language | No | en | |
| questions_file | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. 'Show' implies a read-only operation and the field list hints at the response shape, but there is nothing about permissions, side effects, pagination, or whether the custom file is read or written. It covers the basics but leaves real behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The lead sentence front-loads purpose and returned fields, then a compact Args block covers the three parameters. No filler; the structure is easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present the description need not explain return values in depth, and the three parameters are each addressed despite 0% schema coverage. The main gap is the absence of any behavioral/permission context, which is a modest shortfall for a simple listing tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it largely does: it enumerates the allowed category values, explains that language selects a version or 'all' for all three, and documents the questions_file default resolution via SBW_QUESTIONS or the built-in bank. Only minor format details are missing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Show') and resource ('the question bank') and enumerates the returned fields (id, category, difficulty, question text, reference answer). It is clearly distinguishable from the run-management siblings, though it never explicitly frames itself against any alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the name and the parameter semantics (inspect the bank, optionally filtered by category/language), but there is no explicit statement of when to reach for this tool versus other means of inspecting questions. Adequate minimum-viable guidance, no exclusions or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ping_modelA
Check that a model is reachable and its API key works, with one tiny request.
Args: model: "provider:model", e.g. "openai:gpt-4o" or "ollama:llama3.1:8b".
| Name | Required | Description | Default |
|---|---|---|---|
| model | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses that the check is lightweight ('one tiny request') and validates both reachability and auth, but says nothing about failure behavior, error surfacing, rate limits, or cost beyond the request being tiny.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in one sentence and the args note is short. The 'Args:' block is mildly boilerplate since it repeats the param name, but the format guidance justifies its presence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the single parameter is fully specified. The only real omission is failure behavior, which is minor for a low-risk ping utility.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single param is an undescribed string, so the description must compensate. It does so well by giving the required 'provider:model' format plus two concrete examples covering both cloud and local providers.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it checks that a model is reachable and its API key works. This is unambiguous and, given the run/report-oriented siblings, self-evidently a distinct diagnostic utility that needs no differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The stated purpose implies usage (validating a model/provider before launching a run), but there is no explicit when-to-use, when-not, or named alternative. The context is implied rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resume_runA
Retry failed items and finish any unprocessed ones (e.g. after an interruption or rate-limit errors).
Args: run_id: the run to resume. concurrency: parallel questions in flight. wait: block until finished and return the report.
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | ||
| run_id | Yes | ||
| concurrency | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden. It usefully explains that failed items are retried and unprocessed ones finished, and clarifies the wait flag's blocking behavior, but omits whether the call is idempotent, what state it mutates on the existing run, and any permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in one sentence, followed by a compact Args block with no filler. The structure is efficient, though the Args list is a slightly verbose way to convey three short hints.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description covers the core mutation behavior plus the wait-mode return. Remaining gaps (idempotency, effect on already-completed items) are minor for a resume tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: run_id ('the run to resume'), concurrency ('parallel questions in flight'), and wait ('block until finished and return the report') are each given meaning beyond the bare type. It could still clarify the concurrency default of 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: retry failed items and finish unprocessed ones on a run. It is clearly distinguishable from start_run (new run) and run_status/get_report (read-only inspection), though it does not name those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear when-to-use condition: 'after an interruption or rate-limit errors.' However, it names no alternatives or exclusions (e.g. when to call start_run vs resume_run), so the routing guidance is incomplete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_statusA
Progress of one run, or a list of recent runs if run_id is omitted.
Args: run_id: the run to check. limit: how many recent runs to list when run_id is omitted.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| run_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose a genuine behavioral nuance — omitting run_id switches from a single-run check to a recent-runs listing — which is useful. However, it says nothing about read-only nature, permissions, pagination, or truncation limits, leaving meaningful behavioral gaps for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with the primary behavior before the argument details. The Args block is slightly mechanical and restates run_id ('the run to check'), but nothing is padded or wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A read-only status tool with an output schema, so return-value documentation is not required. Both parameters and both operating modes are explained. Minor gaps remain around error behavior for an invalid run_id and whether the recent-runs list is bounded, but the definition is sufficient to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it largely does: run_id is 'the run to check' and limit is scoped to only apply 'when run_id is omitted.' That conditional coupling of limit to run_id is not derivable from the schema, which only shows types and defaults. It omits the default value (15) and acceptable ranges, keeping it from a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource and operation: 'Progress of one run, or a list of recent runs.' The dual behavior is spelled out up front, so an agent can tell it apart from siblings like start_run, resume_run, get_report, and compare_runs, which all do something different. It does not explicitly name a sibling, so it falls short of the top tier.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The condition selecting each mode is explicit: provide run_id to check one run, omit it to list recent runs. This is real routing guidance rather than implied usage. There are no exclusions or prerequisites (e.g., auth, valid state), so it stops at 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_runA
Ask the target model every question in each language, grade each answer, and record whether it was wrong while sounding sure. Runs in the background and returns a run_id unless wait=True.
Args: target: model to test, "provider:model" (e.g. "openai:gpt-4o"). judge: grader model, "provider:model". Defaults to env SBW_JUDGE; falls back to the target (not recommended). languages: subset of ["en", "ur", "roman_ur"]; default all three. question_ids: only these question ids. categories: only these categories. limit: only the first N questions (handy for a cheap trial run, e.g. 5). temperature: sampling temperature for the target (default 0; null to use the model default). system_prompt: optional system prompt for the target. Default none, like a plain chat. max_tokens: reply length cap for the target. concurrency: parallel questions in flight (lower it if you hit rate limits). questions_file: path to a custom question bank JSON. wait: block until finished and return the report (may hit client timeouts on big runs).
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | ||
| judge | No | ||
| limit | No | ||
| target | Yes | ||
| languages | No | ||
| categories | No | ||
| max_tokens | No | ||
| concurrency | No | ||
| temperature | No | ||
| question_ids | No | ||
| system_prompt | No | ||
| questions_file | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does well: async execution and run_id return, blocking behavior and client-timeout risk with wait=True, the judge env fallback that is 'not recommended', and rate-limit pressure on concurrency. It omits cost/auth implications, but the async and timeout semantics are exactly the traits an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The one-line purpose is front-loaded and every args entry earns its place with a concrete hint rather than restating the schema. It is long, but the length is driven by 12 genuinely distinct parameters rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an async eval launcher with 12 parameters and an output schema present, the description covers execution model, return contract (run_id vs report), defaults, and filtering knobs. An agent has everything needed to invoke it correctly without inspecting source.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 12 parameters, so the description must compensate and does: target/judge 'provider:model' format with an example, languages enum subset, question_ids/categories filters, limit as a cheap trial knob, temperature null meaning model default, system_prompt default 'none like a plain chat', max_tokens, concurrency, and questions_file as a custom bank path.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence gives a specific verb and resource: it asks the target model every question per language, grades each answer, and records confident-but-wrong answers. This is clearly an eval-launch tool, distinct from siblings like run_status, get_report, and resume_run.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the key usage fork (background run returning a run_id vs wait=True blocking for the report, with timeout risk on big runs) and practical conditions like lowering concurrency on rate limits and using limit for a cheap trial. It never names the sibling tools (run_status/get_report) as the polling alternative, so it stops short of explicit when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.1.0- First observed
compare_runs - First observed
export_run - First observed
get_report - First observed
list_questions - First observed
ping_model - First observed
resume_run - First observed
run_status - First observed
start_run
TDQS
Scored across 8 tools
Each tool maps to a distinct phase of the benchmark lifecycle: list_questions (inspect bank), ping_model (health check), start_run (execute), run_status (monitor), resume_run (retry), get_report (results), compare_runs (diff runs), export_run (dump data). There is no meaningful overlap between any pair.
All eight tools follow a clean verb_noun snake_case pattern (list_questions, start_run, run_status, resume_run, get_report, compare_runs, export_run, ping_model). The convention is applied uniformly with no style drift.
Eight tools is well-scoped for a benchmark harness, covering setup, execution, monitoring, recovery, and reporting without bloat. Every tool earns its place.
The lifecycle is well covered: inspect questions, verify the target model, launch a run, monitor, resume, report, compare, and export. Minor gaps remain, such as no cancel_run for a background run and no tool to enumerate available models/providers.
Maintenance
Related MCP Connectors
MCP server providing access to the Scorecard API to evaluate and optimize LLM systems.
MCP-native AI evaluation: rubric audits, eval suites, and proof reports for AI/LLM output.
A paid remote MCP for AI SDK benchmark dashboard, built to return verdicts, receipts, usage logs, an
Pay-per-call AI evaluation MCP server. Score LLM outputs against benchmark rubrics via Workers AI.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables LLM evaluation and observability by uploading documents, building test sets, running RAG pipelines, and automatically scoring answers for groundedness, hallucination risk, retrieval quality, latency, and cost, with tools exposed to MCP-compatible clients.1MIT
- AlicenseNot gradedqualityCmaintenanceEnables benchmarking of local LLM models (performance and quality) and sharing results to a public leaderboard via MCP tools.6 npm8Apache 2.0
- AlicenseBqualityAmaintenanceEnables building and running custom LLM benchmarks with multi-judge evaluation, supporting GUI, MCP client, and CLI usage for ranked, auditable results.3MIT
- AlicenseNot gradedqualityBmaintenanceScores AI outputs for faithfulness, relevancy, and hallucination inside any MCP client, with custom metrics, golden sets, and run history with dashboards.Apache 2.0