sure-but-wrong
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| SBW_JUDGE | No | The judge model to use, written as provider:model, e.g. anthropic:claude-sonnet-4-5. | |
| XAI_API_KEY | No | API key for xAI provider. | |
| GROQ_API_KEY | No | API key for Groq provider. | |
| SBW_DATA_DIR | No | Directory where results and exports are stored. | |
| SBW_QUESTIONS | No | Path to a custom questions JSON file. | |
| CUSTOM_API_KEY | No | Optional API key for a custom OpenAI-compatible server. | |
| GEMINI_API_KEY | No | API key for Gemini provider. Can be used instead of GOOGLE_API_KEY. | |
| GOOGLE_API_KEY | No | API key for Gemini provider. Can be used instead of GEMINI_API_KEY. | |
| OPENAI_API_KEY | No | API key for OpenAI provider. | |
| CUSTOM_BASE_URL | No | Base URL for a custom OpenAI-compatible server. | |
| MISTRAL_API_KEY | No | API key for Mistral provider. | |
| OLLAMA_BASE_URL | No | Optional override for the Ollama provider base URL. Other providers can similarly be overridden with <PROVIDER>_BASE_URL. | |
| DEEPSEEK_API_KEY | No | API key for DeepSeek provider. | |
| TOGETHER_API_KEY | No | API key for Together provider. | |
| ANTHROPIC_API_KEY | No | API key for Anthropic provider. | |
| OPENROUTER_API_KEY | No | API key for OpenRouter provider. |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| list_questionsA | Show the question bank (id, category, difficulty, question text, reference answer). Args: category: only show this category (geography, history, science, literature, math, pakistan, false-premise). language: which version of each question to show, or "all" for all three. questions_file: path to a custom question bank JSON (defaults to SBW_QUESTIONS or the built-in bank). |
| ping_modelA | Check that a model is reachable and its API key works, with one tiny request. Args: model: "provider:model", e.g. "openai:gpt-4o" or "ollama:llama3.1:8b". |
| start_runA | Ask the target model every question in each language, grade each answer, and record whether it was wrong while sounding sure. Runs in the background and returns a run_id unless wait=True. Args: target: model to test, "provider:model" (e.g. "openai:gpt-4o"). judge: grader model, "provider:model". Defaults to env SBW_JUDGE; falls back to the target (not recommended). languages: subset of ["en", "ur", "roman_ur"]; default all three. question_ids: only these question ids. categories: only these categories. limit: only the first N questions (handy for a cheap trial run, e.g. 5). temperature: sampling temperature for the target (default 0; null to use the model default). system_prompt: optional system prompt for the target. Default none, like a plain chat. max_tokens: reply length cap for the target. concurrency: parallel questions in flight (lower it if you hit rate limits). questions_file: path to a custom question bank JSON. wait: block until finished and return the report (may hit client timeouts on big runs). |
| run_statusA | Progress of one run, or a list of recent runs if run_id is omitted. Args: run_id: the run to check. limit: how many recent runs to list when run_id is omitted. |
| resume_runA | Retry failed items and finish any unprocessed ones (e.g. after an interruption or rate-limit errors). Args: run_id: the run to resume. concurrency: parallel questions in flight. wait: block until finished and return the report. |
| get_reportB | Markdown report: confidently-wrong rate per language, accuracy, confidence when right vs wrong, hedge-word cross-check, language gaps, per-question grid and example answers. Args: run_id: the run to report on. examples: how many confidently-wrong answers to quote. per_question: include the per-question grid. |
| compare_runsA | Side-by-side confidently-wrong % and accuracy % per language for several runs (e.g. different models). Args: run_ids: the runs to compare. |
| export_runB | Write every question, answer, verdict, confidence score and hedge match for a run to a CSV (UTF-8 with BOM so Excel shows Urdu correctly). Returns the file path. Args: run_id: the run to export. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 8 tools
Each tool maps to a distinct phase of the benchmark lifecycle: list_questions (inspect bank), ping_model (health check), start_run (execute), run_status (monitor), resume_run (retry), get_report (results), compare_runs (diff runs), export_run (dump data). There is no meaningful overlap between any pair.
All eight tools follow a clean verb_noun snake_case pattern (list_questions, start_run, run_status, resume_run, get_report, compare_runs, export_run, ping_model). The convention is applied uniformly with no style drift.
Eight tools is well-scoped for a benchmark harness, covering setup, execution, monitoring, recovery, and reporting without bloat. Every tool earns its place.
The lifecycle is well covered: inspect questions, verify the target model, launch a run, monitor, resume, report, compare, and export. Minor gaps remain, such as no cancel_run for a background run and no tool to enumerate available models/providers.