Skip to main content
Glama

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
SBW_JUDGENoThe judge model to use, written as provider:model, e.g. anthropic:claude-sonnet-4-5.
XAI_API_KEYNoAPI key for xAI provider.
GROQ_API_KEYNoAPI key for Groq provider.
SBW_DATA_DIRNoDirectory where results and exports are stored.
SBW_QUESTIONSNoPath to a custom questions JSON file.
CUSTOM_API_KEYNoOptional API key for a custom OpenAI-compatible server.
GEMINI_API_KEYNoAPI key for Gemini provider. Can be used instead of GOOGLE_API_KEY.
GOOGLE_API_KEYNoAPI key for Gemini provider. Can be used instead of GEMINI_API_KEY.
OPENAI_API_KEYNoAPI key for OpenAI provider.
CUSTOM_BASE_URLNoBase URL for a custom OpenAI-compatible server.
MISTRAL_API_KEYNoAPI key for Mistral provider.
OLLAMA_BASE_URLNoOptional override for the Ollama provider base URL. Other providers can similarly be overridden with <PROVIDER>_BASE_URL.
DEEPSEEK_API_KEYNoAPI key for DeepSeek provider.
TOGETHER_API_KEYNoAPI key for Together provider.
ANTHROPIC_API_KEYNoAPI key for Anthropic provider.
OPENROUTER_API_KEYNoAPI key for OpenRouter provider.

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}
prompts
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}

Tools

Functions exposed to the LLM to take actions

NameDescription
list_questionsA

Show the question bank (id, category, difficulty, question text, reference answer).

Args: category: only show this category (geography, history, science, literature, math, pakistan, false-premise). language: which version of each question to show, or "all" for all three. questions_file: path to a custom question bank JSON (defaults to SBW_QUESTIONS or the built-in bank).

ping_modelA

Check that a model is reachable and its API key works, with one tiny request.

Args: model: "provider:model", e.g. "openai:gpt-4o" or "ollama:llama3.1:8b".

start_runA

Ask the target model every question in each language, grade each answer, and record whether it was wrong while sounding sure. Runs in the background and returns a run_id unless wait=True.

Args: target: model to test, "provider:model" (e.g. "openai:gpt-4o"). judge: grader model, "provider:model". Defaults to env SBW_JUDGE; falls back to the target (not recommended). languages: subset of ["en", "ur", "roman_ur"]; default all three. question_ids: only these question ids. categories: only these categories. limit: only the first N questions (handy for a cheap trial run, e.g. 5). temperature: sampling temperature for the target (default 0; null to use the model default). system_prompt: optional system prompt for the target. Default none, like a plain chat. max_tokens: reply length cap for the target. concurrency: parallel questions in flight (lower it if you hit rate limits). questions_file: path to a custom question bank JSON. wait: block until finished and return the report (may hit client timeouts on big runs).

run_statusA

Progress of one run, or a list of recent runs if run_id is omitted.

Args: run_id: the run to check. limit: how many recent runs to list when run_id is omitted.

resume_runA

Retry failed items and finish any unprocessed ones (e.g. after an interruption or rate-limit errors).

Args: run_id: the run to resume. concurrency: parallel questions in flight. wait: block until finished and return the report.

get_reportB

Markdown report: confidently-wrong rate per language, accuracy, confidence when right vs wrong, hedge-word cross-check, language gaps, per-question grid and example answers.

Args: run_id: the run to report on. examples: how many confidently-wrong answers to quote. per_question: include the per-question grid.

compare_runsA

Side-by-side confidently-wrong % and accuracy % per language for several runs (e.g. different models).

Args: run_ids: the runs to compare.

export_runB

Write every question, answer, verdict, confidence score and hedge match for a run to a CSV (UTF-8 with BOM so Excel shows Urdu correctly). Returns the file path.

Args: run_id: the run to export.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

A3.9/5.0

Scored across 8 tools

Disambiguation5/5

Each tool maps to a distinct phase of the benchmark lifecycle: list_questions (inspect bank), ping_model (health check), start_run (execute), run_status (monitor), resume_run (retry), get_report (results), compare_runs (diff runs), export_run (dump data). There is no meaningful overlap between any pair.

Naming Consistency5/5

All eight tools follow a clean verb_noun snake_case pattern (list_questions, start_run, run_status, resume_run, get_report, compare_runs, export_run, ping_model). The convention is applied uniformly with no style drift.

Tool Count5/5

Eight tools is well-scoped for a benchmark harness, covering setup, execution, monitoring, recovery, and reporting without bloat. Every tool earns its place.

Completeness4/5

The lifecycle is well covered: inspect questions, verify the target model, launch a run, monitor, resume, report, compare, and export. Minor gaps remain, such as no cancel_run for a background run and no tool to enumerate available models/providers.

Maintenance

ActivityMaintained
ResponsivenessNo issues