Skip to main content
Glama

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault

No arguments

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}
prompts
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}
experimental
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
list_modelsB

List every provider and model in the catalog, plus each provider's frontier (flagship) model.

Each model carries a `reasoning` flag (supports extended reasoning) and a `tags` list of
curated need/industry tag ids; `tags` in the result gives their labels, one-line "why"
explanations and the date they were verified. Tags are a starting point, not benchmarks.
suggest_modelsC

Suggest sibling models from the same family as the given model id.

list_availabilityA

Return the current model-availability snapshot: live OpenRouter listing status (refreshed at most every 6 hours) plus curated Bedrock/Vertex/Foundry region coverage.

set_policyB

Set the company policy text used to gate prompts before any model is called.

evaluate_promptA

Get pre-run feedback on a prompt's clarity/specificity before running a comparison.

An explicit, separately-triggered LLM call (uses your credentials) — not run
automatically as part of run_comparison. Rate-limited independently from
run_comparison's 3-per-8h budget. Pass `creds` as {"openrouter"?: str,
"bedrock"?: {...}, "vertex"?: {...}, "foundry"?: {...}} to use Amazon Bedrock,
Google Vertex AI, or Microsoft Foundry; a bare `api_key` is treated as an
OpenRouter key. `judge_backend` picks which backend runs the evaluation.
If the operator has set server-side keys for a backend (see https://github.com/thejaredchapman/evalforge-lite/blob/main/docs/hosting-and-server-keys.md), those are used for it automatically and creds for it are not needed.
run_comparisonA

Run a set of test-case prompts against a set of models, scoring each response.

Each test case may include an optional "rubric" (scored by an LLM judge) and/or
"checks" (rule-based checks). Returns per-model grades, cost/latency stats, and an
overall verdict. Rate-limited to 3 calls per 8 hours.
Models are "<catalog id>" (OpenRouter) or "<catalog id>@bedrock" / "<catalog id>@vertex" /
"<catalog id>@foundry"; pass matching creds ({"openrouter"?, "bedrock"?, "vertex"?, "foundry"?})
or a bare OpenRouter api_key.
judge_backend picks which backend runs the judge and policy gate.
At most 4 models. priority (balanced|quality|fastest|cheapest) ranks the results;
repeats (1-3) re-sends each prompt for timing accuracy. Returns suggestions (same
provider and backend only), advice, ranking, and best_for_priority, plus a per-run
`cost` total and, per cell, an `evaluation` (answered/quality/instruction_following/
completeness/helpfulness/safety scores, strengths, weaknesses, reasoning, overall).
If the operator has set server-side keys for a backend (see https://github.com/thejaredchapman/evalforge-lite/blob/main/docs/hosting-and-server-keys.md), those are used for it automatically and creds for it are not needed.
list_runsA

List metadata for the 5 most recent runs, newest first.

get_reportA

Get a PDF report (base64-encoded) for a run. Defaults to the most recent run.

priority (balanced|quality|fastest|cheapest), if given, overrides the priority the
run was made with for the report's priority/best-pick line; otherwise the run's own
priority (from run_comparison) is used.
get_report_csvB

Get a CSV export for a run, one row per (test case x model) cell. Defaults to the most recent run.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

A3.7/5.0

Scored across 9 tools

Disambiguation4/5

Most tools have clearly distinct purposes, but the three model-introspection tools (list_models, suggest_models, list_availability) overlap somewhat, and evaluate_prompt vs run_comparison could be momentarily confused since both invoke LLMs. Descriptions do help differentiate them, especially the explicit pre-run vs full-run distinction.

Naming Consistency5/5

All tools use snake_case with a consistent verb_noun pattern (list_models, suggest_models, set_policy, evaluate_prompt, run_comparison, list_runs, get_report). Even the format variant get_report_csv follows the same convention cleanly.

Tool Count5/5

Nine tools is well-scoped for a model-evaluation server, with each tool earning its place across catalog introspection, policy, evaluation, execution, and reporting. No padding or redundancy in count.

Completeness4/5

Coverage spans discovery, policy, prompt feedback, comparison runs, and report exports (PDF/CSV), which is solid for the domain. Minor gaps: no structured JSON run-detail retrieval or run/policy deletion, and set_policy has no corresponding get_policy to inspect current state.

Maintenance

ActivityMaintained
ResponsivenessNo issues