EvalForge Lite
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| list_modelsB | List every provider and model in the catalog, plus each provider's frontier (flagship) model. |
| suggest_modelsC | Suggest sibling models from the same family as the given model id. |
| list_availabilityA | Return the current model-availability snapshot: live OpenRouter listing status (refreshed at most every 6 hours) plus curated Bedrock/Vertex/Foundry region coverage. |
| set_policyB | Set the company policy text used to gate prompts before any model is called. |
| evaluate_promptA | Get pre-run feedback on a prompt's clarity/specificity before running a comparison. |
| run_comparisonA | Run a set of test-case prompts against a set of models, scoring each response. |
| list_runsA | List metadata for the 5 most recent runs, newest first. |
| get_reportA | Get a PDF report (base64-encoded) for a run. Defaults to the most recent run. |
| get_report_csvB | Get a CSV export for a run, one row per (test case x model) cell. Defaults to the most recent run. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 9 tools
Most tools have clearly distinct purposes, but the three model-introspection tools (list_models, suggest_models, list_availability) overlap somewhat, and evaluate_prompt vs run_comparison could be momentarily confused since both invoke LLMs. Descriptions do help differentiate them, especially the explicit pre-run vs full-run distinction.
All tools use snake_case with a consistent verb_noun pattern (list_models, suggest_models, set_policy, evaluate_prompt, run_comparison, list_runs, get_report). Even the format variant get_report_csv follows the same convention cleanly.
Nine tools is well-scoped for a model-evaluation server, with each tool earning its place across catalog introspection, policy, evaluation, execution, and reporting. No padding or redundancy in count.
Coverage spans discovery, policy, prompt feedback, comparison runs, and report exports (PDF/CSV), which is solid for the domain. Minor gaps: no structured JSON run-detail retrieval or run/policy deletion, and set_policy has no corresponding get_policy to inspect current state.