Skip to main content
Glama

EvalForge Lite

Compare text LLMs side by side. Write a few test prompts, pick up to four models across OpenRouter, Amazon Bedrock, Google Vertex AI and Microsoft Foundry, and EvalForge Lite sends the same prompts to all of them, scores the answers automatically, and shows a leaderboard with letter grades, response time, speed and estimated cost. Use it as a web app or as an MCP server for Claude and other assistants. You bring your own credentials, or the person hosting it keeps them on the server for you.

Documentation site: https://thejaredchapman.github.io/evalforge-lite/

Why use it

  • Test on your own prompts. Choose models from evidence about your tasks, not a generic benchmark.

  • Up to 4 models per run, across 4 backends. X (OpenRouter) and X@bedrock are separate targets, so you can check one model on two platforms in a single run.

  • Automatic grading. A judge model scores each answer against your rubric, and rule checks (contains, regex, json_valid, max_length, available through the API and MCP) add a pass or fail. You get a 0-100 score and a letter grade.

  • A second opinion on every response. Each answer is also evaluated on six criteria (answered, quality, instruction following, completeness, helpfulness, safety) with strengths and weaknesses written out.

  • Speed and cost beside quality. Latency, tokens per second and estimated cost for every model, and a "What matters most?" selector that moves the "Best for ..." badge without a new run.

  • Policy gate. Upload a company policy and prompts that violate it are blocked before any model is called. If the check itself fails, the prompt is blocked.

  • Reports. Download a PDF or a CSV for any of your last five runs.

  • No accounts, no database. Credentials are used for one request and not stored. Nothing is written to disk.

Related MCP server: modelmix

Quick start

1. Run the web app on your computer

Requires Python 3.10 or newer (3.12 recommended; download from https://www.python.org/downloads/) and git.

git clone https://github.com/thejaredchapman/evalforge-lite.git
cd evalforge-lite
python3.12 -m venv venv
source venv/bin/activate
pip install -r requirements.txt
python app.py

Open http://localhost:8000, paste a key for at least one backend (an OpenRouter API key is the quickest: https://openrouter.ai/workspaces/default/keys), add a test case, pick two to four models, and click Run comparison. Runs are limited to 3 per 8 hours per browser session. Full walkthrough: Getting started.

2. Use it from Claude (MCP server)

With uv installed:

uvx evalforge-lite

Add it to Claude Code in one line:

claude mcp add evalforge-lite -- uvx evalforge-lite

Or install the Claude Code plugin, which bundles the same server:

claude plugin marketplace add thejaredchapman/evalforge-lite
claude plugin install evalforge-lite@evalforge

Then ask your assistant to compare models. It gets 9 tools: list_models, suggest_models, list_availability, set_policy, evaluate_prompt, run_comparison, list_runs, get_report, get_report_csv. Details, Claude Desktop config and credential shapes: MCP server.

3. Host it for other people

Deploy with the included render.yaml (gunicorn, one worker) or any host that can run gunicorn --workers 1 --threads 4 --bind 0.0.0.0:$PORT app:app. By default every visitor supplies their own key. Optionally keep provider keys on the server with environment variables and a shared daily cap (50 per 24 hours by default). Keep it at one worker: all state is in memory per process. Full guide: Hosting and server-side keys.

Good to know

  • The app does not read .env by itself. To use values from it, run set -a; source .env; set +a before python app.py.

  • Each run is limited to 4 models, and each browser session gets 3 runs per 8 hours.

  • All state lives in memory and is cleared when the server restarts. See Privacy and limits.

  • Upgrading from an older version and reading the CSV or API fields? See the notes in Troubleshooting and FAQ.

Backends and credentials

Backend

What you provide

OpenRouter

One API key

Amazon Bedrock

A region, plus a Bedrock API key or AWS access keys (optional session token)

Google Vertex AI

A project id and region, plus an access token or service-account JSON

Microsoft Foundry

A resource name and region, plus an API key or Entra ID access token

A separate judge backend setting chooses where the judge and policy gate run. Bedrock, Vertex and Foundry costs are estimates from catalog prices, not your cloud bill. See Backends and credentials.

Documentation

Page

What is in it

Overview

What it is, who it is for, the three ways to use it

Getting started

Install, run, and your first comparison

Web app guide

Every part of the screen, in order

Comparing models

Reading metrics, grades, evaluation, cost and their limits

Backends and credentials

Keys, regions, X@backend targets

MCP server

Install paths, all 9 tools, example prompts

Hosting and server-side keys

Deploying for others, operator-held keys, daily cap

Troubleshooting and FAQ

Common messages, fixes, and notes on CSV/API field changes

Privacy and limits

What data goes where, what is stored, every limit

Test

pytest tests/ -v

Every model and HTTP call is mocked, so the suite needs no API key and makes no network calls.

Contributing

Contributions are welcome: bug reports, model-catalog updates, new checks, docs, and new backends. See CONTRIBUTING.md for setup, tests and the pull request process. When the app shows an error, the popup's Report an issue on GitHub button opens a pre-filled bug report.

License

MIT

Available Tools

9 tools
evaluate_promptA

Get pre-run feedback on a prompt's clarity/specificity before running a comparison.

An explicit, separately-triggered LLM call (uses your credentials) — not run
automatically as part of run_comparison. Rate-limited independently from
run_comparison's 3-per-8h budget. Pass `creds` as {"openrouter"?: str,
"bedrock"?: {...}, "vertex"?: {...}, "foundry"?: {...}} to use Amazon Bedrock,
Google Vertex AI, or Microsoft Foundry; a bare `api_key` is treated as an
OpenRouter key. `judge_backend` picks which backend runs the evaluation.
If the operator has set server-side keys for a backend (see https://github.com/thejaredchapman/evalforge-lite/blob/main/docs/hosting-and-server-keys.md), those are used for it automatically and creds for it are not needed.
ParametersJSON Schema
NameRequiredDescriptionDefault
credsNo
promptYes
api_keyNo
judge_backendNoopenrouter

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and does so well: it discloses that this is an explicit LLM call consuming user credentials, that it is rate-limited independently (3-per-8h context), and that server-side operator keys may be substituted automatically. These are meaningful behavioral traits not derivable elsewhere.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose in the first sentence, then layers behavioral detail. Dense but every sentence earns its place; the credential-shape sentence is heavy but necessary given 0% schema coverage.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-param, no-annotation tool with no output schema, the definition covers purpose, trigger semantics, rate limits, and credential handling. The only mild gap is that it does not describe the shape of the returned feedback, which is minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: it documents the creds dict shape, that a bare api_key is treated as OpenRouter, and that judge_backend selects the evaluation backend. Three of four params gain real meaning beyond the bare schema titles.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+scope: 'Get pre-run feedback on a prompt's clarity/specificity before running a comparison.' It explicitly contrasts itself with run_comparison ('not run automatically as part of run_comparison'), letting an agent distinguish it from the sibling without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when it triggers (separately, before a comparison), that it is not automatic as part of run_comparison, and that it has its own rate-limit budget distinct from run_comparison's 3-per-8h. The alternative and its selecting condition are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_reportA

Get a PDF report (base64-encoded) for a run. Defaults to the most recent run.

priority (balanced|quality|fastest|cheapest), if given, overrides the priority the
run was made with for the report's priority/best-pick line; otherwise the run's own
priority (from run_comparison) is used.
ParametersJSON Schema
NameRequiredDescriptionDefault
run_idNo
priorityNo

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It discloses the return format (base64-encoded PDF) and the latest-run default, but says nothing about permission requirements, behavior on in-progress or missing runs, cost/latency of report generation, or failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and output format in the first sentence, then explains the priority parameter. Slightly awkward line breaking and a redundant parenthetical reference to run_comparison, but no filler sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter, no-output-schema tool, the description covers what is returned, the optional nature of both parameters, and the priority override rule. The remaining gaps (error behavior, run_id format) are modest given the tool's low complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does for the more complex parameter: it supplies the enum values for priority (balanced|quality|fastest|cheapest) and explains the override semantics relative to the run's own priority. run_id receives only the implicit "defaults to most recent run" behavior, leaving its format unspecified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (get a PDF report for a run), gives the output encoding (base64), and names the default scope (most recent run). Together with the sibling get_report_csv, which it implicitly contrasts by format, an agent can identify this tool without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Defaults to the most recent run" implies when the run_id can be omitted, which is useful usage context. However, it never names the alternative (get_report_csv) or states when to prefer PDF over CSV, and gives no preconditions such as requiring a completed run.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_report_csvB

Get a CSV export for a run, one row per (test case x model) cell. Defaults to the most recent run.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idNo

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses output shape (one row per test case x model) and the default-run behavior, but says nothing about whether the result is a file path, URL, or inline text, nor about permissions, size, or failure modes for an unknown run_id.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with no padding; the core action and granularity come first, and the default behavior follows. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple single-param export tool with no output schema, the description covers the essential action, output granularity, and default behavior. It is missing the return mechanics (file vs inline content) and where run_id values come from, which an agent calling it blind would need.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With a single optional parameter and 0% schema description coverage, the description must compensate, and it partially does by explaining that omitting run_id falls back to the most recent run. It does not state the expected run_id format or how to obtain one (e.g., via list_runs), so the semantics are only half-covered.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Get a CSV export for a run') and adds the row granularity ('one row per (test case x model) cell'), which is concrete and useful. It implicitly distinguishes itself from the sibling get_report via the CSV format, but never names or contrasts with that sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Defaults to the most recent run' tells the agent the run_id is optional and what happens when omitted, which is genuine usage context. However, there is no guidance on when to choose this over get_report or how it relates to run_comparison/list_runs, so alternatives are only implied by the format.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_availabilityA

Return the current model-availability snapshot: live OpenRouter listing status (refreshed at most every 6 hours) plus curated Bedrock/Vertex/Foundry region coverage.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does add real value: it distinguishes live (refreshed at most every 6 hours) from curated data, which tells the agent the snapshot can be up to 6 hours stale. It stops short of covering auth requirements or what happens on upstream failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence with the resource and both data sources front-loaded and zero filler. The parenthetical refresh cadence is the one detail that earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-param, no-annotation, no-output-schema read tool, the description covers what categories of data come back and how fresh they are, which is the main thing an agent needs. It could say slightly more about the shape/nesting of the returned snapshot, but the coverage is reasonable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no per-parameter semantics to add; baseline for a parameterless tool is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb+resource ('list_availability') and states precisely what the snapshot contains: live OpenRouter listing status plus curated Bedrock/Vertex/Foundry region coverage. It is clear on its own, but it never differentiates itself from the sibling list_models, which an agent could easily confuse it with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use or when-not-to-use guidance, and no routing to alternatives like list_models or suggest_models. The 6-hour refresh note implicitly signals the freshness trade-off, but the agent must infer that this tool is the one to call for provider/region availability rather than model metadata.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_modelsB

List every provider and model in the catalog, plus each provider's frontier (flagship) model.

Each model carries a `reasoning` flag (supports extended reasoning) and a `tags` list of
curated need/industry tag ids; `tags` in the result gives their labels, one-line "why"
explanations and the date they were verified. Tags are a starting point, not benchmarks.
ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose the returned shape's content: a reasoning flag, tag ids, and tag labels with 'why' explanations and verification dates. It omits any mention of auth, pagination, ordering, or size limits, so the disclosure is partial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose is front-loaded in the first sentence, and the second block explains returned fields with little filler. It is slightly over-explained for a no-argument list tool but not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully documents the shape of results, covering the main completeness gap. What remains missing is routing guidance against sibling tools like suggest_models, which an agent needs to pick correctly in this catalog-oriented toolset.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing to document and the baseline of 4 applies. The description correctly avoids inventing parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('List every provider and model in the catalog') and adds scope detail (frontier flagship model per provider). It does not, however, explicitly contrast itself with the sibling suggest_models, so an agent must infer the difference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit statement of when to use this tool versus suggest_models or list_availability. The line 'Tags are a starting point, not benchmarks' faintly hints that this is a browse-first tool, but no alternative is named and no conditions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_runsA

List metadata for the 5 most recent runs, newest first.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose two real behavioral traits: a hard cap of 5 results and newest-first ordering, which tells the agent this is not a paginated full listing. It does not confirm the read-only nature or say what happens on a fresh workspace with no runs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, zero filler, and the most decision-relevant fact (the 5-result cap) is front-loaded right after the verb.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-argument, no-output-schema listing tool, the description covers what an agent needs: what comes back, how many, and in what order. Minor gaps remain around read-only confirmation and the empty-state behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the rubric the baseline is 4. The description adds the ordering and count semantics, which is all that could be said here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('runs') plus the scope: the 5 most recent, newest first. That is clear and actionable. It stops short of explicitly distinguishing itself from siblings, though none of the listed siblings is another run-listing tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the name and the recency scope, but the description never says when to reach for this versus other tools, nor whether it is meant for quick inspection rather than full enumeration. No alternatives are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_comparisonA

Run a set of test-case prompts against a set of models, scoring each response.

Each test case may include an optional "rubric" (scored by an LLM judge) and/or
"checks" (rule-based checks). Returns per-model grades, cost/latency stats, and an
overall verdict. Rate-limited to 3 calls per 8 hours.
Models are "<catalog id>" (OpenRouter) or "<catalog id>@bedrock" / "<catalog id>@vertex" /
"<catalog id>@foundry"; pass matching creds ({"openrouter"?, "bedrock"?, "vertex"?, "foundry"?})
or a bare OpenRouter api_key.
judge_backend picks which backend runs the judge and policy gate.
At most 4 models. priority (balanced|quality|fastest|cheapest) ranks the results;
repeats (1-3) re-sends each prompt for timing accuracy. Returns suggestions (same
provider and backend only), advice, ranking, and best_for_priority, plus a per-run
`cost` total and, per cell, an `evaluation` (answered/quality/instruction_following/
completeness/helpfulness/safety scores, strengths, weaknesses, reasoning, overall).
If the operator has set server-side keys for a backend (see https://github.com/thejaredchapman/evalforge-lite/blob/main/docs/hosting-and-server-keys.md), those are used for it automatically and creds for it are not needed.
ParametersJSON Schema
NameRequiredDescriptionDefault
credsNo
modelsYes
api_keyNo
repeatsNo
priorityNobalanced
test_casesYes
judge_backendNoopenrouter

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the 3-calls-per-8-hours rate limit, the 4-model cap, creds/api_key handling, the judge_backend role, and fallback to server-side keys. It stops short of stating cost implications or whether the run persists state, so not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded and nearly every sentence carries non-obvious operational detail (limits, cred formats, return contents). The prose is dense and somewhat clipped, which hurts scanability, but there is little outright filler beyond the trailing server-keys paragraph.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter tool with 0% schema coverage and no output schema, the description is remarkably complete: it covers inputs, credential rules, hard limits, and enumerates the returned fields (grades, cost/latency, verdict, suggestions, advice, ranking, best_for_priority, per-cell evaluation). An agent has enough to call it correctly and anticipate results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it largely does: model id formats with @bedrock/@vertex/@foundry suffixes, creds/key shape, repeats range (1-3), priority values (balanced|quality|fastest|cheapest, an undeclared enum), and judge_backend semantics are all defined. test_cases internals (rubric/checks) are only loosely described, leaving a small gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence gives a specific verb and resource: run a set of test-case prompts against a set of models, scoring each response. This is unmistakably a batch-comparison tool. It does not, however, name or differentiate itself from the overlapping sibling evaluate_prompt, so it stays at 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the framing (comparing multiple models with optional rubric/checks), and hard constraints are given (rate limit, max 4 models, credential requirements). But there is no explicit when-to-use vs evaluate_prompt guidance and no stated exclusions, so the agent must infer which tool to pick.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_policyB

Set the company policy text used to gate prompts before any model is called.

ParametersJSON Schema
NameRequiredDescriptionDefault
policy_textYes

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses that the policy gates prompts before model calls, but omits whether this overwrites an existing policy, what permissions are required, and whether the change is reversible.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. It states the action, resource, and effect efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity one-parameter setter, the description gives the core purpose and parameter role. However, with no annotations or output schema, it leaves mutation semantics, permissions, and usage context underexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It identifies policy_text as the company policy text used for prompt gating, which adds conceptual meaning beyond the schema, but provides no format, length, or content constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Set) and resource (company policy text), and adds the runtime effect of gating prompts before model calls. This clearly distinguishes it from the listed siblings, none of which manage policy.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not say when to use this tool, when not to use it, or how it relates to any alternative. Context implies the use case, but there is no explicit guidance or prerequisite information.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

suggest_modelsC

Suggest sibling models from the same family as the given model id.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It implies a read-only lookup but says nothing about whether the model must exist, how many suggestions are returned, ranking/ordering criteria, or any rate/limit behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler; the purpose and the input relationship are conveyed immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description should describe the shape of returned suggestions (list of ids? objects? ordering?), but it does not. For a discovery tool with an undocumented parameter and no annotations, this leaves the agent guessing about results and edge cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single model_id parameter, so the description must compensate. It only implies the parameter is a model identifier and gives no format, namespace, or example (e.g., whether it is a provider-qualified or internal id).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('suggest') and resource ('sibling models from the same family as the given model id'), which is a clear, non-tautological purpose. It does not, however, contrast itself with the sibling list_models, which sounds like it may return related data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus list_models or list_availability, no statement of prerequisites, and no indication of when the tool is inappropriate (e.g., for a model id that has no siblings).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 9 tool updatesv1.0.1
    • First observedevaluate_prompt
    • First observedget_report
    • First observedget_report_csv
    • First observedlist_availability
    • First observedlist_models
    • First observedlist_runs
    • First observedrun_comparison
    • First observedset_policy
    • First observedsuggest_models

TDQS

A3.7/5.0

Scored across 9 tools

Disambiguation4/5

Most tools have clearly distinct purposes, but the three model-introspection tools (list_models, suggest_models, list_availability) overlap somewhat, and evaluate_prompt vs run_comparison could be momentarily confused since both invoke LLMs. Descriptions do help differentiate them, especially the explicit pre-run vs full-run distinction.

Naming Consistency5/5

All tools use snake_case with a consistent verb_noun pattern (list_models, suggest_models, set_policy, evaluate_prompt, run_comparison, list_runs, get_report). Even the format variant get_report_csv follows the same convention cleanly.

Tool Count5/5

Nine tools is well-scoped for a model-evaluation server, with each tool earning its place across catalog introspection, policy, evaluation, execution, and reporting. No padding or redundancy in count.

Completeness4/5

Coverage spans discovery, policy, prompt feedback, comparison runs, and report exports (PDF/CSV), which is solid for the domain. Minor gaps: no structured JSON run-detail retrieval or run/policy deletion, and set_policy has no corresponding get_policy to inspect current state.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    Enables blind, bias-free AI writing evaluations through randomized A/B duels, Bradley-Terry MLE rankings, and preference analytics, while letting Claude delegate heavy drafting to OpenRouter models to save tokens.
    11
    MIT