Skip to main content
Glama

search_models

Shortlist models by budget, context and capability, ranked by measured performance. Use this to answer 'which model should I use for X': it returns benchmark scores alongside price so the trade-off is visible in one call. Sort by a subject (math, coding, science, reading) to find a model good at one thing.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
modeNoWhat the model does. Non-chat modes are priced in other units: check price.unit.
limitNo1-50, default 10
sortByNoRanking basis. Default 'intelligence' (overall). 'buzz' is popularity, not skill.
vendorNoOpenRouter namespace, e.g. anthropic
acceptsNoWhat the model must be able to take in, e.g. image for vision tasks.
outputsNoWhat the model produces. A model that accepts video but writes text is 'text', not 'video': filter on what you need made, not what it can read.
minContextNo
maxInputPriceNoUSD per 1M input tokens
minIntelligenceNoLowest acceptable overall index. Leaders sit near 53; the median is 16.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed2 schema fields changed
    • changedInput schema / properties / mode / description
      Previous value: -"What the model does. Non-chat modes are priced in other units — check price.unit."New value: +"What the model does. Non-chat modes are priced in other units: check price.unit."
    • changedInput schema / properties / outputs / description
      Previous value: -"What the model produces. A model that accepts video but writes text is 'text', not 'video' — filter on what you need made, not what it can read."New value: +"What the model produces. A model that accepts video but writes text is 'text', not 'video': filter on what you need made, not what it can read."
  2. Changed1 schema field changed
    • addedInput schema / properties / mode
      Added value: +{
      +  "description": "What the model does. Non-chat modes are priced in other units — check price.unit.",
      +  "enum": [
      +    "chat",
      +    "embedding",
      +    "rerank",
      +    "audio_transcription",
      +    "audio_speech",
      +    "video_generation"
      +  ],
      +  "type": "string"
      +}
  3. Changed2 schema fields changed
    • addedInput schema / properties / accepts
      Added value: +{
      +  "description": "What the model must be able to take in, e.g. image for vision tasks.",
      +  "enum": [
      +    "text",
      +    "image",
      +    "audio",
      +    "video",
      +    "file"
      +  ],
      +  "type": "string"
      +}
    • addedInput schema / properties / outputs
      Added value: +{
      +  "description": "What the model produces. A model that accepts video but writes text is 'text', not 'video' — filter on what you need made, not what it can read.",
      +  "enum": [
      +    "text",
      +    "image",
      +    "audio",
      +    "video"
      +  ],
      +  "type": "string"
      +}
  4. First observed

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses that results are ranked by measured performance, that benchmark scores and price are returned together, and that sorting by subject changes the ranking emphasis. It does not mention default ordering or pagination, but the core behavior is clearly surfaced.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences each earn their place: the first states the function, the second gives the canonical query and output value, and the third explains subject-based sorting. There is no filler or repetition of the input schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

In the absence of annotations and an output schema, the description covers purpose, ranking behavior, result contents, and a sorting refinement. It does not describe return shape beyond score/price or justify when this tool beats its siblings, but for a stateless search endpoint the essential context is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 89%, so the schema already documents most parameters. The description maps high-level concepts (budget, context, capability, subject sorting) onto parameters but adds no new parameter-level detail; it also does not compensate for minContext lacking a schema description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description opens with a specific verb-resource pair—'Shortlist models by budget, context and capability, ranked by measured performance'—and pins the exact question it answers ('which model should I use for X'). The mention of subject-specific sorting further differentiates it from generic model lookup tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It names the canonical use case: answering model-selection questions where benchmark scores and price should be visible together. It does not explicitly compare against sibling tools like compare_models, estimate_cost, or find_replacement, nor state when not to use it, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.