Skip to main content
Glama

inference

Chat-completion inference (standard messages request shape) against a six-model catalogue (see /api/v1/models). $0.01 charged per call: pay $0.01 exact on Base or Solana, or authorize up to $1.00 with the Base "upto" scheme and still be charged $0.01.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
modelNoModel ID (see /api/v1/models)deepseek/deepseek-v4-flash
streamNoEnable SSE streaming
messagesYesChat messages array

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changed
    • addedInput schema / properties / messages / examples
      Added value: +[
      +  [
      +    {
      +      "content": "Say hello in five words.",
      +      "role": "user"
      +    }
      +  ]
      +]
  2. Changed1 schema field changed
    • removedInput schema / properties / max_tokens
      Removed value: -{
      -  "default": 2000,
      -  "description": "Defensively capped output tokens",
      -  "maximum": 2000,
      -  "type": "integer"
      -}
  3. Changed1 schema field changed
    • addedInput schema / properties / max_tokens
      Added value: +{
      +  "default": 2000,
      +  "description": "Defensively capped output tokens",
      +  "maximum": 2000,
      +  "type": "integer"
      +}
  4. Changed4 schema fields changed
    • removedInput schema / properties / max_tokens
      Removed value: -{
      -  "default": 4096,
      -  "type": "integer"
      -}
    • addedInput schema / properties / messages / items
      Added value: +{
      +  "type": "object"
      +}
    • addedInput schema / properties / model / default
      Added value: +"deepseek/deepseek-v4-flash"
    • addedInput schema / properties / stream
      Added value: +{
      +  "default": false,
      +  "description": "Enable SSE streaming",
      +  "type": "boolean"
      +}
  5. First observed

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose a critical behavioral trait: a $0.01 per-call charge and the payment mechanics (exact $0.01 on Base or Solana, or up to $1.00 via the Base 'upto' scheme with actual charge of $0.01). It omits auth prerequisites, rate limits, and error behavior, but the cost/payment disclosure is substantial and non-obvious.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the purpose before the pricing detail, which is the right ordering. The pricing clause is somewhat dense but every element (amount, chains, scheme) is load-bearing for an agent deciding whether it can pay.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema and no annotations, so the description should ideally cover the return shape and auth; it implies an OpenAI-standard chat shape but never states the response format explicitly. Payment mechanics are well covered, but the response/error surface and the distinction from chat_completion are missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters are already documented in the schema. The description adds only that the request uses the 'standard messages request shape' and references the model catalogue, which is marginal value beyond the schema — baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

It names a specific verb+resource (chat-completion inference) and scopes it to a six-model catalogue pointed to via /api/v1/models. However, it does not distinguish this tool from the highly similar sibling 'chat_completion', leaving an agent unable to tell them apart from the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance, no exclusion criteria, and no routing hint against alternatives. Given siblings like chat_completion and route_task that overlap heavily, the absence of any 'use this instead of X' statement is a real gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.