Skip to main content
Glama

evaluate

Answer typed noul, choice, or score questions about a given state with probability distributions and confidence for routing, classification, scoring, and escalation gating.

Instructions

Jev-compatible decision call. Send a state and typed questions (noul/choice/score); get back answers with probability distributions and confidence. Use for routing, classification, scoring, escalation gating — not for open-ended generation.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
stateYesText or JSON the questions are about
questionsYesMap questionId to a Jev question. Optional conditions v1 defines typed fact extractors, ordered local rules, and a default outcome. Output probabilities are composed from fact marginals under an independence assumption.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden. It discloses the output shape (probabilities plus confidence) and the independence-based composition is left to the schema, which is useful. But it says nothing about backend/model dependencies, latency, cost, determinism, or failure behavior for a non-trivial inference tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with the core operation front-loaded and the exclusion trailing. Efficient and nothing is wasted, but the unexplained jargon 'noul' reduces immediate readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a deeply nested tool with no output schema and no annotations, the description covers the essential picture: inputs, returned probabilities with confidence, and the intended decision-making use cases. The child schema carries the structural detail, so only operational behavior (latency, cost, backend) is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents state and the questions map, including the conditions versioning, supports, and the 4096 combination cap. The description adds only a restatement of the three question types, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (decision call / evaluate) and its inputs (state plus typed questions) and output (answers with probability distributions and confidence). It also carves out a negative scope (not open-ended generation), but never names or contrasts the siblings record_outcome and decision_stats, so an agent must infer routing from the domain term 'Jev-compatible'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete when-to-use contexts (routing, classification, scoring, escalation gating) plus an explicit exclusion ('not for open-ended generation'). That is strong routing guidance, though it stops short of naming an alternative tool for the excluded case.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools