Skip to main content
Glama

Evaluate with Jev

evaluate
Read-onlyIdempotent

Judge content against typed questions and get probabilities instead of prose, for yes/no, choice, or scored outcomes. Use it to avoid parsing model text.

Instructions

Judge some content against one or more typed questions and get back probabilities rather than prose — a yes/no likelihood (noul), a pick from a named set (choice), or a position on ordered levels (score). Use it wherever you would otherwise ask a model for an answer and then parse the reply.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
modelNoModel identifier. Defaults to "jev-latest".
stateYesThe material being judged, and only that. Every question in the call reads this same value and nothing else, so anything the decision depends on has to appear here. Prefer an object with named fields when the input has parts.
questionsYesQuestions keyed by an id of your choosing; answers come back under those same ids. Questions in one call run together over the same state and cannot see each other's answers — split dependent judgments across calls.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnly, openWorld, and idempotent hints. The description adds critical behavioral details: questions in one call cannot see each other's answers, and every question reads only the same state. It also clarifies the output is probabilities, not prose. This goes beyond the annotations and is valuable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core purpose and output types. It includes the key usage guidance without any filler. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with nested objects and no output schema, the description adequately explains the output format and key constraints. It does not mention error handling or the 'model' parameter, but the schema covers the model default. Overall, the information an agent needs to call it correctly is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description does not add new parameter-level semantics beyond what the schema already provides for 'state' and 'questions'. It does explain the overall concept and the split-dependent-judgments rule, but that is more behavioral than parameter-specific. The schema carries the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: judging content against typed questions and returning structured probabilities. It distinguishes three output types (noul, choice, score) and frames the use case as an alternative to parsing raw model replies. This is specific and actionable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage guidance: 'Use it wherever you would otherwise ask a model for an answer and then parse the reply.' Since there are no sibling tools, this is sufficient. It does not mention when not to use it, but the guidance is clear and context-rich.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools