Skip to main content
Glama

jev_evaluate

Read-only

Ask Jev up to 12 typed questions for routing, ranking, or evidence checks on Codex/Gemini results, with text/JSON context and optional files.

Instructions

Ask Jev 1-12 narrow typed questions for routing, ranking or evidence checks on Codex/Gemini results. Supply context as text/JSON plus optional selected files. Confidence is advisory, not proof or authorization. Repeated identical evaluations are cached in this server process. Uses TypeSafe API credits.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
filesNo
contextYes
questionsYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.0.0

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only and open-world traits. The description adds genuinely useful context beyond them: confidence is advisory rather than proof/authorization, identical evaluations are cached in-process, and the call consumes TypeSafe API credits. These are real operational disclosures, though permissions and any rate limits are not detailed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, front-loaded with the core purpose and safely within size limits. The advisory-confidence sentence is slightly tangential but carries useful caveat value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with a deep nested question schema and no output schema, the description covers inputs reasonably but does not explain the return shape or how the three question types differ. An agent can invoke it but must reverse-engineer the typed-question contract from the schema alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the load. It clarifies context as text/JSON, files as optional, and adds a 1-12 question-count constraint not enforced in the schema. However, it does not explain the three question types (choice/noul/score) or the role of instructions vs criteria, leaving the deeply nested oneOf structure largely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: asking Jev 1-12 narrow typed questions for routing, ranking, or evidence checks on Codex/Gemini results. This is clearly distinct from a status check or task delegation. It does not explicitly name bridge_status or delegate_task to differentiate itself, keeping it just below a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage contexts (routing, ranking, evidence checks) and says to supply context plus optional files, but never states when to prefer this over delegate_task or bridge_status, nor any exclusion conditions. Usage is inferable but not explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools