Skip to main content
Glama
teempai

jev-in-codex

by teempai

jev_label

Assign labels to JSONL text records and produce id/label JSONL. Use preset or custom criteria with 2-16 labels; low-confidence results are flagged for review.

Instructions

Label id/text JSONL text records using Jev and create a complete id/label JSONL artifact. Sends records to TypeSafe. Use feedback_theme or a custom question and 2–16 label definitions. One label per record. Never overwrites files. Returns evidence below confidence 0.8 for review. Benchmark gains apply to the documented workloads; validate accuracy on new taxonomies.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
pathYes
policyNoPreset name or {question, criteria: {label: definition}}. Define exclusive labels and an unknown/other label when needed.feedback_theme
output_pathNoNew relative output file; defaults to decisions.jsonl beside the input. Parent must exist.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.3.0

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description reveals multiple behavioral traits beyond the sparse annotations: 'Never overwrites files', 'Sends records to TypeSafe', 'Returns evidence below confidence 0.8 for review', and a caveat about benchmark gains. These add meaningful safety, data-flow, and output-trigger information that the agent could not infer from the schema or annotations alone. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but efficient, front-loading the core purpose and then adding config guidance, safety guarantees, output behavior, and a validation caveat. Each sentence earns its place, though it is slightly long for a simple labeling tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with three parameters, no output schema, and sparse annotations, the description covers purpose, label policy configuration, safety guarantees, and output review conditions. It misses a precise return-format specification, but the phrase 'complete id/label JSONL artifact' and the evidence-below-0.8 note give enough context for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema description coverage at 67%, the description compensates by adding parameter-relevant constraints: '2–16 label definitions', 'Define exclusive labels and an unknown/other label when needed', and 'Never overwrites files' for output behavior. The path parameter remains undocumented, but the description's extra semantic details provide value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Label id/text JSONL text records using Jev and create a complete id/label JSONL artifact.' This clearly distinguishes the tool's purpose and output, making it immediately understandable even without sibling tools for comparison.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides clear context on when to use the tool: 'Use feedback_theme or a custom question and 2–16 label definitions', 'One label per record', and 'Returns evidence below confidence 0.8 for review.' While no alternatives are named because there are no siblings, the guidance on configuration and accuracy validation makes usage expectations clear. It stops short of explicitly stating when not to use the tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools