Skip to main content
Glama
teempai

jev-in-codex

by teempai

Jev in Codex

A Codex plugin for batch labelling with Jev. It turns a JSONL file of text records into a complete labelled file, while Codex reviews uncertain decisions.

One tool: jev_label, with custom question/criteria policies and the feedback_theme preset. The previous capability-selection, context-search and output-triage tools have been removed because their tested workflows did not show a useful improvement. There is no lexical or local-classifier fallback.

Define your own labels for topics, sentiment, document collections or other text decisions. See custom labelling for a complete example, taxonomy design and limits. Each call assigns one label per record.

Why this capability is included

The 0.3 custom-policy benchmark passed separately on action-needed, sentiment, document routing and feedback: 21.5–31.9% less total Codex input and 8.4–23.5% less elapsed time across four paired repetitions per task. Both arms labelled all records correctly. These are 64-record synthetic workloads, not a guarantee for every taxonomy or batch size.

The original 0.2 release benchmark on 64 synthetic feedback messages measured 42.7% less total Codex input and 26.4% less elapsed time, with all labels correct across four paired runs. Tiny batches performed worse in the earlier sweep.

The benchmark report includes the shipped implementation check, prototype results, methods, complete artifacts and negative results. This is exploratory verification on known synthetic development data, not independent confirmation. It does not prove real-world accuracy, general labelling superiority, invoice savings or automatic desktop adoption.

New capabilities, built-in presets and workflow expansions require a passing, reviewable benchmark against plain Codex before being exposed. User-supplied taxonomies use the custom-policy interface; their accuracy and benefit still need validation on the actual data. See the admission policy.

Related MCP server: siftr

Use

Prepare an approved UTF-8 JSONL file inside the configured workspace:

{"id":"feedback-1","text":"Saved reports disappear after reopening. Please preserve them."}
{"id":"feedback-2","text":"Please offer a cheaper plan for occasional users."}

Call:

{"path":"feedback/items.jsonl","policy":"feedback_theme"}

The tool creates feedback/decisions.jsonl, containing every original ID and one label:

{"id":"feedback-1","label":"reliability"}
{"id":"feedback-2","label":"pricing"}

These examples explain the format; a two-record batch is not the benchmarked use case. Labels are reliability, usability, pricing and feature. The response includes the policy, counts, model, request count, output path, and original text for decisions below confidence 0.8. Codex can correct that file after review. Above-threshold confidence does not guarantee correctness.

Use output_path to choose another new relative file. Existing files are never overwritten. Both input and output stay inside the configured workspace; the output parent must exist. Input accepts 1–256 records, unique IDs of up to 100 letters/digits/underscore/hyphen, nonempty text up to 6,000 characters per record, and no extra fields. Files are limited to 1 MiB. These are operational limits, not validated performance ranges.

Install

Ask Codex:

Install https://github.com/teempai/jev-in-codex using docs/INSTALL.md, configure my workspace and TypeSafe key privately, and verify that only jev_label is available. Preserve existing instructions. Do not add global instructions.

See installation and migration. Node.js 22+ is required; ripgrep is no longer required. Build with npm ci --ignore-scripts && npm run check. Set TYPESAFE_API_KEY privately and pass --root or JEV_WORKSPACE_ROOT. The model is pinned to the benchmarked jev-1.13.0; the old JEV_MODEL override is no longer used.

For direct MCP configuration, resolve the Node and built-server paths on your machine:

[mcp_servers.jev]
command = "/absolute/path/to/node"
args = ["/absolute/path/to/jev-in-codex/dist/index.js", "--root", "/absolute/workspace"]
env_vars = ["TYPESAFE_API_KEY"]
tool_timeout_sec = 90

Prefer the plugin installation for its bundled skill. Do not register both copies. The tool writes a file, so normal host approval may apply. Installation does not change global approval policy or persistent global instructions.

Data and failure handling

Every input record and the selected policy are sent over HTTPS to https://api.typesafe.ai/v1/systemone. Jev makes the decisions; local code only validates, batches, serializes and selects uncertain records for review. Up to eight requests run concurrently, with at most eight questions per request and a 28,000-byte body budget. Requests have an eight-second timeout and a twenty-second total inference deadline.

Missing credentials, malformed responses or a failed batch produce an explicit error and no labelled output. No local labels replace Jev results. Existing outputs, traversal, escaping symlinks, common credential paths, invalid UTF-8, binary and oversized input are rejected. Provider error bodies are not exposed. The production server has no content logs, telemetry or persistent cache.

Filename exclusions are best-effort, not secret detection. Use only data authorized for TypeSafe. This local path boundary is not a sandbox against concurrent malicious filesystem changes. Labels are advisory and never authorize external actions. See the testing guide; the older security review covers the retired implementation, not this new writing capability.

Development

Contributing · Benchmark report and reproduction · Capability admission policy

MIT. Inspired by the TypeSafe Jev approach to bounded decisions and fast-jev-compaction; this repository does not include that project's compactor code.

Available Tools

1 tool
jev_labelA

Label id/text JSONL text records using Jev and create a complete id/label JSONL artifact. Sends records to TypeSafe. Use feedback_theme or a custom question and 2–16 label definitions. One label per record. Never overwrites files. Returns evidence below confidence 0.8 for review. Benchmark gains apply to the documented workloads; validate accuracy on new taxonomies.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
policyNoPreset name or {question, criteria: {label: definition}}. Define exclusive labels and an unknown/other label when needed.feedback_theme
output_pathNoNew relative output file; defaults to decisions.jsonl beside the input. Parent must exist.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description reveals multiple behavioral traits beyond the sparse annotations: 'Never overwrites files', 'Sends records to TypeSafe', 'Returns evidence below confidence 0.8 for review', and a caveat about benchmark gains. These add meaningful safety, data-flow, and output-trigger information that the agent could not infer from the schema or annotations alone. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but efficient, front-loading the core purpose and then adding config guidance, safety guarantees, output behavior, and a validation caveat. Each sentence earns its place, though it is slightly long for a simple labeling tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with three parameters, no output schema, and sparse annotations, the description covers purpose, label policy configuration, safety guarantees, and output review conditions. It misses a precise return-format specification, but the phrase 'complete id/label JSONL artifact' and the evidence-below-0.8 note give enough context for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema description coverage at 67%, the description compensates by adding parameter-relevant constraints: '2–16 label definitions', 'Define exclusive labels and an unknown/other label when needed', and 'Never overwrites files' for output behavior. The path parameter remains undocumented, but the description's extra semantic details provide value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Label id/text JSONL text records using Jev and create a complete id/label JSONL artifact.' This clearly distinguishes the tool's purpose and output, making it immediately understandable even without sibling tools for comparison.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides clear context on when to use the tool: 'Use feedback_theme or a custom question and 2–16 label definitions', 'One label per record', and 'Returns evidence below confidence 0.8 for review.' While no alternatives are named because there are no siblings, the guidance on configuration and accuracy validation makes usage expectations clear. It stops short of explicitly stating when not to use the tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.3.0
    • First observedjev_label

TDQS

A4.4/5.0

Scored across 1 tool

Disambiguation5/5

There is only one tool, so there is no possibility of confusing it with another tool. The purpose is singular and well-contained.

Naming Consistency5/5

With a single tool named jev_label, there are no conflicting naming conventions. The name clearly follows the server's prefix and indicates its labeling purpose.

Tool Count3/5

One tool feels thin for a labeling workflow, but it is not an extreme mismatch since the tool covers the core labeling operation end-to-end. Supporting operations like managing label definitions or reviewing outputs would make the set feel more complete.

Completeness4/5

The tool covers the main labeling lifecycle: input records, label definitions, artifact creation, and confidence evidence for low-confidence cases. Minor gaps exist around separate review or correction workflows, but the returned evidence mitigates this.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    Enables frontier coding agents to delegate routine probabilistic judgments to TypeSafe Jev, providing calibrated triage signals for failures, attempts, completion, context ranking, findings, risk, and generic evidence-grounded questions.
    7
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI coding agents to quickly find relevant code, focus reads on important parts of large files, and pick relevant items from long lists, reducing token usage and search time.
    2
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Enables coding agents to rank repository files, functions, and web search results using Jev-based probabilistic scoring, returning locations and confidence scores for the agent to review.
    3
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Enables ranking large workspace search results and classifying English text into caller-defined labels using Jev via OpenRouter, with local secret filtering and spend guards.
    2
    MIT