Skip to main content
Glama

Jev Delegate for Codex and Claude Code

Jev Delegate gives Codex and Claude Code two local MCP tools: search_and_rank ranks large rg result sets, and classify_items assigns caller-defined labels to English text. Local search collects candidates; TypeSafe Jev through OpenRouter judges them. It cannot browse, write code, summarize, or act as a general subagent.

Each installation uses its own OPENROUTER_API_KEY. No key belongs in this repository or an agent prompt. OpenRouter may charge for requests.

Install for Codex

Requires Node.js 20+, npm, rg, and Codex CLI. Get your own OpenRouter key from OpenRouter.

git clone https://github.com/koosbcom/jev-delegate.git
cd jev-delegate
npm ci
npm run check
codex mcp add jev-delegate -- node "$PWD/dist/server.js"
mkdir -p "${CODEX_HOME:-$HOME/.codex}/skills"
cp -R skills/jev-delegate "${CODEX_HOME:-$HOME/.codex}/skills/"

Set the key in the environment that starts Codex. For a Codex CLI session, enter it without putting its value in shell history:

read -s OPENROUTER_API_KEY
export OPENROUTER_API_KEY
codex

For Codex desktop on macOS, set it for the app session, then restart Codex:

read -s OPENROUTER_API_KEY
launchctl setenv OPENROUTER_API_KEY "$OPENROUTER_API_KEY"
unset OPENROUTER_API_KEY

Start a new Codex task after installation. If your MCP client does not expose workspace roots, set JEV_DELEGATE_WORKSPACE_ROOT to the project directory in the Codex process environment. Check registration with codex mcp list.

Related MCP server: openrouter-jev-mcp

Install for Claude Code

Requires Node.js 20+, npm, rg, and Claude Code. Build once:

git clone https://github.com/koosbcom/jev-delegate.git
cd jev-delegate
npm ci
npm run check

Register the MCP server and the skill for every project:

claude mcp add --scope user jev_delegate -- node "$PWD/dist/server.js"
mkdir -p ~/.claude/skills
cp -R skills/jev-delegate ~/.claude/skills/

Or load the repository as a plugin for one session. .claude-plugin/plugin.json bundles the skill and the MCP server:

claude --plugin-dir /path/to/jev-delegate

Export OPENROUTER_API_KEY in the shell that starts claude, as shown above. Claude Code exposes the project directory as the MCP workspace root, so JEV_DELEGATE_WORKSPACE_ROOT is not needed. Check registration with claude mcp list or /mcp. Telemetry goes to the same ~/.codex/state/jev-delegate/usage.jsonl path unless JEV_DELEGATE_TELEMETRY_PATH is set.

Opening this repository itself in Claude Code also offers the project .mcp.json server; it fails to start until npm run build has created dist/.

When to use

  • search_and_rank: more than 20 workspace search hits or roughly 4,000 raw-result tokens. Supply explicit literal or regex patterns and a relevance criterion.

  • classify_items: many English items with 2–20 distinct labels and clear label criteria.

  • Keep small, sensitive, non-English, and open-ended work with the main model.

The server returns confidence, cost metadata, and a bounded raw fallback if Jev fails. Treat decisions as evidence, not proof. The search tool reads only the current workspace, respects ignore rules, excludes common credential paths, and screens snippets for secret-like text. These filters are heuristic; review material before sending it to any external API.

Default spend guards are $0.01 per tool call and $0.25 per day. The server writes local metadata-only telemetry to ~/.codex/state/jev-delegate/usage.jsonl; set JEV_DELEGATE_TELEMETRY=off to disable it. See src/config.ts for optional limits.

Development

npm ci
npm run check

Offline tests and an MCP stdio test run in npm run check. An opt-in live test needs your own key: npm run test:live. The live API path has not yet been verified against a user's OpenRouter key, so API behavior may need adjustment.

License

MIT. See LICENSE.

Available Tools

2 tools
classify_itemsClassify supplied items with JevA
Read-onlyIdempotent

Assign exactly one caller-supplied label to each item using Jev typed decisions. Adds other and insufficient_context labels. Use for English bulk classification, not prose generation.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemsYes
labelsYes
instructionsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
usageYes
statusYes
classificationsYes
fallback_reasonNo

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnlyHint, openWorldHint, idempotentHint, and destructiveHint, so the safety profile is well defined. The description adds that the tool automatically adds 'other' and 'insufficient_context' labels, which is a behavioral trait beyond the annotations. This is useful but minimal; it does not explain output format or edge cases, which is acceptable given the output schema exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. The core action is front-loaded, and the additional label behavior and usage constraints are stated efficiently. Every sentence contributes meaning without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity—nested item objects, a labels mapping, instructions, and an output schema—the description is too sparse. It does not explain how to populate the parameters correctly or what the output represents. While the output schema covers return values, the input semantics are largely unexplained, leaving an agent to infer from parameter names alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage, so the description must compensate. It only hints at 'items' and 'caller-supplied labels' but does not explain that labels is an object mapping label names to descriptions, that items are structured objects with id/text/metadata, or what instructions should contain. This is insufficient for a tool with nested objects and three required parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (assign labels to items), the resource (items), and the constraint of exactly one caller-supplied label per item. It also distinguishes itself from prose generation and from the sibling search_and_rank by being a classification tool. The phrase 'using Jev typed decisions' is specific enough to signal a distinct mechanism.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says 'Use for English bulk classification, not prose generation,' which gives positive and negative usage context. However, it does not mention the sibling tool search_and_rank or provide a when-not-to-use alternative beyond prose generation, so guidance is partial.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_and_rankSearch workspace and rank matches with JevA
Read-onlyIdempotent

Run read-only ripgrep inside the current workspace, then use Jev to return only relevant matches. Use automatically when raw search would exceed 20 candidates or about 4K tokens. English judgments only.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoliteral
top_kNo
patternsYes
exclude_globsNo
include_globsNo
relevance_criterionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
usageYes
statusYes
resultsYes
truncatedYes
evaluated_countYes
fallback_reasonNo
raw_match_countYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the safety profile is covered. The description adds behavioral context beyond annotations: it reveals that the tool runs ripgrep internally, that it uses Jev for relevance ranking, and that it returns only relevant matches. It also discloses the 'English judgments only' limitation. This adds meaningful behavioral transparency beyond the structured annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: the first states the mechanism, the second gives the trigger condition, the third adds a constraint. The most important information (read-only, ripgrep, Jev ranking) is front-loaded. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema, so return values are covered. The description explains the core pipeline (ripgrep + Jev), the trigger condition, and the English-only constraint. It doesn't explain what 'relevant' means in terms of the relevance_criterion parameter, but the parameter name and schema make that reasonably clear. For a 6-parameter tool with an output schema, this is nearly complete. The only gap is a bit more detail on how the relevance_criterion interacts with Jev, but that's a minor omission.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. The description mentions 'patterns' implicitly via 'ripgrep' and 'relevance_criterion' via 'relevance', but it doesn't explain the semantics of mode, top_k, exclude_globs, include_globs, or the relationship between patterns and relevance_criterion. The description adds some context (the tool is a search+rank pipeline) but doesn't fully compensate for the 0% schema coverage. Baseline 3 is appropriate because the description gives a general sense of the parameters but leaves the agent to infer details from the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('run read-only ripgrep'), a resource ('current workspace'), and a distinct post-processing step ('use Jev to return only relevant matches'). It also names the sibling tool 'classify_items' implicitly by contrast? Actually it doesn't name the sibling, but it clearly distinguishes itself from a raw search by describing the ranking step. The verb+resource+scope is specific enough to differentiate from a generic search tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Use automatically when raw search would exceed 20 candidates or about 4K tokens.' This is a clear when-to-use condition. It also implies when not to use it (when raw search is small enough), and the 'English judgments only' note adds a constraint. This is strong usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.1.0
    • First observedclassify_items
    • First observedsearch_and_rank

TDQS

A3.9/5.0

Scored across 2 tools

Disambiguation5/5

The two tools are completely distinct: one focuses on searching and ranking relevant matches from raw search results, while the other assigns labels to items. There is no overlap in purpose, and an agent can easily select the correct tool based on the task.

Naming Consistency4/5

Both tools use an imperative verb phrase, but 'search_and_rank' is a compound verb while 'classify_items' follows a verb_noun pattern. The slight inconsistency is minor and does not hinder readability, but a more uniform pattern (e.g., 'search_and_rank_items') would be ideal.

Tool Count3/5

With only two tools, the server feels thin for a delegation service. However, the narrow scope of search ranking and classification justifies a small count, making it borderline rather than excessive. Each tool serves a clear purpose, so the count is acceptable but leaves little room for additional functionality.

Completeness4/5

The tool surface covers the core functions implied by the server name: search-and-rank and classification. Minor gaps exist (e.g., no summarization or generation tool), but the stated purpose is well-addressed, and an agent can complete common tasks without hitting dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    Enables Codex to rank available tools and skills, rerank code and documentation search results, and triage saved output artifacts using Jev-based relevance scoring while leaving final decisions to Codex.
    1
    6
    MIT
  • A
    license
    B
    quality
    B
    maintenance
    Enables coding agents to obtain probabilistic decisions from Jev AI via OpenRouter for classification, scoring, and validation, with tools like jev_check, jev_classify, jev_score, and jev_evaluate.
    5
    1
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Connect AI agents to Jev AI (jev-ai.pro) for classification, scoring, yes/no checks, advisory action assessment, batched typed decisions and saved judges. Returns structured answers and probabilities using your Jev AI API key.
    6
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables Claude Code or any MCP client to ask TypeSafe's Jev for calibrated, typed judgments (probabilities, choices, scores) instead of prose, with local caching and cost tracking.
    1
    MIT