jev-delegate
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@jev-delegateRank the search results for 'memory leak' by relevance to our issue."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Jev Delegate for Codex and Claude Code
Jev Delegate gives Codex and Claude Code two local MCP tools: search_and_rank ranks large rg result sets, and classify_items assigns caller-defined labels to English text. Local search collects candidates; TypeSafe Jev through OpenRouter judges them. It cannot browse, write code, summarize, or act as a general subagent.
Each installation uses its own OPENROUTER_API_KEY. No key belongs in this repository or an agent prompt. OpenRouter may charge for requests.
Install for Codex
Requires Node.js 20+, npm, rg, and Codex CLI. Get your own OpenRouter key from OpenRouter.
git clone https://github.com/koosbcom/jev-delegate.git
cd jev-delegate
npm ci
npm run check
codex mcp add jev-delegate -- node "$PWD/dist/server.js"
mkdir -p "${CODEX_HOME:-$HOME/.codex}/skills"
cp -R skills/jev-delegate "${CODEX_HOME:-$HOME/.codex}/skills/"Set the key in the environment that starts Codex. For a Codex CLI session, enter it without putting its value in shell history:
read -s OPENROUTER_API_KEY
export OPENROUTER_API_KEY
codexFor Codex desktop on macOS, set it for the app session, then restart Codex:
read -s OPENROUTER_API_KEY
launchctl setenv OPENROUTER_API_KEY "$OPENROUTER_API_KEY"
unset OPENROUTER_API_KEYStart a new Codex task after installation. If your MCP client does not expose workspace roots, set JEV_DELEGATE_WORKSPACE_ROOT to the project directory in the Codex process environment. Check registration with codex mcp list.
Related MCP server: openrouter-jev-mcp
Install for Claude Code
Requires Node.js 20+, npm, rg, and Claude Code. Build once:
git clone https://github.com/koosbcom/jev-delegate.git
cd jev-delegate
npm ci
npm run checkRegister the MCP server and the skill for every project:
claude mcp add --scope user jev_delegate -- node "$PWD/dist/server.js"
mkdir -p ~/.claude/skills
cp -R skills/jev-delegate ~/.claude/skills/Or load the repository as a plugin for one session. .claude-plugin/plugin.json bundles the skill and the MCP server:
claude --plugin-dir /path/to/jev-delegateExport OPENROUTER_API_KEY in the shell that starts claude, as shown above. Claude Code exposes the project directory as the MCP workspace root, so JEV_DELEGATE_WORKSPACE_ROOT is not needed. Check registration with claude mcp list or /mcp. Telemetry goes to the same ~/.codex/state/jev-delegate/usage.jsonl path unless JEV_DELEGATE_TELEMETRY_PATH is set.
Opening this repository itself in Claude Code also offers the project .mcp.json server; it fails to start until npm run build has created dist/.
When to use
search_and_rank: more than 20 workspace search hits or roughly 4,000 raw-result tokens. Supply explicit literal or regex patterns and a relevance criterion.classify_items: many English items with 2–20 distinct labels and clear label criteria.Keep small, sensitive, non-English, and open-ended work with the main model.
The server returns confidence, cost metadata, and a bounded raw fallback if Jev fails. Treat decisions as evidence, not proof. The search tool reads only the current workspace, respects ignore rules, excludes common credential paths, and screens snippets for secret-like text. These filters are heuristic; review material before sending it to any external API.
Default spend guards are $0.01 per tool call and $0.25 per day. The server writes local metadata-only telemetry to ~/.codex/state/jev-delegate/usage.jsonl; set JEV_DELEGATE_TELEMETRY=off to disable it. See src/config.ts for optional limits.
Development
npm ci
npm run checkOffline tests and an MCP stdio test run in npm run check. An opt-in live test needs your own key: npm run test:live. The live API path has not yet been verified against a user's OpenRouter key, so API behavior may need adjustment.
License
MIT. See LICENSE.
Available Tools
2 toolsclassify_itemsClassify supplied items with JevARead-onlyIdempotent
Assign exactly one caller-supplied label to each item using Jev typed decisions. Adds other and insufficient_context labels. Use for English bulk classification, not prose generation.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | ||
| labels | Yes | ||
| instructions | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| usage | Yes | |
| status | Yes | |
| classifications | Yes | |
| fallback_reason | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnlyHint, openWorldHint, idempotentHint, and destructiveHint, so the safety profile is well defined. The description adds that the tool automatically adds 'other' and 'insufficient_context' labels, which is a behavioral trait beyond the annotations. This is useful but minimal; it does not explain output format or edge cases, which is acceptable given the output schema exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The core action is front-loaded, and the additional label behavior and usage constraints are stated efficiently. Every sentence contributes meaning without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity—nested item objects, a labels mapping, instructions, and an output schema—the description is too sparse. It does not explain how to populate the parameters correctly or what the output represents. While the output schema covers return values, the input semantics are largely unexplained, leaving an agent to infer from parameter names alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must compensate. It only hints at 'items' and 'caller-supplied labels' but does not explain that labels is an object mapping label names to descriptions, that items are structured objects with id/text/metadata, or what instructions should contain. This is insufficient for a tool with nested objects and three required parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (assign labels to items), the resource (items), and the constraint of exactly one caller-supplied label per item. It also distinguishes itself from prose generation and from the sibling search_and_rank by being a classification tool. The phrase 'using Jev typed decisions' is specific enough to signal a distinct mechanism.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says 'Use for English bulk classification, not prose generation,' which gives positive and negative usage context. However, it does not mention the sibling tool search_and_rank or provide a when-not-to-use alternative beyond prose generation, so guidance is partial.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_and_rankSearch workspace and rank matches with JevARead-onlyIdempotent
Run read-only ripgrep inside the current workspace, then use Jev to return only relevant matches. Use automatically when raw search would exceed 20 candidates or about 4K tokens. English judgments only.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | literal | |
| top_k | No | ||
| patterns | Yes | ||
| exclude_globs | No | ||
| include_globs | No | ||
| relevance_criterion | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| usage | Yes | |
| status | Yes | |
| results | Yes | |
| truncated | Yes | |
| evaluated_count | Yes | |
| fallback_reason | No | |
| raw_match_count | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the safety profile is covered. The description adds behavioral context beyond annotations: it reveals that the tool runs ripgrep internally, that it uses Jev for relevance ranking, and that it returns only relevant matches. It also discloses the 'English judgments only' limitation. This adds meaningful behavioral transparency beyond the structured annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: the first states the mechanism, the second gives the trigger condition, the third adds a constraint. The most important information (read-only, ripgrep, Jev ranking) is front-loaded. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema, so return values are covered. The description explains the core pipeline (ripgrep + Jev), the trigger condition, and the English-only constraint. It doesn't explain what 'relevant' means in terms of the relevance_criterion parameter, but the parameter name and schema make that reasonably clear. For a 6-parameter tool with an output schema, this is nearly complete. The only gap is a bit more detail on how the relevance_criterion interacts with Jev, but that's a minor omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. The description mentions 'patterns' implicitly via 'ripgrep' and 'relevance_criterion' via 'relevance', but it doesn't explain the semantics of mode, top_k, exclude_globs, include_globs, or the relationship between patterns and relevance_criterion. The description adds some context (the tool is a search+rank pipeline) but doesn't fully compensate for the 0% schema coverage. Baseline 3 is appropriate because the description gives a general sense of the parameters but leaves the agent to infer details from the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('run read-only ripgrep'), a resource ('current workspace'), and a distinct post-processing step ('use Jev to return only relevant matches'). It also names the sibling tool 'classify_items' implicitly by contrast? Actually it doesn't name the sibling, but it clearly distinguishes itself from a raw search by describing the ranking step. The verb+resource+scope is specific enough to differentiate from a generic search tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Use automatically when raw search would exceed 20 candidates or about 4K tokens.' This is a clear when-to-use condition. It also implies when not to use it (when raw search is small enough), and the 'English judgments only' note adds a constraint. This is strong usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.1.0- First observed
classify_items - First observed
search_and_rank
TDQS
Scored across 2 tools
The two tools are completely distinct: one focuses on searching and ranking relevant matches from raw search results, while the other assigns labels to items. There is no overlap in purpose, and an agent can easily select the correct tool based on the task.
Both tools use an imperative verb phrase, but 'search_and_rank' is a compound verb while 'classify_items' follows a verb_noun pattern. The slight inconsistency is minor and does not hinder readability, but a more uniform pattern (e.g., 'search_and_rank_items') would be ideal.
With only two tools, the server feels thin for a delegation service. However, the narrow scope of search ranking and classification justifies a small count, making it borderline rather than excessive. Each tool serves a clear purpose, so the count is acceptable but leaves little room for additional functionality.
The tool surface covers the core functions implied by the server name: search-and-rank and classification. Minor gaps exist (e.g., no summarization or generation tool), but the stated purpose is well-addressed, and an agent can complete common tasks without hitting dead ends.
Maintenance
Related MCP Connectors
Jev-powered decisions, web search, PDF/web to Markdown, summarize. From $0.001, no API key.
Token guard and rate limiter preventing runaway API cost spikes for OpenAI and Anthropic.
Sentiment, toxicity, entity extraction, PII, translation, summary, QA, fraud scoring, safety audit.
Deterministic trust gate for AI output: leaked-secret, prompt-injection & PII in one call.
Related MCP Servers
- AlicenseAqualityBmaintenanceEnables Codex to rank available tools and skills, rerank code and documentation search results, and triage saved output artifacts using Jev-based relevance scoring while leaving final decisions to Codex.16MIT
- AlicenseBqualityBmaintenanceEnables coding agents to obtain probabilistic decisions from Jev AI via OpenRouter for classification, scoring, and validation, with tools like jev_check, jev_classify, jev_score, and jev_evaluate.51MIT
- AlicenseAqualityBmaintenanceConnect AI agents to Jev AI (jev-ai.pro) for classification, scoring, yes/no checks, advisory action assessment, batched typed decisions and saved judges. Returns structured answers and probabilities using your Jev AI API key.6MIT
- AlicenseNot gradedqualityAmaintenanceEnables Claude Code or any MCP client to ask TypeSafe's Jev for calibrated, typed judgments (probabilities, choices, scores) instead of prose, with local caching and cost tracking.1MIT