decide
This server performs bulk automated decisions (e.g., classification, feature detection) on large sets of records using the Jev model, with configurable criteria, confidence thresholds, and review workflows, while keeping results and full audit trails on disk.
Process bulk input files (log lines, JSONL records, or glob-matched UTF-8 files) without loading entire contents into the agent's context.
Define your own question and criteria (label descriptions with custom cutoffs), plus optional context and per-label confidence thresholds.
Route decisions automatically (using Jev) and optionally flag chosen labels or low-confidence results for manual review via
review_labelsandreview_limit.Control concurrency and review size (1–16 parallel workers, 0–100 review items).
Receive summary counts and paths to full JSONL result files (
results.jsonl,review.jsonl,request.json,summary.json) for auditing and follow-up.Handle provider failures gracefully – corrupt responses do not cancel other records; oversized items are sent intact to review.
No file mutation or shell execution – retains content on disk by default, optionally includes bounded previews for review.
Send selected content to TypeSafe (api.typesafe.ai) for decision-making; scope input paths accordingly.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@decideFind feature requests in these 500 app reviews"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
decide
Your agent reasons. Jev sorts.
One MCP tool for bulk decisions over files, logs and JSONL.
Get started · Benchmarks · Tool reference · MIT
Give your agent a path and a rubric. decide sends each record to
Jev, writes decisions locally, and returns
counts and paths. Your agent reviews the exceptions and processes the rest with code.
Files + rubric → Jev → accepted decisions + review queue → your agentKeep bulk input out of the expensive agent's context. Tune review to the quality your task needs; there is no universal “review 5%” setting.
Get started
You need uv and a TypeSafe key. No clone or Git required: this command installs the small release wheel, without the benchmark archive.
uvx --python 3.11 --from https://github.com/alsoleg89/decide/releases/download/v0.1.1/decide_mcp-0.1.1-py3-none-any.whl decide-mcpUse it as your MCP client's launch command. Set TYPESAFE_API_KEY and
DECIDE_ROOT to an absolute data-directory path in the server environment.
Missing or relative roots are rejected before any files are read.
Codex setup · Claude setup · Verify the wheel
Then ask your agent:
Use decide on feedback.jsonl to find feature requests. Pass the path without reading the full file first. Use yes/no criteria, a global cutoff of 0 and a cutoff of 0.9 for no. Review every queued case, save one final label per ID, and verify complete coverage.
{
"question": "Does this review request a new or changed capability?",
"criteria": {"yes": "Requests a new or changed capability", "no": "Does not"},
"source": {"kind": "jsonl", "paths": ["feedback.jsonl"]},
"confidence_threshold": 0,
"confidence_thresholds": {"no": 0.9}
}Each JSONL row is {"id":"review-1","content":"Please add CSV export."}.
The example uses the UX study's routing. Validate the rubric and cutoffs on your
own labeled sample; a confidence score is not a probability of correctness.
Related MCP server: Jev MCP Server
What we measured
Real UX feedback and developer issue triage, with complete output files and all model turns counted. Cost includes Jev, agent reading, tools, retained history, writing and the final answer. These are guided API loops with 25-record batches, not native Codex/Claude sessions or optimized baselines; larger standalone-model batches have not been compared.
Task and reviewer | Total cost reduction | Accuracy: reviewer alone → decide + reviewer | Quality tradeoff |
48.9% | 84.2% → 90.2% | Accuracy, macro-F1, feature precision and recall all improved | |
77.3% | 68.0% → 74.7% | Observed feature recall −5 p.p.; 95% interval −11.01 to +0.89 | |
32.8% | 87.8% → 90.1% | Feature recall fell by one correct request | |
70.4% | 74.7% → 74.7% | Bug and feature recall fell |
Luna used reasoning none. Both Luna policies failed the frozen all-metrics
quality gate. The mini UX accuracy gain has a paired 95% interval of +4 to +8
percentage points; statistics and limitations
are published alongside every prediction, mistake and cost calculation.
Review buys a different error tradeoff. In the mini UX study it raised Jev-only feature recall from 67.5% to 84.7%, while lowering accuracy from 91.4% to 90.2%. More review is not automatically better. Review effect · Losing tasks and baselines
The broader workflow study includes four UX labels and issues from five projects. Results vary by task; sorting labels is not UX synthesis, prioritization or incident diagnosis.
Full audit trail
results.jsonl: every decision, confidence, probabilities, usage and error.review.jsonl: complete original records needing review, including provider failures.request.jsonandsummary.json: rubric, routing settings and run totals.
Files live in DECIDE_ROOT/.decide/<run-id>/. The default tool response has no
input previews. Oversized items go intact to review; corrupt provider responses
do not cancel the other records. decide.py is the entire server.
Selected content is sent to TypeSafe. Scope your input paths accordingly. Limits, timeouts and retry costs
Try your own workload
Label a sample, run the evaluator, and compare errors, recall and complete cost. Share a small, reproducible issue when a task works well or fails; both belong in the benchmark suite.
Available Tools
1 tooldecideA
Bulk decisions via Jev. Supply either inline {id,content} items or source.
source: {kind: lines|jsonl|files, paths: [...]}, relative to DECIDE_ROOT. lines = one log line per item; jsonl = {id,content} records; files = one UTF-8 file per item, supports globs such as src/**/*.py. No shell commands. criteria maps labels to descriptions. confidence_threshold uses Jev's confidence, NOT selected-label probability. confidence_thresholds overrides the cutoff for specified labels; others use confidence_threshold. review_labels always escalate chosen labels (e.g. other). Returns counts and paths to full JSONL results. Content stays on disk by default; opt into bounded previews with review_limit. Read review_path from the beginning: previews are not completed reviews. Does not execute decisions. All selected content is sent to api.typesafe.ai.
| Name | Required | Description | Default |
|---|---|---|---|
| items | No | ||
| source | No | ||
| context | No | ||
| criteria | Yes | ||
| question | Yes | ||
| concurrency | No | ||
| review_limit | No | ||
| review_labels | No | ||
| confidence_threshold | No | ||
| confidence_thresholds | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the annotations by disclosing important behaviors: content is sent to api.typesafe.ai, decisions are not executed, content stays on disk by default, and previews are bounded. It also clarifies threshold semantics ('Jev's confidence, NOT selected-label probability') and that review_labels always escalate chosen labels. No contradiction with the annotations exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and mostly front-loaded, with each line adding a distinct behavior or constraint. The line-break organization aids readability. Some statements are cryptic, such as 'Read review_path from the beginning' and the unexplained reference to Jev, but overall the text is compact and purposeful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 10 parameters, nested objects, and an output schema, the description covers the most error-prone parts: source formats, globs, threshold semantics, side effects, and preview caveats. It does not provide an overarching example or explain question, context, and concurrency, but those are comparatively self-evident. The key operational warnings are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates substantially: it explains items, source kinds, criteria, confidence_threshold, confidence_thresholds, review_labels, and review_limit. However, it omits question, context, and concurrency entirely, and some details like review_path are mentioned without being defined in the schema. Strong but not complete coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with the vague phrase 'Bulk decisions via Jev,' but the surrounding text clarifies that the tool labels/decides on items according to criteria, applies confidence thresholds, and can read from inline items or source files. It is reasonably clear what the tool does, though an explicit verb like 'classify' or 'evaluate with labels' would sharpen it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly explains the two input modes ('Supply either inline {id,content} items or source'), how source kinds differ, and warns against using shell commands. It also advises reading review_path from the beginning because previews are not completed reviews. There are no siblings to distinguish from, so the lack of alternative-selection guidance is acceptable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.0- First observed
decide
TDQS
Scored across 1 tool
With only one tool, there is no possibility of confusion between tools. The single tool's purpose is clearly defined in its description, so an agent cannot misselect among alternatives.
The tool name 'decide' is a single, simple verb with no conflicting conventions. Since there is only one name, there is no inconsistency to evaluate, and it follows a clear imperative style.
The server has exactly one tool, which feels thin for a general-purpose utility. However, the tool itself is quite comprehensive, covering multiple input sources and options, so it is not trivial. The count is on the borderline of being too sparse.
The single tool covers the full decision-making workflow including input handling, criteria, thresholds, and review options. Minor gaps exist such as no dedicated retrieval tool for prior results, but the tool returns paths to full results, making the surface reasonably complete for its stated purpose.
Maintenance
Related MCP Connectors
Jev-powered decisions, web search, PDF/web to Markdown, summarize. From $0.001, no API key.
Verified AI-agent outcomes: secret scanning, JSON cleanup, dedupe, anomaly and schema checks.
Pre-execution governance for AI agents. Deterministic PASS/FAIL/REVIEW verdicts, replayable proof.
- openhelmOAuthai.openhelm
Autonomous cloud agent tasks: real browser + your tools, structured evidence-backed results.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables coding agents to run source-bound evidence checks and bounded batch judgments for classification, extraction, and decision tasks via TypeSafe Jev.999 npmMIT
- AlicenseAqualityBmaintenanceEnables AI assistants to perform ultra-fast, calibrated decision tasks such as boolean evaluation, category selection, scoring, and batch decisions through TypeSafe AI's Jev model.4674 npm1MIT
- AlicenseAqualityBmaintenanceConnect AI agents to Jev AI (jev-ai.pro) for classification, scoring, yes/no checks, advisory action assessment, batched typed decisions and saved judges. Returns structured answers and probabilities using your Jev AI API key.6MIT
- AlicenseAqualityBmaintenanceEnables AI coding agents to offload yes/no, multiple-choice, and scoring questions to TypeSafe's Jev, returning compact confidence-scored answers to save tokens and improve speed.1MIT