Skip to main content
Glama

decide

Your agent reasons. Jev sorts.

One MCP tool for bulk decisions over files, logs and JSONL.

Get started · Benchmarks · Tool reference · MIT

Give your agent a path and a rubric. decide sends each record to Jev, writes decisions locally, and returns counts and paths. Your agent reviews the exceptions and processes the rest with code.

Files + rubric → Jev → accepted decisions + review queue → your agent

Keep bulk input out of the expensive agent's context. Tune review to the quality your task needs; there is no universal “review 5%” setting.

Get started

You need uv and a TypeSafe key. No clone or Git required: this command installs the small release wheel, without the benchmark archive.

uvx --python 3.11 --from https://github.com/alsoleg89/decide/releases/download/v0.1.1/decide_mcp-0.1.1-py3-none-any.whl decide-mcp

Use it as your MCP client's launch command. Set TYPESAFE_API_KEY and DECIDE_ROOT to an absolute data-directory path in the server environment. Missing or relative roots are rejected before any files are read. Codex setup · Claude setup · Verify the wheel

Then ask your agent:

Use decide on feedback.jsonl to find feature requests. Pass the path without reading the full file first. Use yes/no criteria, a global cutoff of 0 and a cutoff of 0.9 for no. Review every queued case, save one final label per ID, and verify complete coverage.

{
  "question": "Does this review request a new or changed capability?",
  "criteria": {"yes": "Requests a new or changed capability", "no": "Does not"},
  "source": {"kind": "jsonl", "paths": ["feedback.jsonl"]},
  "confidence_threshold": 0,
  "confidence_thresholds": {"no": 0.9}
}

Each JSONL row is {"id":"review-1","content":"Please add CSV export."}. The example uses the UX study's routing. Validate the rubric and cutoffs on your own labeled sample; a confidence score is not a probability of correctness.

Related MCP server: Jev MCP Server

What we measured

Real UX feedback and developer issue triage, with complete output files and all model turns counted. Cost includes Jev, agent reading, tools, retained history, writing and the final answer. These are guided API loops with 25-record batches, not native Codex/Claude sessions or optimized baselines; larger standalone-model batches have not been compared.

Task and reviewer

Total cost reduction

Accuracy: reviewer alone → decide + reviewer

Quality tradeoff

1,000 new UX reviews / GPT-4.1 mini

48.9%

84.2% → 90.2%

Accuracy, macro-F1, feature precision and recall all improved

300 OpenCV issues / GPT-4.1 mini

77.3%

68.0% → 74.7%

Observed feature recall −5 p.p.; 95% interval −11.01 to +0.89

Same UX reviews / Luna

32.8%

87.8% → 90.1%

Feature recall fell by one correct request

Same OpenCV issues / Luna

70.4%

74.7% → 74.7%

Bug and feature recall fell

Luna used reasoning none. Both Luna policies failed the frozen all-metrics quality gate. The mini UX accuracy gain has a paired 95% interval of +4 to +8 percentage points; statistics and limitations are published alongside every prediction, mistake and cost calculation.

Review buys a different error tradeoff. In the mini UX study it raised Jev-only feature recall from 67.5% to 84.7%, while lowering accuracy from 91.4% to 90.2%. More review is not automatically better. Review effect · Losing tasks and baselines

The broader workflow study includes four UX labels and issues from five projects. Results vary by task; sorting labels is not UX synthesis, prioritization or incident diagnosis.

Full audit trail

  • results.jsonl: every decision, confidence, probabilities, usage and error.

  • review.jsonl: complete original records needing review, including provider failures.

  • request.json and summary.json: rubric, routing settings and run totals.

Files live in DECIDE_ROOT/.decide/<run-id>/. The default tool response has no input previews. Oversized items go intact to review; corrupt provider responses do not cancel the other records. decide.py is the entire server.

Selected content is sent to TypeSafe. Scope your input paths accordingly. Limits, timeouts and retry costs

Try your own workload

Label a sample, run the evaluator, and compare errors, recall and complete cost. Share a small, reproducible issue when a task works well or fails; both belong in the benchmark suite.

Download v0.1.1 · Local verification · Development

Available Tools

1 tool
decideA

Bulk decisions via Jev. Supply either inline {id,content} items or source.

source: {kind: lines|jsonl|files, paths: [...]}, relative to DECIDE_ROOT. lines = one log line per item; jsonl = {id,content} records; files = one UTF-8 file per item, supports globs such as src/**/*.py. No shell commands. criteria maps labels to descriptions. confidence_threshold uses Jev's confidence, NOT selected-label probability. confidence_thresholds overrides the cutoff for specified labels; others use confidence_threshold. review_labels always escalate chosen labels (e.g. other). Returns counts and paths to full JSONL results. Content stays on disk by default; opt into bounded previews with review_limit. Read review_path from the beginning: previews are not completed reviews. Does not execute decisions. All selected content is sent to api.typesafe.ai.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemsNo
sourceNo
contextNo
criteriaYes
questionYes
concurrencyNo
review_limitNo
review_labelsNo
confidence_thresholdNo
confidence_thresholdsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the annotations by disclosing important behaviors: content is sent to api.typesafe.ai, decisions are not executed, content stays on disk by default, and previews are bounded. It also clarifies threshold semantics ('Jev's confidence, NOT selected-label probability') and that review_labels always escalate chosen labels. No contradiction with the annotations exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense and mostly front-loaded, with each line adding a distinct behavior or constraint. The line-break organization aids readability. Some statements are cryptic, such as 'Read review_path from the beginning' and the unexplained reference to Jev, but overall the text is compact and purposeful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 10 parameters, nested objects, and an output schema, the description covers the most error-prone parts: source formats, globs, threshold semantics, side effects, and preview caveats. It does not provide an overarching example or explain question, context, and concurrency, but those are comparatively self-evident. The key operational warnings are present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description compensates substantially: it explains items, source kinds, criteria, confidence_threshold, confidence_thresholds, review_labels, and review_limit. However, it omits question, context, and concurrency entirely, and some details like review_path are mentioned without being defined in the schema. Strong but not complete coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with the vague phrase 'Bulk decisions via Jev,' but the surrounding text clarifies that the tool labels/decides on items according to criteria, applies confidence thresholds, and can read from inline items or source files. It is reasonably clear what the tool does, though an explicit verb like 'classify' or 'evaluate with labels' would sharpen it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly explains the two input modes ('Supply either inline {id,content} items or source'), how source kinds differ, and warns against using shell commands. It also advises reading review_path from the beginning because previews are not completed reviews. There are no siblings to distinguish from, so the lack of alternative-selection guidance is acceptable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.1.0
    • First observeddecide

TDQS

A4.2/5.0

Scored across 1 tool

Disambiguation5/5

With only one tool, there is no possibility of confusion between tools. The single tool's purpose is clearly defined in its description, so an agent cannot misselect among alternatives.

Naming Consistency5/5

The tool name 'decide' is a single, simple verb with no conflicting conventions. Since there is only one name, there is no inconsistency to evaluate, and it follows a clear imperative style.

Tool Count3/5

The server has exactly one tool, which feels thin for a general-purpose utility. However, the tool itself is quite comprehensive, covering multiple input sources and options, so it is not trivial. The count is on the borderline of being too sparse.

Completeness4/5

The single tool covers the full decision-making workflow including input handling, criteria, thresholds, and review options. Minor gaps exist such as no dedicated retrieval tool for prior results, but the tool returns paths to full results, making the surface reasonably complete for its stated purpose.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    Connect AI agents to Jev AI (jev-ai.pro) for classification, scoring, yes/no checks, advisory action assessment, batched typed decisions and saved judges. Returns structured answers and probabilities using your Jev AI API key.
    6
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Enables AI coding agents to offload yes/no, multiple-choice, and scoring questions to TypeSafe's Jev, returning compact confidence-scored answers to save tokens and improve speed.
    1
    MIT