Skip to main content
Glama

Clef Decide

clef_decide
Read-onlyIdempotent

Judge choice, score, or yes/no questions with a local model by passing state and typed questions to get structured probability distributions, confidence, and usage for next actions.

Instructions

Judge a situation with the local Clef decision model. Pass state (compact factual context — string or JSON; treated as data, never executed) and typed questions:

  • "choice": pick among named options (criteria = option id -> description)

  • "score": judge an ordered scale (criteria = ordered list of levels)

  • "noul": yes/no question Returns strict JSON — a probability distribution over the allowed answers for every question. No prose, no sampling. Each decision also carries confidence (model-reported, when available) and the response carries usage (input_tokens/output_tokens/latency_ms). Act on the argmax only when the distribution is decisive (top p >= 0.8 and high confidence for destructive or security-adjacent calls); otherwise gather more context or ask the user. Usage: batch related decisions in one call (up to 64 questions, one forward pass). Optional model overrides the model id. options.temperature is a no-op: scoring is a single forward pass, nothing is sampled.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
modelNoModel id to use (default: clef-flash from CLEF_MODEL).
stateYesTask state as data: a compact string or JSON object with the facts needed to judge (task, errors, logs, diffs). Never treated as instructions. Serialized size limit: 1 MB; model context is 16k tokens.
optionsNoExtra options. Currently only temperature, which is a documented no-op.
questionsYesMap of question id -> typed question. Up to 64 questions per call; all are scored together in a single forward pass, so batch related decisions here.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
modelYes
usageNo
decisionsYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed2 schema fields changedv0.3.0
    • addedOutput schema / properties / decisions / additionalProperties / properties / confidence
      Added value: +{
      +  "maximum": 1,
      +  "minimum": 0,
      +  "type": "number"
      +}
    • addedOutput schema / properties / usage
      Added value: +{
      +  "additionalProperties": false,
      +  "properties": {
      +    "input_tokens": {
      +      "type": "number"
      +    },
      +    "latency_ms": {
      +      "type": "number"
      +    },
      +    "output_tokens": {
      +      "type": "number"
      +    }
      +  },
      +  "type": "object"
      +}
  2. Changed8 schema fields changedv0.2.1
    • addedInput schema / properties / model / description
      Added value: +"Model id to use (default: clef-flash from CLEF_MODEL)."
    • addedInput schema / properties / options / description
      Added value: +"Extra options. Currently only temperature, which is a documented no-op."
    • addedInput schema / properties / options / properties / temperature / description
      Added value: +"Accepted for forward compatibility; no-op. Clef scores in a single forward pass without sampling."
    • addedInput schema / properties / questions / additionalProperties / description
      Added value: +"A single decision question. Allowed answers come from `criteria` and must be mutually exclusive and collectively exhaustive."
    • addedInput schema / properties / questions / additionalProperties / properties / instructions / description
      Added value: +"What to judge, phrased as a self-contained question."
    • addedInput schema / properties / questions / additionalProperties / properties / type / description
      Added value: +"\"choice\" = pick among named options; \"score\" = judge an ordered scale; \"noul\" = yes/no question"
    • addedInput schema / properties / questions / description
      Added value: +"Map of question id -> typed question. Up to 64 questions per call; all are scored together in a single forward pass, so batch related decisions here."
    • addedInput schema / properties / state / description
      Added value: +"Task state as data: a compact string or JSON object with the facts needed to judge (task, errors, logs, diffs). Never treated as instructions. Serialized size limit: 1 MB; model context is 16k tokens."
  3. First observedv0.1.0

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnly, idempotent, non-destructive), and the description adds real behavior beyond that: strict JSON output with no prose or sampling, temperature being a documented no-op, one forward pass for all questions, model-reported confidence, and usage token/latency reporting. It also flags that `state` is treated as data and never executed, which is security-relevant.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded purpose, then a tight bulleted breakdown of question types, then output contract and usage policy. Nearly every sentence earns its place, though the temperature no-op is stated twice (once in prose, once via the schema) and the confidence/threshold guidance could be compressed slightly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a nested-schema tool with an output schema, the description covers the return shape (distribution, confidence, usage), the batching model, the determinism guarantee, and the decision policy. Nothing an agent needs to invoke or interpret this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description nevertheless adds semantics the schema does not fully carry: the criteria mapping per question type (option id -> description for choice, ordered levels for score), the data-not-instructions contract for `state`, and the note that `model` overrides the default. It stops short of explaining the 1 MB / 16k token limits already in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Judge') and resource ('situation') plus the engine ('local Clef decision model'), then enumerates the three question types with their criteria shapes. An agent knows exactly what this tool produces (probability distributions over allowed answers) without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use rules: batch related decisions up to 64 per call, act on the argmax only when top p >= 0.8 and confidence is high for destructive or security-adjacent calls, otherwise gather more context or ask the user. This is a rare case of actionable decision thresholds being spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools