decision-lite
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| HF_ENDPOINT | No | Hugging Face endpoint, e.g. https://hf-mirror.com, if huggingface.co is unreachable | |
| DECISION_LOG | No | Append-only decision log | ./decisions.jsonl |
| DECISION_MODE | No | `verbalized` | `sampling` | `logprobs` (openai backend) | verbalized |
| DECISION_MODEL | No | Model name (openai backend) | qwen3.5-4b |
| DECISION_PYTHON | No | Python executable for the needle/laya bridge | python |
| DECISION_API_KEY | No | Bearer key (omit for local) | |
| DECISION_BACKEND | No | Backend engine: `needle` | `laya` | `openai` | needle |
| DECISION_SAMPLES | No | Votes per question in sampling mode | 7 |
| DECISION_BASE_URL | No | OpenAI-compatible endpoint (openai backend) | http://127.0.0.1:11434/v1 |
| DECISION_LAYA_MODEL | No | Pin `english` or `multilingual` laya checkpoint instead of script-based routing | auto |
| DECISION_LAYA_WARMUP | No | Preload laya checkpoints at server start (`0`/`off` disables lazy-loading avoidance) | 1 |
| DECISION_TEMPERATURE | No | Sampling temperature | 0.8 |
| DECISION_NEEDLE_GENERATION | No | Needle model generation (3=121M, 2=45M) | 3 |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| evaluateA | Jev-compatible decision call. Send a state and typed questions (noul/choice/score); get back answers with probability distributions and confidence. Use for routing, classification, scoring, escalation gating — not for open-ended generation. |
| record_outcomeB | Attach ground truth to a past evaluate() call for calibration tracking. outcome: boolean applied to all answers, or {questionId: boolean}. |
| decision_statsB | Calibration report: confidence buckets vs observed accuracy from recorded outcomes. Shows whether high-confidence answers are actually right that often. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 3 tools
The three tools map cleanly to distinct phases: making a decision (evaluate), providing feedback (record_outcome), and inspecting calibration (decision_stats). There is no overlap in purpose and each has a clearly bounded role.
evaluate is a bare verb, record_outcome is verb_noun, and decision_stats is a noun phrase, so conventions are mixed. Each name is still readable and self-explanatory, but there is no single predictable pattern.
Three tools form a tight, complete decision-calibration loop with no filler. The count is well matched to the narrow stated purpose.
The evaluate/record_outcome/decision_stats cycle covers the core lifecycle of making, labeling, and auditing decisions. Minor gaps exist (e.g., no way to list or retrieve past evaluate calls), but essential workflows are covered.