Skip to main content
Glama

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
HF_ENDPOINTNoHugging Face endpoint, e.g. https://hf-mirror.com, if huggingface.co is unreachable
DECISION_LOGNoAppend-only decision log./decisions.jsonl
DECISION_MODENo`verbalized` | `sampling` | `logprobs` (openai backend)verbalized
DECISION_MODELNoModel name (openai backend)qwen3.5-4b
DECISION_PYTHONNoPython executable for the needle/laya bridgepython
DECISION_API_KEYNoBearer key (omit for local)
DECISION_BACKENDNoBackend engine: `needle` | `laya` | `openai`needle
DECISION_SAMPLESNoVotes per question in sampling mode7
DECISION_BASE_URLNoOpenAI-compatible endpoint (openai backend)http://127.0.0.1:11434/v1
DECISION_LAYA_MODELNoPin `english` or `multilingual` laya checkpoint instead of script-based routingauto
DECISION_LAYA_WARMUPNoPreload laya checkpoints at server start (`0`/`off` disables lazy-loading avoidance)1
DECISION_TEMPERATURENoSampling temperature0.8
DECISION_NEEDLE_GENERATIONNoNeedle model generation (3=121M, 2=45M)3

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
evaluateA

Jev-compatible decision call. Send a state and typed questions (noul/choice/score); get back answers with probability distributions and confidence. Use for routing, classification, scoring, escalation gating — not for open-ended generation.

record_outcomeB

Attach ground truth to a past evaluate() call for calibration tracking. outcome: boolean applied to all answers, or {questionId: boolean}.

decision_statsB

Calibration report: confidence buckets vs observed accuracy from recorded outcomes. Shows whether high-confidence answers are actually right that often.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

A3.7/5.0

Scored across 3 tools

Disambiguation5/5

The three tools map cleanly to distinct phases: making a decision (evaluate), providing feedback (record_outcome), and inspecting calibration (decision_stats). There is no overlap in purpose and each has a clearly bounded role.

Naming Consistency3/5

evaluate is a bare verb, record_outcome is verb_noun, and decision_stats is a noun phrase, so conventions are mixed. Each name is still readable and self-explanatory, but there is no single predictable pattern.

Tool Count5/5

Three tools form a tight, complete decision-calibration loop with no filler. The count is well matched to the narrow stated purpose.

Completeness4/5

The evaluate/record_outcome/decision_stats cycle covers the core lifecycle of making, labeling, and auditing decisions. Minor gaps exist (e.g., no way to list or retrieve past evaluate calls), but essential workflows are covered.

Maintenance

ActivityMaintained
ResponsivenessNo issues