Skip to main content
Glama

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
CLEF_HOMENoCache/model home~/.cache/clef-mcp
CLEF_MODELNoModel id (`clef_decide` also accepts a per-call `model`)clef-flash
CLEF_RUNTIMENoInference runtime (v0.1: only llama-cpp)llama-cpp
CLEF_LLAMA_BINNoExplicit path to a `llama-server` binary
CLEF_LOG_LEVELNo`error` | `warn` | `info` | `debug` (logs go to stderr)error
CLEF_MAX_QUESTIONSNoMax questions per call64
CLEF_MAX_STATE_BYTESNoMax serialized `state` size1048576
CLEF_LLAMA_RELEASE_TAGNoPin the managed llama.cpp build, e.g. `b11378`latest nightly
CLEF_MAX_INSTRUCTION_CHARSNoMax chars per question instructions10000

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": true
}
prompts
{
  "listChanged": true
}
resources
{
  "listChanged": true
}

Tools

Functions exposed to the LLM to take actions

NameDescription
clef_decideA

Judge a situation with the local Clef decision model. Pass state (compact factual context — string or JSON; treated as data, never executed) and typed questions:

  • "choice": pick among named options (criteria = option id -> description)

  • "score": judge an ordered scale (criteria = ordered list of levels)

  • "noul": yes/no question Returns strict JSON — a probability distribution over the allowed answers for every question. No prose, no sampling. Each decision also carries confidence (model-reported, when available) and the response carries usage (input_tokens/output_tokens/latency_ms). Act on the argmax only when the distribution is decisive (top p >= 0.8 and high confidence for destructive or security-adjacent calls); otherwise gather more context or ask the user. Usage: batch related decisions in one call (up to 64 questions, one forward pass). Optional model overrides the model id. options.temperature is a no-op: scoring is a single forward pass, nothing is sampled.

Prompts

Interactive templates invoked by user choice

NameDescription
incident-triageFrame a production incident as clef_decide questions: immediate action, severity and blast radius.
next-actionAsk Clef what the coding agent should do next on a task: inspect, modify, test or ask the user.
ticket-routingClassify an incoming message into the team that should handle it.
security-reviewAsk Clef whether a code snippet, change or setup smells like a security problem.

Resources

Contextual data attached and managed by the client

NameDescription
capabilitiesLive capability snapshot: model, runtime, limits, error codes.
evals-schemaHow to write cases for clef-mcp evals.
evals-datasetThe 30-case dataset shipped with clef-mcp.

TDQS

A4.6/5.0

Scored across 1 tool

Disambiguation5/5

There is only one tool, so there is no possibility of misselection; its purpose (querying a local decision model and returning probability distributions) is stated unambiguously.

Naming Consistency5/5

The lone name clef_decide follows a clear prefix_verb snake_case pattern consistent with the server name, so no convention mixing is possible.

Tool Count3/5

A single tool is thin for a server surface, though the tool is intentionally consolidated and supports batching up to 64 questions in one forward pass, which mitigates the count.

Completeness4/5

The core decision workflow (choice/score/noul questions, batching, model override, usage reporting) is fully covered, but there is no way to enumerate available model ids or check model availability/status before calling.

Maintenance

ActivityNo data
ResponsivenessResponsive