ember
Serves Cloudflare's Clef-Flash decision model locally (via the pinned upstream revision) as the engine behind the advise tool: given a situation/state and a schema of typed questions (noul yes/no, choice named options, score ordered options), it returns a calibrated probability per option with token usage and latency — advising the calling agent rather than generating text.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@embergiven this stack trace, is the bug in the cache, the parser, or the network?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Brand assets and palette · Provenance and licensing · Project site
A local gut feeling for coding agents. ember runs
Cloudflare's Clef-Flash model on your Apple
Silicon Mac and gives agents one MCP tool, advise: describe a situation, ask typed
questions, and get back a calibrated feeling about every option. It's a little buddy for
judgment calls — it advises; the agent decides.
How it works
Clef is a decision model, not a chat model: it takes a state plus a schema of typed
questions and returns one probability per option, with no text generation. So ember
plugs into agents as a tool, while their reasoning stays on their normal LLM:
The model server (
ember/server.py) loads the model once and stays warm across agent sessions.The MCP server (
ember/mcp_server.py) never loads the model; it starts the model server on the first tool call, so the MCP handshake stays instant.
Names
Thing | Name |
Product, Python import, repository |
|
Distribution ( |
|
CLI |
|
MCP server / command |
|
Tool |
|
Playbook skill / MCP resource |
|
Environment variables |
|
"Clef" always refers to Cloudflare's upstream model, never to this product.
Related MCP server: jev-code
Verified
On a MacBook Pro M4 Max / 128 GB, torch 2.14.1, transformers 5.18.0, mcp 2.3:
Step | Result |
Model load | ~5 s |
Warm request | ~0.9–1.3 s for ~220–360 input tokens |
opencode end to end | ✅ the agent reads the instructions, lists the skill, and calls the tool unprompted |
Install (Apple Silicon)
Requirements
Apple Silicon Mac (M-series) on macOS. Intel Macs and NVIDIA/CUDA are out of scope.
Unified memory above the model's size: 32 GB or more for
flash(9B), 64 GB or more forfull(27B). Only 128 GB has been verified.Disk: about 18 GiB for
flashor 55 GiB forfull, in Hugging Face's shared cache (~/.cache/huggingface). Config, state, and logs live in~/Library/Application Support/ember.Python 3.12, managed by uv.
uv tool install --python 3.12 "gut @ git+https://github.com/shapeandshare/ember"
ember model pull # ~18 GB, resumable, disk-space checked
ember doctor # platform, dependencies, model, and server
ember init --opencode # register with opencode: config, plugin, and skillKeep --python 3.12: uv otherwise picks your newest interpreter, which the pinned
torch/transformers stack is not tested on.
The repository is public, so the install above needs no credentials. To use SSH instead,
install from git+ssh://git@github.com/shapeandshare/ember.
Restart opencode and the agent gains ember_advise. The model server stays lazy —
it starts on the first tool call (or with ember start).
Everything is pinned for reproducibility: flash to the commit verified on MPS (17f0b0a),
full to its release commit (2f3de3d, not yet verified locally), and torch/torchvision to
the tested minor series. Set EMBER_MODEL_DIR to run another weights directory.
Cloud
ember is Apple-Silicon-first. In the cloud, run it on an Apple Silicon host with the memory
above, or on any host with the CPU fallback (EMBER_DEVICE=cpu — float32, roughly twice the
memory, much slower). NVIDIA/CUDA is out of scope. The HTTP server binds to loopback with no
authentication, so exposing it beyond the host needs your own access controls. See
COMPATIBILITY.md and SECURITY.md.
Agent onboarding
Installing the tool is half the job; the other half is making agents want to consult it
at the right moments and read its answers sensibly. ember/agent_kit/ ships that
guidance through every channel each agent actually reads:
Channel | opencode | Claude Code | Codex CLI | How you get it |
MCP server instructions (when to consult, how to ask, how to read answers) | ✅ in the system prompt | ✅ (2 KB cap) | — | built in, nothing to do |
| ✅ via | ✅ | — | built in |
| ✅ | ✅ | ✅ |
|
AGENTS.md / CLAUDE.md policy block | ✅ | ✅ (CLAUDE.md) | ✅ |
|
ember init --opencode # opencode: config entry, plugin, and skill
ember agents install --agent claude # .claude/skills/ember-advise/SKILL.md
ember agents install --agent codex # .agents/skills/... (opencode reads this too)
ember agents show snippet >> AGENTS.md # then edit the project policy at the endThe skill is a playbook, not a reference card: the decision points worth consulting ember about, copy-paste question sets for intent and readiness, failure triage, change risk, routing, and effort, and starting confidence thresholds calibrated from observed model output (re-measured whenever the pinned model revision changes). The snippet ends with a project policy — edit it to wire ember into your own workflow, e.g. "check change risk before every push".
Agent-specific notes:
opencode reads skills from
.opencode/skills,.claude/skills, and.agents/skills, so install one copy per project to avoid duplicate listings.Claude Code: register the server with
claude mcp add ember -- ember-mcp; the tool appears asmcp__ember__advise.Codex CLI support for MCP server instructions and resources is unconfirmed, so rely on the skill and the AGENTS.md snippet there.
CLI
gut is a short alias for every command below.
# lifecycle
ember doctor # platform, dependencies, model, and server
ember serve # run the model server in the foreground
ember start | stop | restart | status | logs
ember config path | show
ember uninstall [--purge-models] # also removes global opencode/skill installs
# models
ember model pull [flash|full] # flash = 9B (default), full = 27B
ember model list | path [name] | rm [name]
# opencode and agents
ember init [--opencode] [--global]
ember agents install [--agent opencode|claude|codex] [--global]
ember agents show instructions|skill|snippet
ember mcp # the MCP stdio server agents launchUsage
Ask the agent in natural language — "Is this bug report urgent, and which team should own
it?" — and it calls ember_advise with a state and typed questions:
{
"model": "clef-flash",
"answers": {
"urgent": { "type": "noul", "noul": 0.967 },
"team": {
"type": "choice",
"choice": "db",
"confidence": 0.9782,
"probabilities": { "db": 0.9782, "frontend": 0.0218 }
}
},
"usage": { "input_tokens": 220, "output_tokens": 0 },
"latency_ms": 930.9
}Question types: noul (yes/no → P(true)), choice (named options), score (ordered options →
expected score plus legend). The model field echoes the upstream model label.
When the evidence is visual, attach images or video frames inline: images is a list of
data:image/png;base64,... URIs (or {"content_type": "image/png", "base64": "..."}
objects), and videos is a list of videos, each a list of frame refs. Remote URLs and local
paths are rejected — the model server never reads host files or fetches URLs for an agent.
Configuration
Settings resolve as CLI flag > environment variable > config file > default. The config file
is JSON at ember config path (keys model, host, port, device, max_length; a
max_length of 0 means the model's own maximum).
Variable | Default | Purpose |
|
| Model server address |
|
|
|
|
|
|
| — | Run weights from this directory instead of the pinned cache |
|
| Token cap per request; |
|
| Where the MCP server sends requests |
|
| Let the MCP server start the model server on demand |
|
| Seconds to wait for the model server to start |
| Application Support | Where the pid file and logs live |
Metrics
While the model server is running it exposes Prometheus metrics at
http://127.0.0.1:8765/metrics (the EMBER_HOST/EMBER_PORT address):
Metric | Type | Meaning |
| counter | advise requests by HTTP status ( |
| histogram | advise request latency by status |
| counter | input tokens processed |
| counter | output tokens produced |
| gauge |
|
The endpoint binds to the same address as the rest of the API (loopback by default), so
it is reachable only there unless you change EMBER_HOST. Ember's metrics live in a
dedicated Prometheus registry, so the standard python_*/process_* collectors are not
included — /metrics shows only the table above.
Development
ember/
cli.py # the `ember` command (alias `gut`)
runtime.py # MPS-safe loader (CPU load → .to("mps")) + Engine
server.py # FastAPI: POST /v1/systemone, GET /health (reports pid)
mcp_server.py # MCP stdio server: advise tool, instructions, ember://guide
process.py # model-server lifecycle (pid file + HTTP health)
models.py # pinned model registry + pull/list/rm
paths.py config.py # platform dirs, config precedence
opencode_config.py opencode_plugin.py # opencode integration
agent_kit/ # what agents read: instructions, ember-advise skill, AGENTS snippet
packages/opencode-plugin/ # npm-ready opencode plugin source
scripts/ # MPS smoke test, MCP end-to-end check, provenance and vault audits
tests/ # pytest suite (unit + model-backed, host-isolated)
vault/ # project memory (Obsidian): decisions, discoveries, session logs
.specify/ # spec-kit; memory/constitution.md governs this repo
AGENTS.md CLAUDE.md # guidelines for agents working on this repo
Makefile # contributor lifecycle (wraps the CLI)make bootstrap # deps + pinned weights + opencode.json + doctor (idempotent)
make start # model server in the background; stop | restart | status | logs
make opencode # this checkout's opencode plugin and skillopencode.json, .opencode/plugins/ember.js, and .opencode/skills/ember-advise/
embed this clone's absolute paths or copy packaged files, so they are gitignored — regenerate
them with make init / make opencode after cloning. .opencode/opencode.json is shared and
committed: it registers the vault MCP server that agents use to read and write vault/
(launch opencode from the repository root).
Make targets
Target | What it does |
| List all targets (default) |
| From a fresh clone: |
| Deps + weights / deps only / pinned weights to |
|
|
| Model-server lifecycle via the CLI |
| Run the MCP server / |
| Full suite / unit tests only / full suite that fails without weights |
| Calibration eval suite: positive + negative recipe cases (loads model) |
| Run the benchmark dataset against the live server; writes |
| Copy the latest run into |
| Render the most recent run as a Markdown table |
| MCP end-to-end check / direct MPS inference |
| Byte-compile / compile + unit tests |
|
|
|
|
| Check |
| Build the Pages site into |
| Caches and build output / weights ( |
Make re-syncs the environment automatically when pyproject.toml or uv.lock changes.
Tests
make test # full suite (~30 s, loads the model once)
make test-fast # unit tests only, no model load (~12 s)
make check # compile + test-fast
make test-strict # full suite; fails (not skips) if weights are missing
make ci # bootstrap + check + test-stricttests/ covers the supported call paths:
GET /healthandPOST /v1/systemoneacrossnoul/choice/score, probability normalization, and request-validation422sMCP tool discovery,
adviseover stdio, actionable errors (server down, model missing, malformed questions), and autostart on a configured port with pid cleanupthe agent kit: instructions size, skill frontmatter, the
ember://guideresourcethe CLI:
agents,init --opencode,status/stopon unused ports,doctor,uninstallprocess safety:
stopnever signals a pid that is not an ember server it started
Safe on a host running other opencode instances. The suite binds random free ports (never
8765), keeps state in temporary directories, stops only the processes it started, and never
invokes the opencode CLI or touches global opencode config.
CI (.github/workflows/ci.yml) runs uv sync --locked, make check, uv build, and an
install smoke of the built wheel on a hosted Apple Silicon runner, and audits the workflows
with zizmor. Model-backed tests are
not run in CI: the ~19 GB fp16 model does not fit the available runners (hosted or the
org's 8 GiB self-hosted VMs), so run make test locally for model-affecting changes.
Benchmark
evals/clef-flash.jsonl is a 264-item benchmark covering the five agent-kit recipes (intent
and readiness, failure triage, change risk, routing, effort and approach) plus four vision
recipes (vision_noul, vision_choice, vision_score, vision_video): 456 scored questions
(184 choice, 152 noul, 88 score, 32 across the four vision recipes), split into
dev (130 items) and test (134). Vision items carry an images or videos field
(base64 data: URIs) alongside state and questions. Each line is
self-contained (state, the recipe's fixed question set, a gold label for every question, and
a rationale), so other implementations can score it without this harness.
make eval-run # or: ember eval run [--split dev|test] [--category <recipe>]
make eval-report # Markdown report of the most recent run
make eval-export # or: ember eval export; the reviewer bundle described below
ember eval report --compare results/<a>_results.json results/<b>_results.jsonember eval export writes results/<run_id>_report/ for external reviewers:
report.html (one self-contained file with inline charts, light and dark themes, a print
layout, and item filters; it loads nothing from the network), report.md with its
figures/*.svg, and data/ with the run's results, trace, and the exact dataset scored.
The report covers context, the system under test, benchmark design, metric definitions with
references, results, calibration, the agent kit's decision rules replayed on every item, a
review card for every miss, limitations, and reproduction pins.
The report gives accuracy with a 95% bootstrap interval, macro-F1, top-label ECE, and Brier
score (choice, noul); ranked probability score and MAE (score); and coverage and accuracy
at the agent kit's act-on-it thresholds. Each run records the git hash, the engine, and the
dataset's SHA-256, and --compare warns when two runs scored different item sets.
Gold labels are judged from state alone against the recipe's option descriptions, and they
are fixed before any run: needs_review is true exactly when risk is Medium or High, and
retry only for flaky failures. A person reviews disagreements; labels are never changed to
match the model. make check validates the dataset (tests/test_eval_benchmark.py).
ember eval reads evals/ and scripts/ from the checkout, so it works only in a clone.
Agent in the loop
The benchmark above scores the model. ember eval agent scores ember the way it is used:
through a coding agent. It gives opencode 54 scripted requests (vague and precise asks,
failing tests, risky and trivial commits, routing and effort questions, and controls), each
in a throwaway git repo, and judges the agent only by what it did (files, commits, tests,
its reply). Each scenario runs under four conditions: none (no ember), mcp (the MCP
server and its instructions), skill (plus the ember-advise skill), and full (plus the
AGENTS.md policy).
ember start # the agent sessions call the warm server
make eval-agent-smoke # 6 scenarios x 4 conditions x 2 models, 1 trial
make eval-agent # or: ember eval agent [--models ...] [--trials 3] [--parallel 8]
ember eval export --agent latest # put the agent results first in the reviewer reportEvery session runs opencode run --pure with a private HOME and XDG directories and no TCP
port, so your opencode config, plugins, sessions, and running instances are untouched. The
provider key is read from the environment or opencode's auth file and passed only to the child
process. The run reports, per model and condition: gold-action accuracy, its paired change
against none, how often the agent consulted ember at decision points (and on controls),
whether it consulted before acting, asked the kit's recipe questions, passed the evidence, and
followed ember's answer, plus cost and time per session. It spends real provider credit and is
never part of make check, make test, or CI.
Caveats
Gated DeltaNet fallback. Qwen3.5's backbone is a hybrid linear-attention model. The optimized CUDA kernels (
causal_conv1d,flash-linear-attention) don't exist for Apple Silicon, so it uses the pure-PyTorch reference path: correct but slower. You'll see two "falling back" log lines — expected.fp16 on MPS.
bfloat16works but is emulated and less battle-tested on MPS, so the loader usesfloat16; CPU mode usesfloat32.Vision works on MPS. Images and video frames run through the same fp16 path as text. Inline them as base64
data:URIs inimages/videos; remote URLs and local paths are rejected on purpose, so an agent cannot make the server read host files or fetch URLs. Text-only inputs skip the vision tower entirely.Numerics. MPS can differ slightly from CUDA/CPU. If calibrated probabilities matter, cross-check with
EMBER_DEVICE=cpu ember restart.device_map={"": "mps"}segfaults with the pinned stack. The loader loads on CPU and then moves the model to MPS — don't "simplify" that away.
Troubleshooting
Start with
ember doctor, thenember logs(~/Library/Application Support/ember/logs/server.log).The agent can't see the tool:
opencode mcp listshould show✓ ember connected."model … is not pulled": run
ember model pull, or pointEMBER_MODEL_DIRat the weights.Qwen3VLVideoProcessor requires Torchvision→ torchvision is a pinned dependency; runuv sync(or reinstall the tool).Any single unimplemented MPS op falls back to CPU (
PYTORCH_ENABLE_MPS_FALLBACK=1, set by the runtime).
License
MIT (see LICENSE). Cloudflare's Clef weights and joint_schema_model.py are Apache-2.0; they
are downloaded from Hugging Face at runtime, not redistributed here.
Available Tools
1 tooladviseA
Get ember's read on a situation, as calibrated probabilities.
ember (Cloudflare's Clef-Flash model, running locally) advises; you decide.
Consult it at bounded decision points: intent, triage, routing, yes/no gates, and
risk, severity, or effort scores. It sees only what you pass, so include every
piece of evidence the call depends on and attach images or video frames as base64
`data:` URIs in `images`/`videos` when pixels are the evidence. For 'choice' the
answer has the leading option, its confidence, and full probabilities; for 'score'
an expected score over the ordered criteria; for 'noul' the probability the
proposition is true.
| Name | Required | Description | Default |
|---|---|---|---|
| input | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does substantial work: it discloses that the model runs locally, that it 'sees only what you pass' (so evidence must be included explicitly), that media must be base64 data: URIs rather than URLs or paths, and what each question type returns. It omits cost/latency/rate-limit and any failure-mode behavior, keeping it out of the top band.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, followed by role, usage contexts, input requirements, and per-type outputs in a logical order. Dense but nearly every sentence carries information; the 'ember advises; you decide' restatement is the only mildly redundant line.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, yet the description helpfully summarizes the shape per question type. Combined with the media-encoding constraints and evidence-inclusion warning, an agent has enough to call this correctly; only edge cases (model selection, media_kwargs tuning) are left entirely to the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is reported as 0%, so the description must compensate, and it does: it explains what belongs in `state` ('include every piece of evidence the call depends on'), how `images`/`videos` must be supplied, and the semantics of the `choice`/`score`/`noul` question types. It adds little about `model` or `media_kwargs`, but the core call-shaping parameters are covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: get a calibrated-probability read on a situation from a named model (ember, Clef-Flash running locally). It also distinguishes the tool's role from the caller's ('ember advises; you decide'), so there is no ambiguity about what the tool returns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Enumerates concrete invocation contexts ('bounded decision points: intent, triage, routing, yes/no gates, and risk, severity, or effort scores'), which is clear when-to-use guidance. It does not state when NOT to use it, and there are no sibling tools to route against, so it falls just short of the top band.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.0- First observed
advise
TDQS
Scored across 1 tool
Only one tool exists, so there is no risk of confusion with other tools; its purpose is clearly stated as getting calibrated probabilistic advice.
The single tool name 'advise' is a clear verb, and no inconsistent conventions exist within the set.
One tool is slightly below the typical 3-15 range, but it is well-scoped for a focused advisory endpoint that handles multiple query modes internally.
The tool covers choice, score, and noul response types, plus image/video evidence, providing a complete surface for its stated advisory purpose.
Maintenance
Related MCP Connectors
Calibrated judgments for text: yes/no probabilities, picks from your options, or scores.
Calibrated decisions for agents: choice, yes/no and score questions answered with probabilities.
Calibrated world model for AI agents. 40 tools: world state, markets, trading. Kalshi + Polymarket.
A paid remote MCP for OpenAI Codex agent coordination MCP, built to return verdicts, receipts, usage
Related MCP Servers
- AlicenseAqualityAmaintenanceEnables agents to get fast, calibrated probabilistic answers from Jev (Typesafe AI) to yes/no, scale, or choice questions about provided material, without using a generative model.115 npm1MIT
- AlicenseNot gradedqualityAmaintenanceProvides coding agents with typed classification, yes/no checks, scoring, ranking, and question-answering tools that return calibrated probabilities for fast, reliable decisions.889 npm41MIT
- AlicenseNot gradedqualityBmaintenanceEnables coding agents to query a locally running Kev decision model through MCP tools, returning calibrated probabilities for typed questions such as yes/no, choice, and score.Apache 2.0
- AlicenseNot gradedqualityAmaintenanceEnables Claude Code or any MCP client to ask TypeSafe's Jev for calibrated, typed judgments (probabilities, choices, scores) instead of prose, with local caching and cost tracking.1MIT