snapdec
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@snapdecclassify these support tickets as billing, bug, or feature request"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
β‘ snapdec
π What is snapdec?
snapdec (Snap Decisions) is a lightweight, zero-configuration Model Context Protocol (MCP) server and Agent Skill that gives AI coding agents a dedicated System 1 fast-thinking engine.
Instead of burning thousands of tokens and 3β5 seconds of latency having frontier LLMs (Claude 3.5 Sonnet, GPT-4o, Gemini 1.5 Pro) deliberate over routine categorical choices, snapdec executes typed micro-decisionsβclassify, check, score, rank, and deterministic project factsβusing specialized local models (Kev, Laya) or free hosted routers in under 50 milliseconds.
π‘ Why Coding Agents Need Snapdec
Coding agents execute hundreds of micro-decisions during multi-turn workflows:
"Is this test error an environmental flake or a code bug?"
"Which 3 files out of 30 in this git diff touch authentication?"
"What severity level is this linter violation?"
"Does this repo use uv, poetry, npm, pnpm, or cargo?"
Sending these trivial questions into a 200,000-token context window bloats costs, slows down the agent loop, and wastes time. snapdec intercepts these tasks, processes them in parallel with calibrated confidence scores, and returns an honest auto | review recommendation.
Related MCP server: jev-mcp
β¨ Key Features
β‘ Sub-50ms Micro-Decisions: Run classification and ranking in 5β50 ms locally on CPU, Apple Silicon (MLX), or CUDA GPU.
π― Calibrated Probabilities & Fail-Closed Abstention: Every decision includes exact confidence scores (
probabilities) and a typed verdict (auto | review). If the model is uncertain, it safely abstains and defers to the host agent.π οΈ Tier-0 Deterministic Facts (0 Tokens): Instant, zero-token repository introspection (
project_facts) detecting test runners, linters, package managers, monorepos, and CI setups.π» Hardware-Aware Auto-Profiler:
snapdec initinspects your exact CPU, RAM, GPU, VRAM, and OS, presenting only the local models your machine can genuinely run.π Zero-Config Agent Integration: Automatically registers with 9+ coding agents with atomic backups and 1-click clean uninstall:
Claude Code
Cursor
Windsurf
VS Code & GitHub Copilot
Codex CLI
Cline / Roo Code
OpenCode
Google Antigravity
π Always-On, Zero Idle Drain: Lightweight stdio MCP shim starts in <1s. Daemon starts lazily on first tool call and automatically resurrects dead backends. Zero background battery drain when idle.
π 100% Private & Local Loopback: Local servers bind strictly to
127.0.0.1. API keys are stored in user-owned state with restricted permissions (0600) and never leak to agent configs. Zero telemetry.π Safe, Non-Intrusive Updates:
snapdec updatechecks PyPI with a 24-hour cache. Never installs silently or modifies agent files without user consent.
π Quickstart (60 Seconds)
1. Install snapdec
Install using uv (recommended), pipx, or standard pip:
# Recommended: isolated tool installation via uv
uv tool install snapdec
# Or via pipx
pipx install snapdec
# Or standard python pip
pip install snapdec2. Run the Interactive Setup Wizard
Run snapdec init to profile your system, choose your backend, and auto-wire all detected coding agents:
snapdec initThe wizard scans your hardware and displays an honest, benchmarked menu tailored to your machine:
Recommended for this machine (13th Gen i5 Β· 16 GB RAM Β· Intel UHD Β· Windows 11)
LOCAL β free Β· private Β· offline
[1] Laya EN (421M) β DI ~0 zero-shot (specialize-first base)
5β15 ms GPU/Apple Β· 50β450 ms CPU Β· setup: ~2 GB (recommended)
[2] Laya multilingual (322M) β DI ~0 zero-shot Β· 100+ languages
[3] Decision 2.0 Eos 0.8B β card: JevArena 53.9 Β· transfer 50.3 (vllm-sr)
setup: ~3 GB Β· best Decision 2.0 fit for this machine
CPU: ~7 s/question (bench) [slow on CPU]
[4] Decision 2.0 Kai 0.6B β card: JevArena 48.6 Β· smallest (~2 GB)
CPU: ~6 s/question (bench) [slow on CPU]
[5] Kev 0.8B β DI 23.3 Β· OOD acc 0.65
40β80 ms CUDA Β· fast on Apple (MLX) Β· CPU: seconds/question
setup: ~5 GB [slow on CPU]
HOSTED β API key Β· best accuracy
[6] OpenRouter β free keys (openrouter.ai/keys) Β· typesafe/jev-router Β· Jev DI 54.0
[7] TypeSafe Jev β native /v1/systemone Β· Jev DI 54.0 (best known) Β· paid per call
[8] Other /v1/systemone URL
Choice [1]:Headless / CI Mode: You can also initialize non-interactively:
snapdec init --yes --backend local # Auto-select best local model snapdec init --yes --api-key sk-or-v1-xxxx # Auto-detect provider & model
π How It Works
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Coding Agents (Claude Code, Cursor, ...) β
ββββββββββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββ
β stdio MCP (thin shim, <1s startup)
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β snapdec daemon (127.0.0.1 IPC) β
β Lazy-start Β· Health check Β· Auto-resurrect β
ββββββββββββββββ¬ββββββββββββββββββββββββββββββββ¬βββββββββββββββ
β β
βΌ βΌ
[ Local Inference Engine ] [ Hosted System One API ]
β’ Kev 0.8B / 4B / 9B / 27B β’ OpenRouter (jev-router)
β’ Laya EN / Multilingual β’ TypeSafe Jev
β’ Decision 2.0 (vllm-sr) β’ Custom /v1/systemone
β’ ONNX / PyTorch / MLXEvery response returned to the agent includes calibrated metadata and a discreet status footer:
{
"results": [
{"id": "t1", "label": "bug", "p": 0.94, "decision": "auto"},
{"id": "t2", "label": "infra", "p": 0.61, "decision": "review"}
],
"summary": {
"items": 2,
"auto": 1,
"review": 1
},
"status": "Β· snapdec 0.5.0 Β· kev-0.8b Β· 38 ms Β· 1/2 auto Β· ~210 tok offloaded"
}The auto | review Philosophy (Safe Abstention)
decision: "auto": The model's confidence exceeds the calibrated threshold. The agent can act immediately in batch without asking the user or second-guessing.decision: "review": The model's confidence is below threshold or evidence is ambiguous. The agent falls back to inspecting the problem directly.Advisory Only:
snapdecis strictly an advisory decision engine. It is never used to automatically approve destructive actions (deletions, pushes, migrations).
π§° Available MCP Tools
When snapdec is registered, agents gain access to 6 specialized tools:
Tool | Tier | Latency | Tokens | Description |
| 0 | <1 ms | 0 | Instant inspection of test runner, linter, package manager, monorepo layout, and CI configuration. Fully deterministic. |
| 1 | 15β50 ms | Offloaded | Multi-class categorizer. Takes a list of items and candidate classes, returning labels with calibrated probabilities. |
| 1 | 10β40 ms | Offloaded | Fast boolean verification ( |
| 1 | 15β45 ms | Offloaded | Ordinal rating on a calibrated scale of 2β10 levels (e.g. risk assessment, severity rating, priority). |
| 1 | 20β60 ms | Offloaded | Evaluates candidates against a query, returning relevance ranking and best-match recommendations. |
| 1 | 20β50 ms | Offloaded | Direct structured question answering over short context state. |
Parallel Fan-Out: Multi-item requests are automatically fanned out concurrently in parallel batches so small-context local models never truncate or bottleneck on large collections.
π» CLI Usage & Mirrors
All MCP tools have direct CLI counterparts for terminal workflows, shell scripts, and CI pipelines:
# Deterministic repository inspection (Tier 0)
snapdec project-facts .
# Categorize errors or logs from stdin (Tier 1)
echo '{"items":[{"id":"1","text":"ConnectionResetError during upload"}],
"classes":{"network":"transient socket error","bug":"code bug"}}' \
| snapdec classify --input -
# Diagnostic health check
snapdec doctor --live
# Update to latest version
snapdec update
# View hardware-gated model catalog
snapdec models
# Run local accuracy and latency benchmark suite
snapdec bench --backend mockπ€ Supported Coding Agents
snapdec init and snapdec agents add automatically configure all installed agent environments:
Agent | Config Path / Mechanism | Status |
Claude Code | CLI integration ( | β Auto-configured |
Cursor |
| β Auto-configured |
Windsurf |
| β Auto-configured |
VS Code / Copilot |
| β Auto-configured |
Codex CLI |
| β Auto-configured |
Cline / Roo Code |
| β Auto-configured |
OpenCode |
| β Auto-configured |
Antigravity | Workspace & user agent custom rules | β Auto-configured |
To export standard configuration for any other MCP-compliant client:
snapdec agents print-snippetπ Models & Benchmarks
snapdec supports both local open weights and hosted router endpoints:
Model | Parameters | Hardware / Runtime | Latency | Accuracy (DI / OOD) | Notes |
Laya EN | 421M | CPU / DirectML / Apple Silicon | 5β15 ms | Base zero-shot | Ultra-lightweight, 2 GB footprint |
Laya Multilingual | 322M | CPU / DirectML / Apple Silicon | 5β15 ms | Base zero-shot | 100+ languages supported |
Kev 0.8B | 0.8B | CUDA / Apple Silicon (MLX) | 40β80 ms | DI 23.3 Β· OOD 0.65 | Recommended for Apple Silicon & GPUs |
Kev 4B | 4.0B | 16 GB+ VRAM GPU | 60β120 ms | DI 38.0 | High-accuracy local model |
Kev 9B | 9.0B | 24 GB+ VRAM GPU | 80β180 ms | DI 41.0 | Heavyweight local specialist |
Kev 27B | 27.0B | 80 GB+ GPU / 96GB+ Mac | 150β350 ms | DI 52.3 | Near-frontier decision intelligence |
Decision 2.0 Kai | 0.6B | CPU (slow) / CUDA / Apple (CPU, slow) | ~5.8 s CPU Β· 4.9 ms GPU (card) | card: JevArena 48.6 (β ) | Smallest of the family |
Decision 2.0 Eos | 0.8B | CPU (slow) / CUDA / Apple (CPU, slow) | ~7.4 s CPU Β· 6.0 ms GPU (card) | card: JevArena 53.9 (β ) | Starred on CPU/Windows; beats Kev-0.8B on card |
Decision 2.0 Sol | 2B | CPU (slow) / CUDA / Apple (32 GB+) | not benched Β· 7.2 ms GPU (card) | card: JevArena 52.1 (β ) | Fits 16 GB RAM, marked slow on CPU |
Decision 2.0 Nox | 4B | CPU (very slow) / CUDA / Apple (32 GB+) | not benched Β· 12.9 ms GPU (card) | card: JevArena 63.6 (β ) | Needs β₯20 GB RAM |
imajev 2B | 2.2B | CPU (slow, 12 GB+ RAM) / Apple (MLX, 8 GB+) | p50 14.8 s / p95 35.1 s CPU | board: JevBench hard 60.4 Β· Img JevBench 68.72 #6 (β‘) | Only variant a 16 GB PC runs |
imajev 4B | 4.3B | CPU (very slow, 24 GB+) / Apple (MLX, 16 GB+) | bench pending | board: JevBench 67.37 #1 Β· Img 76.39 #1 Β· DecisionBench 79.65 #3 (β‘) | Board #1 text decision model |
imajev 9B | 9.4B | CPU (48 GB+) / Apple (MLX, 32 GB+) | bench pending | board: JevBench hard 69.4 (β‘) | Heavyweight Qwen3.5 specialist |
TypeSafe Jev | Hosted | Native | ~120 ms | DI 54.0 (Reference) | Best known decision intelligence |
OpenRouter | Hosted |
| ~150 ms | Jev DI 54.0 | Free API key tier available |
DI (Decision Intelligence) benchmarks cited from the official Kev 1.0 test suite. Measured local performance available in docs/benchmarks/.
(β ) Decision 2.0 numbers are from the vendor's model cards (vllm-sr, 2026-10) on their own JevArena index β a different scale from the held-out breadth-v1 numbers above, so the two are never cross-compared (ADR-0008/0010). Measured CPU latency lives in docs/benchmarks/. On Apple Silicon, Decision 2.0 runs via the plain CPU path (no MLX build yet; MPS unvalidated), so Kev 0.8B (MLX) stays the recommended fast local pick there.
(β‘) imajev numbers are JevBench board standings (2026-09) and repo runs on their own indices β a third scale, never cross-compared with kev's breadth-v1 or the vllm-sr card (ADR-0008/0011). Measured CPU latency lives in docs/benchmarks/. On Apple Silicon imajev runs a real MLX fast path (fp16), unlike Decision 2.0 there. Downloads show live byte + speed progress bars.
π Security, Privacy & Reliability
Strict Loopback Binding: Local model servers bind only to
127.0.0.1. No external ports are ever opened.Protected Secrets: API keys are saved with strict
0600file permissions in~/.local/state/snapdec/or read from environment variables. They are never written into agent configuration files.Fail-Closed Design (NFR-4): If a backend crashes, drops connection, or times out, snapdec returns a valid envelope with
decision: "review". It never throws an unhandled exception or interrupts your agent session.Zero Telemetry: No usage stats, prompts, code snippets, or user data are ever tracked or phoned home.
π οΈ CLI Command Reference
Command | Description |
| Interactive system setup wizard (auto-detects hardware and agents). |
| Comprehensive health check of daemon, backend, and agent registrations. |
| Check PyPI and upgrade snapdec installation safely. |
| Manage the background decision daemon. |
| Manage coding agent integrations. |
| List all runnable models matching current hardware. |
| Run the decision accuracy and latency benchmark suite. |
| Deterministic repository fact extraction. |
| CLI classifier reading JSON from stdin. |
| CLI verification reading JSON from stdin. |
| CLI ordinal scoring tool reading JSON from stdin. |
| CLI ranker reading JSON from stdin. |
| Cleanly reverses all agent integrations and removes configurations. |
β Frequently Asked Questions (FAQ)
Frontier LLMs incur input token costs for every turn in their conversation history. In agent loops, repeatedly passing long logs, diffs, and lists into a 200k context window to ask small classification or boolean questions consumes significant tokens and compute. snapdec offloads these discrete questions to a local model or fast router, returning only the concise answer and confidence score.
Yes. When you choose a local model (such as Laya or Kev), all dependencies, weights, and runtimes run on your local machine on 127.0.0.1. No internet connection is required after initial model download.
The MCP shim features built-in self-healing: if the daemon is stopped or crashes, the shim automatically re-spawns it on the next incoming tool call. If the backend fails to recover, it returns a safe decision: "review" envelope so the agent continues operating without crashing.
Simply run:
snapdec uninstallThis restores all agent configuration files from their original backups and cleans up registered skills.
π Architecture & Decisions
Architecture Design: SYSONE_ARCHITECTURE.md
Architectural Decision Records (ADRs):
π License
Distributed under the Apache-2.0 License. See LICENSE for details.
snapdec routes to and credits Kev, Laya, TypeSafe Jev, the vllm-sr Decision 2.0 family, and the imajev family (pinned @ ccf586d4). Model weights are distributed under their respective Apache-2.0 licenses.
This server cannot be deployed
Maintenance
Related MCP Connectors
Your team's shipping standards, org map and delivery metrics, inside your coding agent.
Shared control plane for AI coding agents β tasks, memory, decisions, file locks. 12 tools.
Code intelligence platform for AI agents. 20 tools for architecture, security & impact analysis.
- vibsyncOAuthcom.vibsync
One shared brain for your AI coding agents: team memory, agent Q&A, tasks, and file claims.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceEnables coding agents to scout, rank, and preflight software work before implementation, returning evidence-backed ACT, VERIFY, or SKIP decisions for issues and pull requests.98 npm2MIT
- AlicenseAqualityBmaintenanceEnables coding agents to make cheap, fast probabilistic decisions on every turn, with tools for coding-loop checks, review, verification, screening untrusted input, and ranking candidates.6470 npm61MIT
- AlicenseNot gradedqualityAmaintenanceProvides coding agents with typed classification, yes/no checks, scoring, ranking, and question-answering tools that return calibrated probabilities for fast, reliable decisions.889 npm44MIT
- AlicenseAqualityBmaintenanceEnables coding agents to query files, logs, and search results through TypeSafe Jev's typed, probabilistic answers, so they retrieve only the needed conclusion instead of raw context. Supports classification, scoring, yes/no checks, extraction, page ranking, and injection screening.11470 npmMIT