Skip to main content
Glama

⚑ snapdec


πŸ“Œ What is snapdec?

snapdec (Snap Decisions) is a lightweight, zero-configuration Model Context Protocol (MCP) server and Agent Skill that gives AI coding agents a dedicated System 1 fast-thinking engine.

Instead of burning thousands of tokens and 3–5 seconds of latency having frontier LLMs (Claude 3.5 Sonnet, GPT-4o, Gemini 1.5 Pro) deliberate over routine categorical choices, snapdec executes typed micro-decisionsβ€”classify, check, score, rank, and deterministic project factsβ€”using specialized local models (Kev, Laya) or free hosted routers in under 50 milliseconds.

πŸ’‘ Why Coding Agents Need Snapdec

Coding agents execute hundreds of micro-decisions during multi-turn workflows:

  • "Is this test error an environmental flake or a code bug?"

  • "Which 3 files out of 30 in this git diff touch authentication?"

  • "What severity level is this linter violation?"

  • "Does this repo use uv, poetry, npm, pnpm, or cargo?"

Sending these trivial questions into a 200,000-token context window bloats costs, slows down the agent loop, and wastes time. snapdec intercepts these tasks, processes them in parallel with calibrated confidence scores, and returns an honest auto | review recommendation.


Related MCP server: jev-mcp

✨ Key Features

  • ⚑ Sub-50ms Micro-Decisions: Run classification and ranking in 5–50 ms locally on CPU, Apple Silicon (MLX), or CUDA GPU.

  • 🎯 Calibrated Probabilities & Fail-Closed Abstention: Every decision includes exact confidence scores (probabilities) and a typed verdict (auto | review). If the model is uncertain, it safely abstains and defers to the host agent.

  • πŸ› οΈ Tier-0 Deterministic Facts (0 Tokens): Instant, zero-token repository introspection (project_facts) detecting test runners, linters, package managers, monorepos, and CI setups.

  • πŸ’» Hardware-Aware Auto-Profiler: snapdec init inspects your exact CPU, RAM, GPU, VRAM, and OS, presenting only the local models your machine can genuinely run.

  • πŸ”Œ Zero-Config Agent Integration: Automatically registers with 9+ coding agents with atomic backups and 1-click clean uninstall:

    • Claude Code

    • Cursor

    • Windsurf

    • VS Code & GitHub Copilot

    • Codex CLI

    • Cline / Roo Code

    • OpenCode

    • Google Antigravity

  • πŸ”‹ Always-On, Zero Idle Drain: Lightweight stdio MCP shim starts in <1s. Daemon starts lazily on first tool call and automatically resurrects dead backends. Zero background battery drain when idle.

  • πŸ”’ 100% Private & Local Loopback: Local servers bind strictly to 127.0.0.1. API keys are stored in user-owned state with restricted permissions (0600) and never leak to agent configs. Zero telemetry.

  • πŸ”„ Safe, Non-Intrusive Updates: snapdec update checks PyPI with a 24-hour cache. Never installs silently or modifies agent files without user consent.


πŸš€ Quickstart (60 Seconds)

1. Install snapdec

Install using uv (recommended), pipx, or standard pip:

# Recommended: isolated tool installation via uv
uv tool install snapdec

# Or via pipx
pipx install snapdec

# Or standard python pip
pip install snapdec

2. Run the Interactive Setup Wizard

Run snapdec init to profile your system, choose your backend, and auto-wire all detected coding agents:

snapdec init

The wizard scans your hardware and displays an honest, benchmarked menu tailored to your machine:

Recommended for this machine (13th Gen i5 Β· 16 GB RAM Β· Intel UHD Β· Windows 11)

  LOCAL β€” free Β· private Β· offline
    [1] Laya EN (421M)  β€”  DI ~0 zero-shot (specialize-first base)
        5–15 ms GPU/Apple Β· 50–450 ms CPU Β· setup: ~2 GB        (recommended)
    [2] Laya multilingual (322M)  β€”  DI ~0 zero-shot Β· 100+ languages
    [3] Decision 2.0 Eos 0.8B  β€”  card: JevArena 53.9 Β· transfer 50.3 (vllm-sr)
        setup: ~3 GB Β· best Decision 2.0 fit for this machine
        CPU: ~7 s/question (bench)                           [slow on CPU]
    [4] Decision 2.0 Kai 0.6B  β€”  card: JevArena 48.6 Β· smallest (~2 GB)
        CPU: ~6 s/question (bench)                           [slow on CPU]
    [5] Kev 0.8B  β€”  DI 23.3 Β· OOD acc 0.65
        40–80 ms CUDA Β· fast on Apple (MLX) Β· CPU: seconds/question
        setup: ~5 GB                                             [slow on CPU]
  HOSTED β€” API key Β· best accuracy
    [6] OpenRouter  β€”  free keys (openrouter.ai/keys) Β· typesafe/jev-router Β· Jev DI 54.0
    [7] TypeSafe Jev  β€”  native /v1/systemone Β· Jev DI 54.0 (best known) Β· paid per call
    [8] Other /v1/systemone URL

Choice [1]:

Headless / CI Mode: You can also initialize non-interactively:

snapdec init --yes --backend local             # Auto-select best local model
snapdec init --yes --api-key sk-or-v1-xxxx     # Auto-detect provider & model

πŸ” How It Works

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚             Coding Agents (Claude Code, Cursor, ...)        β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚ stdio MCP (thin shim, <1s startup)
                               β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                 snapdec daemon (127.0.0.1 IPC)              β”‚
β”‚       Lazy-start Β· Health check Β· Auto-resurrect            β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
               β”‚                               β”‚
               β–Ό                               β–Ό
    [ Local Inference Engine ]     [ Hosted System One API ]
    β€’ Kev 0.8B / 4B / 9B / 27B     β€’ OpenRouter (jev-router)
    β€’ Laya EN / Multilingual       β€’ TypeSafe Jev
    β€’ Decision 2.0 (vllm-sr)       β€’ Custom /v1/systemone
    β€’ ONNX / PyTorch / MLX

Every response returned to the agent includes calibrated metadata and a discreet status footer:

{
  "results": [
    {"id": "t1", "label": "bug", "p": 0.94, "decision": "auto"},
    {"id": "t2", "label": "infra", "p": 0.61, "decision": "review"}
  ],
  "summary": {
    "items": 2,
    "auto": 1,
    "review": 1
  },
  "status": "Β· snapdec 0.5.0 Β· kev-0.8b Β· 38 ms Β· 1/2 auto Β· ~210 tok offloaded"
}

The auto | review Philosophy (Safe Abstention)

  • decision: "auto": The model's confidence exceeds the calibrated threshold. The agent can act immediately in batch without asking the user or second-guessing.

  • decision: "review": The model's confidence is below threshold or evidence is ambiguous. The agent falls back to inspecting the problem directly.

  • Advisory Only: snapdec is strictly an advisory decision engine. It is never used to automatically approve destructive actions (deletions, pushes, migrations).


🧰 Available MCP Tools

When snapdec is registered, agents gain access to 6 specialized tools:

Tool

Tier

Latency

Tokens

Description

project_facts

0

<1 ms

0

Instant inspection of test runner, linter, package manager, monorepo layout, and CI configuration. Fully deterministic.

classify

1

15–50 ms

Offloaded

Multi-class categorizer. Takes a list of items and candidate classes, returning labels with calibrated probabilities.

check

1

10–40 ms

Offloaded

Fast boolean verification (yes | no | uncertain) against provided evidence.

score

1

15–45 ms

Offloaded

Ordinal rating on a calibrated scale of 2–10 levels (e.g. risk assessment, severity rating, priority).

rank

1

20–60 ms

Offloaded

Evaluates candidates against a query, returning relevance ranking and best-match recommendations.

ask

1

20–50 ms

Offloaded

Direct structured question answering over short context state.

Parallel Fan-Out: Multi-item requests are automatically fanned out concurrently in parallel batches so small-context local models never truncate or bottleneck on large collections.


πŸ’» CLI Usage & Mirrors

All MCP tools have direct CLI counterparts for terminal workflows, shell scripts, and CI pipelines:

# Deterministic repository inspection (Tier 0)
snapdec project-facts .

# Categorize errors or logs from stdin (Tier 1)
echo '{"items":[{"id":"1","text":"ConnectionResetError during upload"}],
       "classes":{"network":"transient socket error","bug":"code bug"}}' \
  | snapdec classify --input -

# Diagnostic health check
snapdec doctor --live

# Update to latest version
snapdec update

# View hardware-gated model catalog
snapdec models

# Run local accuracy and latency benchmark suite
snapdec bench --backend mock

πŸ€– Supported Coding Agents

snapdec init and snapdec agents add automatically configure all installed agent environments:

Agent

Config Path / Mechanism

Status

Claude Code

CLI integration (claude mcp add)

βœ… Auto-configured

Cursor

~/.cursor/mcp.json

βœ… Auto-configured

Windsurf

~/.codeium/windsurf/mcp_config.json

βœ… Auto-configured

VS Code / Copilot

.vscode/mcp.json / user settings

βœ… Auto-configured

Codex CLI

~/.codex/config.toml

βœ… Auto-configured

Cline / Roo Code

cline_mcp_settings.json

βœ… Auto-configured

OpenCode

~/.opencode/mcp.json

βœ… Auto-configured

Antigravity

Workspace & user agent custom rules

βœ… Auto-configured

To export standard configuration for any other MCP-compliant client:

snapdec agents print-snippet

πŸ“Š Models & Benchmarks

snapdec supports both local open weights and hosted router endpoints:

Model

Parameters

Hardware / Runtime

Latency

Accuracy (DI / OOD)

Notes

Laya EN

421M

CPU / DirectML / Apple Silicon

5–15 ms

Base zero-shot

Ultra-lightweight, 2 GB footprint

Laya Multilingual

322M

CPU / DirectML / Apple Silicon

5–15 ms

Base zero-shot

100+ languages supported

Kev 0.8B

0.8B

CUDA / Apple Silicon (MLX)

40–80 ms

DI 23.3 Β· OOD 0.65

Recommended for Apple Silicon & GPUs

Kev 4B

4.0B

16 GB+ VRAM GPU

60–120 ms

DI 38.0

High-accuracy local model

Kev 9B

9.0B

24 GB+ VRAM GPU

80–180 ms

DI 41.0

Heavyweight local specialist

Kev 27B

27.0B

80 GB+ GPU / 96GB+ Mac

150–350 ms

DI 52.3

Near-frontier decision intelligence

Decision 2.0 Kai

0.6B

CPU (slow) / CUDA / Apple (CPU, slow)

~5.8 s CPU Β· 4.9 ms GPU (card)

card: JevArena 48.6 (†)

Smallest of the family

Decision 2.0 Eos

0.8B

CPU (slow) / CUDA / Apple (CPU, slow)

~7.4 s CPU Β· 6.0 ms GPU (card)

card: JevArena 53.9 (†)

Starred on CPU/Windows; beats Kev-0.8B on card

Decision 2.0 Sol

2B

CPU (slow) / CUDA / Apple (32 GB+)

not benched Β· 7.2 ms GPU (card)

card: JevArena 52.1 (†)

Fits 16 GB RAM, marked slow on CPU

Decision 2.0 Nox

4B

CPU (very slow) / CUDA / Apple (32 GB+)

not benched Β· 12.9 ms GPU (card)

card: JevArena 63.6 (†)

Needs β‰₯20 GB RAM

imajev 2B

2.2B

CPU (slow, 12 GB+ RAM) / Apple (MLX, 8 GB+)

p50 14.8 s / p95 35.1 s CPU

board: JevBench hard 60.4 Β· Img JevBench 68.72 #6 (‑)

Only variant a 16 GB PC runs

imajev 4B

4.3B

CPU (very slow, 24 GB+) / Apple (MLX, 16 GB+)

bench pending

board: JevBench 67.37 #1 Β· Img 76.39 #1 Β· DecisionBench 79.65 #3 (‑)

Board #1 text decision model

imajev 9B

9.4B

CPU (48 GB+) / Apple (MLX, 32 GB+)

bench pending

board: JevBench hard 69.4 (‑)

Heavyweight Qwen3.5 specialist

TypeSafe Jev

Hosted

Native /v1/systemone

~120 ms

DI 54.0 (Reference)

Best known decision intelligence

OpenRouter

Hosted

typesafe/jev-router

~150 ms

Jev DI 54.0

Free API key tier available

DI (Decision Intelligence) benchmarks cited from the official Kev 1.0 test suite. Measured local performance available in docs/benchmarks/.

(†) Decision 2.0 numbers are from the vendor's model cards (vllm-sr, 2026-10) on their own JevArena index β€” a different scale from the held-out breadth-v1 numbers above, so the two are never cross-compared (ADR-0008/0010). Measured CPU latency lives in docs/benchmarks/. On Apple Silicon, Decision 2.0 runs via the plain CPU path (no MLX build yet; MPS unvalidated), so Kev 0.8B (MLX) stays the recommended fast local pick there.

(‑) imajev numbers are JevBench board standings (2026-09) and repo runs on their own indices β€” a third scale, never cross-compared with kev's breadth-v1 or the vllm-sr card (ADR-0008/0011). Measured CPU latency lives in docs/benchmarks/. On Apple Silicon imajev runs a real MLX fast path (fp16), unlike Decision 2.0 there. Downloads show live byte + speed progress bars.


πŸ”’ Security, Privacy & Reliability

  • Strict Loopback Binding: Local model servers bind only to 127.0.0.1. No external ports are ever opened.

  • Protected Secrets: API keys are saved with strict 0600 file permissions in ~/.local/state/snapdec/ or read from environment variables. They are never written into agent configuration files.

  • Fail-Closed Design (NFR-4): If a backend crashes, drops connection, or times out, snapdec returns a valid envelope with decision: "review". It never throws an unhandled exception or interrupts your agent session.

  • Zero Telemetry: No usage stats, prompts, code snippets, or user data are ever tracked or phoned home.


πŸ› οΈ CLI Command Reference

Command

Description

snapdec init

Interactive system setup wizard (auto-detects hardware and agents).

snapdec doctor

Comprehensive health check of daemon, backend, and agent registrations.

snapdec update

Check PyPI and upgrade snapdec installation safely.

snapdec daemon [start|stop|status|logs]

Manage the background decision daemon.

snapdec agents [list|add|remove|print-snippet]

Manage coding agent integrations.

snapdec models [--all]

List all runnable models matching current hardware.

snapdec bench [--backend mock]

Run the decision accuracy and latency benchmark suite.

snapdec project-facts [path]

Deterministic repository fact extraction.

snapdec classify --input -

CLI classifier reading JSON from stdin.

snapdec check --input -

CLI verification reading JSON from stdin.

snapdec score --input -

CLI ordinal scoring tool reading JSON from stdin.

snapdec rank --input -

CLI ranker reading JSON from stdin.

snapdec uninstall

Cleanly reverses all agent integrations and removes configurations.


❓ Frequently Asked Questions (FAQ)

Frontier LLMs incur input token costs for every turn in their conversation history. In agent loops, repeatedly passing long logs, diffs, and lists into a 200k context window to ask small classification or boolean questions consumes significant tokens and compute. snapdec offloads these discrete questions to a local model or fast router, returning only the concise answer and confidence score.

Yes. When you choose a local model (such as Laya or Kev), all dependencies, weights, and runtimes run on your local machine on 127.0.0.1. No internet connection is required after initial model download.

The MCP shim features built-in self-healing: if the daemon is stopped or crashes, the shim automatically re-spawns it on the next incoming tool call. If the backend fails to recover, it returns a safe decision: "review" envelope so the agent continues operating without crashing.

Simply run:

snapdec uninstall

This restores all agent configuration files from their original backups and cleans up registered skills.


πŸ“œ Architecture & Decisions


πŸ“„ License

Distributed under the Apache-2.0 License. See LICENSE for details.

snapdec routes to and credits Kev, Laya, TypeSafe Jev, the vllm-sr Decision 2.0 family, and the imajev family (pinned @ ccf586d4). Model weights are distributed under their respective Apache-2.0 licenses.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables coding agents to scout, rank, and preflight software work before implementation, returning evidence-backed ACT, VERIFY, or SKIP decisions for issues and pull requests.
    98 npm
    2
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Enables coding agents to make cheap, fast probabilistic decisions on every turn, with tools for coding-loop checks, review, verification, screening untrusted input, and ranking candidates.
    6
    470 npm
    61
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Enables coding agents to query files, logs, and search results through TypeSafe Jev's typed, probabilistic answers, so they retrieve only the needed conclusion instead of raw context. Supports classification, scoring, yes/no checks, extraction, page ranking, and injection screening.
    11
    470 npm
    MIT