local-llm-mcp
local-llm-mcp is an MCP server that lets a cloud/calling model hand work to a local LLM worker, returning a scrubbed or digested result so private data stays off-cloud and context stays small.
Run shell commands locally and get a digest —
local_llm_runexecutes commands on the host and returns a summary instead of raw output (logs, tests, git, listings, etc.).Delegate tasks over material —
local_llm_delegateprocesses inline text, files/directories, or command output: summarise, extract, parse, calculate, draft, or research.Read exact slices of raw output —
local_llm_artifactfetches stored raw output by line or character range, still scrubbed.Switch between privacy modes —
local_llm_set_modetoggles PII mode (fully sanitized placeholders) and ASSIST mode (only secrets scrubbed).Turn the server on only with user consent —
local_llm_enableshows a dialog before any work is delegated.Control what personal data may be shown —
local_llm_disclosurerecords whether identity/number values may appear in clear in ASSIST mode.Compact the worker's memory —
local_llm_compactforces the worker to compress its running conversation summary.Check session state and savings —
local_llm_statusreports mode, session identity, turn counts, placeholder counts, compaction info, and token/dollar savings estimates.Keep private values as reusable placeholders — emails, credentials, names, etc. are replaced by stable tokens like
[EMAIL-1], which the server re-expands server-side when reused.Enforce a command safety policy — deny dangerous commands, apply an allow list, and ask the user before running anything else.
Verbatum quoting — retrieve code/config/error text byte-for-byte when you need to reproduce it exactly.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@local-llm-mcpSummarize this shell output and replace any private values with placeholders."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
local-llm-mcp
A local worker model for your cloud model. local-llm-mcp is an
MCP server that sits between a calling model
(Claude Code, or any MCP client) and an LLM running on your own network. The caller
hands work over; the local model does it; the caller gets back a result that is either
sanitized (private values replaced by stable placeholders) or digested (long
raw output reduced to what the caller needs). The local model keeps its own running
memory of the conversation and compacts it whenever the caller compacts its own.
It exists for two reasons:
Privacy. Some material must never reach a cloud model: credentials, keys, personal mail, contacts, addresses, account numbers, the contents of
.envfiles. In PII mode the caller is told to delegate every such step; the worker reads the real data and the caller only ever sees[EMAIL-1],[PERSON-2],[SECRET-3].Context economy. A cloud model's context is expensive and finite. Most of what a shell command prints is noise to the caller. In ASSIST mode the caller is told to route any context-free step whose raw output would exceed the digest it needs: logs, test runs, listings,
gitoutput,grepresults, file reads, calculations.
The server is deployment-agnostic: endpoint, key, detection rules, private terms and attribution are all configuration. Nothing about any particular network lives in the package.
Contents
Related MCP server: Doc Sanitizer MCP Server
How it works
caller (Claude Code) local-llm-mcp (stdio child) local model
──────────────────── ─────────────────────────── ───────────
tool call ──────────────────────────────► 1. rehydrate placeholders in task/command
2. gather material (run command / read files)
3. register detectable values in the vault
4. PII mode: entity pass ─────────────────────► "which private
register what the worker found values are here?"
5. build prompt: rules + session memory +
task + material ───────────────────────────► answer
6. draft revises itself? one finalize call ───► committed answer
7. SCRUB the answer (rules · secret shapes ·
identity shapes · private terms · every vault value)
8. store artifact (raw) + turn (scrubbed)
digest + trailer ◄──────────────────────── 9. return
...
PreCompact hook ── unix socket ──────────► compact own memory (background) ─────────────► new summaryEvery value that goes to the worker or the shell is rehydrated; every value that leaves the server is scrubbed. That asymmetry is the whole design.
There is a second path beside the digest: verbatim mode. A digest paraphrases, and a
caller that needs code or configuration text cannot act on a paraphrase. With
verbatim=true the worker is shown numbered lines and asked only which line ranges the
task refers to; the server then copies those lines out of the material exactly (still
scrubbed on the way out). The model's judgement chooses the region; it never rewrites a
character of it.
Turning it on — nothing happens without the user's say-so
The server never starts intercepting on its own. At connect time the model sees only the
offer: what the server can do and that it is off until the user turns it on. The first time
the model calls a working tool — or calls local_llm_enable to offer it — the server asks the
user directly through MCP elicitation (Claude Code shows a dialog naming the tool that was
attempted, never its arguments): turn it on for this session, turn it on in PII mode, or keep it
off. Only then does anything run; the result of that first call carries the full operating rules.
off is remembered for the session; further calls return a refusal and the model is told not to ask again. A cancelled dialog refuses that call and may ask again later (three times at most).
A client that cannot show a dialog is refused, with the text telling the model how the user can enable the server:
LOCAL_LLM_MCP_ARM=onin that client's MCP configuration for the server (the user's standing approval for that client), or the turn on switch on the session's row in the admin app. Those two, and the dialog, are the only ways the gate is ever passed.local_llm_statusis the one tool that answers while off (it shows the gate's state).The dialog is default-safe by construction, whatever the client does with it: the answer field is required, has no default, and lists off first, so a client that auto-accepts a form (empty, with defaults, or with its first option) yields "not answered" or "off" — never "on". Only an explicit on or on_pii chosen by the user arms the server.
The command policy — what may actually run
Turning the server on is one decision; what it then executes is another. Every command a caller
hands to local_llm_run (or to local_llm_delegate's command) runs with your privileges,
and a client in a bypass or auto permission mode never shows you a per-tool prompt — so the
server holds its own line, in three layers:
Deny shapes — refused outright, nothing run, the reason returned to the caller: privilege escalation (
sudo,doas,su,pkexec), power state (shutdown,reboot,halt), filesystem writes (mkfs*,fdisk,parted,wipefs,dd of=), recursiverm, recursivechmod/chown, package installs,shred, a pipe fromcurl/wgetinto a shell, a write to a real device under/dev/, and any write under/etc,/bootor~/.ssh(reading them is fine). A command containing a placeholder that would rehydrate to a secret is refused too — a secret is never handed to a shell.Your allow list — one glob per line in
~/.config/local-llm-mcp/run-allow.txt(LOCAL_LLM_MCP_RUN_ALLOW). These run without a question: the standing approval for the shapes you delegate every day.git log* journalctl* ls* rg* pytest*Every segment of a command line must match an entry, so approving
ls*does not approvels; something-else. Approvinggit log*does not approvegit push*: for a multiplexer (git, docker, systemctl, uv, npm…) the subcommand is part of the shape.Everything else asks. The same MCP elicitation the gate uses, showing you the exact command and its working directory: refuse / run once / always allow this shape. "Always" appends the shape to the allow file, so the file stays the record of what you approved and you can edit it by hand. Like the turn-on dialog it lists refuse first and has no default, so a client that auto-accepts a form can never approve a command.
No allow entry, and no dialog answer, can exempt a deny shape — layer 1 is checked first,
always. A client that cannot show a dialog gets layer 1 and a trailer line saying it was not
asked. Every decision is stamped on the turn and in the trailer (policy: allow-list (ls*),
policy: asked:run, policy: refused: recursive rm), counted in local_llm_status.run_policy,
and shown as a chip on the session's row in the admin app.
LOCAL_LLM_MCP_RUN_POLICY chooses how much of this runs: ask (default), allow (layer 1
only — never ask), or off (no policy at all).
This is deliberately not a sandbox. It reads command position, quoting and redirection well enough to judge shapes; it does not follow a script it invokes, a variable it expands, or an alias. It is a second lock on a door you already opened.
The two modes and when the caller calls
The server tells the caller, in its MCP instructions (returned at initialize and
again by local_llm_set_mode), how often and under what circumstances to call. The
active mode's rule is written first.
PII mode | ASSIST mode | |
Route here | every step that might touch private data: names, home/postal addresses, phone numbers, personal email, family details, account/card/bank numbers, credentials, API keys, tokens, vault contents, | every step that needs no conversational context and whose raw output would exceed the digest: a direct call returning more than ~40 lines / 2 KB — logs, journals, test/build/lint runs, listings, |
Do directly | anything with no private data | calls that print little or nothing ( |
What comes back | fully sanitized: placeholders for every private value; secrets are used server-side and never returned | the result edited for clarity and usefulness; only secret shapes are scrubbed |
When unsure | delegate | keep exact counts/sorts in the command ( |
The mode is a server setting (LOCAL_LLM_MCP_MODE) and can be switched per
conversation with local_llm_set_mode. Claude Code truncates server instructions at
2048 characters; both texts fit.
Tools
All tools return text; every result ends with a trailer:
— local-llm · pii · turn t_ef33d4 · ref a_1606e7eb · rc 0 · raw 4471 chars · finalized · local · 1.8s · scrubbed 2 EMAIL, 1 SECRETturn is the id in the session memory, ref the stored raw material, rc the exit
code (for commands), finalized marks a draft folded into a committed answer.
local_llm_run
Run a shell command on the host (bash -c, this user's privileges, stdin closed,
output capped) and receive a digest of its output. Raw output is kept as an artifact.
Parameter | Type | Default | Meaning |
| string | required | the command; placeholders ( |
| string | outcome, errors/warnings verbatim, key values | what to report |
| string | server cwd | working directory |
| int |
| kill after |
| int |
| soft digest budget |
| bool |
| copy the matching output lines byte for byte instead of digesting — only for text you will reproduce or edit; never for questions, counts, summaries or listings. If nothing matches, the turn is digested instead and the trailer says so |
A timeout or non-zero exit is reported inside the digest and in the trailer.
local_llm_delegate
Hand a task to the worker over material you name: inline text, files or directories
(listed; ~ expands; missing or unreadable paths are reported inline), and/or the output
of a shell command.
Parameter | Type | Default | Meaning |
| string | required | what to do — what to extract, thresholds, output shape |
| string |
| inline text |
| list[string] |
| files or directories to read |
| string |
| shell command whose output is added to the material |
| string |
| base for relative paths and the command |
| int | server default | soft answer budget |
| bool |
| copy the matching lines byte for byte instead of answering — only for text you will reproduce or edit; never for questions, counts, summaries or listings |
Without material the worker answers from the task and its session memory. In verbatim
mode the result is one or more source lines a-b: blocks with numbered lines; ranges
that did not fit the budget are named in the trailer so you can fetch them with
local_llm_artifact.
local_llm_artifact
Return an exact slice of a stored raw output by ref, scrubbed (fully in PII mode,
secrets only in ASSIST mode). Line mode: line_start / line_end (1-based, inclusive,
numbered output; line_end defaults to 200 lines). Character mode: offset / limit
(default 4000, max 20000). The trailer says where the next slice starts.
local_llm_status
Mode, session key, the caller's session id, model, endpoint, turn counts, context size, compaction count and time, placeholder counts by kind, artifact count, the scrubbed summary, control-socket path, observer name.
The block also carries tokens_saved: the estimate for this session and for all sessions (see Tokens saved).
local_llm_enable
local_llm_enable() offers the server to the user: the server shows the dialog and returns the
operating rules if the user turns it on, or a refusal. The model should call it once, when a step
would print far more than it needs or would touch private data, and never again in a session where
the user said no. Every other working tool asks the same question on its first use.
local_llm_disclosure
local_llm_disclosure(identity="open"|"masked"|"ask"|"keep", numbers="open"|"masked"|"keep", reason="")
records what this ASSIST session may show in clear, for when the user has said so in the
conversation; identity="ask" resets the question. Returns the resulting state. Secrets are
never shown; PII mode masks everything regardless. See Disclosure.
local_llm_compact
Compact now. Normally automatic; useful when a client has no PreCompact hook.
local_llm_set_mode
Switch pii ↔ assist for the rest of the conversation; returns the instructions for
the new mode so the caller learns the rule in-context.
The privacy boundary
Placeholders
Every detected private value maps to a placeholder that is stable for the whole
conversation: [EMAIL-1] is the same address every time it appears, in every
result, artifact slice and summary. Kinds: PERSON, ADDRESS, PHONE, EMAIL,
SSN, CARD, ACCOUNT, SECRET, PII.
The map lives in the session vault (placeholders.json, mode 0600) and is never
returned. Placeholders are re-expanded server-side when the caller reuses them:
"reply to [EMAIL-1]" reaches the worker with the real address; curl -H 'Authorization: Bearer [SECRET-1]' … runs with the real token. The caller never holds
the value.
Detection — layered, deterministic
Layer | What it catches | Source |
shape rules | emails, phone numbers, national ids, payment cards (Luhn-checked), IBANs, key prefixes, PEM headers | built in, or your rules file ( |
secret shapes | PEM blocks, JWTs, bearer headers, well-known token prefixes, | built in |
identity shapes (PII mode) | what identifies a person by shape alone, no seed data: street addresses and PO boxes, | built in; bounded regexes, low milliseconds per 100 KB; |
private terms | the operator's own literal names, addresses, numbers — an unlabelled name has no shape | terms file ( |
entity pass (PII mode) | whatever private values the worker itself finds in the material: names, addresses, dates of birth, credentials it recognises in context | one extra local call per delegation; only exact substrings of the material are accepted, never a label or a variable name |
known values | every value already in the vault, exact-matched | automatic |
answer pass (while identity is masked) | after the scrub, the worker is asked what private values the OUTBOUND text still holds — every result, every compaction summary and every artifact slice, in PII mode and in ASSIST until the user opens identity; the text is short, so this covers it whole, unlike the material pass (first two chunks); anything found is registered and the text re-scrubbed; the trailer says | one short local call per outbound text |
Order of operations matters:
Everything detectable in incoming material is registered in the vault before the worker sees it. An echo in the answer is then caught by exact match even in a context no rule would recognise.
The worker is instructed to write placeholders itself and never quote a secret.
The answer is scrubbed (rules · secret shapes · identity shapes · terms · every vault value), then re-scanned; anything still detectable becomes
[REDACTED]. In PII mode the worker is then asked what private values the scrubbed answer still contains, and those are masked too (the answer pass). A privacy boundary cannot rest on a model choosing to comply, so step 2 is help, not the guarantee; the deterministic layers and the answer pass are.
Disclosure — what may come back in clear
Mode never decides where work runs: everything runs on the local worker in both modes, and a delegation by shape ("read this PDF", "list the accounts") lands here either way. Mode decides what comes back to the caller. Three tiers:
Tier | Kinds | PII mode | ASSIST mode |
secrets | keys, tokens, passwords, PEM blocks | masked | masked, always |
numbers | SSN, card, account and id numbers, dates of birth | masked | masked by default ( |
identity | person, address, email, phone | masked | open once the user says so |
The first time identity or number values appear in an ASSIST session, the server asks the
user directly through MCP elicitation (Claude Code shows a dialog; the worker keeps working
meanwhile): switch the session to PII mode, continue with identity values shown, or
show everything except secrets. The answer is kept for the session, shown in
local_llm_status, stamped into that result's trailer (disclosure: asked → continue) and
listed in the admin app. Cancelling the dialog masks that result and asks again next time,
three times at most. Choosing "switch" also vaults the values already written into the
session's memory, so they leave masked from then on.
Bypasses, in order of directness: the dialog itself; local_llm_disclosure(identity=…, numbers=…), which the caller invokes when the user has said what they want in the
conversation ("just show me the names"), so no dialog is needed; and the deployment default
LOCAL_LLM_MCP_ESCALATION: ask (default), auto (no dialog: mask and tell the caller,
also what a client without elicitation gets) or off (identity values pass as they always
did). Nothing opens secrets.
The high-entropy rule is PII-only, so git SHAs and ids stay exact in ASSIST digests. In
ASSIST, everything detectable is still registered in the session vault (it is local), which is
what makes a later switch retroactive and lets the caller reuse [CARD-1] in a follow-up
command that needs the real digits.
The endpoint must be local
LOCAL_LLM_MCP_BASE_URL must resolve to a loopback, RFC 1918 private or CGNAT
(100.64.0.0/10, e.g. a tailnet) address. Anything else is refused at startup, and
there is no override. This is the one setting that is not the operator's to relax.
Session memory and compaction
The server is a stdio child of the caller and is told nothing about which conversation it serves. Two small hooks close that gap; the server copes without them.
Identity. The SessionStart hook walks its own process ancestry to the Claude
Code process and writes <state>/pidmap/<pid> = the session id; the server walks
its ancestry and adopts the mapped id, migrating a pid-keyed directory onto it. A
resumed conversation (new process, same id) lands on the same memory. Because the hook
can fire before the server has opened its socket, a pid-keyed session re-checks the
map on every tool call until it is named.
Memory. Each tool call appends a turn: kind, task, source, a scrubbed excerpt of
what the worker saw, the scrubbed result, ref, sizes, timing. Every prompt to the
worker carries the compacted summary plus as many recent turns as fit
(LOCAL_LLM_MCP_CONTEXT_CHARS), rehydrated for the worker's eyes only.
Compaction. The PreCompact hook sends {"op":"compact"} to the control socket
($XDG_RUNTIME_DIR/local-llm-mcp/<pid>.sock, 0600). The server queues a compaction
and answers immediately, so the caller's own compaction is never delayed; the worker
then rewrites its summary (goals, decisions, facts with exact identifiers, placeholders
in use, open items) under a lock, and the summary is scrubbed like everything else. The
server also compacts on its own when uncompacted turns exceed
LOCAL_LLM_MCP_AUTO_COMPACT_CHARS, and local_llm_compact does it on demand.
Control socket ops: compact (queue), session (announce the session id),
status. One JSON object per line in, one out.
Tokens saved
Every counted turn records what the server gathered on the caller's behalf (command output, files) and what it returned (the digest). From those two sizes the server estimates what the caller did not have to spend, per session and across all sessions:
Quantity | Estimate | Why |
avoided input | tokens(gathered) − tokens(returned) | the caller read the digest instead of the raw output; inline material it pasted was already in its context and counts for nothing |
avoided output | tokens(returned) × | for a delegation the worker wrote the answer the caller would otherwise have generated (credit 1.0); a digest of command output (0) and verbatim quoting (0) replace reading, not writing |
carried | avoided input × | every token that enters the caller's context is re-sent on each later model call until compaction; priced at the model's cache-read rate where it has one, else at the input rate |
Token counts are exact where the worker can count: after each result is returned, the server asks
the worker's own tokenize endpoint (vLLM's POST /tokenize, or llama.cpp's) for the token count
of what it gathered and of what it returned, in the background, and records both in the ledger. A
worker with no such endpoint, or a turn that outlives the process, falls back to characters ÷
chars_per_token; the report says how many turns were measured. The counts are the worker
tokenizer's, not the caller model's, so each price row may carry a tokenizer_factor: that
model's tokens per character relative to the worker's, applied at pricing time only (the ledger
stays in worker tokens). Seeded at 1.3 for the Claude 4.7+ tokenizer family from Anthropic's own
note that it yields roughly 30 % more tokens, 1.0 (assumed equal) everywhere else; editable per
row in the admin app and the override file. The status block and the admin table show which
factor priced the headline. Measure it instead of assuming it: tools/measure_tokenizer.py
counts public text only (this repository's README, source and price table, plus synthesized
listings, process tables and journal lines) with the worker's tokenizer and with Anthropic's
token-count endpoint, prints the ratio per text class, and with --write sets each row's factor
in the override file from the measurement. Nothing private ever leaves the machine for this.
Dollars are list prices per 1M tokens for flagship models of every major provider, read
from the providers' own pricing pages and stamped with the date they were checked
(local_llm_mcp/prices.json). They are an equivalent, not an invoice: on a subscription the
figure is the usage kept out of the plan's limits; batch, flex and long-context tiers are
ignored (the context_note names them).
Where it shows:
every result trailer carries
saved ≈ N tok(avoided input + avoided output for that turn);local_llm_statusreturns atokens_savedblock — this session, all sessions, the headline model's dollar figure, and a per-model table;the admin app's Tokens saved tab: all-time tiles, by model, by session, and an editor for the assumptions and prices.
Everything is configurable and applies retroactively, because the ledger stores characters,
not tokens: chars_per_token (only for unmeasured turns; default 3.8 — deliberately conservative: Anthropic states that Claude 4.7
and later use a tokenizer producing roughly 30 % more tokens for the same text, so for those models the
true figure is nearer 3, and the estimate understates the saving), context_reuse_calls (default 8), the output
credits, and the headline model (caller_model, or LOCAL_LLM_MCP_CALLER_MODEL). Edits go to
the override file LOCAL_LLM_MCP_PRICES (default ~/.config/local-llm-mcp/prices.json),
which merges over the package defaults by model id, can add models or disable one
("enabled": false), and is hot-reloaded. The admin tab writes that file; "reset" removes it.
Prices go stale: the tab, the status block and --check flag the table once its checked date is
older than stale_after_days (default 30), so re-check the providers' pages and save.
The cross-session ledger is <state>/savings.jsonl, append-only, one line per counted turn;
it survives session purges. Sessions recorded before the ledger existed can be added from their
context logs with local-llm-mcp-admin --backfill-savings (or the button on the tab); the
gathered size of those older turns is inferred from the turn's raw size and source. Idempotent.
Install and wire into Claude Code
Requirements: Python 3.12+, a local OpenAI-compatible chat endpoint (vLLM, llama.cpp server, LM Studio, Ollama's OpenAI route, …).
uv tool install git+https://github.com/lealvona/local-llm-mcp # or pipx; or uv pip install into a venv you manage
mkdir -p ~/.config/local-llm-mcp
cat > ~/.config/local-llm-mcp/env <<'EOF2'
LOCAL_LLM_MCP_MODE=assist
LOCAL_LLM_MCP_BASE_URL=http://127.0.0.1:8000/v1
LOCAL_LLM_MCP_MODEL=local
LOCAL_LLM_MCP_API_KEY_FILE=$HOME/.config/local-llm-mcp/model.key
EOF2
chmod 600 ~/.config/local-llm-mcp/env
local-llm-mcp --check # prints effective config, rules source, session key, and probes the model
claude mcp add --scope user local-llm -- local-llm-mcpHooks — add to ~/.claude/settings.json (both entries run the same stdlib-only
script, exit 0 always, cost milliseconds):
{
"hooks": {
"SessionStart": [{"hooks": [{"type": "command", "command": "/usr/bin/python3 /path/to/local-llm-mcp-hook.py", "timeout": 5}]}],
"PreCompact": [{"hooks": [{"type": "command", "command": "/usr/bin/python3 /path/to/local-llm-mcp-hook.py", "timeout": 5}]}]
}
}hooks/local-llm-mcp-hook.py ships in the repo; copy it wherever you keep hooks.
Verify with claude mcp list (should say Connected) and, in a new session, the
local_llm_status tool.
Other MCP clients: run local-llm-mcp over stdio. Without the hooks, sessions are
keyed by the parent process id and compaction is manual or automatic-by-size.
Using it from other MCP clients
The server is a plain stdio MCP server; nothing in it requires Claude Code. Every client below was exercised against this code (protocol clients with the reference tooling, agents with a real model turn) — what changes from client to client is only what you lose without Claude Code's hooks, listed after the table.
Client | Register | Verified |
Claude Code |
| full: hooks, dialog, compaction trigger |
Codex CLI |
| agent turn: tools called and digested; it does declare elicitation. Note the non-interactive caveat below |
Hermes Agent | in | agent turn: Hermes called |
Kimi Code CLI |
| agent turn: |
opencode |
| see below |
Claude Desktop, Cursor, Windsurf and most others |
| same wire protocol as the Inspector run below |
MCP Inspector (reference client) |
|
|
Python SDK |
|
|
Pass configuration the same way for every client: put it in the dotenv the server reads
(~/.config/local-llm-mcp/env), or in the client's per-server env.
What you lose without Claude Code's hooks, and what replaces it:
Session identity. Claude Code's SessionStart hook maps the conversation id onto the server; elsewhere the session is keyed by the client process that spawned the server (
pid-<pid>), one per conversation, or byLOCAL_LLM_MCP_SESSIONif the client sets it. The pidmap is read only when Claude Code is the process that spawned the server, so another harness launched from inside a Claude Code session gets its own session rather than inheriting that conversation's vault, memory and answer to the gate. The session records which client it belongs to (clientinlocal_llm_statusand the admin sessions table).Compaction trigger. The PreCompact hook tells the worker to compact when the caller does; without it the worker still compacts itself at
LOCAL_LLM_MCP_AUTO_COMPACT_CHARSand onlocal_llm_compact.The connection instructions. Some clients never show the model the
instructionsfield ofinitialize. The same text is therefore also an MCP prompt (local_llm_instructions) and a resource (local-llm://instructions), and the routing rules are written into the tool descriptions, which every client shows. For a harness that reads an instructions file (AGENTS.md, a system prompt), paste the ASSIST or PII text fromlocal_llm_status.A non-interactive run cannot answer a dialog. A client may declare elicitation and still auto-decline every question when it has no user in front of it —
codex execdoes, and so does Claude Code's print mode. That is the right direction to fail: the gate stays off and no command runs. For an unattended run, give that client its standing approvals explicitly in its own server config —LOCAL_LLM_MCP_ARM=on, plus the command shapes inLOCAL_LLM_MCP_RUN_ALLOW(orLOCAL_LLM_MCP_RUN_POLICY=allow). Both are needed: they are separate decisions, and each dialog is refused on its own.The turn-on, disclosure and command dialogs need a client that supports MCP elicitation (Claude Code and Codex do; Kimi Code 0.37 does not). Without it, the server refuses to start until the user sets
LOCAL_LLM_MCP_ARM=onin that client's server configuration (their explicit approval); disclosure falls back toauto(masked, with a note telling the model to calllocal_llm_disclosurewhen the user decides); and the command policy falls back to its deny shapes alone, with a trailer line saying the user was not asked — so in such a client the allow list is the only per-command control, and it is worth filling in. Measured on Kimi Code 0.37: refused as designed, then withARM=onand one allow entry it returned the digest and the full trailer,policy: allow-list (echo*)included.
One server process serves one conversation; that is the isolation model behind the vault. Do
not put it behind a multi-user MCP gateway (a shared mcpo instance serving a chat UI, say): every
user would share one placeholder map, and local_llm_run runs shell commands as the user that
owns the process.
Configuration reference
The process environment wins over the dotenv file
(~/.config/local-llm-mcp/env, or LOCAL_LLM_MCP_ENV_FILE); the file accepts
KEY=VALUE, export KEY=VALUE, quotes, # comments and $HOME.
Variable | Default | Meaning |
|
|
|
|
| OpenAI-compatible endpoint; must be local |
|
| model name sent to the endpoint |
| unset | a second local worker used when the primary is unreachable or answers 5xx (must be local too) |
| primary model | model name at the fallback |
| unset | its key |
|
| seconds the fallback is asked first after the primary fails |
| unset | bearer for the endpoint (file: one line, keep it 0600) |
|
| pass |
|
| default digest budget; |
|
| ceiling on |
|
| seconds per model call |
|
| material above this is processed map-reduce style |
|
| concurrent chunk calls |
|
| cap on gathered material (head 70 % + tail 30 % kept, marker in between) |
|
| budget for summary + recent turns in each prompt |
|
| scrubbed excerpt of material kept per turn |
|
| compaction target |
|
| self-compaction threshold on uncompacted turns |
|
| default command timeout, seconds |
|
| the command policy: |
|
| commands that run without asking, one glob per line; the "always" answer appends here |
|
| treat every bare 10-digit run as a phone number (default: only formatted numbers or ones near phone words) |
|
| PII mode: ask the worker for the private values before answering |
|
| the opt-in gate: |
|
| first identity/number values in an ASSIST session: |
|
| whether account/card/id numbers and dates of birth are shown in ASSIST before the user decides |
|
| seconds to wait for the disclosure dialog before masking that result |
|
| PII mode: the identity-shape layer (addresses, labelled names, dates of birth, labelled id numbers) |
|
| override file for the tokens-saved estimate (assumptions + prices; the admin app writes it) |
| unset | headline model for the dollar figure (defaults to the prices file's |
| built-in | JSON rules file (see below) |
|
| terms file (see below) |
|
| sessions, pidmap |
|
| control sockets; AF_UNIX paths cap at 108 bytes, keep it short |
| unset |
|
| unset | webhook target (also available to your own observer) |
| unset | directory added to |
| unset | force a session key (tests, scripts) |
|
| admin app listener |
| unset | bearer token the admin API requires (set it whenever the bind is not loopback — a non-loopback bind with neither this nor the override below set refuses to start) |
|
| start anyway on a non-loopback bind with no token |
|
| stderr logging |
CLI: local-llm-mcp [--mode pii|assist] [--session KEY] [--check].
Rules file
LOCAL_LLM_MCP_RULES points at a JSON file that replaces the built-in shape
rules. The schema is deliberately compatible with a privacy-routing rules file another
tool might already maintain — other top-level keys are ignored, so one file can serve
several tools:
{
"regex_rules": [
{"id": "email", "pattern": "[A-Za-z0-9._%+-]{1,64}@[A-Za-z0-9.-]{1,255}\\.[A-Za-z]{2,24}"},
{"id": "phone", "pattern": "(?<![\\d.-])(?:\\+\\d{1,3}[ .-]?)?(?:\\(\\d{3}\\)|\\d{3})[ .-]?\\d{3}[ .-]?\\d{4}(?![\\d.-])"},
{"id": "credit_card", "pattern": "\\b(?:\\d[ -]?){12,18}\\d\\b", "luhn": true},
{"id": "employee_id", "pattern": "\\bEMP-\\d{6}\\b", "enabled": true}
]
}The rule id decides the placeholder kind: email→EMAIL, phone→PHONE,
us_ssn→SSN, credit_card→CARD, iban→ACCOUNT; ids containing key, token,
secret, private, password or credential→SECRET; mail/phone/card/ssn/
account/bank/routing by substring; anything else→PII. The file is hot-reloaded
on change. Keep quantifiers bounded — an unbounded email pattern backtracks
quadratically on long unbroken runs.
Built-in acceptance filters apply to any rules file: a bare ten-digit phone match in
non-strict mode needs formatting or a nearby phone word (timestamps and ids are not
phone numbers); a credit_card match shaped like YYYYMMDD-HHMMSS is a build id, not
a card.
Private terms
{
"terms": [
{"kind": "PERSON", "values": ["Jane Q. Example", "J. Example"]},
{"kind": "ADDRESS", "values": ["12 Example Lane"]},
{"kind": "PHONE", "values": ["+1 555 010 0000"]},
{"kind": "PII", "values": []}
]
}Case-insensitive, whole-word, longest first, hot-reloaded. Put the names, addresses and numbers of the people this deployment protects here; keep usernames and hostnames out (they appear in every path). This is the deterministic floor for identity that no pattern can infer; the entity pass adds what the worker finds on top.
Observers
Observers receive scalar-only events — tool name, turn id, ref, sizes, seconds, exit code, scrub and entity counts, session key — never task text, material or results. Three options:
unset — no observer;
LOCAL_LLM_MCP_OBSERVER=webhookwithLOCAL_LLM_MCP_OBSERVER_URL— one JSON object per event:{"event": "start"|"event"|"end", ...};LOCAL_LLM_MCP_OBSERVER=mymodule:MyObserverwithLOCAL_LLM_MCP_OBSERVER_PATH=/dir— your own subclass, kept outside this package:
# /dir/mymodule.py
from local_llm_mcp.observers import Observer
class MyObserver(Observer):
name = "mine"
async def start(self, info: dict) -> None: ... # session, claude_session_id, mode, model, endpoint
def event(self, span: str, kind: str, attrs: dict) -> None: ... # must not block
async def end(self, state: str = "done") -> None: ...
@property
def run_id(self) -> str: return "" # shown in local_llm_statusA broken or unreachable observer never fails a tool call; the server logs and continues.
Admin app
local-llm-mcp-admin serves a small single-page app plus a JSON API for the PII layer —
the parts the caller never sees:
Tab | What it shows (glance) | What it does (delve) |
Overview | counts: terms by kind, sessions and how many are live, placeholders by kind, rules source, entity pass | each tile opens its tab |
Terms | the private terms by kind, values masked (hover or reveal values) | add a term; remove one; harvest candidates from pasted text with the local model and tick the ones to keep |
Sessions | one row per conversation: live dot, mode, last seen, turns, compactions, placeholder counts, artifacts, size | open a row: the placeholder vault (placeholder → kind → value), compacted memory, recent turns, artifacts; purge placeholders, delete artifacts, delete the session — refused while a server holds it |
Scrub tester | paste text, pick a mode | the text as it would leave the server, every detected span highlighted by kind, and the placeholders that would be minted |
Rules · Config | the shape rules in force and their source; the effective configuration (key shown as set/unset) | read-only |
local-llm-mcp-admin # http://127.0.0.1:8631/ (loopback, no token)
LOCAL_LLM_MCP_ADMIN_BIND=0.0.0.0 LOCAL_LLM_MCP_ADMIN_PORT=8631 \
LOCAL_LLM_MCP_ADMIN_TOKEN_FILE=~/.config/local-llm-mcp/admin.token local-llm-mcp-adminIt reads the same dotenv, state directory and terms file as the server (terms are
hot-reloaded, so an edit here applies to running servers at once). It shows private
values to whoever reaches the port: it binds to loopback by default, and when bound
wider it should sit behind a firewall and a token (LOCAL_LLM_MCP_ADMIN_TOKEN or
_TOKEN_FILE; the page asks once per tab) — a non-loopback bind with no token set
refuses to start rather than serving the vault to whoever finds the port; set
LOCAL_LLM_MCP_ADMIN_ALLOW_NO_TOKEN=1 if you genuinely want that (a LAN you already
trust, say) and understand what it means. No external assets; ?demo=1 renders the
page with sample data and no server, for a look at the UI.
Keeping a deployment separate from the code
This package carries nothing about any network. A deployment is:
Layer | Where | Contents |
general code | this repo | the package, hook script, tests, tools |
runtime | a venv you install the package into ( | re-installed to update |
instance config |
| dotenv, key files (0600, never committed), |
A minimal sync script for a checkout-based deployment:
#!/usr/bin/env bash
set -euo pipefail
REPO=~/src/local-llm-mcp; INST=~/.local/share/local-llm-mcp
[[ -d $INST/venv ]] || uv venv --python 3.12 "$INST/venv"
uv pip install --python "$INST/venv/bin/python" --reinstall-package local-llm-mcp "$REPO"
install -m 755 "$REPO/hooks/local-llm-mcp-hook.py" "$INST/hooks/local-llm-mcp-hook.py"
"$INST/venv/bin/local-llm-mcp" --checkPoint Claude Code at $INST/venv/bin/local-llm-mcp and the hook copy; edit the repo,
run the script, and new sessions pick up the change. Version the config directory
privately if you like (ignore *.key).
The git gate — nothing personal reaches a commit, a message, or a push
The separation above is enforced at the repository, not by care. tools/install-hooks.sh
points a clone at the tracked hooks in tools/githooks/, and every one of them calls
tools/publiccheck.py:
Hook | Refuses |
| staged content or file names carrying a private network address, a home-directory path, a non-example e-mail, a tailnet name, or a denylisted term; then |
| the same terms in the message; any attribution trailer ( |
| every commit the push would publish, checked in full — identity, message, file names and the whole tree at that commit — so a commit made with |
| nothing; with |
The denylist is a private file the repository never sees — one term per line in
$LOCAL_LLM_MCP_DENYLIST (default ~/.config/local-llm-mcp/publiccheck-denylist.txt):
your user name, host names, project names, model aliases, domains, anything that identifies you or
your network. Terms match case-insensitively on word boundaries, and a space in a term matches a
space, underscore or hyphen, so one entry covers Blue Kettle, Blue-Kettle and blue_kettle.
Hits are reported by line number, never by value, so a refusal can be pasted anywhere.
⚠️ Keep it current: the shape checks catch addresses, but only the denylist catches NAMES. A host
or project name has no shape a regex can find, so a name that is not on your list will pass every
gate — this is the one place the gate depends on you. When publiccheck runs without the list — on
CI, where the file does not exist — it says so instead of reporting "clean", because half the gate
did not run there. python tools/publiccheck.py --all on a machine that has the list is the real
proof, and it reads annotated tags as well as commits.
State on disk
~/.local/state/local-llm-mcp/
pidmap/<pid> caller pid → session id (written by the SessionStart hook)
sessions/<session>/ 0700
meta.json mode, counts, timestamps, identity
context.jsonl turns (scrubbed); a "compaction" turn marks each summary
summary.md the compacted memory (scrubbed)
placeholders.json the vault: placeholder ↔ value, 0600 — never leaves the host
artifacts/a_xxxxxxxx.txt raw material per turn, 0600 (served only through the scrub)
$XDG_RUNTIME_DIR/local-llm-mcp/<pid>.sock control socket, 0600, removed on exitSessions are never deleted automatically; rm -r a session directory to forget it.
Testing and verification
uv sync
.venv/bin/python -m pytest -q # scrubber, vault, rules file, dotenv, entity registration, verbatim helpers, admin API
.venv/bin/python tools/smoke.py # end to end over stdio against your configured model
.venv/bin/python tools/leakcheck.py ~/.secrets/app.env # a REAL secrets file through PII mode; asserts no value leaks
.venv/bin/python tools/piicheck.py # adversarial tasks over SYNTHETIC private data, PII mode + ASSIST default: nothing may come back
ANTHROPIC_API_KEY=… .venv/bin/python tools/measure_tokenizer.py [--write] # tokenizer factors from PUBLIC text, never from turns
python tools/publiccheck.py --all # every commit in the history: nothing personal, no stray identity or trailerThe smoke client exercises: initialize (instructions carry the mode), tool listing, status, a calculation, a command digest, a PII delegation (the email and password never come back), placeholder rehydration through a shell command (an uppercased echo comes back re-scrubbed), artifact slicing, compaction over the control socket exactly as the hook does it, recall of a fact after compaction, verbatim quoting of a function out of a 190-line file, line-mode artifact slicing, a command supplied to delegate, and a mode switch.
leakcheck.py reads a file you name only to build the list of values to guard, runs
a fresh PII-mode server against it, and prints the digest only if none of those values
— not even a 12-character fragment — appears in it.
Threat model and limits
Protects against: private values in material (files, command output, pasted text) reaching the caller, including when the worker paraphrases, uppercases, or quotes them; private values surviving into summaries and artifact slices; and — through the command policy — a caller running a destructive command, or one built around a secret, on a client whose own approval prompts are turned off.
Assumes: the host and the local model endpoint are trusted; the caller is not malicious toward the operator (it is instructed, not sandboxed); the operator lists the identities to protect or accepts the entity pass's recall; and that what they put on the allow list is what they meant to approve.
Does not protect against: a caller that reads private files with its own tools instead of delegating (the instructions say not to; nothing enforces it — put a hook in front of those tools if you need enforcement); names and addresses that are neither in the terms file nor recognised by the worker; secrets shaped like ordinary words with no label; side channels such as the length of a rehydrated value (a command can measure it); the worker's own logs at the endpoint. The command policy is not a sandbox — it judges shapes, not shell semantics, so it does not follow a script it invokes, a variable it expands, or an alias; a command the user approves runs as the user.
Known behaviours: bare ten-digit numbers in command output are ids unless
formatted or near a phone word (LOCAL_LLM_MCP_STRICT_PII=1 flips that); in PII mode,
long mixed alphanumeric runs (git SHAs included) become [SECRET-n]; a variable name
is never treated as a secret, its value always is; a no-thinking model that revises in
the open is folded into a committed answer by a second short call.
Troubleshooting
Symptom | Cause / fix |
| read Claude's log: |
| the endpoint must be local; that is the design |
empty answers / | keep |
the caller ignores the mode rules | instructions are truncated at 2048 chars by Claude Code; check |
| the SessionStart hook is not installed or not on that path; the server still works, keyed by pid, and adopts the id as soon as the hook's mapping appears |
compaction never triggers on the caller's compaction | the PreCompact hook is missing, or the socket dir differs between hook and server ( |
| set |
the digest hides variable names as | a rules file rule is matching the name; names shaped |
everything is |
|
Logs go to stderr (Claude Code captures them); LOCAL_LLM_MCP_LOG_LEVEL=DEBUG for
observer traffic.
FAQ
Why placeholders instead of [REDACTED]? So the caller can keep working: it can
address [EMAIL-1], compare [PERSON-1] and [PERSON-2], and pass [SECRET-1] into
a command — all without holding the value. Referential integrity survives
sanitization.
When should I use verbatim mode? Whenever you will reproduce the text rather than act on a summary of it: a function to copy, a config block to edit, an error to quote. The worker only chooses the lines; the bytes come from the source. A digest is for when you need to know what something says; verbatim is for when you need to have it.
Why does the worker get its own memory instead of the caller's? The caller's context is what we are trying to keep small and clean. The worker's memory is local, cheap, and rehydrated only for the worker; the caller receives a summary only through the scrub.
Why one extra call in PII mode (the entity pass)? Because a regex knows the shape of an email, not a person's name. Asking the worker "what private values are here?" and registering the exact substrings it names turns the model's judgement into a deterministic guarantee — the scrub catches those strings regardless of what the final answer says.
Why refuse non-local endpoints outright? The server's purpose is that private material never leaves the network. A configuration mistake must not be able to defeat that.
Can I use a cloud model as the worker over a tunnel? If it resolves to a private address the check passes, but you would be defeating the purpose. Don't.
Does it work without Claude Code? Yes, with any MCP client over stdio. The hooks are Claude Code specific; without them, sessions are keyed by parent pid and compaction is by size or on demand.
Why isn't it on PyPI? Because publishing a package is a support commitment, and the
author isn't taking one. Install it from git — the line above works, and uv tool install git+https://github.com/lealvona/local-llm-mcp does too. The repository is MIT: fork it,
package it, publish it under your own name if you want a release you can rely on.
License
MIT — see LICENSE.
Available Tools
8 toolslocal_llm_artifactRead a slice of a stored raw outputARead-onlyIdempotent
Return an exact slice of the raw material behind an earlier result, by its artifact ref (a_xxxxxxxx from a trailer): by line (line_start/line_end, numbered) or by character (offset/limit). Use when the digest left out something you need exactly. In PII mode the slice is scrubbed (placeholders); in ASSIST mode only secrets are scrubbed.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | Artifact ref from a result trailer, e.g. a_1f2e3d4c. | |
| limit | No | Maximum characters to return (character mode). | |
| offset | No | Character offset to start from (character mode). | |
| line_end | No | Last line to return, inclusive (0 = line_start + 199). | |
| line_start | No | First line to return, 1-based (line mode; 0 = character mode). |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool read-only and idempotent, and the description adds meaningful behavior beyond that: line vs character slicing, exact-range semantics, and mode-dependent scrubbing in PII/ASSIST modes. This gives an agent valuable expectations about what the returned content will actually contain.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no fluff. The core action and artifact ref format are front-loaded, slicing modes are grouped, and the usage trigger and scrubbing caveat are stated compactly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a rich output schema, full parameter coverage, and clear annotations, the description covers what remains: when to invoke it, how refs are obtained, how slicing modes work, and how output is transformed in different modes. Nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds useful grouping: line_start/line_end select line mode while offset/limit select character mode. It also clarifies the relationship between the two modes, which the individual schema fields only imply.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb and resource: 'Return an exact slice of the raw material behind an earlier result' via an artifact ref. It is inherently distinguishable from siblings like run/status, but it does not explicitly name any sibling or contrast itself with another tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an explicit usage condition: 'Use when the digest left out something you need exactly.' This is clear context, but it does not list alternatives or state when not to use the tool, so it falls short of a full when/when-not comparison.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
local_llm_compactCompact the worker's running context nowAIdempotent
Ask the worker to compress its running memory of this conversation into a fresh summary. Normally automatic (a PreCompact hook triggers it when your own context is compacted, and the server compacts on its own when its context grows large); call it only if you compacted manually and the hook is absent.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | Why (recorded in the context log). | requested by caller |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide idempotentHint=true and destructiveHint=false; the description adds useful context about automatic triggers (PreCompact hook and server-side compaction) and clarifies the tool's role as a manual fallback. It is consistent with annotations and gives the agent the behavioral nuance it needs, though it could slightly expand on what 'compact' means for future context fidelity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the first sentence states the action, and the second sentence supplies necessary usage constraints. There is no filler or repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple input schema, existing output schema, and clear annotations, the description covers the essential context: what the tool does, when it is normally triggered, and when manual invocation is required. Nothing critically missing for an agent to decide whether and how to call it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single optional 'reason' parameter, so the schema already explains its meaning ('Why (recorded in the context log).'). The description does not need to add parameter-level detail, and the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('compress') and resource ('the worker's running memory of this conversation'), making the tool's function unmistakable. It also distinguishes this tool from siblings by framing it as a manual fallback for automatic compaction, which is clear without needing to inspect other tool schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when not to call the tool ('Normally automatic') and gives the precise condition for calling it ('only if you compacted manually and the hook is absent'). This is concrete, actionable guidance that prevents unnecessary invocations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
local_llm_delegateDelegate a task with inline material and/or filesARead-only
Hand a task to the local worker model over material you name: inline text (material), files or directories to read (paths; directories are listed), and/or the output of a shell command (command). Use for reading or summarizing files and documents, research and synthesis over provided text, calculations, format transformations, parsing, boilerplate drafting, and — in PII mode — anything that touches private data (names, addresses, credentials, personal mail; the worker reads it, you receive placeholders such as [PERSON-1] that you can reuse in later calls). Set verbatim=true ONLY when you will reproduce or edit the text itself (code, config, an error with its stack): the worker locates the lines and the server quotes them. Leave it false for anything to be answered, counted, summarised or listed. The worker also carries its own running memory of this conversation, so follow-up tasks can refer to earlier results by turn id (t_xxxxxx) or artifact ref (a_xxxxxxxx).
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Base directory for relative paths and the command. | |
| task | Yes | What to do, precisely: what to extract, what to report, thresholds, output shape. | |
| paths | No | Files or directories to read as material (~ expands). Optional. | |
| command | No | Shell command whose output is added to the material (bash -c; placeholders expanded server-side). Optional. | |
| material | No | Inline text to work on (pasted output, a document, data). Optional. | |
| verbatim | No | Copy the matching lines byte for byte instead of answering (the worker only locates them). Only for text you must reproduce or edit (a function, a config block, an error with its stack). NOT for questions, counts, summaries or listings: those need the digest, which is the default. | |
| max_output_chars | No | Soft budget for the answer (0 = server default). |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds behavioral detail beyond the annotations: it explains the delegation to the local worker model, the handling of PII with placeholders, the verbatim mode behavior ('the worker only locates them'), and the worker's running memory. It explicitly notes that command output is 'added to the material' and that max_output_chars is a 'soft budget,' all consistent with the read-only, non-destructive hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but tightly organized: a one-sentence definition, a use-case list, a dedicated explanation of the verbatim flag, and a closing note on memory references. Each sentence adds necessary guidance without redundancy, and the flow from general purpose to specific parameter semantics is logical and easy to follow.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Combined with the fully covered schema (100% parameter descriptions) and annotations (read-only, non-destructive), the description fills all gaps an agent might encounter. It covers special cases such as PII placeholders, verbatim behavior, shell command expansion, and the ability to refer to prior turn outputs. The presence of an output schema further completes the picture, so no critical guidance is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
While the schema already describes each parameter, the tool description enriches their meaning. For example, verbatim is explained with a clear rule: 'Set verbatim=true ONLY when you will reproduce or edit the text itself... Leave it false for anything to be answered, counted, summarised or listed.' Material, paths, and command are given contextual framing ('inline text to work on', 'files or directories to read as material', 'shell command whose output is added to the material').
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's core function: 'Hand a task to the local worker model over material you name...' and enumerates the input types (inline text, files/directories, shell command output). The title 'Delegate a task with inline material and/or files' reinforces the specific action and resource, distinguishing it from sibling run/artifact/status tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use it: 'Use for reading or summarizing files and documents, research and synthesis over provided text, calculations, format transformations, parsing, boilerplate drafting, and — in PII mode — anything that touches private data...' It also details when to set verbatim=true vs false and mentions the ability to reference earlier results via turn/artifact IDs, giving the agent clear decision criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
local_llm_disclosureSet what private values may be shown in ASSIST modeAIdempotent
Record the user's decision on what this ASSIST session may show in clear, when the user has said so in the conversation (otherwise the server asks the user itself the first time personal data appears). identity = names, addresses, emails, phones; numbers = account/card/id/SSN numbers and dates of birth (masked by default). Secrets are never shown in either mode. identity='ask' resets the question. Returns the resulting state.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | What the user said, in a few words. Recorded in the session log. | |
| numbers | No | 'open', 'masked' or 'keep'. | keep |
| identity | No | 'open', 'masked', 'ask' (reset) or 'keep'. | keep |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the annotations, explaining that the tool records a decision, that personal data is masked by default, that secrets are never shown, that identity='ask' resets the earlier question, and that the tool returns the resulting state. These are valuable behavioral details not present in readOnlyHint, idempotentHint, or destructiveHint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core purpose, followed by necessary parameter semantics and behavioral edge cases. Every sentence contributes useful information; there is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity, the presence of an output schema, and strong annotations, this description covers the when, what, parameter meanings, privacy defaults, reset behavior, and return value. Nothing critical is missing for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even though schema coverage is 100%, the description adds real meaning to the parameters: it defines what 'identity' and 'numbers' concretely cover (names, addresses, emails, phones; account/card/id/SSN numbers and DOB), notes the masked default, and explains the special 'ask' reset value. This is significant added value beyond the enum patterns.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Record the user's decision on what this ASSIST session may show in clear.' This clearly distinguishes the tool from LLM execution, delegation, artifact, and mode-setting siblings, even without naming them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an explicit usage condition: use this when the user has stated a privacy decision in conversation; otherwise the server asks the user itself. It also explains the reset behavior for identity='ask'. It does not explicitly contrast with sibling tools, but the context is clear enough for an agent to select it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
local_llm_enableOffer to turn this server on for the session (the user decides)AIdempotent
Ask the user whether to turn local-llm-mcp on for this session. The server shows the user a dialog; nothing is delegated unless they say yes. Call it once when a step would print far more than you need or would touch private data. If the user declined, do not call it again unless they ask. Returns the operating rules when the user turns it on, or a refusal otherwise.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate readOnlyHint=false, idempotentHint=true, and destructiveHint=false. The description adds behavioral context: 'The server shows the user a dialog; nothing is delegated unless they say yes' and 'Returns the operating rules when the user turns it on, or a refusal otherwise.' This clarifies the interactive nature and consent requirement. It does not contradict any annotation, though it does not address the idempotency hint explicitly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences and front-loaded: the first sentence states the purpose, the second explains the interactive behavior, and the third gives the conditional usage and return outcome. No wasted words; every sentence adds distinct value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no parameters and an output schema (though details aren't shown). The description covers when to call, the interactive nature, and the return behavior. It explains the condition for not calling again. An agent has enough context to decide when and how to use this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are no parameters, so the schema covers all (100%). The description does not need to explain parameters. It adds value by describing the interaction and return behavior, but since there are no parameters, a baseline of 4 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb and resource: 'Ask the user whether to turn local-llm-mcp on for this session.' It is specific about the action and target. However, it does not explicitly differentiate from sibling tools by name, although the name and description make the intent obvious. It does not say 'use this instead of X', so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to call it: 'Call it once when a step would print far more than you need or would touch private data.' It also gives a when-not: 'If the user declined, do not call it again unless they ask.' This provides clear usage boundaries and is actionable for an agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
local_llm_runRun a command here and receive a digest of its outputA
Execute a shell command on this host (bash -c, this user's privileges, stdin closed) and have the local worker model digest the output for you. You receive the digest plus a trailer (turn id, artifact ref, exit code, raw size); the raw output is stored as an artifact you can slice with local_llm_artifact. Use this instead of your own shell tool whenever the command would print more than you need to read (logs, tests, builds, listings, git output, grep results); rule of thumb: anything over ~40 lines / 2 KB of output belongs here, while commands that print little or nothing run directly. Placeholders such as [SECRET-1] in the command are expanded server-side before execution and never appear in the result.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Working directory for the command. Default: the server's own cwd. | |
| task | No | What to report from the output. Default: outcome, every error/warning verbatim, key values. | |
| command | Yes | The shell command to run, e.g. 'journalctl -u nginx --since -1h' or 'pytest -q'. | |
| verbatim | No | Copy the matching output lines byte for byte instead of digesting (the worker only locates them). Only for text you must reproduce or edit (a function, a config block, an error with its stack). NOT for questions, counts, summaries or listings: those need the digest, which is the default. | |
| timeout_s | No | Kill the command after this many seconds (0 = server default). | |
| max_output_chars | No | Soft budget for the digest (0 = server default). |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
All four annotation hints are false, so the description carries the full behavioral burden — and it pays it in full: execution environment (bash -c, this user's privileges, stdin closed), return shape (digest plus a trailer with turn id, artifact ref, exit code, raw size), raw output persisted as a sliceable artifact, and server-side secret expansion where placeholders 'never appear in the result.' Nothing contradicts the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences with no filler, and the core action plus execution context are front-loaded before usage routing and the secret-handling caveat. It is on the longer side, but for a six-parameter tool with digest, artifact, and secret mechanics, each sentence carries distinct decision-relevant information, so the density is earned.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a high-complexity tool (arbitrary shell execution with side-effect and security implications), and the description covers execution environment, alternative routing, return/artifact flow, and secret handling; an output schema exists to document return values. Minor gaps remain — no explicit statement about non-zero exit handling or how it relates to local_llm_delegate — though the trailer's exit-code field implies graceful handling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with unusually strong per-parameter descriptions (verbatim even explains when not to use it), establishing a baseline of 3. The description adds one schema-absent semantic: placeholders such as [SECRET-1] in the command are expanded server-side and stripped from results, which meaningfully extends the command parameter's value space beyond the schema's examples.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise action — 'Execute a shell command on this host (bash -c, this user's privileges, stdin closed) and have the local worker model digest the output for you' — naming the verb, resource, and the differentiating mechanism (digestion by the local worker model). It also explicitly references the sibling local_llm_artifact for the stored raw output, so an agent can distinguish this tool from the sibling set without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit routing: 'Use this instead of your own shell tool whenever the command would print more than you need to read', backed by a concrete rule of thumb ('anything over ~40 lines / 2 KB of output belongs here, while commands that print little or nothing run directly') and examples (logs, tests, builds, listings, git output, grep results). The verbatim parameter schema description adds a further exclusion ('NOT for questions, counts, summaries or listings'), making when-to-use versus when-not-to-use unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
local_llm_set_modeSwitch between PII and ASSIST modeAIdempotent
Switch the server's mode for the rest of this conversation and return the instructions that apply in the new mode. 'pii': delegate everything that may touch private data; results are fully sanitized. 'assist': delegate anything context-free whose output would be larger than its digest; only secrets are scrubbed.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | Yes | 'pii' or 'assist'. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses meaningful behavior beyond the annotations: the mode change persists for the rest of the conversation, calling it returns the instructions for the new mode, and it details the sanitization behavior for each mode. The annotations only indicate idempotence and non-destructiveness; the description adds the stateful and return-value context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loads the primary action and result, and packs each mode's behavior into concise clauses. There is no filler or repetition of schema details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one parameter, a full schema, an output schema, and siblings covering execution, this description fully equips an agent to invoke it correctly. It covers the persistent side effect, the return value, and the precise mode semantics without needing to explain outputs already described by the output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter already has 100% schema description coverage, but the description adds substantial semantic value beyond the schema's simple
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Switch the server's mode' and explicitly names the two modes, 'pii' and 'assist'. It clearly distinguishes this tool as the mode-setter among siblings like local_llm_run, local_llm_delegate, and local_llm_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear, actionable context for when each mode is appropriate: use 'pii' for anything touching private data and 'assist' for context-free outputs larger than their digest. It does not explicitly name sibling tools or state when not to use set_mode, but the intended usage is well implied by the persistence wording: 'for the rest of this conversation.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
local_llm_statusShow the worker's session stateARead-onlyIdempotent
Mode, session identity, worker (primary/fallback, which is active), turns since compaction, context size, compaction count, placeholder counts, artifact count, and the tokens-saved estimate for this session and all sessions.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds useful context about the scope ('this session and all sessions') and the active worker distinction, but it discloses no additional behavioral traits such as latency, freshness, or side effects beyond what annotations already provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that front-loads the most important category (mode) and then lists the remaining state fields without filler or repetition. Every listed item contributes to the agent's understanding of what the tool returns.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the zero-parameter interface, the read-only annotations, and the presence of an output schema, the description is sufficient for an agent to understand the tool's role and returned contents. It could be slightly more explicit about intended use cases, but nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and an empty input schema, so the description has no parameter documentation burden. With schema coverage effectively at 100%, the baseline of 4 applies; there is nothing missing that the description would need to explain.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The title and description clearly state this tool shows the worker's session state, and the description enumerates the exact contents of that state (mode, session identity, worker, compaction counts, tokens saved). This is specific enough to distinguish it from the sibling action-oriented tools like local_llm_run, local_llm_compact, and local_llm_set_mode.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this is the inspection/status tool, but it never explicitly states when to use it instead of the sibling tools or whether it is meant as a precursor to actions like compaction or delegation. Usage context is implied by the name and title rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
- Added
local_llm_enable
7 tool updates
v0.1.0- First observed
local_llm_artifact - First observed
local_llm_compact - First observed
local_llm_delegate - First observed
local_llm_disclosure - First observed
local_llm_run - First observed
local_llm_set_mode - First observed
local_llm_status
TDQS
Scored across 8 tools
Most tools have clearly distinct purposes—delegate handles text/file analysis, run digests shell output, artifact slices raw material, and status reports state. However, delegate and run both accept commands, and enable vs set_mode both touch activation/mode, creating occasional ambiguity.
The shared local_llm_ prefix gives the set a strong family resemblance, but the second word is inconsistent: nouns (artifact, disclosure, status), verbs (compact, delegate, enable, run), and one verb-noun (set_mode) are mixed rather than following a single pattern.
8 tools is a well-scoped size for a local LLM proxy server. Each tool maps to a distinct part of the workflow—opt-in, mode control, delegation, shell digestion, artifact access, memory compaction, disclosure settings, and status—without redundancy.
The set covers the full lifecycle: enabling the server, switching modes, delegating over text/files/commands, retrieving raw artifacts, compacting memory, and checking status. Minor gaps like explicit cancellation or artifact list/delete are workable since calls are synchronous digests and status exposes artifact counts.
Maintenance
Related MCP Connectors
MCP-Native LLM Orchestration Agent
Authenticated LLM MCP Agent
Private-by-default, local-first memory/context/task orchestrator for MCP apps and agents.
Connect MCP clients to 2,000+ AI models without managing provider API keys.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceA local-first LLM routing MCP server that keeps sensitive data on your own models, with fail-closed privacy and manager-worker delegation, exposing route and complete tools to any MCP client.MIT
- AlicenseNot gradedqualityDmaintenanceA local, containerized MCP server that uses a local LLM to sanitize documents by removing or transforming PII before content is sent to public LLM services.MIT
- AlicenseAqualityAmaintenanceLocal pseudonymisation MCP server that detects PII in text, replaces it with opaque tokens before sending to cloud LLMs, and restores tokens afterward.2199 npm2MIT
- AlicenseNot gradedqualityBmaintenanceMCP server that delegates mechanical tasks like summarization, classification, extraction, and drafting to a local Llama.cpp LLM, serving as a cost-optimization layer while Claude handles reasoning and quality control.MIT