local-llm-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| LOCAL_LLM_MCP_MODE | No | pii or assist. | assist |
| LOCAL_LLM_MCP_MODEL | No | Model name sent to the endpoint. | local |
| LOCAL_LLM_MCP_RULES | No | JSON rules file (replaces built-in shape rules). | |
| LOCAL_LLM_MCP_PRICES | No | Override file for the tokens-saved estimate (assumptions + prices). | ~/.config/local-llm-mcp/prices.json |
| LOCAL_LLM_MCP_SHAPES | No | PII mode: the identity-shape layer (addresses, labelled names, dates of birth, labelled id numbers). | 1 |
| LOCAL_LLM_MCP_API_KEY | No | Bearer token for the OpenAI-compatible endpoint. | |
| LOCAL_LLM_MCP_SESSION | No | Force a session key (tests, scripts). | |
| LOCAL_LLM_MCP_BASE_URL | No | OpenAI-compatible endpoint; must be local (loopback, private, or CGNAT). | http://127.0.0.1:8000/v1 |
| LOCAL_LLM_MCP_ENV_FILE | No | Path to the dotenv file (default ~/.config/local-llm-mcp/env). | |
| LOCAL_LLM_MCP_OBSERVER | No | webhook, or module:Class. | |
| LOCAL_LLM_MCP_SOCK_DIR | No | Control sockets directory (default $XDG_RUNTIME_DIR/local-llm-mcp, else <state>/sock). | |
| LOCAL_LLM_MCP_THINKING | No | Pass enable_thinking to the chat template. | 0 |
| LOCAL_LLM_MCP_LOG_LEVEL | No | Stderr logging level. | INFO |
| LOCAL_LLM_MCP_STATE_DIR | No | Sessions, pidmap. | ~/.local/state/local-llm-mcp |
| LOCAL_LLM_MCP_ADMIN_BIND | No | Admin app listener bind address. | 127.0.0.1 |
| LOCAL_LLM_MCP_ADMIN_PORT | No | Admin app listener port. | 8631 |
| LOCAL_LLM_MCP_ESCALATION | No | First identity/number values in an ASSIST session: ask (dialog), auto (mask and tell), or off. | ask |
| LOCAL_LLM_MCP_STRICT_PII | No | Treat every bare 10-digit run as a phone number. | 0 |
| LOCAL_LLM_MCP_ADMIN_TOKEN | No | Bearer token the admin API requires (set it whenever the bind is not loopback). | |
| LOCAL_LLM_MCP_CHUNK_CHARS | No | Material above this is processed map-reduce style. | 24000 |
| LOCAL_LLM_MCP_ENTITY_PASS | No | PII mode: ask the worker for the private values before answering. | 1 |
| LOCAL_LLM_MCP_LLM_TIMEOUT | No | Seconds per model call. | 300 |
| LOCAL_LLM_MCP_API_KEY_FILE | No | Path to a file containing the endpoint bearer token (one line, keep it 0600). | |
| LOCAL_LLM_MCP_CALLER_MODEL | No | Headline model for the dollar figure in token savings. | |
| LOCAL_LLM_MCP_OBSERVER_URL | No | Webhook target. | |
| LOCAL_LLM_MCP_CONTEXT_CHARS | No | Budget for summary + recent turns in each prompt. | 12000 |
| LOCAL_LLM_MCP_EXCERPT_CHARS | No | Scrubbed excerpt of material kept per turn. | 1500 |
| LOCAL_LLM_MCP_OBSERVER_PATH | No | Directory added to sys.path to import your observer. | |
| LOCAL_LLM_MCP_PRIVATE_TERMS | No | Terms file for private values. | ~/.config/local-llm-mcp/private_terms.json |
| LOCAL_LLM_MCP_SUMMARY_CHARS | No | Compaction target. | 6000 |
| LOCAL_LLM_MCP_ASSIST_NUMBERS | No | Whether account/card/id numbers and dates of birth are shown in ASSIST mode before the user decides. | masked |
| LOCAL_LLM_MCP_DIALOG_TIMEOUT | No | Seconds to wait for the disclosure dialog before masking that result. | 600 |
| LOCAL_LLM_MCP_FALLBACK_MODEL | No | Model name at the fallback. | |
| LOCAL_LLM_MCP_MAX_TOKENS_CAP | No | Ceiling on max_tokens. | 8192 |
| LOCAL_LLM_MCP_COMMAND_TIMEOUT | No | Default command timeout, seconds. | 120 |
| LOCAL_LLM_MCP_PARALLEL_CHUNKS | No | Concurrent chunk calls. | 4 |
| LOCAL_LLM_MCP_ADMIN_TOKEN_FILE | No | Path to a file containing the admin bearer token. | |
| LOCAL_LLM_MCP_FALLBACK_API_KEY | No | Bearer token for the fallback endpoint. | |
| LOCAL_LLM_MCP_MAX_OUTPUT_CHARS | No | Default digest budget; max_tokens is derived from it. | 2000 |
| LOCAL_LLM_MCP_FAILOVER_COOLDOWN | No | Seconds the fallback is asked first after the primary fails. | 120 |
| LOCAL_LLM_MCP_FALLBACK_BASE_URL | No | A second local worker used when the primary is unreachable or answers 5xx (must be local too). | |
| LOCAL_LLM_MCP_AUTO_COMPACT_CHARS | No | Self-compaction threshold on uncompacted turns. | 60000 |
| LOCAL_LLM_MCP_MATERIAL_MAX_CHARS | No | Cap on gathered material (head 70% + tail 30% kept, marker in between). | 400000 |
| LOCAL_LLM_MCP_FALLBACK_API_KEY_FILE | No | Path to a file containing the fallback endpoint bearer token. |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| local_llm_runA | Execute a shell command on this host (bash -c, this user's privileges, stdin closed) and have the local worker model digest the output for you. You receive the digest plus a trailer (turn id, artifact ref, exit code, raw size); the raw output is stored as an artifact you can slice with local_llm_artifact. Use this instead of your own shell tool whenever the command would print more than you need to read (logs, tests, builds, listings, git output, grep results); rule of thumb: anything over ~40 lines / 2 KB of output belongs here, while commands that print little or nothing run directly. Placeholders such as [SECRET-1] in the command are expanded server-side before execution and never appear in the result. |
| local_llm_delegateA | Hand a task to the local worker model over material you name: inline text (material), files or directories to read (paths; directories are listed), and/or the output of a shell command (command). Use for reading or summarizing files and documents, research and synthesis over provided text, calculations, format transformations, parsing, boilerplate drafting, and — in PII mode — anything that touches private data (names, addresses, credentials, personal mail; the worker reads it, you receive placeholders such as [PERSON-1] that you can reuse in later calls). Set verbatim=true ONLY when you will reproduce or edit the text itself (code, config, an error with its stack): the worker locates the lines and the server quotes them. Leave it false for anything to be answered, counted, summarised or listed. The worker also carries its own running memory of this conversation, so follow-up tasks can refer to earlier results by turn id (t_xxxxxx) or artifact ref (a_xxxxxxxx). |
| local_llm_artifactA | Return an exact slice of the raw material behind an earlier result, by its artifact ref (a_xxxxxxxx from a trailer): by line (line_start/line_end, numbered) or by character (offset/limit). Use when the digest left out something you need exactly. In PII mode the slice is scrubbed (placeholders); in ASSIST mode only secrets are scrubbed. |
| local_llm_statusA | Mode, session identity, worker (primary/fallback, which is active), turns since compaction, context size, compaction count, placeholder counts, artifact count, and the tokens-saved estimate for this session and all sessions. |
| local_llm_compactA | Ask the worker to compress its running memory of this conversation into a fresh summary. Normally automatic (a PreCompact hook triggers it when your own context is compacted, and the server compacts on its own when its context grows large); call it only if you compacted manually and the hook is absent. |
| local_llm_set_modeA | Switch the server's mode for the rest of this conversation and return the instructions that apply in the new mode. 'pii': delegate everything that may touch private data; results are fully sanitized. 'assist': delegate anything context-free whose output would be larger than its digest; only secrets are scrubbed. |
| local_llm_enableA | Ask the user whether to turn local-llm-mcp on for this session. The server shows the user a dialog; nothing is delegated unless they say yes. Call it once when a step would print far more than you need or would touch private data. If the user declined, do not call it again unless they ask. Returns the operating rules when the user turns it on, or a refusal otherwise. |
| local_llm_disclosureA | Record the user's decision on what this ASSIST session may show in clear, when the user has said so in the conversation (otherwise the server asks the user itself the first time personal data appears). identity = names, addresses, emails, phones; numbers = account/card/id/SSN numbers and dates of birth (masked by default). Secrets are never shown in either mode. identity='ask' resets the question. Returns the resulting state. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
| local_llm_instructions | When and how to call this server (the same text as the connection instructions; for clients that do not surface them). |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
| local_llm_instructions | Current mode and the rules for calling this server. |
TDQS
Scored across 8 tools
Most tools have clearly distinct purposes—delegate handles text/file analysis, run digests shell output, artifact slices raw material, and status reports state. However, delegate and run both accept commands, and enable vs set_mode both touch activation/mode, creating occasional ambiguity.
The shared local_llm_ prefix gives the set a strong family resemblance, but the second word is inconsistent: nouns (artifact, disclosure, status), verbs (compact, delegate, enable, run), and one verb-noun (set_mode) are mixed rather than following a single pattern.
8 tools is a well-scoped size for a local LLM proxy server. Each tool maps to a distinct part of the workflow—opt-in, mode control, delegation, shell digestion, artifact access, memory compaction, disclosure settings, and status—without redundancy.
The set covers the full lifecycle: enabling the server, switching modes, delegating over text/files/commands, retrieving raw artifacts, compacting memory, and checking status. Minor gaps like explicit cancellation or artifact list/delete are workable since calls are synchronous digests and status exposes artifact counts.