Skip to main content
Glama

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
LOCAL_LLM_MCP_MODENopii or assist.assist
LOCAL_LLM_MCP_MODELNoModel name sent to the endpoint.local
LOCAL_LLM_MCP_RULESNoJSON rules file (replaces built-in shape rules).
LOCAL_LLM_MCP_PRICESNoOverride file for the tokens-saved estimate (assumptions + prices).~/.config/local-llm-mcp/prices.json
LOCAL_LLM_MCP_SHAPESNoPII mode: the identity-shape layer (addresses, labelled names, dates of birth, labelled id numbers).1
LOCAL_LLM_MCP_API_KEYNoBearer token for the OpenAI-compatible endpoint.
LOCAL_LLM_MCP_SESSIONNoForce a session key (tests, scripts).
LOCAL_LLM_MCP_BASE_URLNoOpenAI-compatible endpoint; must be local (loopback, private, or CGNAT).http://127.0.0.1:8000/v1
LOCAL_LLM_MCP_ENV_FILENoPath to the dotenv file (default ~/.config/local-llm-mcp/env).
LOCAL_LLM_MCP_OBSERVERNowebhook, or module:Class.
LOCAL_LLM_MCP_SOCK_DIRNoControl sockets directory (default $XDG_RUNTIME_DIR/local-llm-mcp, else <state>/sock).
LOCAL_LLM_MCP_THINKINGNoPass enable_thinking to the chat template.0
LOCAL_LLM_MCP_LOG_LEVELNoStderr logging level.INFO
LOCAL_LLM_MCP_STATE_DIRNoSessions, pidmap.~/.local/state/local-llm-mcp
LOCAL_LLM_MCP_ADMIN_BINDNoAdmin app listener bind address.127.0.0.1
LOCAL_LLM_MCP_ADMIN_PORTNoAdmin app listener port.8631
LOCAL_LLM_MCP_ESCALATIONNoFirst identity/number values in an ASSIST session: ask (dialog), auto (mask and tell), or off.ask
LOCAL_LLM_MCP_STRICT_PIINoTreat every bare 10-digit run as a phone number.0
LOCAL_LLM_MCP_ADMIN_TOKENNoBearer token the admin API requires (set it whenever the bind is not loopback).
LOCAL_LLM_MCP_CHUNK_CHARSNoMaterial above this is processed map-reduce style.24000
LOCAL_LLM_MCP_ENTITY_PASSNoPII mode: ask the worker for the private values before answering.1
LOCAL_LLM_MCP_LLM_TIMEOUTNoSeconds per model call.300
LOCAL_LLM_MCP_API_KEY_FILENoPath to a file containing the endpoint bearer token (one line, keep it 0600).
LOCAL_LLM_MCP_CALLER_MODELNoHeadline model for the dollar figure in token savings.
LOCAL_LLM_MCP_OBSERVER_URLNoWebhook target.
LOCAL_LLM_MCP_CONTEXT_CHARSNoBudget for summary + recent turns in each prompt.12000
LOCAL_LLM_MCP_EXCERPT_CHARSNoScrubbed excerpt of material kept per turn.1500
LOCAL_LLM_MCP_OBSERVER_PATHNoDirectory added to sys.path to import your observer.
LOCAL_LLM_MCP_PRIVATE_TERMSNoTerms file for private values.~/.config/local-llm-mcp/private_terms.json
LOCAL_LLM_MCP_SUMMARY_CHARSNoCompaction target.6000
LOCAL_LLM_MCP_ASSIST_NUMBERSNoWhether account/card/id numbers and dates of birth are shown in ASSIST mode before the user decides.masked
LOCAL_LLM_MCP_DIALOG_TIMEOUTNoSeconds to wait for the disclosure dialog before masking that result.600
LOCAL_LLM_MCP_FALLBACK_MODELNoModel name at the fallback.
LOCAL_LLM_MCP_MAX_TOKENS_CAPNoCeiling on max_tokens.8192
LOCAL_LLM_MCP_COMMAND_TIMEOUTNoDefault command timeout, seconds.120
LOCAL_LLM_MCP_PARALLEL_CHUNKSNoConcurrent chunk calls.4
LOCAL_LLM_MCP_ADMIN_TOKEN_FILENoPath to a file containing the admin bearer token.
LOCAL_LLM_MCP_FALLBACK_API_KEYNoBearer token for the fallback endpoint.
LOCAL_LLM_MCP_MAX_OUTPUT_CHARSNoDefault digest budget; max_tokens is derived from it.2000
LOCAL_LLM_MCP_FAILOVER_COOLDOWNNoSeconds the fallback is asked first after the primary fails.120
LOCAL_LLM_MCP_FALLBACK_BASE_URLNoA second local worker used when the primary is unreachable or answers 5xx (must be local too).
LOCAL_LLM_MCP_AUTO_COMPACT_CHARSNoSelf-compaction threshold on uncompacted turns.60000
LOCAL_LLM_MCP_MATERIAL_MAX_CHARSNoCap on gathered material (head 70% + tail 30% kept, marker in between).400000
LOCAL_LLM_MCP_FALLBACK_API_KEY_FILENoPath to a file containing the fallback endpoint bearer token.

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}
prompts
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}
experimental
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
local_llm_runA

Execute a shell command on this host (bash -c, this user's privileges, stdin closed) and have the local worker model digest the output for you. You receive the digest plus a trailer (turn id, artifact ref, exit code, raw size); the raw output is stored as an artifact you can slice with local_llm_artifact. Use this instead of your own shell tool whenever the command would print more than you need to read (logs, tests, builds, listings, git output, grep results); rule of thumb: anything over ~40 lines / 2 KB of output belongs here, while commands that print little or nothing run directly. Placeholders such as [SECRET-1] in the command are expanded server-side before execution and never appear in the result.

local_llm_delegateA

Hand a task to the local worker model over material you name: inline text (material), files or directories to read (paths; directories are listed), and/or the output of a shell command (command). Use for reading or summarizing files and documents, research and synthesis over provided text, calculations, format transformations, parsing, boilerplate drafting, and — in PII mode — anything that touches private data (names, addresses, credentials, personal mail; the worker reads it, you receive placeholders such as [PERSON-1] that you can reuse in later calls). Set verbatim=true ONLY when you will reproduce or edit the text itself (code, config, an error with its stack): the worker locates the lines and the server quotes them. Leave it false for anything to be answered, counted, summarised or listed. The worker also carries its own running memory of this conversation, so follow-up tasks can refer to earlier results by turn id (t_xxxxxx) or artifact ref (a_xxxxxxxx).

local_llm_artifactA

Return an exact slice of the raw material behind an earlier result, by its artifact ref (a_xxxxxxxx from a trailer): by line (line_start/line_end, numbered) or by character (offset/limit). Use when the digest left out something you need exactly. In PII mode the slice is scrubbed (placeholders); in ASSIST mode only secrets are scrubbed.

local_llm_statusA

Mode, session identity, worker (primary/fallback, which is active), turns since compaction, context size, compaction count, placeholder counts, artifact count, and the tokens-saved estimate for this session and all sessions.

local_llm_compactA

Ask the worker to compress its running memory of this conversation into a fresh summary. Normally automatic (a PreCompact hook triggers it when your own context is compacted, and the server compacts on its own when its context grows large); call it only if you compacted manually and the hook is absent.

local_llm_set_modeA

Switch the server's mode for the rest of this conversation and return the instructions that apply in the new mode. 'pii': delegate everything that may touch private data; results are fully sanitized. 'assist': delegate anything context-free whose output would be larger than its digest; only secrets are scrubbed.

local_llm_enableA

Ask the user whether to turn local-llm-mcp on for this session. The server shows the user a dialog; nothing is delegated unless they say yes. Call it once when a step would print far more than you need or would touch private data. If the user declined, do not call it again unless they ask. Returns the operating rules when the user turns it on, or a refusal otherwise.

local_llm_disclosureA

Record the user's decision on what this ASSIST session may show in clear, when the user has said so in the conversation (otherwise the server asks the user itself the first time personal data appears). identity = names, addresses, emails, phones; numbers = account/card/id/SSN numbers and dates of birth (masked by default). Secrets are never shown in either mode. identity='ask' resets the question. Returns the resulting state.

Prompts

Interactive templates invoked by user choice

NameDescription
local_llm_instructionsWhen and how to call this server (the same text as the connection instructions; for clients that do not surface them).

Resources

Contextual data attached and managed by the client

NameDescription
local_llm_instructionsCurrent mode and the rules for calling this server.

TDQS

A4.3/5.0

Scored across 8 tools

Disambiguation4/5

Most tools have clearly distinct purposes—delegate handles text/file analysis, run digests shell output, artifact slices raw material, and status reports state. However, delegate and run both accept commands, and enable vs set_mode both touch activation/mode, creating occasional ambiguity.

Naming Consistency3/5

The shared local_llm_ prefix gives the set a strong family resemblance, but the second word is inconsistent: nouns (artifact, disclosure, status), verbs (compact, delegate, enable, run), and one verb-noun (set_mode) are mixed rather than following a single pattern.

Tool Count5/5

8 tools is a well-scoped size for a local LLM proxy server. Each tool maps to a distinct part of the workflow—opt-in, mode control, delegation, shell digestion, artifact access, memory compaction, disclosure settings, and status—without redundancy.

Completeness4/5

The set covers the full lifecycle: enabling the server, switching modes, delegating over text/files/commands, retrieving raw artifacts, compacting memory, and checking status. Minor gaps like explicit cancellation or artifact list/delete are workable since calls are synchronous digests and status exposes artifact counts.

Maintenance

ActivityMaintained
ResponsivenessNo issues