Enables MCP clients to scan local GGUF models, estimate VRAM and suggest GPU offload layers, manage llama-server lifecycle, and proxy OpenAI-format chats with idle auto-unload.
LLM deployment planner: given a model and a GPU, answers will it fit, will it hit your SLO, and what will it cost. Sizes VRAM and KV-cache from the model's real architecture, and labels every number measured, estimated, or unknown.
Enables engineers to estimate LLM costs, check GPU VRAM fit, audit MCP configs for security risks, and analyze network device config changes and compliance—all locally, with no account or telemetry.
MCP server for LLM quantization. Compress any HuggingFace model to GGUF, GPTQ, or AWQ format. 6 tools: info, check, recommend, quantize, evaluate, push. Self-contained Python server — no external CLI needed.
Enables AI agents and developers to dynamically tokenize session context windows and compact dialogue history into structured outputs via MCP, CLI, or Python client. Runs on pure standard library Python with no external dependencies for deterministic, low-latency execution.