F
licenseNot graded
qualityB
maintenanceMCP server that lets Claude Code delegate trivial, standalone questions to a local Gemma 4 model via llama.cpp, saving paid tokens.
-
No user-submitted related servers found.
Scored across 19 tools
Each tool targets a distinct operation or resource: chat vs. completion vs. infill vs. embed vs. rerank, plus model and server management. No ambiguous overlaps.
All tools follow the 'llama_' prefix with lowercase snake_case names, using verbs or verb_noun patterns consistently.
19 tools is slightly above the typical 3-15 range, but each tool serves a clear purpose covering model interaction, management, and server control, so it remains reasonable.
The tool set covers the full lifecycle: chat/completion/infill/embed/rerank, tokenization, model loading/management, server health and metrics, and LoRA adapter control. No obvious gaps.