Nanites
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| NANITES_HOME | No | Storage root for profiles, database, and logs. | ~/.nanites |
| NANITES_UI_PORT | No | Port for the local dashboard. | 4700 |
| NANITES_HF_FETCH | No | Set to 1 to allow Hugging Face network lookups. | off |
| NANITES_AUTOSTART_UI | No | Set to 0 to disable automatic dashboard orchestration. | 1 |
| NANITES_VISION_ROOTS | No | Allowed roots for image paths, separated by semicolons. | current directory |
| NANITES_LMS_API_TOKEN | No | LM Studio API auth token, used when a profile has no token set. | |
| NANITES_LMSTUDIO_MODELS_DIR | No | Directory used to measure free disk space for model downloads. | ~/.lmstudio/models |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
| prompts | {
"listChanged": true
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| nanites_pingA | Liveness probe for the Nanites MCP server. Returns server identity and transport health. |
| list_modelsA | List the models known to the active profile's LM Studio endpoint. Returns a trimmed shape by default; pass verbose for the full model objects. |
| get_loaded_modelA | Return the models currently loaded into GPU/CPU on the active profile's LM Studio endpoint (loaded_instances non-empty). |
| load_modelA | Load a model into the active profile's LM Studio endpoint. model_id is the model key (the identifier shown by list_models); optional load params (context_length, flash_attention, ...) may be passed. |
| unload_modelC | Unload a loaded model instance by its instance_id from the active profile's LM Studio endpoint. |
| chatA | Send a multi-message chat to a loaded model instance. System messages become the system prompt; output passes through the reply validator/cleaner and includes a validation field describing anything stripped. |
| download_modelC | Ask the active profile's LM Studio endpoint to download a model by HF source id, optionally pinned to a quantization. |
| get_download_statusA | Poll the download progress for a job_id returned by download_model. |
| read_registryA | Read model-registry entries (roles, scores, best params, last tested) for a profile, optionally filtered to one model_id. Trimmed by default; verbose includes timestamps. |
| write_registry_entryA | Insert or update one registry entry for a profile/model_id. Omitted entry fields keep their existing values. |
| share_test_resultsA | Copy approved test results from source_profile into profile when both point at the SAME LM Studio instance (endpoint fingerprint: normalized URL + auth presence). Opt-in sharing to avoid redundant regimen runs; different endpoints are refused and stay isolated. Nothing is auto-finalized — a later regimen/finalize aggregates the copied results without re-running the model. |
| create_profileC | Create a named profile. Machine specs, endpoint, pricing, test_plan_ref and ntfy are optional and resolve to documented defaults; concurrency is derived from machine specs via the guardrail tiers. |
| switch_profileA | Make a named profile the active one. All subsequent model tools use its endpoint until switched again. |
| update_profileA | Partially update a named profile (machine specs, endpoint, pricing, ntfy, inference effort/ceiling, theme, tool grant). Omitted fields keep their current values. Used by /nanites-effort to set the active profile's effort. |
| list_profilesA | List profile names. Pass verbose for full profile objects. |
| get_active_profileA | Return the currently active profile, or null if none has been switched to yet. |
| get_first_run_statusA | Report whether Nanites needs first-run initialization: true exactly when zero profiles exist. The orchestrator calls this once at session start; when true, it runs the /nanites-new-profile flow. |
| list_test_unitsA | List the registered test units for a profile (the default regimen is registered automatically on first use). |
| validate_test_unitA | Validate a test-unit object against the schema/validator rules without persisting it. Returns { ok, issues }. |
| register_test_unitB | Validate and persist a test unit for a profile. Rejects invalid units with the issues listed; duplicate ids are refused. |
| run_test_regimenA | Test one model against the profile's registered test units (the default regimen is auto-registered on first use). Loads the model, runs deterministic units inline with a logged param search, runs orchestrator_judged units to pending, unloads exactly once, and writes the registry entry when nothing is pending. Returns only pending_unit_ids, never raw output. |
| get_pending_judgmentsB | Fetch cleaned raw output plus rubric and prompt context for pending orchestrator-judged units (all, or a requested subset) for the orchestrator to read and judge in the same turn. |
| submit_test_judgmentA | Record an orchestrator judgment for one pending unit. Nothing becomes final in the registry until user_approved is true; a submission with user_approved false records the judgment without finalizing that score. |
| run_sub_agentA | Delegate one bounded task to a local model. Resolves the model from the profile's registry by roles (or an explicit model_id), acquires a model respecting the profile's concurrency tier (reuse already-loaded, evict on sequential tiers, refuse at parallel capacity), runs the brief once, cleans the reply, unloads exactly once if it loaded the model, and logs exactly one token-usage entry for cost tracking. |
| start_sub_agent_jobA | Queue a sub-agent job on the profile's async job FIFO and return its job_id immediately (non-blocking — for long-horizon work on slow/big models instead of a tool call that blocks for minutes). Same validated inputs as run_sub_agent; when the profile's concurrency tier is at capacity the job waits queued rather than erroring. Poll with get_sub_agent_job_status. |
| get_sub_agent_job_statusA | Poll a sub-agent job started with start_sub_agent_job. Returns { job_id, status: queued|running|done|error, result? }; result is present for done (shaped like run_sub_agent's response) and error (structured { code, message, retryable }). |
| start_btw_chatA | Start a /nanites-btw working-memory chat for a profile: replaces the active chat row and wipes its old transcript (the compaction caches persist), enqueues compaction of the given host-session transcript as an async job (returns immediately — it can never stall or eject a running job), pins and holds a context_qa model when the job completes, and answers initial_question inline if it finishes inside the ~5s grace window. Returns the dashboard deep link (open it in the preview) plus the job handle and status. |
| diff_untestedA | List downloaded LLM models on the profile's endpoint that have no registry entry yet (the Workflow #3 discovery step). Returns a trimmed per-model shape. |
| run_untested_sweepB | Workflow #3 end to end: diff downloaded LLM models against the registry and run Workflow #1 (run_test_regimen) on each unregistered model, sequentially. |
| download_and_waitB | Workflow #4 download half: ask LM Studio to download a model by HF source id, then poll download status with exponential backoff until completed/failed/paused (or a poll ceiling). |
| download_and_testB | Workflow #4 end to end: download_and_wait, then on completion trigger Workflow #1 (run_test_regimen) on the newly downloaded model automatically. A failed/paused/gave-up download returns without testing. |
| filter_by_guardrailA | Shortlist candidate models (e.g. Hugging Face search results the host gathered via its HF connector) against the profile's machine-spec guardrail tier. Excludes models far outside the tier's recommended size range with an explicit reason; never silently suggests. |
| get_cost_saved_reportA | Report tokens and estimated USD saved by delegating work to local models, from real logged sub-agent usage over a period (all/day/week/month) at the profile's input/output rates. Local delegation costs ~$0, so saved_usd is the orchestrator-equivalent cost not spent. |
| check_adaptationA | Check whether a profile's use case diverges from nanites-default. When it does, returns a user prompt asking whether to draft custom test units, plus the default-plan notice that reusing the default plan may not be well-calibrated. |
| register_adapted_unitsA | Validate a batch of authored test units through the Phase 4 validator and register the ones that pass; each rejected unit is returned with its validation issues surfaced, never silently dropped or force-registered. |
| system_health_checkA | Check LM Studio health for a profile: endpoint reachability (with a one-shot lms server start autostart recovery and recheck), free disk space for downloads, and stuck-loaded-model detection. Returns an overall status (healthy/degraded/down) plus per-check sub-fields. |
| send_ntfyA | Fire-and-forget push notification to the profile's ntfy topic (public-server default resolution per profile config). A failed push never fails the underlying operation; it is logged server-side only. |
| nanites_addProviderKeyC | Add an API key for a cloud provider (Cloudflare, OpenRouter, OmniRoute, Generic). |
| nanites_removeProviderKeyC | Remove an API key from a cloud provider. |
| nanites_listProviderKeysA | List all API keys for a cloud provider. |
| nanites_toggleProviderKeyC | Enable or disable an API key for a cloud provider. |
| nanites_discoverProviderModelsC | Auto-discover available models from a cloud provider. |
| nanites_listProviderModelsC | List cached models from a cloud provider. |
| nanites_registerProviderModelC | Register a discovered model for use with a cloud provider. |
| nanites_deregisterProviderModelB | Deregister a registered model from a cloud provider. |
| nanites_showProviderErrorsC | Show recent errors from cloud provider calls. |
| nanites_setProviderEnabledB | Enable or disable a cloud provider globally. |
| nanites_getProviderConfigA | Get provider preference order and per-provider settings. |
| nanites_setProviderPreferenceOrderC | Set the global provider preference order for cloud routing. |
| seed_provider_modelsA | Bulk-register the canonical Cloudflare agentic models (canonical manifest), role-tag them in the registry, and write the default role pins. Idempotent; unknown model ids are refused before anything is written. |
| set_role_pinA | Pin a role to a preferred (provider, model). Auto-route by pin on later sub-agent runs; falls back to dynamic when the pinned target is unusable. |
| list_role_pinsB | List the role->model pins for a profile (default active). |
| delete_role_pinB | Remove a role pin so that role falls back to dynamic model selection. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
| nanites-new-profile | Create a new profile from the user's stated fields and make it active. Pass every field the user supplied through unchanged; omit nothing they gave you. |
| nanites-switch-profile | Make a named profile the active one. |
| nanites-profiles | List the known profiles and which one is active. |
| nanites-cost-saved | Report tokens and estimated USD saved by local delegation for the active profile, over a period. |
| nanites-effort | Set the active profile's effort level (low/medium/high) and optional output-token ceiling. Effort drives the inference planner's reasoning + token-budget decisions for every subsequent sub-agent call. |
| nanites-dynamic-model | Toggle whether run_sub_agent hot-loads the best registry match for a role (on) or uses whatever models the user has loaded in LM Studio (off). |
| nanites-seed-agents | Bulk-register the canonical Cloudflare agentic models (plus llama-3.2-11b-vision) for a profile: catalog registration with manifest capabilities, registry role-tagging, and default role pins. Idempotent. |
| nanites-pin | List, set, or delete a preferred (provider, model) pin for a role on a profile. Pins auto-route sub-agent runs for that role; the fallback ladder takes over when the pinned target is unusable. |
| nanites-vision | Flip a profile's vision_capable flag. When off, the profile does no image work at all — delegation of image analysis to vision-capable models is disabled for it. |
| nanites-btw | Compact this session's context into a /nanites-btw working-memory chat for a profile and open it in the dashboard chat mode (btw-spec-v2 §5). One new active chat per profile; the previous chat's visible transcript is reset but the compaction caches persist. |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 53 tools
Most tools are clearly distinct (e.g., list_models vs get_loaded_model, run_sub_agent vs start_sub_agent_job), but a few pairs could be confused: nanites_deregisterProviderModel vs nanites_removeProviderKey (deregister vs remove), and get_download_status vs download_and_wait overlap in polling behavior. Overall, the core model/test workflows are well separated.
Naming is mixed: some tools use snake_case with domain prefix (nanites_*), others use plain verbs (chat, load_model, create_profile), and a few use different conventions (get_cost_saved_report, seed_provider_models). The pattern is not uniform, making it harder to predict tool names, though each individual name is readable.
With 53 tools, the surface is both wide and deep, covering profiles, providers, models, testing, jobs, workflows, and notifications. While each domain has its own cluster, the sheer number exceeds what an agent can easily navigate, and some tools (e.g., send_ntfy, get_cost_saved_report) feel peripheral. A more focused set of ~25-35 would be more manageable.
The set covers the full lifecycle for models (list/load/unload/download/test), profiles (create/update/switch/list), providers (add/remove/list/toggle keys, set preference), and test management (validate/register/run/judge). Minor gaps exist: no explicit tool to delete a profile or a test unit, and no tool to list provider errors beyond 'show' (which is read-only). Overall, core workflows are well-covered.