Skip to main content
Glama
README.md
# Nanites

**An MCP server that lets Claude Code and Claude Desktop delegate bounded, disposable
work to your own local models — so you stop paying frontier tokens for grunt work.**

Nanites is the delegation layer, not the orchestrator. Claude Code/Desktop stays your
orchestrator on a paid frontier model and handles anything requiring real judgment.
Nanites hands the cheap, bounded, throwaway work — scanning a codebase, summarizing a
file, drafting a test, answering a side question — to a local LM Studio model or a
cheap cloud model, and returns the result.

It ships as a Claude Code plugin (slash commands + a delegation skill + a live
dashboard) and as a plain MCP server you can register with any MCP-capable harness.

---

## Table of contents

- [Why it exists](#why-it-exists)
- [Requirements](#requirements)
- [Install](#install)
- [Getting started](#getting-started)
- [How it works](#how-it-works)
- [The dashboard](#the-dashboard)
- [Configuration](#configuration)
- [Tool reference](#tool-reference)
- [Slash commands](#slash-commands)
- [Development](#development)
- [Known limitations](#known-limitations)

---

## Why it exists

Delegating work to a small local model is a good idea. Delegating it *blindly* — without
knowing which model is good at what — is how you end up with worse results and a bigger
bill. Nanites builds the part that is tedious:

- **A model registry.** Each model gets a `performance_score` (1-100), plus rolling
  `avg_load_ms` / `avg_response_ms`, recomputed from your real logged runs. The score is a
  speed/stability signal, not an oracle — the skill teaches Claude to weigh it alongside
  the test-regimen results and your own approvals.
- **A test regimen.** A built-in suite of modular test units runs against a model.
  Deterministic ones (valid JSON, exact label) are scored automatically; the rest come
  back to *you* for judgment — never auto-filed as passing.
- **Guardrail tiers keyed to your VRAM.** A 4 GB card and a 24 GB card get different
  concurrency advice, and the advice comes with a stated reason so Claude can explain it.
- **A cost-saved report.** Every delegation is logged with token counts and a computed
  saving, so the claim is auditable rather than vibes.

## Requirements

| Requirement | Notes |
|---|---|
| **Node.js 22.5+** | Hard requirement. Nanites uses the built-in `node:sqlite` module, which does not exist on Node 20. |
| **LM Studio** | Only for local models. The app talks to LM Studio's local HTTP server (`http://localhost:1234` by default). |
| **Claude Code or Claude Desktop** | For the plugin experience. Any MCP-capable client works for the raw server. |
| A cloud provider key | Optional. Cloudflare / OpenRouter / OmniRoute / any OpenAI-compatible endpoint. |

> No `lms` CLI is required. Nanites will *try* `lms server start` as a recovery step if
> LM Studio is unreachable, but it is best-effort and non-blocking.

## Install

### As a Claude Code plugin (recommended)

```bash
git clone https://github.com/ADn-001/NANITES_MCP.git
cd NANITES_MCP
npm install
npm run build
```

Point Claude Code at the plugin directory for this session:

```bash
claude --plugin-dir ./plugin/nanites
```

To install it permanently as a marketplace plugin:

```bash
claude plugin marketplace add ADn-001/NANITES_MCP
claude plugin install nanites@nanites
```

The plugin ships its compiled server and its runtime dependencies, so a
marketplace install needs no build step.

### As a plain MCP server

Register it with any MCP client. For Claude Code, add to your project `.mcp.json`:

```json
{
  "mcpServers": {
    "nanites": {
      "command": "node",
      "args": ["/absolute/path/to/NANITES_MCP/dist/index.js"]
    }
  }
}
```

Run `npm run build` first — `dist/index.js` does not exist until you do.

## Getting started

### 1. Create a profile

A profile describes *your machine* — its VRAM, the LM Studio endpoint, pricing, and what
you want to delegate. This is what makes the guardrail and scoring advice specific.

In Claude Code, say something like:

> Create a Nanites profile called `workstation` with 24 GB VRAM and 32 GB RAM, pointing at
> my LM Studio at http://localhost:1234.

Or call the tool directly: `create_profile` with a name, optional `machine_specs`, and an
optional `endpoint`. If you omit `machine_specs`, a conservative baseline is used
(4 GB VRAM — the safe default that keeps you sequential).

### 2. Point it at your endpoint

If your LM Studio requires an API token, set it on the profile (`endpoint.auth_token`),
or export `NANITES_LMS_API_TOKEN` to cover profiles that leave it null. If LM Studio is
down when you start, run `system_health_check` — it attempts a one-shot `lms server start`
and rechecks before reporting `down`.

### 3. Load a model and delegate something small

> What models do I have? → `list_models`
> Load the small coding model. → `load_model`
> Summarize what this repo's build script does. → `run_sub_agent`

The first `run_sub_agent` is where the machinery engages: the model is hot-loaded if
needed, the inference gate serializes access for your profile's tier, the reply passes
through the cleaner (strips reasoning tags, catches degeneration loops), and the run is
logged so the score can be recomputed next time.

### 4. Build up the registry

Once you have a few models:

- `run_test_regimen` — score a model against the built-in units
- `get_pending_judgments` — pull back the units that need your call
- `submit_test_judgment` — record your score and approve it
- `write_registry_entry` — persist a model and its parameters

After that, `run_sub_agent` picks better models automatically because the registry knows
what each one is actually good at.

### 5. Check what you saved

> How much have I saved? → `get_cost_saved_report`

Reads real logged token counts and computes the not-spent cost at your profile's rates.

## How it works

**Transport.** stdio. The MCP server is a child process; Claude Code owns the pipe.

**Storage.** Everything lives under `NANITES_HOME` (default `~/.nanites`):
- `profiles/*.json` — small, human-editable, rarely written
- `nanites.db` — SQLite (WAL mode) for the registry, call logs, events, test results

**Concurrency.** Your profile's VRAM determines a guardrail tier, which determines a
`(max_parallel_models × num_parallel)` pair. Forced-sequential tiers (< 12 GB) also route
through a per-profile **inference gate** — one in-flight inference at a time — so
overlapping sub-agents queue instead of fighting over a single model slot.

**Providers.** Local LM Studio, or cloud (Cloudflare Workers AI / OpenRouter / OmniRoute /
any OpenAI-compatible endpoint). Cloud runs get a sandboxed filesystem tool loop that
Nanites executes in-process, confined to a configured root.

**Sanitization.** Anything returned to the orchestrator is scrubbed: no reasoning-tag
leakage, no absolute filesystem paths, no usernames, no raw upstream error bodies. Loop
detection in generated text cuts the degenerate tail and flags it rather than silently
shipping it.

## The dashboard

A local web dashboard runs as a *separate* process sharing the same database.

```bash
npm run ui      # http://127.0.0.1:4700
```

Views: Live Execution, Registry, Hardware, Cost, Health, Settings. Live sub-agent output
streams in over SSE by polling the events table — no cross-process IPC.

### Themes

Three switchable profiles, set in Settings and stored on the active profile:

| Theme | Character |
|---|---|
| **Retro Instrument** (default) | E-ink paper-and-ink with a single orange signal, hard 1px borders, corner brackets, a faint dither texture, and a day/night toggle. System monospace only. |
| **Phosphor Terminal** | Green CRT, angular and gritty, with the cursor-evading skull in the header. |
| **Modern Minimal** | Clean dark, quiet typography, sans headings. |

The dashboard binds `127.0.0.1` by default. A Broadcast setting can expose it on the LAN
for a phone or tablet; that mode is **read-only** unless the caller presents the
per-boot LAN token (sent as `X-Nanites-Lan-Token`), and Broadcast warns before enabling
because it exposes local data to the network.

## Configuration

Everything is environment variables — **there is no `.env` loader.** Export them into your
shell (or your MCP server's `env` block) before starting.

| Variable | Purpose | Default |
|---|---|---|
| `NANITES_HOME` | Storage root | `~/.nanites` |
| `NANITES_UI_PORT` | Dashboard port | `4700` |
| `NANITES_LMS_API_TOKEN` | LM Studio auth token (used when the profile has none) | — |
| `NANITES_LMSTUDIO_MODELS_DIR` | Where to measure free disk for downloads | `~/.lmstudio/models` |
| `NANITES_VISION_ROOTS` | Allowed roots for image paths (`;`-separated) | current directory |
| `NANITES_AUTOSTART_UI` | Set `0` to disable dashboard orchestration | `1` |
| `NANITES_HF_FETCH` | Set `1` to allow Hugging Face network lookups | off |

## Tool reference

Nanites exposes the real 53-tool surface, grouped by what you would use it for:

**Models and inference** — `list_models`, `get_loaded_model`, `load_model`, `unload_model`,
`chat`, `download_model`, `get_download_status`, `download_and_wait`, `download_and_test`

**Delegation** — `run_sub_agent`, `start_sub_agent_job`, `get_sub_agent_job_status`,
`start_btw_chat`, `share_test_results`, `get_cost_saved_report`

**Registry** — `read_registry`, `write_registry_entry`, `diff_untested`,
`run_untested_sweep`, `filter_by_guardrail`, `seed_provider_models`, `set_role_pin`,
`list_role_pins`, `delete_role_pin`

**Profiles** — `create_profile`, `switch_profile`, `update_profile`, `list_profiles`,
`get_active_profile`, `get_first_run_status`

**Testing and scoring** — `list_test_units`, `validate_test_unit`, `register_test_unit`,
`run_test_regimen`, `get_pending_judgments`, `submit_test_judgment`, `check_adaptation`,
`register_adapted_units`

**Cloud providers** — `nanites_addProviderKey`, `nanites_removeProviderKey`,
`nanites_listProviderKeys`, `nanites_toggleProviderKey`, `nanites_discoverProviderModels`,
`nanites_listProviderModels`, `nanites_registerProviderModel`,
`nanites_deregisterProviderModel`, `nanites_showProviderErrors`,
`nanites_setProviderEnabled`, `nanites_getProviderConfig`,
`nanites_setProviderPreferenceOrder`

**Health and notifications** — `system_health_check`, `send_ntfy`

Every tool takes schema-validated input and returns a structured envelope —
`{code, message, retryable, details?}` — never a raw throw or stack trace.

## Slash commands

Available once the plugin is loaded:

`/nanites-new-profile`, `/nanites-switch-profile`, `/nanites-profiles`,
`/nanites-models`, `/nanites-registry`, `/nanites-untested`, `/nanites-cost-saved`,
`/nanites-effort`, `/nanites-dynamic-model`, `/nanites-pin`, `/nanites-vision`,
`/nanites-seed-agents`, `/nanites-health`, `/nanites-btw`

`/nanites-btw` opens a side-conversation with a model mid-task — useful when you want to
ask "wait, what does that function do?" without derailing the main thread.

## Development

```bash
npm install
npm run build        # tsc + copy frontend to dist/ui, skill, and plugin server bundle
npm test             # full suite, mocked — no LM Studio or network required
npm run typecheck    # tsc --noEmit
```

The test suite is fully self-contained: 141 test files run against a mock LM Studio
server and a scratch `NANITES_HOME`, so nothing touches your real data or the network.
CI runs typecheck + tests on every push.

Optional live checks (require a running LM Studio with a model loaded):

```bash
npm run live-smoke
npm run live-feature
```

## Known limitations

- **Node 22.5+ is required.** This is a hard floor imposed by `node:sqlite`, not a
  preference. Node 20 will not run it.
- **There is no `.env` loader.** If you rely on a `.env` file in your current workflow,
  you will need to export variables into the environment instead.
- **The cloud filesystem grant executes real writes** when a profile enables it. It is
  confined to a configured root and `write_file` requires explicit opt-in, but it is a
  real capability, not a simulation.
- **`command-runner` integration carries a shell grant.** Keep `tools.enabled: false`
  unless a profile genuinely needs it.
- **Provider API keys are stored in plaintext** in `nanites.db`. The file is created with
  `0600` permissions where the OS supports it, but this is not encrypted at rest.
- **The dashboard's Broadcast mode exposes local data to the LAN.** It is read-only
  without a token, but the token is printed to the console at startup.
- **`live-smoke` timing is sensitive to LM Studio contention.** A loaded machine can
  exceed the script's budget even when everything is working.

## License

[MIT](LICENSE) © 2026 Adnan Shelim. Use it, fork it, ship it commercially.

## Acknowledgements

Built against the [LM Studio REST API](https://lmstudio.ai/docs/app/api) and the
[Model Context Protocol](https://modelcontextprotocol.io) TypeScript SDK.

TDQS

B3.2/5.0

Scored across 53 tools

Disambiguation4/5

Most tools are clearly distinct (e.g., list_models vs get_loaded_model, run_sub_agent vs start_sub_agent_job), but a few pairs could be confused: nanites_deregisterProviderModel vs nanites_removeProviderKey (deregister vs remove), and get_download_status vs download_and_wait overlap in polling behavior. Overall, the core model/test workflows are well separated.

Naming Consistency3/5

Naming is mixed: some tools use snake_case with domain prefix (nanites_*), others use plain verbs (chat, load_model, create_profile), and a few use different conventions (get_cost_saved_report, seed_provider_models). The pattern is not uniform, making it harder to predict tool names, though each individual name is readable.

Tool Count2/5

With 53 tools, the surface is both wide and deep, covering profiles, providers, models, testing, jobs, workflows, and notifications. While each domain has its own cluster, the sheer number exceeds what an agent can easily navigate, and some tools (e.g., send_ntfy, get_cost_saved_report) feel peripheral. A more focused set of ~25-35 would be more manageable.

Completeness4/5

The set covers the full lifecycle for models (list/load/unload/download/test), profiles (create/update/switch/list), providers (add/remove/list/toggle keys, set preference), and test management (validate/register/run/judge). Minor gaps exist: no explicit tool to delete a profile or a test unit, and no tool to list provider errors beyond 'show' (which is read-only). Overall, core workflows are well-covered.

Maintenance

ActivityMaintained
ResponsivenessNo issues