Nanites
Allows using Cloudflare Workers AI as a cloud model provider for delegated inference tasks, including provider key management for Cloudflare-backed models.
Enables Hugging Face network lookups for model information when enabled, supporting model discovery in the registry.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@NanitesUse a local model to summarize this repo's README"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Nanites
An MCP server that lets Claude Code and Claude Desktop delegate bounded, disposable work to your own local models — so you stop paying frontier tokens for grunt work.
Nanites is the delegation layer, not the orchestrator. Claude Code/Desktop stays your orchestrator on a paid frontier model and handles anything requiring real judgment. Nanites hands the cheap, bounded, throwaway work — scanning a codebase, summarizing a file, drafting a test, answering a side question — to a local LM Studio model or a cheap cloud model, and returns the result.
It ships as a Claude Code plugin (slash commands + a delegation skill + a live dashboard) and as a plain MCP server you can register with any MCP-capable harness.
Table of contents
Related MCP server: Local Worker MCP
Why it exists
Delegating work to a small local model is a good idea. Delegating it blindly — without knowing which model is good at what — is how you end up with worse results and a bigger bill. Nanites builds the part that is tedious:
A model registry. Each model gets a
performance_score(1-100), plus rollingavg_load_ms/avg_response_ms, recomputed from your real logged runs. The score is a speed/stability signal, not an oracle — the skill teaches Claude to weigh it alongside the test-regimen results and your own approvals.A test regimen. A built-in suite of modular test units runs against a model. Deterministic ones (valid JSON, exact label) are scored automatically; the rest come back to you for judgment — never auto-filed as passing.
Guardrail tiers keyed to your VRAM. A 4 GB card and a 24 GB card get different concurrency advice, and the advice comes with a stated reason so Claude can explain it.
A cost-saved report. Every delegation is logged with token counts and a computed saving, so the claim is auditable rather than vibes.
Requirements
Requirement | Notes |
Node.js 22.5+ | Hard requirement. Nanites uses the built-in |
LM Studio | Only for local models. The app talks to LM Studio's local HTTP server ( |
Claude Code or Claude Desktop | For the plugin experience. Any MCP-capable client works for the raw server. |
A cloud provider key | Optional. Cloudflare / OpenRouter / OmniRoute / any OpenAI-compatible endpoint. |
No
lmsCLI is required. Nanites will trylms server startas a recovery step if LM Studio is unreachable, but it is best-effort and non-blocking.
Install
As a Claude Code plugin (recommended)
git clone https://github.com/ADn-001/NANITES_MCP.git
cd NANITES_MCP
npm install
npm run buildPoint Claude Code at the plugin directory for this session:
claude --plugin-dir ./plugin/nanitesTo install it permanently as a marketplace plugin:
claude plugin marketplace add ADn-001/NANITES_MCP
claude plugin install nanites@nanitesThe plugin ships its compiled server and its runtime dependencies, so a marketplace install needs no build step.
As a plain MCP server
Register it with any MCP client. For Claude Code, add to your project .mcp.json:
{
"mcpServers": {
"nanites": {
"command": "node",
"args": ["/absolute/path/to/NANITES_MCP/dist/index.js"]
}
}
}Run npm run build first — dist/index.js does not exist until you do.
Getting started
1. Create a profile
A profile describes your machine — its VRAM, the LM Studio endpoint, pricing, and what you want to delegate. This is what makes the guardrail and scoring advice specific.
In Claude Code, say something like:
Create a Nanites profile called
workstationwith 24 GB VRAM and 32 GB RAM, pointing at my LM Studio at http://localhost:1234.
Or call the tool directly: create_profile with a name, optional machine_specs, and an
optional endpoint. If you omit machine_specs, a conservative baseline is used
(4 GB VRAM — the safe default that keeps you sequential).
2. Point it at your endpoint
If your LM Studio requires an API token, set it on the profile (endpoint.auth_token),
or export NANITES_LMS_API_TOKEN to cover profiles that leave it null. If LM Studio is
down when you start, run system_health_check — it attempts a one-shot lms server start
and rechecks before reporting down.
3. Load a model and delegate something small
What models do I have? →
list_modelsLoad the small coding model. →load_modelSummarize what this repo's build script does. →run_sub_agent
The first run_sub_agent is where the machinery engages: the model is hot-loaded if
needed, the inference gate serializes access for your profile's tier, the reply passes
through the cleaner (strips reasoning tags, catches degeneration loops), and the run is
logged so the score can be recomputed next time.
4. Build up the registry
Once you have a few models:
run_test_regimen— score a model against the built-in unitsget_pending_judgments— pull back the units that need your callsubmit_test_judgment— record your score and approve itwrite_registry_entry— persist a model and its parameters
After that, run_sub_agent picks better models automatically because the registry knows
what each one is actually good at.
5. Check what you saved
How much have I saved? →
get_cost_saved_report
Reads real logged token counts and computes the not-spent cost at your profile's rates.
How it works
Transport. stdio. The MCP server is a child process; Claude Code owns the pipe.
Storage. Everything lives under NANITES_HOME (default ~/.nanites):
profiles/*.json— small, human-editable, rarely writtennanites.db— SQLite (WAL mode) for the registry, call logs, events, test results
Concurrency. Your profile's VRAM determines a guardrail tier, which determines a
(max_parallel_models × num_parallel) pair. Forced-sequential tiers (< 12 GB) also route
through a per-profile inference gate — one in-flight inference at a time — so
overlapping sub-agents queue instead of fighting over a single model slot.
Providers. Local LM Studio, or cloud (Cloudflare Workers AI / OpenRouter / OmniRoute / any OpenAI-compatible endpoint). Cloud runs get a sandboxed filesystem tool loop that Nanites executes in-process, confined to a configured root.
Sanitization. Anything returned to the orchestrator is scrubbed: no reasoning-tag leakage, no absolute filesystem paths, no usernames, no raw upstream error bodies. Loop detection in generated text cuts the degenerate tail and flags it rather than silently shipping it.
The dashboard
A local web dashboard runs as a separate process sharing the same database.
npm run ui # http://127.0.0.1:4700Views: Live Execution, Registry, Hardware, Cost, Health, Settings. Live sub-agent output streams in over SSE by polling the events table — no cross-process IPC.
Themes
Three switchable profiles, set in Settings and stored on the active profile:
Theme | Character |
Retro Instrument (default) | E-ink paper-and-ink with a single orange signal, hard 1px borders, corner brackets, a faint dither texture, and a day/night toggle. System monospace only. |
Phosphor Terminal | Green CRT, angular and gritty, with the cursor-evading skull in the header. |
Modern Minimal | Clean dark, quiet typography, sans headings. |
The dashboard binds 127.0.0.1 by default. A Broadcast setting can expose it on the LAN
for a phone or tablet; that mode is read-only unless the caller presents the
per-boot LAN token (sent as X-Nanites-Lan-Token), and Broadcast warns before enabling
because it exposes local data to the network.
Configuration
Everything is environment variables — there is no .env loader. Export them into your
shell (or your MCP server's env block) before starting.
Variable | Purpose | Default |
| Storage root |
|
| Dashboard port |
|
| LM Studio auth token (used when the profile has none) | — |
| Where to measure free disk for downloads |
|
| Allowed roots for image paths ( | current directory |
| Set |
|
| Set | off |
Tool reference
Nanites exposes the real 53-tool surface, grouped by what you would use it for:
Models and inference — list_models, get_loaded_model, load_model, unload_model,
chat, download_model, get_download_status, download_and_wait, download_and_test
Delegation — run_sub_agent, start_sub_agent_job, get_sub_agent_job_status,
start_btw_chat, share_test_results, get_cost_saved_report
Registry — read_registry, write_registry_entry, diff_untested,
run_untested_sweep, filter_by_guardrail, seed_provider_models, set_role_pin,
list_role_pins, delete_role_pin
Profiles — create_profile, switch_profile, update_profile, list_profiles,
get_active_profile, get_first_run_status
Testing and scoring — list_test_units, validate_test_unit, register_test_unit,
run_test_regimen, get_pending_judgments, submit_test_judgment, check_adaptation,
register_adapted_units
Cloud providers — nanites_addProviderKey, nanites_removeProviderKey,
nanites_listProviderKeys, nanites_toggleProviderKey, nanites_discoverProviderModels,
nanites_listProviderModels, nanites_registerProviderModel,
nanites_deregisterProviderModel, nanites_showProviderErrors,
nanites_setProviderEnabled, nanites_getProviderConfig,
nanites_setProviderPreferenceOrder
Health and notifications — system_health_check, send_ntfy
Every tool takes schema-validated input and returns a structured envelope —
{code, message, retryable, details?} — never a raw throw or stack trace.
Slash commands
Available once the plugin is loaded:
/nanites-new-profile, /nanites-switch-profile, /nanites-profiles,
/nanites-models, /nanites-registry, /nanites-untested, /nanites-cost-saved,
/nanites-effort, /nanites-dynamic-model, /nanites-pin, /nanites-vision,
/nanites-seed-agents, /nanites-health, /nanites-btw
/nanites-btw opens a side-conversation with a model mid-task — useful when you want to
ask "wait, what does that function do?" without derailing the main thread.
Development
npm install
npm run build # tsc + copy frontend to dist/ui, skill, and plugin server bundle
npm test # full suite, mocked — no LM Studio or network required
npm run typecheck # tsc --noEmitThe test suite is fully self-contained: 141 test files run against a mock LM Studio
server and a scratch NANITES_HOME, so nothing touches your real data or the network.
CI runs typecheck + tests on every push.
Optional live checks (require a running LM Studio with a model loaded):
npm run live-smoke
npm run live-featureKnown limitations
Node 22.5+ is required. This is a hard floor imposed by
node:sqlite, not a preference. Node 20 will not run it.There is no
.envloader. If you rely on a.envfile in your current workflow, you will need to export variables into the environment instead.The cloud filesystem grant executes real writes when a profile enables it. It is confined to a configured root and
write_filerequires explicit opt-in, but it is a real capability, not a simulation.command-runnerintegration carries a shell grant. Keeptools.enabled: falseunless a profile genuinely needs it.Provider API keys are stored in plaintext in
nanites.db. The file is created with0600permissions where the OS supports it, but this is not encrypted at rest.The dashboard's Broadcast mode exposes local data to the LAN. It is read-only without a token, but the token is printed to the console at startup.
live-smoketiming is sensitive to LM Studio contention. A loaded machine can exceed the script's budget even when everything is working.
License
MIT © 2026 Adnan Shelim. Use it, fork it, ship it commercially.
Acknowledgements
Built against the LM Studio REST API and the Model Context Protocol TypeScript SDK.
Available Tools
53 toolschatChatA
Send a multi-message chat to a loaded model instance. System messages become the system prompt; output passes through the reply validator/cleaner and includes a validation field describing anything stripped.
| Name | Required | Description | Default |
|---|---|---|---|
| params | No | ||
| messages | Yes | ||
| timeout_s | No | ||
| instance_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It meaningfully states that system messages become the system prompt and that output passes through a reply validator/cleaner, including a validation field for stripped content. It does not cover error behavior or side effects, but the disclosed behavior goes well beyond a minimal summary.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences with no wasted words. The core action is front-loaded, and the behavioral details about system messages and the validation field each earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of an output schema and annotations, the description does provide useful return-relevant context via the validation field and the loaded-instance prerequisite. However, it omits guidance about generation parameters, timeout semantics, error cases, and how to obtain or reference a loaded model instance, leaving clear gaps for a tool with nested objects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds helpful meaning for the messages parameter (multi-message, system-role behavior), but it leaves instance_id, timeout_s, and the nested params generation settings entirely to their raw schema names and structure. This is insufficient compensation for the schema's lack of descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Send a multi-message chat to a loaded model instance.' It also explains system-message behavior, which helps identify what the tool does. It does not explicitly differentiate from chat-like siblings such as start_btw_chat, but the 'loaded model instance' scope is fairly distinctive.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'to a loaded model instance' implies the prerequisite of having a model loaded, giving some usage context. However, the description does not explicitly state when to prefer this tool over alternatives, when not to use it, or mention sibling tools that might be more appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_adaptationCheck AdaptationA
Check whether a profile's use case diverges from nanites-default. When it does, returns a user prompt asking whether to draft custom test units, plus the default-plan notice that reusing the default plan may not be well-calibrated.
| Name | Required | Description | Default |
|---|---|---|---|
| profile | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the transparency burden. It clearly signals a read-only check and discloses the main behavioral outcome: a returned user prompt about drafting custom test units plus a default-plan notice. It does not explicitly rule out side effects or describe the negative-branch behavior, but 'Check whether' strongly implies inspection.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence that states the core action immediately and then adds the conditional output. It is reasonably concise, though the phrase 'default-plan notice that reusing the default plan...' is slightly redundant and could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the positive branch well: when divergence is detected, the agent should expect a user prompt and a notice. However, it does not describe what happens when the profile does not diverge, which matters because there is no output schema. It also leaves the exact meaning of 'nanites-default' and the required profile identifier format implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only a string type and minLength for 'profile', with 0% description coverage. The description does add some semantic context by referring to 'a profile's use case,' but it still does not clarify whether the value should be a profile name, ID, or path. This is adequate but leaves a clear gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: it checks whether a profile's use case diverges from nanites-default. It also states the conditional output, a user prompt about drafting custom test units plus a default-plan notice. This is specific enough to distinguish it from siblings like validate_test_unit or register_adapted_units.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains what the tool checks but gives no guidance on when to choose it over alternatives, nor any when-not-to-use conditions. Usage must be inferred from the tool's purpose, so there is no explicit routing or exclusion information.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_profileCreate ProfileC
Create a named profile. Machine specs, endpoint, pricing, test_plan_ref and ntfy are optional and resolve to documented defaults; concurrency is derived from machine specs via the guardrail tiers.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| ntfy | No | ||
| tools | No | ||
| pricing | No | ||
| endpoint | No | ||
| use_case | No | ||
| inference | No | ||
| dynamic_model | No | ||
| machine_specs | No | ||
| test_plan_ref | No | ||
| vision_capable | No | ||
| concurrency_override | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must disclose behavior itself; it does add value by stating that optional fields resolve to documented defaults and that concurrency is derived via guardrail tiers. However, it omits effects such as duplicate-name handling, persistence, whether the new profile becomes active, and whether an override can bypass the guardrail derivation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The two sentences are compact and front-loaded with the core purpose before the defaults note. Minor deduction for being too terse relative to the 12-parameter nested schema, though every included clause carries meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite a large nested schema, no output schema, and no annotations, the description leaves important context uncovered: tool routing, effects of creation, key parameter meanings, and the relationship between concurrency_override and derived concurrency. An agent can call it with only the required name but would be guessing on most optional fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it names only 5 of 12 parameters and omits major ones such as tools, inference, use_case, dynamic_model, and vision_capable. It adds useful semantics for the listed fields, but the unlisted parameters remain unexplained and concurrency_override is not reconciled with the statement that concurrency is derived from machine specs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Create') and resource ('a named profile'), and the emphasis on name aligns with the schema's only required field. It does not explicitly distinguish itself from update_profile or switch_profile, but the creation semantics are unambiguous enough to avoid confusion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to choose this tool over update_profile, switch_profile, or list_profiles, and no conditions or prerequisites are stated. The only usage-related information is that several fields are optional and fall back to defaults, which is more parameter behavior than routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_role_pinDelete Role PinB
Remove a role pin so that role falls back to dynamic model selection.
| Name | Required | Description | Default |
|---|---|---|---|
| role | Yes | ||
| profile | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It does state the behavioral outcome (fallback to dynamic selection), but it does not mention whether removal is permanent, idempotent, or what errors may occur. This adds some context but leaves key undefined behaviors.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, front-loaded with the action and consequence, with no wasted words. It is concise and readable, though the lack of parameter detail limits its practical value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 2 parameters, no annotations, no output schema, and zero schema description coverage, the description is too sparse. An agent cannot tell what 'profile' does, what happens after deletion in detail, or how it interacts with related role pin tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description itself does not explain either the required 'role' parameter or the optional 'profile' parameter. The agent gets no additional meaning beyond the bare schema types, which is inadequate for a tool with multiple parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Remove'), a distinct resource ('role pin'), and a clear consequence ('role falls back to dynamic model selection'). It is easily distinguished from sibling tools like set_role_pin and list_role_pins without needing to inspect schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is implied: use this when you want to remove a role pin and revert to dynamic model selection. However, the description does not explicitly mention alternatives or exclusion conditions (e.g., when to use set_role_pin instead).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
diff_untestedDiff Untested ModelsA
List downloaded LLM models on the profile's endpoint that have no registry entry yet (the Workflow #3 discovery step). Returns a trimmed per-model shape.
| Name | Required | Description | Default |
|---|---|---|---|
| profile | Yes | ||
| verbose | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the full behavioral disclosure burden. It does reveal output style ('trimmed per-model shape') and the filtering logic, but it does not explicitly confirm read-only/no-side-effect behavior, error conditions, or endpoint-related failure modes. The 'List' verb implies safety but does not state it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler. The core action is front-loaded, the distinguishing condition follows immediately, and the return-shape note is the only additional sentence. Every phrase earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list-style tool the description is mostly sufficient, but with no output schema and no parameter explanations, the vague 'trimmed per-model shape' and unspecified 'verbose' behavior leave meaningful gaps. An agent could call it, but not with full certainty about the result shape or option effects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only indirectly references the 'profile' parameter via 'profile's endpoint' and says nothing about the 'verbose' parameter. The agent is left to guess what verbose toggles and what the output shape actually contains.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List'), a specific resource ('downloaded LLM models on the profile's endpoint'), and a clear distinguishing condition ('no registry entry yet'). This differentiates it from siblings like list_models and read_registry by scope, and the parenthetical 'Workflow #3 discovery step' anchors its role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides clear context for when to use the tool: the discovery step of Workflow #3, specifically for models lacking a registry entry. It does not explicitly name alternatives or state when not to use it, so it stops short of a 5, but the intended usage is reasonably clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
download_and_testDownload And TestB
Workflow #4 end to end: download_and_wait, then on completion trigger Workflow #1 (run_test_regimen) on the newly downloaded model automatically. A failed/paused/gave-up download returns without testing.
| Name | Required | Description | Default |
|---|---|---|---|
| source | Yes | ||
| profile | Yes | ||
| quantization | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It does well by stating the sequencing (download_and_wait, then run_test_regimen) and the failure condition (failed/paused/gave-up download returns without testing). This goes beyond what the title alone implies, though it does not describe what happens after testing or what the return value looks like.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler and the core behavior is front-loaded. The references to workflow numbers add some ambiguity, but the text is compact and each sentence contributes meaningful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a composed workflow with no annotations, no output schema, and no parameter documentation, the description is incomplete. It explains the high-level sequence and failure behavior but omits parameter semantics, prerequisites, return format, and what happens if the test itself fails.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not mention the three parameters (source, profile, quantization) at all. The agent is left with only the parameter names and types, which is insufficient for correctly selecting and providing values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: it downloads a model and then automatically triggers the test regimen workflow. It references specific workflow names which add context, though it relies on internal knowledge of 'Workflow #4' and 'Workflow #1' that may not be obvious to an agent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool is used when you want the combined download-and-test sequence, but it does not explicitly say when to use it versus the sibling tools download_and_wait, download_model, or run_test_regimen. There is no when-not-to-use guidance or explicit alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
download_and_waitDownload And WaitB
Workflow #4 download half: ask LM Studio to download a model by HF source id, then poll download status with exponential backoff until completed/failed/paused (or a poll ceiling).
| Name | Required | Description | Default |
|---|---|---|---|
| source | Yes | ||
| profile | Yes | ||
| quantization | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full transparency burden. It usefully discloses the polling strategy, exponential backoff, terminal states, and poll ceiling, but it does not mention prerequisites like LM Studio being available, auth requirements, rate limits, or what the tool returns when it exits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single focused sentence with no fluff, and the core action is front-loaded. The only slightly opaque phrase is 'Workflow #4 download half,' but it does not significantly hurt readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 3 undocumented parameters, no annotations, and no output schema, the description omits important context: what profile means, what quantization values are valid, what the return value is, and what the poll ceiling is. The description captures the workflow shape but not enough to call it reliably in all cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain all three parameters. It only clarifies that 'source' is a Hugging Face source id; 'profile' and 'quantization' are left unexplained, with no format, examples, or required-value context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource: ask LM Studio to download a model by HF source id, then poll download status. It is clear enough to distinguish from siblings like download_model or get_download_status, though it does not name them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The behavior 'then poll download status until completed/failed/paused' implies this tool is for callers who need to wait for a download outcome. However, it gives no explicit when-to-use guidance or alternatives, such as 'use download_model if you only want to start the download.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
download_modelDownload ModelC
Ask the active profile's LM Studio endpoint to download a model by HF source id, optionally pinned to a quantization.
| Name | Required | Description | Default |
|---|---|---|---|
| source | Yes | ||
| quantization | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. The description mentions it asks the endpoint to download, but does not specify whether the download is asynchronous, whether it waits for completion, or what happens if the download fails. It also doesn't disclose rate limits, auth requirements, or potential side effects like populating the model registry. The verb 'download' implies a state-changing operation, but without annotations, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence that front-loads the core action and key parameters. There is no wasted wording; it efficiently conveys the tool's purpose. It earns its place by providing enough detail to understand the basic operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (2 parameters, 0% schema coverage), the description is insufficient. It lacks detail on the return behavior (e.g., does it return a status?), whether the download is blocking, and how it relates to other download variants. With no output schema and no annotations, the description does not fully equip an agent to handle errors or interpret results, leaving an agent to infer too much.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema coverage is 0%, with two parameters that have no descriptions. The description explains the 'source' parameter as an HF source id and 'quantization' as a pinning option, but it does not clarify the format of the HF source id (e.g., whether it includes the repo id or a full URL), nor does it specify acceptable values for quantization (e.g., 'q4_K_M' vs 'int4'). This leaves important semantics undocumented beyond the basic purpose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'download' and the resource 'model by HF source id', with a specific optional parameter 'quantization'. It distinguishes itself from sibling tools like 'download_and_wait' and 'download_and_test' by specifying it is asking the active profile's LM Studio endpoint, rather than downloading and then waiting or testing. However, it does not explicitly differentiate from these siblings or mention alternative tools for downloading with different workflows.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context: it says 'Ask the active profile's LM Studio endpoint to download', which indicates it operates on the active profile, but it does not clarify when to use this tool versus alternatives like 'download_and_wait' or 'download_and_test'. It lacks explicit guidance on when to use this tool versus waiting for completion or testing, and does not mention prerequisites like whether the model must be available or if a profile must be active.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
filter_by_guardrailFilter By GuardrailA
Shortlist candidate models (e.g. Hugging Face search results the host gathered via its HF connector) against the profile's machine-spec guardrail tier. Excludes models far outside the tier's recommended size range with an explicit reason; never silently suggests.
| Name | Required | Description | Default |
|---|---|---|---|
| profile | Yes | ||
| candidates | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It discloses key behaviors: excludes models 'far outside the tier's recommended size range' and does so 'with an explicit reason; never silently suggests'. This informs the agent about the tool's transparency and filtering logic. It does not mention side effects or return format, but the disclosed behavior is valuable and specific.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, tightly written, and front-loads the core purpose. Every clause adds value: the example input, the filtering criterion, and the exclusion behavior. There is no redundancy or filler. It is concise yet informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simplicity of the schema (2 params, no nested required fields) and the absence of an output schema, the description is largely complete. It covers what the tool does, the input source, and key behavior. It does not specify the exact return structure (e.g., whether it returns a modified list or a result with reasons), but that is a minor omission for a filter tool. The description is sufficient for an agent to attempt a call confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the conceptual role of 'profile' (machine-spec guardrail tier) and 'candidates' (models to be shortlisted), but it does not elaborate on the structure of the candidates array (e.g., how size_bytes is used) or the format of the profile string. It adds some meaning beyond raw types but leaves gaps that an agent would need to infer.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Shortlist'), a specific resource ('candidate models'), and a clear criterion ('against the profile's machine-spec guardrail tier'). It distinguishes itself from siblings like list_models (just lists) and load_model (loads) by focusing on filtering/shortlisting a set of candidates. The purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a concrete usage scenario ('e.g. Hugging Face search results the host gathered via its HF connector'), which clarifies when to apply it. However, it does not explicitly state when not to use this tool or point to alternative tools. The 'never silently suggests' phrasing hints at its selection behavior but does not provide exclusions. Overall, the context is clear but lacks explicit alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_active_profileGet Active ProfileA
Return the currently active profile, or null if none has been switched to yet.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of behavioral disclosure. It explicitly describes the main behavior—returning the active profile or null if none has been switched to—which gives the agent accurate expectations about the tool's result.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler. It states the primary behavior and the edge-case return value without wasting words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, parameterless getter with no output schema, the description is complete: it explains what is returned and the null case. Nothing else is needed for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema documents everything there is to document about parameters. The description adds no parameter details, which is appropriate and non-penalized; the baseline for zero parameters is 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Return' and the resource 'currently active profile,' and also specifies the null case. This distinguishes it from sibling tools like list_profiles and switch_profile, which handle different operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The context is clear: an agent should use this tool when it needs the currently active profile. It doesn't explicitly name alternatives or exclusions, but for a zero-parameter getter, the usage context is sufficiently obvious and unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_cost_saved_reportGet Cost Saved ReportA
Report tokens and estimated USD saved by delegating work to local models, from real logged sub-agent usage over a period (all/day/week/month) at the profile's input/output rates. Local delegation costs ~$0, so saved_usd is the orchestrator-equivalent cost not spent.
| Name | Required | Description | Default |
|---|---|---|---|
| period | No | ||
| profile | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It adds meaningful behavioral context: data comes from real logged usage, local delegation costs ~$0, and saved_usd is the orchestrator-equivalent cost not spent. It does not mention read-only behavior or edge cases, but the core semantics are transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler. The main purpose is front-loaded, and the clarifying sentence about local delegation cost and saved_usd meaning earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description covers the essential context: source of data, period options, profile-dependent rates, and the economic meaning of saved_usd. It does not specify the exact return shape, but for a simple report tool this is adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does: it names the period values (all/day/week/month) and explains that profile determines the input/output rates used for the calculation. It does not define the profile string format, but enough meaning is added for correct invocation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Report') and resource ('tokens and estimated USD saved') with clear scope: real logged sub-agent usage over a period at the profile's rates. It is precise enough to distinguish from all sibling tools, none of which share this reporting purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it (after sub-agent usage, to see savings) but provides no explicit when-to-use guidance, exclusions, or alternatives. The sibling list contains related sub-agent tools, but the description does not differentiate this report from them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_download_statusGet Download StatusA
Poll the download progress for a job_id returned by download_model.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral disclosure burden. It does reveal the core trait — that this is a polling operation intended for non-blocking, repeated status checks rather than a blocking wait. However, it does not disclose terminal states (how the agent recognizes completion or failure), whether the result is consumed by polling, or error behavior for invalid/expired job_ids.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Twelve words, front-loaded with the operative verb, and every phrase earns its place: 'Poll' states the mode, 'download progress' the resource, 'job_id returned by download_model' the parameter provenance. Zero filler or repetition of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter polling tool the core usage is adequately covered, but with no output schema and no annotations the description should arguably state what the response contains (progress percentage? status enum?) so an agent knows when to stop polling. This is a real gap: a poller needs a terminal condition to recognize completion.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate — and it does meaningfully. The phrase 'returned by download_model' tells the agent exactly where to obtain the job_id value, which is the single most important semantic fact about the parameter and is completely absent from the schema. It could add format hints, but provenance is the critical gap filled.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Poll'), a precise resource ('download progress'), and anchors the parameter provenance ('job_id returned by download_model'). This immediately distinguishes it from download_model (which starts the job), download_and_wait, and download_and_test without requiring an agent to inspect sibling schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description establishes a clear usage context: this is a polling follow-up to download_model, implying the agent should call it repeatedly after initiating a download. It stops short of naming alternatives or stating when-not-to-use (e.g., preferring download_and_wait for blocking scenarios), but the context is clear enough that no exclusions would be strictly required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_first_run_statusGet First Run StatusA
Report whether Nanites needs first-run initialization: true exactly when zero profiles exist. The orchestrator calls this once at session start; when true, it runs the /nanites-new-profile flow.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It fully defines the semantics ('true exactly when zero profiles exist') and explains the caller's intended use, making it clear this is a read-only status check. No side effects or hidden mutations are implied.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler; the core meaning is front-loaded in the first sentence and the orchestration context occupies only the second. Every word contributes useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter status tool with no output schema and no annotations, the description tells the agent exactly what the boolean means, when to call the tool, and what action to take based on the result. No critical information is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema fully covers the input space (trivially 100% coverage). The description adds no parameter-specific detail, which is appropriate for a parameterless tool, so the baseline score of 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise predicate: reports true exactly when zero profiles exist, making the tool's function unambiguous. The verb 'report' and resource 'Nanites first-run initialization' are specific, and the description frames it as a session-start check rather than a configuration or management operation, distinguishing it from sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says the orchestrator calls this once at session start, giving a clear invocation context. It also prescribes the follow-up action when true ('runs the /nanites-new-profile flow'). It does not mention when not to use it or name alternative tools, but none are obvious, so the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_loaded_modelGet Loaded ModelA
Return the models currently loaded into GPU/CPU on the active profile's LM Studio endpoint (loaded_instances non-empty).
| Name | Required | Description | Default |
|---|---|---|---|
| verbose | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must handle safety and behavior. 'Return' implies read-only, which is helpful, but it doesn't disclose if it can fail when no models are loaded or any side effects. The parenthetical about loaded_instances non-empty hints at return condition but is minimal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, concise and front-loaded with the core purpose. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only getter with one optional parameter, the description is almost sufficient. However, the 'verbose' parameter is undocumented, and the return format is not specified (no output schema). It's adequate for a simple call but could benefit from explaining the verbose behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter 'verbose' is not described in the schema (0% coverage). The description doesn't mention it at all, so the agent has no idea what verbose does. However, having only one parameter, the description's omission is a significant gap, but the baseline for no schema coverage is 4, and the description adds no value here, so a 3 might be fair. Given it's a simple boolean, I'd give 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool returns currently loaded models from the LM Studio endpoint, distinguishing it from list_models (likely all available) and load_model/unload_model. Specific resource and action are clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies use when you need to know what's loaded, but doesn't explicitly contrast with load_model or unload_model or mention when not to use. Among siblings, list_models is a potential alternative but no differentiation is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_pending_judgmentsGet Pending JudgmentsB
Fetch cleaned raw output plus rubric and prompt context for pending orchestrator-judged units (all, or a requested subset) for the orchestrator to read and judge in the same turn.
| Name | Required | Description | Default |
|---|---|---|---|
| profile | Yes | ||
| model_id | Yes | ||
| unit_ids | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It communicates that the tool fetches data for reading and judging, which suggests a non-destructive operation, and it discloses key content categories (raw output, rubric, prompt context). It does not disclose possible side effects, pagination, ordering, or how 'pending' is determined.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence and is reasonably compact, front-loading the main action. Some jargon ('cleaned raw output', 'orchestrator-judged units') adds density, but there is no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with three parameters, no annotations, and no output schema, the description leaves key gaps: the meaning of `profile` and `model_id`, the format of the returned data, how subsets are expressed, and what 'cleaned raw output' actually contains. An agent would likely need external context to invoke this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaning for the optional subset concept ('all, or a requested subset') but never names or explains `unit_ids`. The required `profile` and `model_id` parameters are entirely unexplained, leaving the agent to guess their roles.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Fetch') and a specific resource: cleaned raw output plus rubric and prompt context for pending orchestrator-judged units. It clearly implies a read-oriented tool for judgment work, but it does not explicitly contrast with siblings like submit_test_judgment, so differentiation is somewhat implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for the orchestrator to read and judge in the same turn' implies when to use the tool, and 'all, or a requested subset' gives a scoping choice. However, there is no explicit guidance about when to prefer this tool over alternatives or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_sub_agent_job_statusGet Sub-Agent Job StatusA
Poll a sub-agent job started with start_sub_agent_job. Returns { job_id, status: queued|running|done|error, result? }; result is present for done (shaped like run_sub_agent's response) and error (structured { code, message, retryable }).
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states that the tool is for polling (implying a non-destructive read operation) and describes the return structure, including statuses and the presence of result for terminal states. This gives the agent enough to understand the expected behavior without additional annotation support.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, using one sentence to state the tool's purpose and return format. It front-loads the critical action (poll) and resource (sub-agent job), then immediately provides the response structure. Every word is informative with no redundancy, making it highly efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple polling tool with one parameter and no output schema, the description is complete. It covers the when-to-use (referencing start_sub_agent_job), the return format, and the possible statusesahanincluding result shape for done and error cases. An agent can correctly invoke this tool without needing additional information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must compensate. The description does not explicitly explain the job_id parameter beyond referencing a job started by start_sub_agent_job, which implies the job_id value. However, the schema already specifies the parameter name and type, so the description adds some context but not detailed semantics. A baseline of 3 is appropriate given the single parameter and the inferred meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: polling a sub-agent job. It specifies the exact resource (sub-agent job) and the action (poll status), and it distinguishes itself from the sibling tool start_sub_agent_job by referencing it directly. The description also details the return structure, making it unambiguous what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use this tool: to poll a job started with start_sub_agent_job. It does not list alternatives, but given the context of sibling tools, start_sub_agent_job is the only related one, and the description's reference to it provides clear usage context. There are no exclusions, but the usage direction is precise.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_modelsList ModelsA
List the models known to the active profile's LM Studio endpoint. Returns a trimmed shape by default; pass verbose for the full model objects.
| Name | Required | Description | Default |
|---|---|---|---|
| verbose | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full behavioral disclosure burden. It discloses the key output behavior: 'Returns a trimmed shape by default; pass verbose for the full model objects.' It does not state side-effect absence or error behavior, but for a read-only list operation the default/verbose disclosure covers the main decision an agent faces.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero filler, the verb+resource front-loaded in the first sentence. The second sentence earns its place by explaining the default behavior and the parameter's effect.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one optional boolean, no output schema, and no annotations, the description adequately covers scope and return-shape behavior. Minor omissions such as what the trimmed shape contains and behavior when the endpoint is unreachable are low-stakes for a simple list operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: 'pass verbose for the full model objects' plus the default trimmed shape gives the boolean real meaning beyond the bare schema type. It stops short of specifying which fields appear in each shape, but the semantic distinction is clear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'List the models known to the active profile's LM Studio endpoint.' The scope phrase distinguishes it from sibling tools like nanites_listProviderModels and get_loaded_model, which concern provider-discovered or currently loaded models.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a usage context by scoping the result set to the active profile's LM Studio endpoint, which separates it from provider-discovery tools. However, it never names alternatives or states when not to use this tool, so an agent must infer routing from the sibling list rather than receiving explicit guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_profilesList ProfilesA
List profile names. Pass verbose for full profile objects.
| Name | Required | Description | Default |
|---|---|---|---|
| verbose | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It doesn't explicitly state that this is a read-only operation, nor does it disclose any side effects, return format, or pagination. However, as a list operation, the absence of any mutation hints is neutral, and the description does reveal that verbose changes the output detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler. The primary action and the parameter effect are front-loaded. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a list tool with one optional parameter and no output schema, the description covers the essential behavior: what it lists and how the flag alters the output. It lacks details like error conditions or ordering, but these are minor for this simple operation. The description is complete enough for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only a boolean 'verbose' parameter with no description (0% schema coverage). The description fully compensates by explaining that passing verbose yields full profile objects, while default behavior lists profile names. This is exactly the kind of clarification an agent needs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action ('List') on a specific resource ('profile names'), and distinguishes it from sibling tools like create_profile, switch_profile, and get_active_profile. The phrase 'profile names' vs 'full profile objects' also hints at a scoping detail that sets it apart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it (to list profiles) but provides no explicit exclusions or alternatives. It doesn't say 'use this when you need to enumerate all profiles' or mention any sibling for comparison. For a simple list tool, the usage is fairly obvious, but it lacks explicit guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_role_pinsList Role PinsB
List the role->model pins for a profile (default active).
| Name | Required | Description | Default |
|---|---|---|---|
| profile | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description itself must convey safety and behavior. 'List' implies a read-only operation and 'default active' adds useful behavior, but it does not disclose response format, whether an invalid profile produces an error, or any ordering/filtering behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no wasted words. It communicates the operation, the resource, the profile scope, and the default behavior efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read tool with one optional parameter, the description is mostly sufficient for invocation. However, with no output schema and no mention of return values or error behavior, an agent is left to guess what the tool actually returns beyond a vague idea of a pin listing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema only defines 'profile' as a string with minLength 1 and 0% description coverage. The description adds the key meaning that the parameter is optional and defaults to the active profile, but it does not clarify what identifier values are expected (name vs. id) or how it relates to switch_profile.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb ('List') and a specific resource ('role->model pins'), and notes the profile scope. It is easy to distinguish from sibling mutation tools like set_role_pin and delete_role_pin, though it does not explicitly name any alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance about when to use this tool versus alternatives such as list_models or the pin mutation tools. The 'default active' note gives some context, but the description does not state when this tool is the right choice or when another tool should be used.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_test_unitsList Test UnitsA
List the registered test units for a profile (the default regimen is registered automatically on first use).
| Name | Required | Description | Default |
|---|---|---|---|
| profile | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It does add a useful behavioral nuance about automatic default-regimen registration, which helps the agent understand what appears in the list. However, it does not explicitly confirm that this is a read-only operation or disclose any other behavioral traits like profile existence checks or potential errors.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that gets straight to the point: 'List the registered test units for a profile'. The parenthetical adds a relevant behavioral note without any fluff. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool with one parameter and no output schema, the description covers the primary purpose and a useful edge case. However, it lacks mention of what a 'profile' is or where to obtain it, and it doesn't hint at the return format (e.g., array of unit names). Given the broader toolset includes profile-related siblings like `list_profiles`, a bit more context would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It identifies 'profile' as the scope but provides no additional detail beyond that—no format (ID vs. name), no source, no validation rules. The single parameter is minimally explained, leaving the agent to infer what a 'profile' is from the tool name and context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('List'), a resource ('test units'), and a scope ('for a profile'), making the tool's purpose immediately clear. It is easily distinguished from sibling tools like `register_test_unit` or `run_test_regimen` because it names the read-only listing behavior explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool (to see registered test units) and adds the nuance that the default regimen is auto-registered, but it does not explicitly contrast with alternatives or state when not to use it. There is no mention of using this instead of `validate_test_unit` or `run_test_regimen`, leaving the guidance somewhat implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
load_modelLoad ModelA
Load a model into the active profile's LM Studio endpoint. model_id is the model key (the identifier shown by list_models); optional load params (context_length, flash_attention, ...) may be passed.
| Name | Required | Description | Default |
|---|---|---|---|
| params | No | ||
| model_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It does convey the primary effect (loading into the endpoint) and that params are optional, but it does not mention side effects on the previously loaded model, prerequisites (e.g., model being downloaded), or whether the call blocks until the load completes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no fluff: the first states the action, the second clarifies the key parameter and optional nature of the rest. Every clause earns its place and the most important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with a nested params object and no output schema, the description is adequate but incomplete. It covers the core action and model_id source, but omits prerequisites (e.g., endpoint running, model downloaded), potential side effects, and any indication of the response or failure behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds crucial meaning for model_id by defining it as the key from list_models, and it indicates that params are optional with examples (context_length, flash_attention). However, it does not explain the semantics of the remaining nested parameters (num_experts, eval_batch_size, etc.), leaving significant gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Load') and a precise resource ('a model into the active profile's LM Studio endpoint'), clearly distinguishing this tool from siblings such as list_models, get_loaded_model, and unload_model. It also ties the action to the active profile, which is a concrete context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage context by explaining that model_id is the key shown by list_models, implying the workflow of listing models first. It does not explicitly list excluded alternatives, but the tool's purpose is distinct enough and the connection to list_models gives practical guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
nanites_addProviderKeyAdd Provider KeyC
Add an API key for a cloud provider (Cloudflare, OpenRouter, OmniRoute, Generic).
| Name | Required | Description | Default |
|---|---|---|---|
| api_key | Yes | ||
| provider | Yes | ||
| account_id | No | ||
| gateway_url | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It only implies that the tool mutates state by adding a key. It does not disclose overwrite behavior, validation checks, authentication needs, side effects on existing keys, or failure modes, leaving the agent without critical safety information.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tight sentence with no wasted words. It front-loads the core action and provider list, making it easy to parse. For the amount of content it contains, it is impeccably concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has no annotations, no output schema, and four under-documented parameters, this description is severely incomplete. It does not mention prerequisite conditions, return value, error behavior, or the meaning of optional fields, so an agent cannot reliably predict the tool's behavior or handle edge cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it fails to explain any parameters. It merely repeats the provider enum values already in the schema. The purposes of api_key, account_id, and gateway_url are entirely unexplained, especially the optional parameters, leaving the agent to guess their meaning and correct usage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (Add), the resource (API key for a cloud provider), and enumerates the supported providers. This distinguishes it from siblings like nanites_removeProviderKey and nanites_listProviderKeys, leaving no ambiguity about the tool's function.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance is given on when to use this tool versus the many related provider-key tools. The verb 'Add' implies usage for creating a new key, but there is no mention of prerequisites, such as needing to add a key before using a provider, or when to prefer remove, list, or toggle alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
nanites_deregisterProviderModelDeregister Provider ModelB
Deregister a registered model from a cloud provider.
| Name | Required | Description | Default |
|---|---|---|---|
| model_id | Yes | ||
| provider | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the full burden of behavioral disclosure. It conveys that the operation mutates registration state, but says nothing about side effects, reversibility, failure conditions, or what happens if the model is currently in use.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tightly worded sentence with no filler. The action and object are front-loaded, making it efficient for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a state-changing tool with no annotations and no output schema, this is too thin. An agent cannot determine what deregistration entails, how to confirm success, or what errors to expect, so it must rely on schema and sibling names to fill critical gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description adds essentially no meaning beyond the word 'model' and 'cloud provider'. The schema supplies the provider enum, but the description does not clarify model_id format, how to obtain it, or provider-specific nuances.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Deregister'), the object ('a registered model'), and the scope ('from a cloud provider'). This clearly distinguishes the tool from siblings like nanites_registerProviderModel, nanites_listProviderModels, and nanites_setProviderEnabled.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus related alternatives such as nanites_setProviderEnabled, nanites_removeProviderKey, or unload_model. The phrase 'a registered model' implies a prerequisite, but it is not explicit and no ordering or idempotency guidance is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
nanites_discoverProviderModelsDiscover Provider ModelsC
Auto-discover available models from a cloud provider.
| Name | Required | Description | Default |
|---|---|---|---|
| provider | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, but it only states that models are auto-discovered. It does not disclose whether this is a read-only query, whether it mutates or registers models in storage, whether network calls occur, or whether provider credentials are required.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler or redundancy. It earns its place by adding 'auto', 'available', and 'from a cloud provider' beyond the title, though it is perhaps too terse to carry the full explanatory load.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description should explain what a caller receives and what side effects or preconditions exist. It does neither, leaving important operational ambiguity for a tool that sits among registry-management siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not compensate. The only parameter, 'provider', is backed by an enum, but the description simply says 'from a cloud provider' without explaining distinctions among cloudflare, openrouter, omniroute, and generic, leaving genuine ambiguity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Auto-discover') and resource ('available models from a cloud provider'), which goes beyond a mere restatement of the title. It does not explicitly differentiate from sibling tools like nanites_listProviderModels or seed_provider_models, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus related siblings such as nanites_listProviderModels, seed_provider_models, or list_models. The description also omits prerequisites like provider configuration or API key requirements.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
nanites_getProviderConfigGet Provider ConfigA
Get provider preference order and per-provider settings.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It only says 'Get', which implies read-only, but it does not explicitly state that it does not modify anything, does not require authentication, or what happens if no configuration exists. For a getter, this is minimal but not fully transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that front-loads the action and the content. There is no wasted language. It is appropriately sized for a tool with no parameters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple getter with no parameters and no output schema, the description is adequate but minimal. It tells what the tool returns but does not mention any prerequisites, error conditions, or the exact structure of the response. Given the lack of annotations and output schema, an agent might benefit from more detail on the return shape or potential side effects, but the simplicity of the tool mitigates the gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, so the description needs to add nothing about parameters. However, it does add meaning by specifying what the returned config includes ('preference order and per-provider settings'), which is more than just 'get config'. Since schema coverage is trivially 100% and there are no params, the baseline is 4, and the description adds a bit of useful detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Get') on a specific resource ('provider config') and details what it returns ('preference order and per-provider settings'). This clearly distinguishes it from sibling tools like nanites_setProviderPreferenceOrder, which is a setter, and other config-related tools. It is not a tautology.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives. It does not mention any conditions, exclusions, or mention related tools like nanites_setProviderPreferenceOrder. The agent is left to infer that this is the getter counterpart to the setter, but no explicit routing is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
nanites_listProviderKeysList Provider KeysA
List all API keys for a cloud provider.
| Name | Required | Description | Default |
|---|---|---|---|
| provider | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It clearly indicates a read-only listing operation and the 'all' quantifier adds scope, but it does not disclose output shape, redaction of secrets, or error behavior. This is acceptable for a simple list but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, no filler, front-loaded verb and resource. Every word contributes to the meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple and the description covers its core purpose, but without an output schema or annotations it does not state what a successful response looks like (e.g., full keys vs metadata, pagination). An agent could call it correctly but might be surprised by the return shape.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It ties the lone parameter to the cloud provider ('for a cloud provider'), but does not explain the enum values or requiredness, which are left to the schema. This is minimal but sufficient for one simple parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the verb 'List' and the resource 'API keys', making the operation unambiguous. The 'all' and 'for a cloud provider' scope clearly distinguishes it from sibling key-management tools such as nanites_addProviderKey and nanites_removeProviderKey.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, when-not-to-use, or alternative tools are mentioned. The verb 'List' implies a retrieval scenario, but the description does not say to prefer this over other list tools or explain when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
nanites_listProviderModelsList Provider ModelsC
List cached models from a cloud provider.
| Name | Required | Description | Default |
|---|---|---|---|
| provider | Yes | ||
| registered_only | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry behavioral information. It mentions 'cached' which implies it reads from a local cache rather than calling the provider, a useful behavioral trait. However, it doesn't disclose error handling, rate limits, authentication, or whether the result is ordered or filtered. For a simple list operation, this is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that is easy to parse and front-loads the core action. However, its brevity comes at the cost of missing parameter and usage details, so it is appropriately sized for the information it does convey.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no parameter descriptions, the description is incomplete. It doesn't explain the meaning of the enum values, what 'registered_only' does, or what the return format looks like. For a tool with 2 parameters and a sibling set with similar names, more context is needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain the parameters. It does not mention 'provider' or 'registered_only' at all, leaving the agent to rely solely on the schema enum and boolean type. No additional meaning is added.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb-resource pair: list cached models from a cloud provider. It adds 'cached' which distinguishes it from other listing tools like list_models or nanites_discoverProviderModels, but does not explicitly differentiate beyond that, so not a perfect 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives like nanites_discoverProviderModels or list_models. The description only states what it does, not when to prefer it. Given the large sibling set, this is a significant gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
nanites_pingNanites PingA
Liveness probe for the Nanites MCP server. Returns server identity and transport health.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Because no annotations are provided, the description carries the full behavioral burden. It does disclose the tool's read-only nature implicitly through 'Liveness probe' and specifies the return content (server identity and transport health). It does not describe failure behavior, timeouts, or the exact shape of the response, but for a zero-parameter ping this is reasonably transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with no filler. It front-loads the core purpose ('Liveness probe') and immediately states what the caller receives, making it easy to scan and understand.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple zero-parameter liveness check with no output schema, the description adequately explains the tool's purpose and high-level return values. It is not fully complete because it lacks detail on the exact response format and does not clarify how it differs from the overlapping `system_health_check` sibling, but the missing detail is not critical for invoking this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero properties, and schema description coverage is 100%, so there are no parameters to document. The description correctly needs to add nothing about parameters, matching the baseline for a zero-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool is a 'Liveness probe for the Nanites MCP server' and that it 'Returns server identity and transport health.' This gives a specific verb and resource. However, it does not explicitly distinguish itself from the similarly-named sibling `system_health_check`.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Liveness probe' strongly implies this tool is used to check whether the server is alive and healthy. However, there is no explicit guidance about when to prefer this over sibling tools like `system_health_check`, and no exclusions or alternative usage instructions are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
nanites_registerProviderModelRegister Provider ModelC
Register a discovered model for use with a cloud provider.
| Name | Required | Description | Default |
|---|---|---|---|
| model_id | Yes | ||
| provider | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full responsibility for behavioral disclosure. It does not state whether registration is persistent, whether it overwrites existing entries, whether provider credentials are required, or what side effects occur. 'Register' only implies mutation, not the operational details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. It communicates the core action and object directly, making it appropriately concise for a simple registration tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and minimal parameter detail, the description is too thin for an agent to confidently call the tool correctly. It leaves key operational context—such as valid provider states, overwrite behavior, and expected postconditions—unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description adds only minimal context: model_id is a 'discovered model' and provider is a 'cloud provider.' The provider enum helps, but model_id semantics remain ambiguous, and no relationship between the two parameters is explained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear action ('Register') and a specific object ('a discovered model') plus the context ('for use with a cloud provider'). This distinguishes it from deregistration, listing, and discovery tools, though it does not explicitly contrast it with sibling tools like seed_provider_models.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit when-to-use guidance, prerequisites, or alternatives. The only implicit cue is 'discovered model,' which suggests this should happen after discovery, but the agent is left to infer that from the wording alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
nanites_removeProviderKeyRemove Provider KeyC
Remove an API key from a cloud provider.
| Name | Required | Description | Default |
|---|---|---|---|
| key_id | Yes | ||
| provider | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It only states 'Remove an API key' which implies a destructive action, but it does not disclose potential side effects (e.g., whether associated models are also deregistered, whether the action is reversible, required permissions, or confirmation mechanisms). The description does not contradict any annotations since none exist, but it offers no behavioral depth beyond the verb 'remove'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single clear sentence with no wasted words. It is appropriately short for a simple operation and front-loads the core intent. However, it is under-specified in terms of parameters and usage, but that is captured in other dimensions. The structure itself is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations, no output schema, and zero parameter descriptions, the description is incomplete. It does not mention any return value, error conditions, idempotency, or whether the key must be currently linked to a provider. An agent cannot fully anticipate the tool's behavior or side effects, making the description inadequate for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, meaning the description must explain the parameters. It does not. The description does not clarify what 'key_id' refers to, how to obtain it, or how the 'provider' enum relates to key removal. The agent must rely solely on parameter names and the enum values, which is insufficient for correct invocation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: 'Remove an API key from a cloud provider.' This is a specific verb-resource pairing that immediately distinguishes it from sibling tools like nanites_addProviderKey, nanites_listProviderKeys, and nanites_toggleProviderKey. No ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus its alternatives. It does not mention that this is the inverse of addProviderKey, nor does it explain any prerequisites (e.g., whether the key_id must exist or whether provider must be active). The agent must infer usage context from the name and siblings, which is insufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
nanites_setProviderEnabledSet Provider EnabledB
Enable or disable a cloud provider globally.
| Name | Required | Description | Default |
|---|---|---|---|
| enabled | Yes | ||
| provider | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It merely states the action (enable/disable) without revealing side effects, persistence, permission requirements, or whether changes affect existing sessions. It is minimally transparent beyond the core action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, succinct sentence with no filler. It is perfectly concise and immediately states the tool's function.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool, the description is adequate but not complete. It tells what it does but omits any context about when to use it, potential side effects, or how it interacts with other provider management operations. Given no annotations and no output schema, more context would improve an agent's ability to use it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It implies that 'provider' selects the cloud provider and 'enabled' toggles its state, but it does not explain the enumerated provider values (e.g., cloudflare, openrouter) or any parameter constraints. The description adds only basic meaning already inferable from the parameter names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb (enable/disable) and a clear resource (a cloud provider) with a global scope, which distinguishes it from sibling tools like toggling provider keys or setting preference order. It directly states what the tool does without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives, nor any prerequisites or conditions for calling it. There is no mention of when not to use it or how it relates to other provider management tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
nanites_setProviderPreferenceOrderSet Provider Preference OrderC
Set the global provider preference order for cloud routing.
| Name | Required | Description | Default |
|---|---|---|---|
| order | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It doesn't state that this operation persists globally, whether it overrides existing order, if it affects ongoing routing immediately, or any side effects on other providers. 'Global' hints at scope but not consequences; an agent cannot anticipate what happens to current routing until after invocation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, compact sentence that states the core action and scope. It is appropriately short for a simple one-parameter tool. However, it could be slightly more informative without being verbose, e.g., by mentioning 'order values must be from the supported provider list' or 'affects routing globally', but the structural efficiency is good.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has one parameter and no output schema, the description should suffice for straightforward invocation, but it lacks critical context: what the order array represents (priority ranking), the effect on routing, and any constraints (e.g., must include all providers or at least one). The sibling context shows many configuration tools, but this description doesn't explain how this order interacts with other settings like enabled providers. An agent might not know if setting an order alone is sufficient or if additional steps are needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there are no enum annotations on the parameter, though the schema defines an array of strings with an enum list in the items. The description does not explain what values should be in the order, the expected order of preference (e.g., first is most preferred?), or whether all providers must be listed or a subset is allowed. The agent must read the schema and infer semantics from the enum names, so the description adds no meaning beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Set') and resource ('global provider preference order'), and it distinguishes from siblings by indicating it sets the global order (not per-provider or per-key operations). However, it doesn't specify what 'provider preference order' means practically or how it affects routing, leaving some ambiguity for an agent; it's clear enough to be a minimal viable purpose but not distinct enough to avoid confusion with configuration tools like nanites_setProviderEnabled.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives. With many nanites_* configuration tools in the siblings, an agent could confuse this with setting provider keys or enabling/disabling providers. No context is given on the typical use case (e.g., during initial setup or to change routing priorities), so the agent must infer usage from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
nanites_showProviderErrorsShow Provider ErrorsC
Show recent errors from cloud provider calls.
| Name | Required | Description | Default |
|---|---|---|---|
| filter | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It implies a read-only operation ('show') but doesn't explicitly state that no side effects occur, nor does it mention any permissions, rate limits, or output format. The description is too minimal to convey meaningful behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with no filler. It front-loads the verb and resource. While it is under-specified, the conciseness itself is good—there is no wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one optional parameter, no output schema, and no annotations, the description should at least explain the filter's purpose and the nature of the errors returned. It does neither. The tool is simple, but the description leaves critical operational details unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter 'filter' is completely undocumented in the schema, and the description makes no mention of it whatsoever. With 0% schema description coverage, the description must compensate but fails entirely, leaving the agent with no understanding of what values the filter accepts or how it affects results.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'show' and the resource 'errors from cloud provider calls,' making the tool's purpose unmistakable. It distinguishes itself from siblings by focusing on errors, which no other nanites tool mentions. However, it doesn't explicitly contrast with any sibling, so it doesn't fully capitalize on differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. It doesn't mention any prerequisites, context, or exclusions. An agent receives no help in deciding between this and other nanites_* tools that might also relate to provider status.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
nanites_toggleProviderKeyToggle Provider KeyC
Enable or disable an API key for a cloud provider.
| Name | Required | Description | Default |
|---|---|---|---|
| key_id | Yes | ||
| enabled | Yes | ||
| provider | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It reveals the core action ('enable or disable') but does not mention side effects, idempotency, whether disabling a key invalidates existing usage, or any authorization requirements. This is thin for a state-changing operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler. It communicates the core purpose efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with three required parameters, no output schema, and no annotations, the description is too minimal to be fully actionable. It lacks usage context, parameter clarification, and behavioral details, leaving an agent to infer too much.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not compensate by explaining what key_id refers to, how provider values map, or what the enabled boolean controls beyond the obvious. The parameter names and enum provide some self-evident meaning, but the description adds no real semantic value over the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action ('Enable or disable') on a specific resource ('API key for a cloud provider'), which distinguishes it from sibling tools that add, remove, or list keys, or enable providers. However, it does not explicitly note that the key must already exist or contrast itself with the similarly named nanites_setProviderEnabled.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to use this tool versus alternatives like nanites_addProviderKey, nanites_removeProviderKey, or nanites_setProviderEnabled. The description gives no context about prerequisites or expected use cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_registryRead RegistryA
Read model-registry entries (roles, scores, best params, last tested) for a profile, optionally filtered to one model_id. Trimmed by default; verbose includes timestamps.
| Name | Required | Description | Default |
|---|---|---|---|
| profile | Yes | ||
| verbose | No | ||
| model_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the burden of behavioral disclosure. It adds useful behavior by stating that output is 'Trimmed by default; verbose includes timestamps,' and the verb 'Read' implies no mutation. However, it does not disclose error behavior, whether a missing profile/model_id is tolerated, or output shape limitations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences front-load the core action and resource, then add optional filtering and output-format details. No filler or redundant restatement of the tool name appears.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read operation with no output schema, the description covers the purpose, parameters, and key behavioral nuance (trimming/verbose). It could be more complete by noting what happens when no entries match or when the profile does not exist, but nothing essential to invoking the call correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description compensates by assigning meaning to all three parameters: profile selects the registry context, model_id filters to one entry, and verbose expands timestamps. It adds semantic value beyond the bare string/boolean schema types, though requiredness of profile is left implicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Read'), a concrete resource ('model-registry entries'), and lists the entry fields (roles, scores, best params, last tested). It states the profile scope and optional model_id filtering, making it clearly distinct from siblings like write_registry_entry.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly identifies what data the tool operates on and the available scoping (profile, optional model_id), which tells an agent when this is the right read call. It does not explicitly name alternatives or exclusion criteria, but the read-versus-write contrast with siblings is implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
register_adapted_unitsRegister Adapted UnitsA
Validate a batch of authored test units through the Phase 4 validator and register the ones that pass; each rejected unit is returned with its validation issues surfaced, never silently dropped or force-registered.
| Name | Required | Description | Default |
|---|---|---|---|
| units | Yes | ||
| profile | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full disclosure burden, and it delivers solid value: 'each rejected unit is returned with its validation issues surfaced, never silently dropped or force-registered' is a genuine behavioral commitment about failure handling that an agent would otherwise not know. It does not cover error behavior on partial registration failure or rate limits, but the core mutation semantics (validate-then-register) are well disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with the core action front-loaded before the behavioral guarantee. No wasted words, though the syntax is slightly run-on; every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core flow (validate, register passing, return rejected with issues) well for a mutation tool. But it leaves the required `profile` parameter unexplained and gives no indication of return format beyond the rejected-units note. Moderate complexity with a loose `units` schema warrants a bit more, so this is adequate but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for both parameters. It explains `units` as 'authored test units' in a batch, but the required `profile` parameter is never mentioned or explained — an agent cannot infer what profile means or how it affects registration. A 0% coverage schema with an unspecified required parameter is a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Validate ... test units ... register the ones that pass') with a concrete context (batch, Phase 4 validator). The batch scope clearly differentiates it from sibling singles like validate_test_unit and register_test_unit, so an agent can tell them apart without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this is the batch variant by saying 'a batch of ... test units', which hints at when to reach for it over the single-unit siblings. However, it never explicitly names the alternatives or states when NOT to use this tool, leaving the routing decision to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
register_test_unitRegister Test UnitB
Validate and persist a test unit for a profile. Rejects invalid units with the issues listed; duplicate ids are refused.
| Name | Required | Description | Default |
|---|---|---|---|
| unit | Yes | ||
| profile | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of behavioral disclosure. It explicitly reveals that invalid units are rejected with issues and that duplicate ids are refused, giving an agent concrete error-path and idempotency expectations. It does not cover permissions or side effects, but core behavior is meaningfully disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately compact and front-loads the verb and resource. The phrase 'with the issues listed' is slightly ambiguous because no list appears in the description, but otherwise every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with a nested object parameter, no annotations, no output schema, and zero schema coverage, this description is insufficient for correct invocation. Missing details include the unit's shape, validation criteria, where the duplicate id comes from, and the return/error format.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clarifies that 'profile' is the target and 'unit' is the test unit, and it implies an id exists for duplicate checks, but it does not explain unit structure, required fields, or id semantics. An agent cannot construct a valid 'unit' from this text alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Validate and persist a test unit for a profile.' This clearly distinguishes it from sibling tools like validate_test_unit by emphasizing persistence, and the rejection/duplicate behavior adds further specificity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives such as validate_test_unit or register_adapted_units. The purpose is implied, but there are no explicit exclusions, prerequisites, or routing conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_sub_agentRun Sub-AgentA
Delegate one bounded task to a local model. Resolves the model from the profile's registry by roles (or an explicit model_id), acquires a model respecting the profile's concurrency tier (reuse already-loaded, evict on sequential tiers, refuse at parallel capacity), runs the brief once, cleans the reply, unloads exactly once if it loaded the model, and logs exactly one token-usage entry for cost tracking.
| Name | Required | Description | Default |
|---|---|---|---|
| task | No | ||
| brief | Yes | ||
| roles | No | ||
| effort | No | ||
| images | No | ||
| profile | Yes | ||
| model_id | No | ||
| provider | No | ||
| output_schema | No | ||
| reasoning_budget | No | ||
| output_schema_name | No | ||
| system_prompt_override | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it delivers: it discloses model resolution, concurrency-tier behavior (reuse, evict, refuse), single execution, reply cleaning, exact-once unload, and exactly one token-usage log entry. This is unusually transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence, but it is front-loaded with the core purpose and every clause adds behavioral detail. It is somewhat long, but not bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite rich behavioral detail, the tool has 12 parameters, no annotations, and no output schema. The description omits semantics for most parameters and does not describe return values or error conditions, leaving significant gaps for an agent trying to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains profile, roles, model_id, and brief, but 12 parameters exist and many (effort, images, provider, output_schema, reasoning_budget, system_prompt_override, task) receive no semantic guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Delegate one bounded task to a local model.' It clearly distinguishes this tool from siblings like start_sub_agent_job or chat by emphasizing a one-shot, bounded, local-model delegation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'one bounded task' provides clear context for when to use this tool. It does not explicitly name alternatives or exclusions, but the bounded-task framing and detailed execution flow make the intended use apparent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_test_regimenRun Test RegimenA
Test one model against the profile's registered test units (the default regimen is auto-registered on first use). Loads the model, runs deterministic units inline with a logged param search, runs orchestrator_judged units to pending, unloads exactly once, and writes the registry entry when nothing is pending. Returns only pending_unit_ids, never raw output.
| Name | Required | Description | Default |
|---|---|---|---|
| profile | Yes | ||
| model_id | Yes | ||
| provider | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it delivers richly. It discloses the full side-effect chain: loads the model, runs deterministic units inline with a logged param search, moves orchestrator_judged units to pending, unloads exactly once, and writes a registry entry when nothing is pending. It even clarifies the return contract: only pending_unit_ids, never raw output. This is exceptional transparency for a stateful mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place. The main purpose is front-loaded in the first sentence, and the following sentences pack precise behavioral details without filler. The description is dense but not bloated, and every clause adds decision-relevant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema and no annotations, the description is nearly complete for invoking the tool correctly. It explains what the tool does, the order of operations, the side effects (registry write, unload), and the exact return value. The only minor omission is explicit provider parameter guidance, but the core execution model and postconditions are fully specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It gives meaningful context for profile ('profile's registered test units') and model_id ('one model'), but it never mentions the optional 'provider' parameter or explains how it influences the run. The description partially compensates but leaves the third parameter semantically unaddressed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Test one model against the profile's registered test units'. It then distinguishes this from sibling operations by detailing the execution pipeline: deterministic units inline, orchestrator_judged units to pending, unloads exactly once, and registry writes. This clearly separates it from tools like validate_test_unit, run_untested_sweep, or load_model.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when the tool is appropriate: when you want to run a full test regimen for one model against a profile's registered test units. It also hints at lifecycle behavior (loads, unloads once) that helps an agent understand the operational context. However, it does not explicitly name alternatives or state when not to use it, so it stops short of full usage differentiation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_untested_sweepRun Untested SweepB
Workflow #3 end to end: diff downloaded LLM models against the registry and run Workflow #1 (run_test_regimen) on each unregistered model, sequentially.
| Name | Required | Description | Default |
|---|---|---|---|
| profile | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden. It discloses the two main steps and sequential execution, but omits side effects (e.g., whether tested models are registered afterward), error handling, resource usage, or long-running nature. Some behavioral context is present, but significant gaps remain.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One concise sentence that front-loads the workflow label and states the actions clearly. The references to Workflow #1 and #3 are slightly jargon-heavy but do not waste words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite being a complex workflow tool, the description lacks any explanation of the 'profile' parameter, expected return value, or behavioral side effects. With no output schema or annotations, the agent has insufficient information to call it correctly and interpret results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has zero description coverage for the single required parameter 'profile', and the description does not mention or explain it. The agent is left to infer what profile means and how it affects the sweep, making this a critical gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description specifies an end-to-end workflow with concrete actions: diff downloaded models against the registry, then run run_test_regimen on each unregistered model sequentially. It clearly differentiates from sibling tools like diff_untested and run_test_regimen by indicating it combines them into a full sweep.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Workflow #3 end to end' signals that this is the complete sweep, implying use when the entire process is needed rather than individual steps. However, it does not explicitly name alternatives or state when not to use the tool, so the guidance is clear but not fully explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
seed_provider_modelsSeed Provider ModelsA
Bulk-register the canonical Cloudflare agentic models (canonical manifest), role-tag them in the registry, and write the default role pins. Idempotent; unknown model ids are refused before anything is written.
| Name | Required | Description | Default |
|---|---|---|---|
| profile | No | ||
| provider | No | ||
| model_ids | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses idempotency and that unknown model IDs are refused before any write, which is valuable. However, it does not clarify whether existing role pins are overwritten or only set if absent, nor does it mention any destructive effects on the registry, authentication requirements, or rate limits. For a write operation, this is a partial disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that front-loads the core purpose and packs in the key behavioral notes (idempotent, refusal of unknown IDs). It is efficient and avoids redundancy, though it could be split into two sentences for better readability. Overall it earns a high score for brevity without losing important detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 3 parameters, no output schema, and no annotations, the description leaves significant gaps. It does not explain the meaning of each parameter, what 'default role pins' are, what happens on success (return value?), or how the provider parameter affects the operation. An agent would need to inspect the schema or make assumptions to call it correctly, making the description incomplete for safe invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain the parameters, but it does not mention any of them. The parameter names (profile, provider, model_ids) give some hint, but there is ambiguity: the description says 'canonical Cloudflare agentic models', yet the provider enum includes openrouter, omniroute, and generic. The description fails to clarify how the provider parameter interacts with the canonical manifest, leaving the agent to guess at semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('Bulk-register'), a specific resource (canonical Cloudflare agentic models), and the subsequent actions (role-tag them, write default role pins). It distinguishes this from sibling tools like nanites_registerProviderModel (singular registration) and nanites_discoverProviderModels (discovery), making its unique purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it (bulk registration of the canonical manifest) and contrasts with the singular sibling via the word 'Bulk'. It also adds idempotency and refusal behavior, which helps an agent decide if this is the right operation. However, it does not explicitly name alternatives or state when not to use it, so it stops short of full guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
send_ntfySend NotificationA
Fire-and-forget push notification to the profile's ntfy topic (public-server default resolution per profile config). A failed push never fails the underlying operation; it is logged server-side only.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | ||
| message | Yes | ||
| profile | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and delivers meaningfully: it discloses that a failed push never fails the underlying operation, is logged server-side only, and has fire-and-forget async semantics. This is genuinely valuable behavioral context beyond what the schema reveals. Minor omissions like rate limits or auth requirements keep it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences with the core purpose front-loaded and zero filler. Every clause adds information: the async nature, the topic resolution source, and the failure behavior. Nothing could be trimmed without losing value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequate for a fire-and-forget notification: failure behavior and topic resolution are covered, and no output schema is needed for an async fire-and-forget call. However, the unexplained 'tags' parameter and absence of any usage guidance leave real gaps that the 0% schema coverage makes harder to ignore.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, but it only clarifies 'profile' (topic resolution per profile config). 'message' is self-evident by name, while 'tags' — an array of strings with no schema detail — is left entirely unexplained. The description doesn't add meaning for two of the three parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Fire-and-forget push notification to the profile's ntfy topic.' This is precise and clearly distinguishes it from all 51 siblings, none of which are notification-related. An agent can identify this tool's role instantly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'fire-and-forget' phrasing implies when it's appropriate (async notification where you don't await a result), but there's no explicit comparison to alternatives or when-not-to-use guidance. No sibling notification tool exists, so differentiation isn't critical, but prerequisites or conditions for use are never stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_role_pinSet Role PinA
Pin a role to a preferred (provider, model). Auto-route by pin on later sub-agent runs; falls back to dynamic when the pinned target is unusable.
| Name | Required | Description | Default |
|---|---|---|---|
| role | Yes | ||
| profile | No | ||
| model_id | Yes | ||
| provider | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses that this is a persistent configuration change (pins a role), that it affects future sub-agent runs, and that it falls back to dynamic routing when the pinned target is unusable. Some behavioral details remain uncovered, such as whether this overwrites an existing pin or whether provider/model validity is checked, but the description is strong for a configuration tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no wasted words. The core action is front-loaded, and the behavioral note about fallback earns its place. Perfectly sized for what it communicates.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple config-pinning tool with 4 params and no output schema, the description covers the main outcome and follow-on behavior. It could mention what happens if the pin already exists or how to view pins (list_role_pins), but these are minor gaps given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it provides limited detail about the parameters beyond their names. The term 'preferred (provider, model)' maps loosely to provider and model_id, and 'a role' maps to role, but profile is not mentioned at all. The enum for provider is left to the schema. The description adds some interpretive context but doesn't fully compensate for the lack of schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Pin') and identifies the resource ('a role to a preferred (provider, model)'). It clearly states the purpose and even distinguishes itself from dynamic routing behavior. It doesn't explicitly name sibling alternatives like list_role_pins or delete_role_pin, but the purpose is still clear enough.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when this tool is used ('Pin a role') and what happens on later sub-agent runs ('Auto-route by pin'). It also explains the fallback behavior when the pinned target is unusable. This gives clear usage context without needing to reference sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_btw_chatStart /nanites-btw ChatA
Start a /nanites-btw working-memory chat for a profile: replaces the active chat row and wipes its old transcript (the compaction caches persist), enqueues compaction of the given host-session transcript as an async job (returns immediately — it can never stall or eject a running job), pins and holds a context_qa model when the job completes, and answers initial_question inline if it finishes inside the ~5s grace window. Returns the dashboard deep link (open it in the preview) plus the job handle and status.
| Name | Required | Description | Default |
|---|---|---|---|
| profile | No | ||
| messages | No | ||
| initial_question | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full burden and excels: it discloses the destructive effect ("replaces the active chat row and wipes its old transcript"), the async non-blocking guarantee ("returns immediately — it can never stall or eject a running job"), model pinning behavior, and the ~5s grace window for inline answers. This is exemplary behavioral disclosure with no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core verb and resource, and every clause earns its place — there is zero filler. It is structured as one long run-on sentence with multiple parenthetical asides, which could be split into separate sentences for readability, but the density is justified by the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations, no output schema, and 0% parameter documentation, the description is remarkably complete: side effects, async semantics, timing bounds, and the return payload (dashboard deep link, job handle, status) are all covered. Minor gaps remain — error/failure behavior of the compaction job and what happens when initial_question misses the grace window — but nothing blocks a correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it largely does: profile is grounded as "for a profile," messages as "the given host-session transcript" that gets compacted, and initial_question as the prompt "answered inline" within the grace window. It falls short only by not flagging that all three parameters are optional and by not stating default behaviors for omitted values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-plus-resource statement — "Start a /nanites-btw working-memory chat for a profile" — and the follow-on behavioral list (row replacement, transcript wipe, compaction enqueue, context_qa pinning, grace window) makes it unmistakably distinct from siblings like chat or start_sub_agent_job. An agent can tell what this tool is for without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The use context is clearly established: this is the operation for opening a /nanites-btw working-memory chat with transcript compaction and model pinning, which is a distinct scenario from the sibling chat tool. However, no explicit when-not-to-use or alternative routing is stated (e.g., a note that ordinary conversation should use chat), so it stops short of full guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_sub_agent_jobStart Sub-Agent JobA
Queue a sub-agent job on the profile's async job FIFO and return its job_id immediately (non-blocking — for long-horizon work on slow/big models instead of a tool call that blocks for minutes). Same validated inputs as run_sub_agent; when the profile's concurrency tier is at capacity the job waits queued rather than erroring. Poll with get_sub_agent_job_status.
| Name | Required | Description | Default |
|---|---|---|---|
| task | No | ||
| brief | Yes | ||
| roles | No | ||
| effort | No | ||
| profile | Yes | ||
| model_id | No | ||
| provider | No | ||
| output_schema | No | ||
| output_schema_name | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it delivers the key async traits: non-blocking, FIFO ordering, queued-when-at-capacity rather than erroring, and same validation as run_sub_agent. It does not cover failure modes (e.g., invalid profile, job rejection) or persistence, which keeps it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences with no filler: primary action and return, use case and capacity behavior, and polling instruction. The most decision-relevant facts (non-blocking, returns job_id) are front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with no annotations and no output schema, the description covers the action, return value, async semantics, concurrency-tier behavior, and follow-up polling — surprisingly complete on behavior. It falls short only on parameter definitions and error semantics, which are partially mitigated by the run_sub_agent reference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 9 parameters, so the description must compensate, but it only says 'Same validated inputs as run_sub_agent' — a delegation rather than an explanation of what each parameter means. The schema's names, types, and enums are all an agent has to go on for fields like task, roles, effort, output_schema, and provider.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Queue'), a precise resource ('a sub-agent job on the profile's async job FIFO'), and the immediate return ('return its job_id immediately'). It explicitly contrasts with the blocking sibling run_sub_agent, so an agent can distinguish them without inspecting either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use guidance is present: 'for long-horizon work on slow/big models instead of a tool call that blocks for minutes.' It names the alternative (run_sub_agent) for the blocking case and prescribes the follow-up tool ('Poll with get_sub_agent_job_status'), plus explains capacity behavior (waits queued rather than erroring).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
submit_test_judgmentSubmit Test JudgmentA
Record an orchestrator judgment for one pending unit. Nothing becomes final in the registry until user_approved is true; a submission with user_approved false records the judgment without finalizing that score.
| Name | Required | Description | Default |
|---|---|---|---|
| score | Yes | ||
| profile | Yes | ||
| unit_id | Yes | ||
| model_id | Yes | ||
| user_notes | No | ||
| user_approved | Yes | ||
| orchestrator_notes | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses the critical non-finalizing behavior when user_approved is false, which is a key side-effect distinction. It doesn't mention whether the tool can be called multiple times, overwrite existing judgments, or require pending status, but the core behavioral trait is well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the primary action, and the critical caveat about user_approved is placed immediately after. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter tool with no output schema and no annotations, the description covers the essential behavior and the one parameter with complex semantics (user_approved). It doesn't describe return values or error conditions, but the core decision an agent needs—what this tool does and what finalizes—is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the semantic meaning of user_approved (finalization gate) and implies that score, unit_id, model_id, profile, and orchestrator_notes are the judgment components. It doesn't explain user_notes, but the schema's type info plus the description's context covers most parameters adequately.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Record an orchestrator judgment'), the target ('one pending unit'), and the key distinction from finalization. It distinguishes itself from sibling tools like get_pending_judgments and validate_test_unit by focusing on recording a judgment rather than reading or validating.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool: when there is a pending unit and an orchestrator judgment to record. It explains the user_approved flag's role in finalization, which guides usage. However, it doesn't explicitly name alternatives or state when not to use it, though the sibling list makes the context fairly clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
switch_profileSwitch ProfileA
Make a named profile the active one. All subsequent model tools use its endpoint until switched again.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose a key side effect: the active profile persists and changes the endpoint used by all subsequent model tools. It does not mention failure behavior or profile validation, but the core behavioral trait is clearly communicated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action, and the second sentence adds a necessary consequence. There is no filler and no repetition of schema details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter stateful setter, the description covers the action, scope, and persistence. It does not describe return values or error handling, but there is no output schema and those details are less critical for this kind of tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is only one parameter, `name`, and schema description coverage is 0%, so the description must compensate. It explains that the name identifies a profile, but it does not explicitly say the profile must already exist or how it relates to create_profile/list_profiles. The meaning is adequate but not fully elaborated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Make a named profile the active one.' It also states the lasting consequence ('All subsequent model tools use its endpoint until switched again'), which distinguishes it from read-only siblings like get_active_profile and from create_profile/update_profile.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear operational context: switching is a persistent, global change that affects subsequent model tool calls until another switch. It does not explicitly name alternatives or when-not-to-use conditions, so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
system_health_checkSystem Health CheckA
Check LM Studio health for a profile: endpoint reachability (with a one-shot lms server start autostart recovery and recheck), free disk space for downloads, and stuck-loaded-model detection. Returns an overall status (healthy/degraded/down) plus per-check sub-fields.
| Name | Required | Description | Default |
|---|---|---|---|
| profile | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses a notable side effect: a one-shot lms server start autostart recovery and recheck, which is a behavioral trait. However, it does not disclose whether other checks have side effects, error handling, or timeouts, so transparency is partial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core purpose, and lists the checks succinctly without extraneous detail. Every word serves to inform the agent about scope and return.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description must explain return values. It mentions an overall status (healthy/degraded/down) and per-check sub-fields, but does not detail the sub-field structure or possible values. It also omits prerequisites or potential error conditions, leaving gaps for an agent to infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has zero description coverage for the single 'profile' parameter. The description clarifies that the check is scoped to a profile, but it does not explain what a profile is, its format, or any constraints. This adds minimal meaning beyond the parameter name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (check) and resource (LM Studio health for a profile), and lists distinct health checks (endpoint reachability, disk space, stuck-loaded-model detection) with a clear return structure. This distinguishes it from related sibling tools like nanites_ping or get_loaded_model.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool is for a health overview of a profile, but it does not explicitly state when to use it versus alternatives or when not to use it. No exclusions or sibling references are provided, so usage guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
unload_modelUnload ModelC
Unload a loaded model instance by its instance_id from the active profile's LM Studio endpoint.
| Name | Required | Description | Default |
|---|---|---|---|
| instance_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It implies a state-changing operation (unloading) but does not disclose potential side effects like freeing memory, whether the operation is reversible, or whether it affects other instances. It also doesn't mention errors when instance_id is not loaded. This is a significant gap for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with zero fluff, front-loading the verb and resource. It is appropriately concise for a simple operation with one parameter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with no output schema and no annotations, the description is minimal. It lacks critical context such as how to find instance_id (likely from list_models or get_loaded_model), what to expect in response, and error handling. Given the tool's low complexity, a slightly richer description would be sufficient, but the current one leaves gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has only one parameter (instance_id) with no description, and schema coverage is 0%. The description mentions instance_id but does not explain what an instance_id is or how to obtain it. At 0% coverage, the description should compensate, and it only barely does by mentioning 'loaded model instance' – leaving the agent uncertain about the ID format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (unload a model) and the target resource (a loaded model instance by instance_id) from the active profile's LM Studio endpoint. It distinguishes itself from load_model and get_loaded_model by specifying unload semantics, but it could be more explicit about the distinction from 'get_loaded_model' which is about querying, not modifying.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like load_model or get_loaded_model. It doesn't state prerequisites (e.g., model must be loaded, instance_id must be valid) or consequences. An agent would have to infer usage from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_profileUpdate ProfileA
Partially update a named profile (machine specs, endpoint, pricing, ntfy, inference effort/ceiling, theme, tool grant). Omitted fields keep their current values. Used by /nanites-effort to set the active profile's effort.
| Name | Required | Description | Default |
|---|---|---|---|
| ntfy | No | ||
| tools | No | ||
| pricing | No | ||
| profile | Yes | ||
| endpoint | No | ||
| use_case | No | ||
| inference | No | ||
| dynamic_model | No | ||
| machine_specs | No | ||
| test_plan_ref | No | ||
| vision_capable | No | ||
| concurrency_override | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses the partial-update semantics and the fact that it mutates a profile. However, it doesn't disclose whether changes are reversible, whether the active profile is affected immediately, or whether any validation/restart is needed. The partial-update disclosure is valuable but the mutation consequences are under-specified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, both information-dense. The first sentence front-loads the verb, resource, and field scope; the second adds the critical partial-update semantics and a concrete usage example. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter mutation tool with no annotations and no output schema, the description is reasonably complete but leaves gaps: it doesn't mention return values, error conditions, or whether the update applies to the active profile. The partial-update semantics and field enumeration are strong, but an agent still lacks information about what happens after the update.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It lists the field categories (machine specs, endpoint, pricing, ntfy, inference effort/ceiling, theme, tool grant) which maps to the 12 parameters, and clarifies that 'profile' is the named target. It doesn't explain nested structures like tools.integrations or concurrency_override, but the field list gives an agent a solid map of what can be updated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('partially update'), a clear resource ('named profile'), and enumerates the updatable fields (machine specs, endpoint, pricing, ntfy, inference effort/ceiling, theme, tool grant). It also distinguishes itself from sibling tools like create_profile and switch_profile by emphasizing partial update semantics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states the partial-update behavior ('Omitted fields keep their current values'), which is critical usage guidance. It also names one concrete use case ('Used by /nanites-effort to set the active profile's effort'). However, it doesn't explicitly say when NOT to use it or name alternatives like create_profile/switch_profile.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_test_unitValidate Test UnitA
Validate a test-unit object against the schema/validator rules without persisting it. Returns { ok, issues }.
| Name | Required | Description | Default |
|---|---|---|---|
| unit | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the non-persisting side effect and the return shape ({ ok, issues }), which is essential behavioral information. It does not mention error handling or edge cases, but for a validation tool this is sufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler, front-loads the core action, and includes the key differentiator and return type. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the opaque input schema (no defined properties) and no output schema, the description should explain more about the expected structure of 'unit' and the meaning of 'issues'. It is too sparse for an agent to reliably construct a valid input or interpret results beyond 'ok'.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has zero description coverage for the sole 'unit' parameter, and the description merely says 'test-unit object' without adding any field details, format, or examples. The agent cannot determine what properties the unit object must have, so the description fails to compensate for the opaque schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Validate'), a specific resource ('test-unit object'), and the criterion ('against the schema/validator rules') while explicitly noting the non-persisting behavior, which differentiates it from register_test_unit. It also gives the return shape, making the tool's purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'without persisting it' provides a clear usage context: use this tool to check validity before committing. However, it does not explicitly name the alternative or state when not to use it, so it falls just short of full guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
write_registry_entryWrite Registry EntryA
Insert or update one registry entry for a profile/model_id. Omitted entry fields keep their existing values.
| Name | Required | Description | Default |
|---|---|---|---|
| entry | Yes | ||
| profile | Yes | ||
| model_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral disclosure burden. It discloses the crucial upsert behavior and, more importantly, that omitted entry fields retain existing values, which prevents an agent from assuming full replacement. It does not cover return values or error/side-effect behavior, but the core mutation semantics are clearly disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with no redundant wording. The primary operation is front-loaded, and the important partial-update nuance is stated immediately after, so every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a write tool with no annotations and no output schema, the description provides the core purpose and partial-update behavior but omits expected return value, failure modes, and how to confirm the write via read_registry. It is adequate for basic invocation but not fully complete for an agent operating independently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does add meaningful semantics by identifying profile and model_id as the entry key and by explaining that entry is a partial set of updatable fields. However, it does not explain the meaning of individual entry fields, leaving some semantics to schema names and types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb pair ('insert or update') and names the exact resource ('one registry entry for a profile/model_id'), making the operation unmistakable. This clearly distinguishes it from read_registry and other registry-related sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for writing registry entries, but it never explicitly says when to choose this over read_registry or other alternatives. There is no exclusion or alternative-routing guidance, so the usage context is clear but incomplete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
53 tool updates
v1.0.0- First observed
chat - First observed
check_adaptation - First observed
create_profile - First observed
delete_role_pin - First observed
diff_untested - First observed
download_and_test - First observed
download_and_wait - First observed
download_model - First observed
filter_by_guardrail - First observed
get_active_profile - First observed
get_cost_saved_report - First observed
get_download_status - First observed
get_first_run_status - First observed
get_loaded_model - First observed
get_pending_judgments - First observed
get_sub_agent_job_status - First observed
list_models - First observed
list_profiles - First observed
list_role_pins - First observed
list_test_units - First observed
load_model - First observed
nanites_addProviderKey - First observed
nanites_deregisterProviderModel - First observed
nanites_discoverProviderModels - First observed
nanites_getProviderConfig - First observed
nanites_listProviderKeys - First observed
nanites_listProviderModels - First observed
nanites_ping - First observed
nanites_registerProviderModel - First observed
nanites_removeProviderKey - First observed
nanites_setProviderEnabled - First observed
nanites_setProviderPreferenceOrder - First observed
nanites_showProviderErrors - First observed
nanites_toggleProviderKey - First observed
read_registry - First observed
register_adapted_units - First observed
register_test_unit - First observed
run_sub_agent - First observed
run_test_regimen - First observed
run_untested_sweep - First observed
seed_provider_models - First observed
send_ntfy - First observed
set_role_pin - First observed
share_test_results - First observed
start_btw_chat - First observed
start_sub_agent_job - First observed
submit_test_judgment - First observed
switch_profile - First observed
system_health_check - First observed
unload_model - First observed
update_profile - First observed
validate_test_unit - First observed
write_registry_entry
TDQS
Scored across 53 tools
Most tools are clearly distinct (e.g., list_models vs get_loaded_model, run_sub_agent vs start_sub_agent_job), but a few pairs could be confused: nanites_deregisterProviderModel vs nanites_removeProviderKey (deregister vs remove), and get_download_status vs download_and_wait overlap in polling behavior. Overall, the core model/test workflows are well separated.
Naming is mixed: some tools use snake_case with domain prefix (nanites_*), others use plain verbs (chat, load_model, create_profile), and a few use different conventions (get_cost_saved_report, seed_provider_models). The pattern is not uniform, making it harder to predict tool names, though each individual name is readable.
With 53 tools, the surface is both wide and deep, covering profiles, providers, models, testing, jobs, workflows, and notifications. While each domain has its own cluster, the sheer number exceeds what an agent can easily navigate, and some tools (e.g., send_ntfy, get_cost_saved_report) feel peripheral. A more focused set of ~25-35 would be more manageable.
The set covers the full lifecycle for models (list/load/unload/download/test), profiles (create/update/switch/list), providers (add/remove/list/toggle keys, set preference), and test management (validate/register/run/judge). Minor gaps exist: no explicit tool to delete a profile or a test unit, and no tool to list provider errors beyond 'show' (which is read-only). Overall, core workflows are well-covered.
Maintenance
Related MCP Connectors
AI routing, memory, guardrails, and governance. Routes across Claude, GPT, Gemini.
Enterprise AI Control Plane: governance, guardrails, spend tracking, compliance & smart routing.
One MCP endpoint for Claude, GPT & Gemini: 100+ tools + no-code connectors + agent workers.
Your AI Agent's Infrastructure Layer. Connect Claude, Copilot, Codex, or ChatGPT to 200+ managed open source services. Start databases, pipelines, and applications through natural language.
Related MCP Servers
- AlicenseBqualityDmaintenanceEnables coding agents like Claude Code and Codex to offload boilerplate generation, summarization, and other bounded text tasks to local or cheap cloud LLMs, keeping the frontier agent in charge of judgment and code edits.93MIT
- AlicenseBqualityCmaintenanceDelegates heavy, repetitive, and verifiable tasks like PDF extraction, code analysis, and log processing to a local LLM to reduce token consumption for frontier AI models, while keeping decision-making with the main AI.8MIT
- AlicenseNot gradedqualityBmaintenanceEnables Claude Code to hand off bulk, mechanical, read-heavy tasks to a local model, including agentic loops that can read, write, and run commands sandboxed at zero cloud token cost.MIT
- FlicenseAqualityBmaintenanceEnables Claude Desktop to delegate specialized tasks like security scanning and code review to external AI agents from multiple providers, with dynamic agent registration, flexible pipelines, and safety controls.9-