smol-mcp
This server exposes MCP tools for creating and managing smolVM machines (microVMs) from OCI images, on either a local smolvm fleet or the smol cloud.
List and inspect machines:
list-machinesandget-machinereturn machine id, name, state, CPUs, memory, network, image, pid, and creation time.Create and start machines:
create-machinecreates a machine from an OCI image, lets you set CPU/memory/name/env/egress/network/allow-lists, and waits until it is ready.Run commands:
run-commandruns a shell command or argv in a machine, with workdir, stdin, env, and timeout; returns stdout, stderr, exit code, and timeout/truncation flags.Run one-shot machines:
run-oncecreates a machine, runs a single command, then deletes it even on timeout.Transfer files:
read-fileandwrite-filelet you read and write files inside machines, with utf8 or base64 encoding.Stop and delete machines:
stop-machinestops a running machine;delete-machinedeletes one by name.Tail logs:
machine-logsreturns console log lines from a machine.Pull images:
pull-imagepulls an OCI image into a running machine's local cache.Most tools accept a
targetargument to chooselocalorcloud, defaulting tolocalin this schema.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@smol-mcpcreate a temporary cloud machine and rununame -aon it"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
smol-mcp
An MCP server for smol machines: it gives an AI agent a virtual machine to work in. The agent creates a machine from an OCI image, runs commands in it, reads and writes files, branches it, tails its log and deletes it, all through the tools its client already knows how to call.
Who it is for. Anyone running an agent that should not run commands on the
laptop it is running on. The machine is a real microVM, not a container on
your host, and every command a tool executes goes through the machine API into
a guest; this server spawns exactly one host process, smolvm serve, and only
when a local tool is called.
Built on the official TypeScript SDK (@modelcontextprotocol/sdk, pinned
exactly), with two transports, stdio and Streamable HTTP.
The two targets, and the three modes
Machines come from one of two fleets:
localissmolvm serveon the host this server runs on, reached over a Unix socket. It needssmolvmand a hypervisor. Nothing is billed.cloudis the smol cloud REST API at$SMOL_CLOUD_URL, withAuthorization: Bearer $SMOL_CLOUD_TOKEN. Machines run in the service, and they cost money for as long as they exist.
Which of them a process serves is decided once, at startup, from
SMOL_MCP_TARGETS or, when that is unset, from whether a cloud token is
configured:
Mode | When | What the tools look like |
| no cloud token, or | No |
|
| No |
| a cloud token is configured, or |
|
There is no default in both mode on purpose: a machine on one fleet is
invisible on the other and the two bill differently, so a defaulted argument
sends work to the wrong fleet silently. A client that supports elicitation is
asked once, on its first tool call, which fleet the session is for; an answer
of local or cloud narrows the session and removes the argument from every
tool, which the server reports as a tools/listChanged. A client without
elicitation, or one that declines, keeps the required argument.
The server also carries an instructions string naming the fleets this
process reaches and what each cannot do, and a smol://targets resource with
the same facts as JSON, including whether a served fleet is usable on this
host at all.
The two APIs agree about almost nothing. A machine is a name locally and a
mach-... id on cloud; the image is a string locally and a tagged source
object on cloud; env is a list of {name, value} locally and a map on cloud;
exec takes workdir/timeoutSecs locally and cwd/timeoutSeconds on
cloud; list returns {machines: [...]} locally and a bare array on cloud.
All of that is normalised in the two clients, so a tool call has the same
shape whichever target answers it.
Related MCP server: qemu-mcp-server
Tools
Tool | Targets | Notes |
| local, cloud | |
| local, cloud | Cloud resolves a name to an id with one extra list call. |
| local, cloud | Starts the machine and waits until commands run in it. |
| local, cloud | Returns |
| local, cloud | Create, start, exec, delete. Deletes the machine even on timeout. |
| local, cloud | Takes |
| local, cloud | Waits for readiness first, so the file is not written under a mount that later hides it. |
| local, cloud | Starts a stopped machine and waits until commands run in it. |
| local, cloud | Copies a running branchable machine into a new child, memory and disks included. Local needs Linux or macOS, not Windows. |
| local, cloud | Start it again with |
| local, cloud | |
| local, cloud | Local is the guest console. Cloud is the machine's event log from the control plane: what happened to the machine, not what ran in it. |
| local | The cloud control plane pulls the image itself at create. |
Branching copies a running machine, memory and all, so a child starts from
exactly where its parent was. The source has to be made a branch source when
it is created (branchable: true on create-machine), and neither target can
turn that on for a machine that already exists. The two ask for it in
different places, which the server handles: locally it is a query parameter on
the start, on cloud a field on the create. branch-machine on the local
target needs Linux or macOS; smolvm serve on Windows does not support it.
On cloud, a command, a read or a write in a stopped machine starts it and
leaves it running. startedMachine in the result says when that happened;
stop-machine stops it again.
Following a log is a resource subscription: read
smol://machine/{target}/{name}/logs, subscribe to it, and the server tells
you when there is more. machine-logs with the cursor from the last call
returns what has arrived since.
A machine whose name starts with mcp- is ephemeral: the session that created
it records it and deletes it when that session ends, on either target. A name
you choose yourself persists. The local record is a file in the runtime
directory, so the next start can clean up after a crashed one; the cloud
record is in memory, and a process that dies with entries in it leaves them to
the control plane's own idle stop and TTL.
Install and run
The package is smol-mcp on npm, with two bins: smol-mcp (stdio) and
smol-mcp-http (Streamable HTTP). A client can run it without installing
anything:
npx -y smol-mcpor install it once, globally or into a project:
npm install -g smol-mcpFrom source, for working on it:
git clone https://github.com/smol-machines/smol-mcp
cd smol-mcp
npm install
npm run buildsmolvm is only needed for the local target, and the serve is started on the
first local call rather than at connect, so a client that only ever names
cloud never launches a hypervisor.
Connecting a client
The stdio transport is the default: the client runs smol-mcp and speaks
MCP on its stdin and stdout. The blocks below use npx -y smol-mcp, which
fetches the package on first use; with a global install the command is
smol-mcp and there are no arguments, and from a source checkout it is node
with the absolute path to dist/cli.js.
Claude Code (.mcp.json), Claude Desktop (claude_desktop_config.json) and
Cursor (.cursor/mcp.json) all take the same shape:
{
"mcpServers": {
"smol": {
"command": "npx",
"args": ["-y", "smol-mcp"],
"env": { "SMOL_CLOUD_TOKEN": "..." }
}
}
}OpenCode (opencode.json) uses its own key names:
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"smol": {
"type": "local",
"command": ["npx", "-y", "smol-mcp"],
"enabled": true,
"environment": { "SMOL_CLOUD_TOKEN": "..." }
}
}
}The top-level key is mcp rather than mcpServers, a local server is
type: "local", the command and its arguments are one array, and the
environment is environment rather than env.
Drop SMOL_CLOUD_TOKEN and the server serves the local fleet only. Keep it
and it serves both, and asks once which one the session is for; see the target
modes below.
What has been driven end to end is the MCP SDK's own Client, over both
transports, by the two scripts in scripts/, and OpenCode 1.18.30, which
ran the worked example below over stdio and a machine on another host over
HTTP, from the opencode.json printed above with no changes. The other three
configurations are read from each client's documentation and not driven.
Two computers: the agent here, the machines there
The supported shape for an agent on computer A driving machines on computer B
is to run this server on B, the machine that has smolvm, and connect to
it over the HTTP transport. The local machine API has no authentication of its
own and stays on a Unix socket on B; what crosses the network is this server's
own authenticated endpoint.
The same thirteen tools are served over the official SDK's Streamable HTTP transport by a second bin. stdio is the default and is untouched by it: no port, no token, no listener.
On B:
SMOL_MCP_AUTH_TOKEN=$(openssl rand -hex 16) \
npx -y --package smol-mcp smol-mcp-http --host 0.0.0.0 --port 8080 --path /mcpOn A, one entry:
{
"mcpServers": {
"smol": {
"type": "http",
"url": "http://b.local:8080/mcp",
"headers": { "Authorization": "Bearer <the token you generated>" }
}
}
}It refuses to start without
SMOL_MCP_AUTH_TOKEN, and every request needs it. These tools create and run virtual machines, so a listener without a token is a remote shell.The token check runs before the path check, so an unauthenticated request learns nothing, not even where the endpoint is.
DNS rebinding protection is on, with the Host and Origin allow-lists below.
The token is read from
authorization: Bearer ...orx-smol-mcp-token. The second header is there for a deployment that puts something else inauthorizationbefore the request reaches this server.The bind is
127.0.0.1by default. Publishing the port is a decision to make in the open, not the consequence of a default, and off loopback you should also setSMOL_MCP_HTTP_ALLOWED_HOSTSto the name A dials.An idle session is closed and its ephemeral machines deleted; a session with a tool call still running is never closed, however long the call takes.
This is plain HTTP. Over anything but a network you trust, put it behind TLS.
Replies are plain JSON rather than an SSE frame per response (
enableJsonResponse), because the reply travels through a proxy whose buffering is not ours and no tool here streams.Each session gets its own server instance, and ending the session (an HTTP DELETE) deletes that session's ephemeral machines, the way stdin EOF does on stdio. A session cleans up only what it created, and the
smolvm servethe sessions share is stopped by the last one to let go of it, not by whichever one started it.
Hosting it in a smol machine
Two shapes, both run end to end; see Verified.
Agent inside the machine. npm install smol-mcp in the guest (or npm pack
here and upload the tarball through POST /v1/machines/{id}/exec with the
base64 on stdin, which is how it was first proven), and speak stdio to
smol-mcp there. The account key goes
in the machine's create-time env and is in the guest process environment.
The local target is unavailable inside a guest and says so on the first call.
Agent outside the machine. Publish port 8080 at create, run
smol-mcp-http on 0.0.0.0, and connect with the SDK's
StreamableHTTPClientTransport to https://<name>-<hash>.apps.smolmachines.com/mcp,
the ingress URL the machine record's url field carries once it is ready.
get-machine and list-machines report that field, as url; it is null on
the local target and until something listens on the published port.
This shape needs open egress anyway, to reach the smol cloud API.
A worked example
An agent borrowing a machine to check out a repository, run its tests, read
the result and take a generated file away. Five calls, on the local target,
run exactly as shown against smolvm 1.14.5 and pasted back verbatim.
1. A machine. network: "open" because the image is pulled from a
registry and that pull happens inside the guest; see the egress section below.
$ create-machine {"target":"local","name":"mcp-walkthrough","image":"python:3.12-alpine","cpus":2,"memoryMb":1024,"network":"open"} [3.1s]
{
"machine": { "id": "mcp-walkthrough", "name": "mcp-walkthrough", "state": "running",
"cpus": 2, "memoryMb": 1024, "network": "open", "pid": 36349 },
"ephemeral": true,
"ready": true
}ready: true means a command has already run in it, so the next call does not
have to wait for the guest.
2. Check out the work.
$ run-command {"target":"local","name":"mcp-walkthrough","command":"apk add --no-cache git >/dev/null && pip install --quiet pytz && git clone --depth 1 --quiet https://github.com/dbader/schedule /workspace/schedule && echo cloned","timeoutSecs":120} [3.3s]
{
"stdout": "cloned\n",
"stderr": "WARNING: Running pip as the 'root' user can result in broken permissions ...",
"exitCode": 0,
"truncated": false,
"timedOut": false,
"overflow": [],
"startedMachine": false
}exitCode comes from the guest: a failing command is a result, not a tool
error, so the agent reads the code rather than catching an exception.
3. Run the tests, keeping a report in the machine.
$ run-command {"target":"local","name":"mcp-walkthrough","command":"cd /workspace/schedule && python -m unittest test_schedule 2>&1 | tee /workspace/report.txt | tail -3","timeoutSecs":120} [0.2s]
{
"stdout": "Ran 81 tests in 0.020s\n\nOK\n",
"stderr": "",
"exitCode": 0,
"truncated": false,
"timedOut": false,
"overflow": [],
"startedMachine": false
}4. Take the generated file.
$ read-file {"target":"local","name":"mcp-walkthrough","path":"/workspace/report.txt"} [0.0s]
{
"path": "/workspace/report.txt",
"content": ".....................................................................\n---------------------------------------------------------------\nRan 81 tests in 0.020s\n\nOK\n",
"encoding": "utf8",
"size": 180,
"offset": 0,
"bytes": 180,
"eof": true,
"startedMachine": false
}size is the whole file and bytes is what this call returned, so eof: true
says there is nothing after it. A bigger file comes back in pages: pass
offset and length and read until eof.
5. Give it back.
$ delete-machine {"target":"local","name":"mcp-walkthrough"} [0.1s]
{ "deleted": "mcp-walkthrough" }The delete is not strictly needed. The name carries the mcp- prefix, so the
machine is ephemeral and the session would have deleted it at exit anyway.
The whole sequence took 6.7 seconds.
Defaults, and why each one is what it is
Settable by env var, or by a JSON config file at
$XDG_CONFIG_HOME/smol-mcp/config.json (or $SMOL_MCP_CONFIG). Precedence is
env over file over default. An unknown key in the config file is an error
rather than a silent no-op, because the local API's own habit of accepting and
dropping unknown fields is exactly the failure this avoids.
Setting | Env | Default | Why |
minimum smolvm |
| 1.14.0 | Checked against |
memory |
| 2048 MiB | Enough for an interpreter and a build; small enough to boot several at once. |
cpus |
| 2 | One core leaves nothing for the guest agent while a command runs. |
exec timeout |
| 120 s | Long enough for a package install, short enough that a hung command does not hold a tool call open. |
exec timeout ceiling |
| 900 s | The longest |
output truncation |
| 64 KiB per stream | A tool result is read by a model with a context budget. Past this the result keeps the head and the tail, says how many bytes fell between them, and writes the whole stream into the machine so it can be read back. Truncation is reported, never silent. |
readiness timeout |
| 120 s | Covers a cold image pull inside the guest. |
egress default |
|
|
|
run-once egress |
|
| The same, for |
ephemeral TTL |
| 3600 s | Sent as |
ephemeral idle stop |
| 900 s | Sent as |
machine prefix |
|
| The marker that makes a machine ephemeral. |
log tail |
| 100 lines | What |
log poll |
| 2 s | How often a subscribed log resource is checked for new lines. Neither log route pushes, so following is a poll. |
HTTP bind |
|
| HTTP transport only. Publishing the listener is a decision, not a default. |
HTTP port |
| 8080 | HTTP transport only. What a smol machine publishes by default. |
HTTP path |
|
| HTTP transport only. Everything else on the listener is a 404. |
HTTP token |
| none | HTTP transport only, and required: it refuses to start without one. |
HTTP idle session |
| 1800 s | A session with no request in flight and none for this long is closed, and its ephemeral machines with it. A running tool call keeps its session alive however long it takes. |
HTTP Host allow-list |
| the loopback names of the bound port | Comma separated. Off loopback the name is not knowable here, so name it or the Host check does nothing and the startup log says so. |
HTTP Origin allow-list |
| none | Comma separated. Empty means no browser origin is expected; a request carrying one is refused only when this names some other value. |
Security defaults, stated as reasons
The local API has no authentication. Anything that can reach it can create a VM on this host. So the server listens on a Unix socket in its own runtime directory (mode 0700) by default, and refuses outright to start a serve on a non-loopback address.
Nothing runs on the host outside a VM. Every command a tool executes goes through the machine API into a guest. The server spawns exactly one host process,
smolvm serve, and only when a local tool is called.The token is read from the environment or the config file, and never written anywhere. It is not logged, not echoed in an error,
.envis in.gitignore, and thesmolvm servechild is spawned with a minimal environment that carries neither token.run-onceon cloud denies egress by default, and never publishes a port. See the network section.The HTTP transport authenticates every request and refuses to start without a token. Not a warning, not a default token: the process exits.
Egress is off unless the call asks for it
The intent is that a machine an agent asked for cannot reach the internet
because nobody said otherwise. create-machine and run-once both default to
no egress, on both targets. SMOL_MCP_NETWORK_DEFAULT and
SMOL_MCP_RUN_ONCE_NETWORK move that default for an operator who wants it
open; the tool arguments network, allowHosts and allowCidrs move it per
call.
The local trap, and it is why the default has to be opt-out per create. An image is pulled from a registry inside the guest, so a machine with no egress cannot start from one. The API refuses it at create time:
image 'alpine' must be pulled from a registry, but this machine has no
network, so the pull can never succeed. Add --net (or publish a port with -p,
or set an egress policy with --allow-cidr/--allow-host). To keep the machine
network-isolated, supply the image locally instead: `docker save alpine |
smolvm machine create --image - ...`This server passes that message through and adds one line naming the argument that grants the exception for that create. An egress allow-list is accepted and is honoured on the pull, which sounds like the better answer and is a trap: the blob CDN host is not knowable in advance. Allow-listing docker.io's documented hosts produced
dial tcp: lookup production.cloudfront.docker.com ...: no such hostbecause the CDN name is neither docker.io nor any of the hosts the docs
name. So a registry image on the local target realistically needs
network: "open" for the create that pulls it.
Which means the blocked default is a default, not a guarantee. The
server's instructions tell the agent exactly that: pass network: "open" for
a registry image on local. An agent following them will open egress on its own
whenever it wants an image it does not have. If you need isolation you cannot
talk an agent out of, supply the image locally (docker save ... | smolvm machine create --image -) so nothing has to be pulled, or give an allow-list
that names what the workload may reach. The default stops a machine reaching
the internet by accident; it does not stop an agent asking.
Cloud. The control plane pulls the image, so the guest never needs the
registry and a blocked machine starts normally. One API fact shapes how the
deny is spelled: {"mode": "allowCidrs", "cidrs": []} is refused with
HTTP 400 allowCidrs network mode requires at least one CIDR or host. So the
deny goes out as an allow-list of 192.0.2.0/24, RFC 5737 TEST-NET-1,
reserved for documentation and routed nowhere. A cloud machine that publishes
a port cannot also block egress, and create-machine refuses that
combination rather than sending it.
What this server does not expose
Both APIs are larger than this tool vocabulary. Everything below is a
deliberate omission with a reason, and src/parity.ts carries the same list
as data: the parity check reports each skipped route with its reason and
reports anything in neither list as unknown, so a capability added in a
release shows up as a decision to make rather than an omission nobody noticed.
The tools are hand written and not generated from either spec, because three published entries do not describe the running service: the cloud snapshot route is documented as a 200 and answers 501, the cloud export entry carries no request body at all, and the local spec has no checkpoint route while the product has the feature. The published cloud OpenAPI also lists neither the fork route nor the checkpoint routes, and both exist and answer.
Local, against smolvm serve openapi on v1.14.5
Not exposed | Why |
export | Needs a |
checkpoint | Absent from the local spec entirely. It also restores only on the same OS and architecture, and macOS restore with host mounts has an open defect. |
branch release | Releases a held fork-pool slot; this server has no pool vocabulary. |
sync | Synchronises staged mounts without stopping the machine; mounts here are a create-time argument. |
resize | Expand only, and no tool asks for a machine to grow after it exists. |
| Exec results are returned whole, not streamed. |
fork pools, rollout executors | Fleet and batch surfaces rather than agent ones. |
Volumes on the local target are the mounts argument on create-machine,
which attaches a host directory. There is no separate volume object.
Cloud
Not exposed | Why |
export | The spec entry has no request body and nothing anywhere tests it, so there is no shape to code against. |
snapshot | Documented as a 200 and answers 501, telling the caller to export instead. Never call it. |
volumes | Not built on the service. |
checkpoints | All four routes exist. On 2026-09-09, capture answered 502 three times with a libkrun permission error in the node's checkpoint staging directory, and delete answered 502 on storage cleanup. No tool ships until the service side works. |
fork batch, lineage | Batch branching and branch ancestry, which this vocabulary does not express. |
sessions | A session keeps a working directory and environment across execs; |
code | A code-oriented surface with no counterpart on the local target. |
connect | The authenticated bridge answers |
Exec results are returned whole on both targets. Output past the budget keeps its head and its tail and the whole stream is written back into the machine, which the result names; nothing streams a command as it runs.
Constraints this server is built around
Each one has a test.
Constraint | Covered by | |
a | A local file upload must wait until the workload container is running, or it lands in the agent's namespace and is hidden once the container mounts over the path. |
|
b |
|
|
c | Local parity is stated by the paths this server calls, not by |
|
d | The local create field is |
|
e | The cloud connect bridge answers | Not a unit test: |
f | A server hosted behind an ingress cannot assume |
|
g |
|
|
h | A cloud machine created |
|
Two more, from the same source:
run-onceuses the plain path only: create, start, exec, delete. Never--oci-cache, neverinit(broken by smol-machines/smolvm#1192 and #1193).The settled cloud bill comes from
DELETE /v1/machines/{id}?includeUsage=true. A mid-life/usageread is a documented lower bound.
Tests
npm run lint # eslint, plus a check that no em or en dash is tracked
npm run typecheck # src and test
npm run test:unit # no smolvm, no network, no keyIntegration is opt-in, so a clean clone on a host with no hypervisor and no
key can still run npm test. Those three commands are also the whole CI gate,
for the same reason: booting a machine needs a hypervisor and the cloud suite
needs an account to bill, and neither belongs on a pull request.
# local: needs smolvm on PATH or in SMOLVM
SMOL_MCP_IT=1 npm run test:integration
# cloud: additionally needs SMOL_CLOUD_TOKEN and SMOL_CLOUD_URL
SMOL_MCP_IT=1 npm run test:cloudThe cloud suite reads GET /v1/account before it creates anything and after
every test, and stops the moment period spend passes its own ceiling. It
deletes what it makes in the test that makes it, and its afterAll asserts no
mcp- machine is left on the fleet.
Verified
What was actually run, on what, rather than what should work. Everything below is from one round on 2026-09-09 and 2026-09-10, on the tree this section ships with.
Three hosts. A Mac on Apple Silicon (node 25.9.0), a Windows 11 laptop
(build 10.0.26200.0, node 24.19.0), and Ubuntu 24.04 aarch64 in a Lima VM with
/dev/kvm (node 20.19.5). smolvm 1.14.5 on all three, from the
darwin-arm64, windows-x86_64 and linux-arm64 release archives.
Keyless, on any host. npm run lint, npm run typecheck and
npm run test:unit: 11 files, 173 tests passing, 14 skipped. Identical counts
on macOS and on Linux, and within half a second on the clock. Those three are
also the whole CI gate, because the integration suites need a hypervisor and an
account and neither belongs on a pull request.
Local, on real machines. SMOL_MCP_IT=1 npm run test:integration from an
isolated HOME: 10 passing, the cloud file skipped, on macOS in 23 s and on
Linux under Lima in 167 s, the same tests either way. scripts/smoke.mjs local
listed all thirteen tools and ran a command on both. The worked example above
is a real run, pasted back, and a model driving a real MCP client ran the same
five calls unaided. The branch flow was run end to end on macOS: a machine
created branchable, a file written into it, branch-machine into a child in
0.4 s against 2.7 s for a create, the child reading the parent's file, and a
non-branchable source refused with the API's own message.
Windows. The local target works there: this server starts smolvm serve
on loopback TCP, boots Linux microVMs and runs commands, files and logs in
them, driven both from a client on the same box and from another computer.
branch-machine is refused there with the reason, because a machine started
branchable on that platform has no control socket to branch from.
Two computers over wifi, both directions. The Windows laptop serving and
the Mac driving it, then the Mac serving and the laptop driving it. On the
same runs: a request with no token and a request with a wrong token both
refused with the same 401, including on a path that is not the endpoint, so
the endpoint is not revealed to an unauthenticated caller; a Host outside the
allow-list refused with 403; a session idle past
SMOL_MCP_HTTP_SESSION_IDLE_SECS closed and its machine deleted, while a
session holding a 70 s call through the same limit returned its output and
stayed usable; and netstat on the serving host showed smolvm bound to
loopback only, with the MCP port the single listener on the LAN address.
A real client, driven headless. OpenCode 1.18.30, configured exactly as
the opencode.json above, ran the worked example over stdio and a
create-plus-run-plus-delete against the Windows laptop over HTTP with the
bearer header. Both were tool calls a model chose from the schemas, not
scripted requests.
The package. npm pack gives an 87781 byte tarball carrying dist,
LICENSE, README.md and package.json and nothing from src or test.
Installed into a clean directory on macOS and on Windows, both bins in
node_modules/.bin serve all thirteen tools and boot a real machine.
Cloud, against the live API. SMOL_MCP_IT=1 npm run test:cloud: 4 passing,
26 micros of spend, nothing left on the fleet. Driven through this server over
stdio with SMOL_MCP_TARGETS=cloud: a create with a published port and open
egress, the blocked-plus-port combination refused before anything was sent,
write-file and read-file through the documented files route,
branch-machine into a child that read the parent's file, and machine-logs
from the events route with the cursor coming back empty on the second call.
The hosted shape above was rebuilt and driven: this server installed into a
cloud machine through its own write-file, serving on a published port, with
a client on the Mac creating, running a command in and deleting a second
machine through the ingress URL. Every machine was deleted and
GET /v1/machines was empty afterwards; the whole cloud half cost 449 micros.
Not verified. Claude Code, Claude Desktop and Cursor were not driven: their
configurations above are read from each client's documentation, and the three
claude runs attempted here never reached the server because that CLI had no
usable login on the host. The browser Origin gate was exercised only with
SMOL_MCP_HTTP_ALLOWED_ORIGINS empty, where a request carrying an origin is
accepted, which is what that setting means and not a test of refusing. Cloud
checkpoints failed on the service side and no tool ships for them; see what
this server does not expose.
How this relates to the smolmachines SDK
If you are writing a Node or Python application rather than connecting an
agent, use the embedded smolmachines SDK instead. It runs the local engine
in process, does not need smolvm serve, and reaches smol cloud through the
same Machine API. This server exists for the other case: a client that already
speaks MCP and wants tools rather than a library. Its local target rides on
smolvm serve because that API is the one an out-of-process server can talk
to on every platform the CLI supports.
Traps
Only one
smolvm servecan run per host: it binds127.0.0.1:10081for the guest rollout ingress. If a start fails withAddress already in use, setSMOL_LOCAL_URLto the running serve's listen address; this server will then use it and leave it running.smolvmis a wrapper script thatexecssmolvm-bin, sopkill -f "smolvm serve"does not match a running serve. Look forsmolvm-bin, or for whatever holds port 10081.A failed local start leaves the machine behind in
createdstate. It has to be deleted;run-oncedoes that on every path.On the cloud API, a 400 does not mean the body was not JSON. An empty
cidrsis a 400 with a valid JSON body, so status alone does not separate a parse failure from a validation one.The cloud files route takes the path as a suffix, with no leading slash:
PUT /v1/machines/{id}/files/workspace/app.py, and the same forGET. The published schema lists the route with only{id}, so a deployment that predates the suffix answers 404 andread-fileandwrite-filefall back to exec with base64. The fallback meets the exec response cap, so a read it cut is refused rather than returned short.The connect bridge is GET and HEAD only.
POSTto/v1/machines/{id}/connect/{port}/...is 405 withallow: GET,HEAD, so no MCP client can speak through it. Use the ingress URL in the machine record'surl. A trailing slash on the bridge (connect/8080/) is a 404 whatever the method, which is a different failure from the 405.The ingress needs the account key. Without
authorization: Bearer <key>it answers 401, so a server behind an ingress cannot useauthorizationfor a token of its own. That is whatx-smol-mcp-tokenis for.urlandreadyare bothnull/falseuntil something listens on the published port. Start the server first, then wait on readiness, or the wait can never end.ports[].hostPortis allocated long before either.npm packrefuses to overwrite an existing tarball in--pack-destination, and it writes to~/.npm/_cacacheeven for a pack. In a sandbox that denies either, the failure is an npm error in the middle of a pipeline; delete the old tarball and pass--cache.
Available Tools
11 toolscreate-machineA
Create a machine from an OCI image, start it, and wait until commands run in it.
| Name | Required | Description | Default |
|---|---|---|---|
| cmd | No | Workload command. Default keeps the container alive (sleep loop). Local only; the cloud create request has no such field. | |
| env | No | Environment variables | |
| cpus | No | ||
| name | No | Machine name. Omitted: an ephemeral mcp-<id> name, deleted when this server exits. A name without the mcp- prefix persists. | |
| image | Yes | OCI image reference, e.g. alpine or python:3.12-alpine | |
| start | No | Start and wait for readiness (default true) | |
| target | No | Which fleet to talk to. local is smolvm serve on this host; cloud is smol cloud at SMOL_CLOUD_URL. | local |
| network | No | Egress mode. Default open, except run-once on cloud which is blocked. Local: a blocked machine whose image still has to be pulled from a registry is refused by the API; pass open or an allow-list for that create. Cloud: blocked is sent as an allow-list of an unroutable range. | |
| memoryMb | No | ||
| allowCidrs | No | Egress allow-list of CIDR ranges. Overrides network. | |
| allowHosts | No | Egress allow-list of hostnames. Overrides network. Local: sent as allowedHosts. Cloud: sent inside the same cidrs list the published schema names, alongside allowCidrs. |
Output Schema
| Name | Required | Description |
|---|---|---|
| ready | Yes | |
| machine | Yes | |
| ephemeral | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry disclosure, and it does communicate the create/start/blocking-wait behavior. It does not mention ephemeral-vs-persistent naming, cloud/local target behavior, or that start can be disabled with start=false, so the disclosure is incomplete but not misleading at the default behavior level.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no filler, presenting a clear sequence: create, start, wait. Every segment earns its place and the core operation is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is an 11-parameter tool with local/cloud targets, network modes, and lifecycle implications, yet the description gives only the core flow. The rich input schema and output schema compensate substantially, but the tool-level context is still thinner than the complexity warrants.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is high (82%), so the schema already explains most parameters. The tool description adds no extra parameter semantics beyond 'OCI image' and 'commands', which map to image and cmd but are already documented in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb+resource ('create a machine from an OCI image') and adds behavioral detail (start it, wait until commands run). It separates the tool from obvious lifecycle siblings like stop-machine and delete-machine, though it does not explicitly distinguish it from run-once.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is implied: you want a machine created from an image, started, and ready for commands. It gives no explicit when-not-to-use guidance or alternative routing (e.g., run-once for one-shot tasks), so an agent must infer the boundary between this and run-once.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete-machineA
Delete a machine, running or not.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Machine name | |
| target | No | Which fleet to talk to. local is smolvm serve on this host; cloud is smol cloud at SMOL_CLOUD_URL. | local |
Output Schema
| Name | Required | Description |
|---|---|---|
| deleted | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral burden. It usefully discloses that running machines can be deleted, but it does not state permanence, cascading resource effects, or confirmation requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler, redundant phrasing, or repeated keywords from the name. Every word adds information, especially the qualifier 'running or not.'
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple destructive operation with two well-documented parameters and an output schema, the description covers the essential action and state condition. It could be stronger with a note that deletion is permanent or irreversible, but the current wording is minimally adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents both parameters, including the target enum and default; the description adds no parameter-level meaning. With 100% schema parameter coverage, this is an acceptable baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The verb 'Delete' plus the object 'machine' clearly identifies the operation, and 'running or not' adds a meaningful scope qualifier. It does not explicitly distinguish itself from the sibling 'stop-machine', though the destructive verb makes the contrast implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'running or not' gives useful context that this operation applies regardless of machine state, which helps an agent decide when deletion is allowed. However, it provides no explicit guidance about when to prefer stop-machine versus delete-machine.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get-machineB
Get one machine's state and resources.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Machine name | |
| target | No | Which fleet to talk to. local is smolvm serve on this host; cloud is smol cloud at SMOL_CLOUD_URL. | local |
Output Schema
| Name | Required | Description |
|---|---|---|
| id | Yes | |
| pid | Yes | |
| cpus | Yes | |
| name | Yes | |
| image | Yes | |
| state | Yes | |
| network | Yes | |
| memoryMb | Yes | |
| createdAt | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure, but it does not explicitly state that the operation is read-only, has no side effects, or any auth/rate-limit implications. 'Get' implies retrieval, but does not add beyond the tool name or state that the machine is not modified, which is a notable gap for an unannotated tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, direct sentence that communicates the core function without fluff. It is slightly under-specified for a tool with two inputs, but it is not terse to the point of confusion and earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has a knowledgeable output schema and fully documented parameters, so the description does not need to explain return values. However, the absence of any usage context (e.g., when to call this vs. an alternative) makes it less complete for an agent facing several similar sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already provides meaning for 'name' and 'target'. The description does not add parameter semantics beyond what the schema gives, and therefore appropriately sits at the baseline of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Get'), a specific resource ('one machine'), and the nature of the result ('state and resources'). The singular 'one machine' cleanly distinguishes it from sibling list-machines, and 'state and resources' separates it from machine-logs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives such as list-machines or machine-logs. There is no mention of scenarios for single-machine inspection, no exclusions, and no conditional guidance for choosing a target (local vs cloud).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list-machinesB
List machines on the target.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | Which fleet to talk to. local is smolvm serve on this host; cloud is smol cloud at SMOL_CLOUD_URL. | local |
Output Schema
| Name | Required | Description |
|---|---|---|
| machines | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden. It only states that machines are listed and does not mention read-only behavior, authentication, connectivity, pagination, or any side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. Every word contributes to identifying the operation and scope.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple optional-parameter list operation with an output schema present and a well-documented parameter, the description is nearly sufficient. It lacks behavioral details, but the low complexity and rich schema make it mostly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the target parameter is fully documented with its enum values and default. The description itself adds no extra parameter meaning beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List') and resource ('machines') and identifies a scope ('on the target'). It is clearly distinct from siblings like get-machine, create-machine, and run-command.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance about when to use this tool versus alternatives such as get-machine or machine-logs. There is no explicit when-to-use/when-not-to-use context, though the parameter schema explains the local/cloud distinction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
machine-logsB
Tail the machine's console log. Local only.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Machine name | |
| tail | No | ||
| target | No | Which fleet to talk to. local is smolvm serve on this host; cloud is smol cloud at SMOL_CLOUD_URL. | local |
Output Schema
| Name | Required | Description |
|---|---|---|
| lines | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are not provided, so the description carries the full burden of behavioral disclosure. The description only says 'Tail the machine's console log. Local only.' It does not disclose that tailing is a continuous streaming operation (potentially long-running), whether it requires specific permissions, how it handles log rotation, or what the output format is. For a tool that likely streams data, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise at six words, and it is front-loaded with the primary action ('Tail the machine's console log') followed by the scope restriction ('Local only'). Every word earns its place; there is no fluff or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that there is an output schema and the tool complexity is moderate (tail command with 3 params), the description is incomplete. It does not explain what the output looks like (even though an output schema exists, it may not fully capture streaming behavior), nor does it address how the tail parameter affects the call. For a tail operation, one would expect details on streaming behavior or termination conditions, which are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 67%. The schema already documents the 'name' parameter (machine name) and 'target' parameter (local vs cloud). The description adds the qualifier 'Local only' which reinforces that the target is local, but it does not add new meaning to the 'tail' parameter, which lacks a description in the schema. With 67% coverage, the description partially compensates but does not fully clarify the 'tail' parameter's meaning (number of lines? bytes?).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (tail) and the resource (machine's console log), and specifies 'Local only' to distinguish from cloud targets. This is a specific verb+resource pairing that goes beyond a mere restatement of the name. However, it does not explicitly differentiate from sibling tools like read-file or get-machine, though for an agent the action of tailing is distinct enough from those.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions 'Local only' which gives a hint about when to use this tool (for local machines), but it does not explicitly state when not to use it or mention alternatives among siblings. It lacks clear guidance on when to prefer this over read-file or get-machine. The sibling tools are numerous, so the absence of explicit exclusions makes this only minimally adequate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pull-imageC
Pull an image into a running machine's local cache. Local only.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Machine name | |
| image | Yes | ||
| target | No | Which fleet to talk to. local is smolvm serve on this host; cloud is smol cloud at SMOL_CLOUD_URL. | local |
Output Schema
| Name | Required | Description |
|---|---|---|
| os | Yes | |
| size | Yes | |
| digest | Yes | |
| reference | Yes | |
| layerCount | Yes | |
| architecture | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the action and the 'local only' scope, but it does not disclose side effects (e.g., whether existing cache entries are overwritten), network requirements, authentication needs, or failure modes. For a pull operation this is a notable gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that is concise and front-loads the core action. It avoids unnecessary detail, though it could have included a bit more behavioral context without sacrificing brevity. It is appropriately sized for a simple pull operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (3 parameters, an enum, and an output schema) and the complete absence of annotations, the description is too sparse. It does not explain return values, error conditions, or the distinction between local and cloud behavior beyond the one-word 'Local only'. An agent would have to rely on the schema and runtime feedback to understand the full behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema describes 'name' and 'target' but leaves 'image' undocumented. The description adds no parameter-level detail beyond what the schema already provides, and it fails to compensate for the missing 'image' description. With only 67% schema coverage, the description should have clarified the image parameter but did not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('pull') and resource ('image') with a destination ('into a running machine's local cache'). It also adds a scoping note ('Local only') that distinguishes it from potential cloud operations. However, since no sibling tool performs pulling, the differentiation is implicit rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides minimal context about when to use this tool: it implies it is for pulling into a running machine's local cache. There is no explicit guidance on when not to use it or mention of alternatives, even though the target parameter distinguishes local vs cloud. The agent is left to infer that this is the only tool for image pulling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read-fileB
Read a file from a machine.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Machine name | |
| path | Yes | Absolute path inside the machine | |
| target | No | Which fleet to talk to. local is smolvm serve on this host; cloud is smol cloud at SMOL_CLOUD_URL. | local |
| encoding | No | utf8 |
Output Schema
| Name | Required | Description |
|---|---|---|
| path | Yes | |
| size | Yes | |
| content | Yes | |
| encoding | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It merely says 'Read' without explaining side effects (e.g., read-only), error behavior, permission requirements, or how encoding affects the response. The schema hints at target and encoding, but the description itself is silent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with zero waste, and the core action is front-loaded. It is appropriately brief given the rich input schema and the presence of an output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description plus schema define the required parameters and target, and an output schema exists to explain return values. However, there is no mention of edge cases like file-not-found, permission issues, or encoding's practical effect on the response, leaving noticeable gaps for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 75%, and the schema already documents name, path, and target. The description adds no parameter-specific meaning. Encoding lacks an explanation, but its enum values are somewhat self-explanatory, so the schema carries the load adequately.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Read a file from a machine') with a clear verb and resource. It is distinct from siblings like write-file and machine-logs, though it does not explicitly name an alternative or contrast with related tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use or alternatives are mentioned. The description only says what the tool does, not when to prefer it over run-command, get-machine, or machine-logs. There is no guidance on selecting local vs cloud target or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run-commandC
Run a command in a running machine. exitCode comes from the guest; a failing command is not an error.
| Name | Required | Description | Default |
|---|---|---|---|
| env | No | Environment variables | |
| name | Yes | Machine name | |
| stdin | No | ||
| target | No | Which fleet to talk to. local is smolvm serve on this host; cloud is smol cloud at SMOL_CLOUD_URL. | local |
| command | Yes | argv array, or a string run with sh -c | |
| workdir | No | ||
| timeoutSecs | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| stderr | Yes | |
| stdout | Yes | |
| exitCode | Yes | |
| timedOut | Yes | |
| truncated | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds genuine behavioral nuance: 'exitCode comes from the guest; a failing command is not an error.' This is valuable beyond the schema because it tells agents to treat a non-zero exit code as data, not failure. However, it does not explain stdout/stderr delivery, behavior for a stopped machine, or command serialization.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no redundancy. The core action is stated in the first sentence, and the subtle exit-code behavior is given in the second. Every word contributes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values are not required in the description. The description captures the essential operation and the exit-code nuance, but gives no guidance on selecting this tool over siblings and no parameter hints. Not wholesale but has clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description offers no parameter-level information. Schema description coverage is only 57% (the schema covers env, command, target), so the description would need to add meaning for ambiguous parameters like stdin, workdir, and timeoutSecs, but it does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Run a command') and resource ('in a running machine'). This clearly separates it from machine creation/deletion tools, and the phrase 'running machine' distinguishes it from likely one-shot sibling 'run-once', though that distinction is implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to use this tool versus alternatives like 'run-once' or prerequisites such as 'machine must be running'. The description shows the intended context in passing but does not explicitly state when not to use the command.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run-onceA
Create a throwaway machine from an image, run one command, and delete the machine even on timeout.
| Name | Required | Description | Default |
|---|---|---|---|
| env | No | Environment variables | |
| cpus | No | ||
| image | Yes | ||
| stdin | No | ||
| target | No | Which fleet to talk to. local is smolvm serve on this host; cloud is smol cloud at SMOL_CLOUD_URL. | local |
| command | Yes | argv array, or a string run with sh -c | |
| network | No | Egress mode. Default open, except run-once on cloud which is blocked. Local: a blocked machine whose image still has to be pulled from a registry is refused by the API; pass open or an allow-list for that create. Cloud: blocked is sent as an allow-list of an unroutable range. | |
| workdir | No | ||
| memoryMb | No | ||
| allowCidrs | No | Egress allow-list of CIDR ranges. Overrides network. | |
| allowHosts | No | Egress allow-list of hostnames. Overrides network. Local: sent as allowedHosts. Cloud: sent inside the same cidrs list the published schema names, alongside allowCidrs. | |
| timeoutSecs | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| stderr | Yes | |
| stdout | Yes | |
| machine | Yes | |
| exitCode | Yes | |
| timedOut | Yes | |
| truncated | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses the guarantee of deletion even on timeout, which is a strong behavioral trait. It also implies resource creation and termination. It does not mention failure modes, such as what happens if creation fails or if the command fails, but the core lifecycle is transparent. The description adds significant value beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence that front-loads the primary action and lifecycle. Every element serves a purpose: 'throwaway' signals no persistent state, 'run one command' specifies the scope, and 'delete even on timeout' highlights a key guarantee. No filler; it's efficient and easily parseable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (12 params, nested objects, output schema), the description provides essential lifecycle semantics but leaves many parameter-specific details to the schema. Since an output schema exists, return values are documented. The description covers the unique behavioral aspect (deletion) and the core workflow, which is sufficient for an agent to select and invoke it correctly. It does not list replacement logic or edge cases, but these are not always necessary. With the schema and output schema, this is reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%, not high, so the description must compensate. Many parameters (env, cpus, image, stdin, target, command, network, workdir, memoryMb, allowCidrs, allowHosts, timeoutSecs) are already described in the schema, but the description adds critical value by noting that 'allowCidrs' and 'allowHosts' override network, and that 'blocked' egress behaves differently on local vs cloud. This is crucial context not in the schema. Even with some schema coverage, this extra clarification is valuable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('run'), resource ('machine'), and lifecycle semantics ('create', 'run one command', 'delete'). It distinguishes itself from siblings by emphasizing the throwaway nature and the deletion guarantee even on timeout. The description is dense but conveys the core workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for one-off commands without cleanup, but it doesn't explicitly contrast with run-command or create-machine. It mentions deletion on timeout, which hints at when to prefer this tool (need for cleanup), but does not provide explicit 'use X instead' guidance. Given five siblings, more explicit routing would improve clarity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stop-machineA
Stop a running machine (it can be started again with create-machine's name or the CLI).
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Machine name | |
| target | No | Which fleet to talk to. local is smolvm serve on this host; cloud is smol cloud at SMOL_CLOUD_URL. | local |
Output Schema
| Name | Required | Description |
|---|---|---|
| id | Yes | |
| pid | Yes | |
| cpus | Yes | |
| name | Yes | |
| image | Yes | |
| state | Yes | |
| network | Yes | |
| memoryMb | Yes | |
| createdAt | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of behavioral disclosure. It does disclose the key reversible outcome ('can be started again'), which is useful, but it does not mention what happens if the machine is already stopped, whether resources are released, or any side effects beyond the stop operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One concise sentence front-loads the core action and adds a relevant parenthetical about restartability. There is no wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple lifecycle command with a fully documented schema and output schema present, the description is nearly complete. It lacks only minor details like behavior when the machine is already stopped, but nothing essential is missing for a correct call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents both 'name' and 'target'. The description adds no additional parameter semantics, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Stop') and resource ('running machine'), and the parenthetical clarifies that stopping is reversible. This clearly distinguishes it from delete-machine and create-machine without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description indicates this tool is for stopping a running machine, and it names create-machine as the way to start it again. It doesn't explicitly state when NOT to use this tool, but the context and alternative are clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
write-fileA
Write a file into a machine. Waits until the workload container runs so the file is not lost.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Machine name | |
| path | Yes | Absolute path inside the machine | |
| target | No | Which fleet to talk to. local is smolvm serve on this host; cloud is smol cloud at SMOL_CLOUD_URL. | local |
| content | Yes | ||
| encoding | No | utf8 |
Output Schema
| Name | Required | Description |
|---|---|---|
| path | Yes | |
| size | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It discloses the waiting behavior, which is a useful behavioral trait beyond the schema. However, it doesn't mention whether the file is overwritten if it exists, permissions, or any side effects. The description adds some value but could disclose more about the operation's effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, clear and to the point. The key behavior (waiting) is front-loaded after the primary action. No fluff, every word adds meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (write a file) and the presence of an output schema, the description covers timing behavior, which is the critical gap. It doesn't mention overwrite semantics or error conditions, but for a basic write operation, this is fairly complete. The output schema likely covers return values, so no need to explain those.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema coverage is 60%, meaning two parameters (content and encoding) lack descriptions in the schema. The description does not add semantics for these, but the schema already covers name, path, and target well. Since content and encoding are straightforward, the description doesn't need to add much. The description does add the nuance that path should be 'absolute', which is helpful beyond the schema's 'Absolute path' description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the verb 'Write a file' and the resource 'into a machine', clearly identifying the primary action. It distinguishes from siblings like read-file and run-command. However, it doesn't explicitly note that this is for a remote machine vs local, but the context of 'machine' and sibling names make it clear enough.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions 'Waits until the workload container runs so the file is not lost', providing a timing context that implies when this is safe to use. It doesn't explicitly state when not to use it or compare to alternatives like run-command for writing files. The guidance is implied but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
11 tool updates
v0.1.0- First observed
create-machine - First observed
delete-machine - First observed
get-machine - First observed
list-machines - First observed
machine-logs - First observed
pull-image - First observed
read-file - First observed
run-command - First observed
run-once - First observed
stop-machine - First observed
write-file
TDQS
Scored across 11 tools
Each tool targets a specific resource and action: machine lifecycle, command execution, file transfer, and image caching are cleanly separated. run-command and run-once may seem similar at first, but the descriptions clearly distinguish running in an existing machine from a throwaway one-shot workflow.
Tool names mostly follow a consistent lowercase hyphenated verb-noun pattern (list-machines, get-machine, create-machine, read-file, write-file). Minor deviations like machine-logs (noun-noun) and run-once (verb-adverb) slightly break the pattern but do not cause confusion.
With 11 tools, the server is well-scoped for managing machines, running commands, transferring files, and pulling images. Each tool adds meaningful capability without redundancy.
The surface covers the core machine lifecycle (create, list, get, stop, delete), command execution, file transfer, image pulling, and logs. Minor gaps exist, such as no explicit machine update/resize or image listing, but agents can accomplish the main workflows without dead ends.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
MCP server for building and testing AI agents with multi-model experimentation and insights.
MCP server for AI agents to plan, verify, and deploy Cloudflare-native apps.
Nifty's MCP server — exposes tasks, projects, messages, and files as tools for AI agents.
A registry of AI agent tools — MCP servers, APIs, CLIs, SDKs — kept current by automated ingestion.
Related MCP Servers
- FlicenseAqualityDmaintenanceAn MCP server for managing Incus virtual machines through structured tools for command execution, file management, and snapshot operations. It enables AI agents to puppeteer VMs on a masternode by wrapping the Incus CLI.9-
- AlicenseNot gradedqualityDmaintenanceAn MCP server that gives AI agents direct control over QEMU virtual machines.281MIT
- AlicenseNot gradedqualityCmaintenanceA free, open-source MCP server that gives AI agents real filesystem, terminal, git, and process control over your machine for inspection, diagnosis, and repair.MIT
- AlicenseNot gradedqualityCmaintenanceA self-hosted MCP server that gives AI agents controlled access to a machine: filesystem, shell, background processes, git, web fetching and persistent key-value memory.GPL 3.0