Skip to main content
Glama

mcp-drone-os

An Arch-based, LAN-only disposable agent runner. The project builds a bootable USB image with hardened SSH, per-agent Unix identities, isolated SMB shares, an MCP bridge, and a local control-plane dashboard.

Architecture

Run one coordinator on the laptop where the host agent lives. Booted machines are drones. The host agent connects to one local MCP server; the coordinator keeps fleet state and later maintains one persistent SSH control connection per drone. Drones run one-shot jobs or durable services.

host agent -> local mcp-drone-mcp -> coordinator state -> SSH -> drones

Job submission is batched. A twelve-variant experiment is one MCP call, not twelve independent shell conversations. A drone advertises available slots and capabilities; placement can be explicit or capability-aware.

Related MCP server: RelayMe

Features

This repository contains the contracts and safe host-side tooling needed for the first image profile:

  • mcp-drone-os validate validates an agent manifest without network access.

  • mcp-drone-os devices lists removable block devices.

  • mcp-drone-os write requires --yes, an exact device confirmation, and a removable-disk check before writing.

  • mcp-drone-os events appends and reads the legacy JSONL event format.

  • mcp-drone-mcp provides a local JSON-RPC stdio MCP gateway.

  • mcp-drone-agent provides a persistent JSON-lines worker protocol for launching, inspecting, and reading logs from transient systemd tasks.

  • mcp_drone_os.ssh_transport.SSHDrone keeps one SSH worker stream alive and uses OpenSSH connection multiplexing for repeated operations.

  • The coordinator stores drones, tasks, artifacts, and audit events in SQLite.

  • The MCP gateway exposes registration, batch submission, status, bounded logs, artifact listing/fetching, cancellation, reconciliation, persistence promotion, and user-event recording.

  • run_jobs accepts many variants in one call. Each job can select a drone; otherwise placement uses available slots. Jobs support one-shot work, durable services, and long-lived provisioning work.

  • Each drone uses a persistent JSON-lines SSH worker stream plus OpenSSH connection multiplexing. The host agent sends coarse batch operations rather than starting a new SSH process for every command.

  • Drones use per-agent Unix accounts, transient systemd units with CPU, memory, and process limits, and rootless Podman/Quadlet for durable services.

  • Enrollment supports agent keys, a separate controller key, and an OpenSSH host certificate signed by a coordinator-held CA. Password and root SSH are disabled.

  • A preconfigured image boots without a framebuffer login: it creates declared agent accounts, installs the controller key, derives a unique mcp-drone-<NIC> hostname, records IP/CPU/memory in /run/mcp-drone/identity, and advertises _ssh._tcp over mDNS.

  • runner/profile/ is an ArchISO profile with the minimal runtime package set, SSH hardening, first-boot enrollment hook, and systemd target.

  • tools/test-image.sh boots a built image in QEMU with SSH forwarded to port

Install and use

python3 -m venv .venv
. .venv/bin/activate
pip install -e '.[test]'
mcp-drone-os validate examples/manifest.json
mcp-drone-os devices

# Start the local MCP gateway for an MCP client
mcp-drone-mcp

# The coordinator can also be inspected directly
mcp-drone-os fleet --state /var/lib/mcp-drone/coordinator.db

Launch mcp-drone-mcp as a local stdio MCP server for the host agent. The dashboard is optional and defaults to localhost; view it remotely through SSH port forwarding. Use TLS and a network boundary before exposing it on a LAN.

USB-only operation

After the ISO is built, the runtime does not need this repository, Python packages, or an internet connection. Boot the USB and use the bundled commands:

/usr/local/bin/mcp-drone-os fleet --state /var/lib/mcp-drone/coordinator.db
/usr/local/libexec/mcp-drone-mcp

An MCP client on another computer can use the USB-hosted coordinator over SSH by launching /usr/local/libexec/mcp-drone-mcp as its SSH stdio command. The dashboard is available through ssh -L 8787:127.0.0.1:8787 ...; it is not exposed on the LAN by default. Enrollment and remote multi-drone jobs still need LAN connectivity to the other machines, but the USB itself is the complete control-plane/runtime artifact.

Build an image after installing Arch’s archiso package:

tools/build-image.sh out
tools/test-image.sh out/mcp-drone-os-0.1.0-x86_64.iso

For a dedicated machine, boot the ISO and run the guarded persistent installer from the live shell. It repartitions the selected whole disk, creates a LUKS2-encrypted root, installs the runner filesystem, and configures BIOS and UEFI GRUB boot paths:

/usr/local/bin/mcp-drone-install \
  --device /dev/sda --confirm-device /dev/sda --yes

The installer asks for the device path a second time and prompts for the disk passphrase. Verify the target carefully: this operation destroys all existing partitions on that disk. The live image remains the recommended disposable mode when a machine should be reset by rebooting.

The image configures wired Ethernet and USB Ethernet for DHCP through systemd-networkd. Wi-Fi credentials are intentionally not baked into the generic image; configure the appropriate network profile on machines that need wireless operation.

To bake a runner enrollment bundle into the image, set both MCP_DRONE_MANIFEST and MCP_DRONE_ENROLLMENT. The enrollment JSON must reference /etc/mcp-drone/manifest.json; the builder copies the selected manifest to that path. Enrollment tokens are secrets and should not be committed to the repository.

To make a USB immediately accept work from one controller, also provide its public key at build time. It is copied into the image but is not committed:

MCP_DRONE_MANIFEST=manifest.json \
MCP_DRONE_ENROLLMENT=enrollment.json \
MCP_DRONE_CONTROLLER_PUBLIC_KEY="$HOME/.ssh/id_ed25519.pub" \
tools/build-image.sh out-ready

The enrollment manifest declares the agent IDs that become drone-<id> Unix accounts. Discover a runner with avahi-browse -rt _ssh._tcp, then register its advertised hostname and SSH account with the coordinator. register_drone accepts ssh_user for hosts whose SSH username differs from the default mcp-control.

For an offline/preconfigured runner, start with examples/enrollment.json; add coordinator_url and a one-time token when the coordinator is available on the LAN.

Enrollment URLs must use HTTPS by default. For a temporary isolated lab only, allow_insecure: true can be set in the enrollment bundle. The dashboard supports --tls-cert and --tls-key when it must be reachable from a LAN; otherwise keep its localhost default and use SSH forwarding.

For authenticated host enrollment, configure the coordinator with MCP_DRONE_ENROLL_TOKEN and MCP_DRONE_SSH_CA_KEY. The drone generates its own host key on first boot and receives an OpenSSH host certificate; the CA private key never goes into the image.

SMB is intentionally optional. Install samba in a custom profile, generate /etc/samba/mcp-drone-shares.conf with mcp_drone_os.shares.render_shares, then enable mcp-drone-smb.service. Agent homes remain accessible over SFTP without enabling Samba.

Parallel example: 12 experiments on 3 machines

Register three runners with four slots each. Submit one batch through MCP:

{
  "agent_id": "research-agent",
  "mode": "once",
  "defaults": {"limits": {"cpu_quota": "200%", "memory_max": "2G", "tasks_max": 256}},
  "jobs": [
    {"command": ["./run-test", "--variant", "0"]},
    {"command": ["./run-test", "--variant", "1"]},
    {"command": ["./run-test", "--variant", "2"]}
  ]
}

Continue the array through variant 11. The coordinator persists all tasks, dispatches up to the available slots, and fills slots as work finishes. Use fleet_snapshot or reconcile_running for progress, then task_artifacts and fetch_artifact to collect results. The laptop handles coordination and metadata; the three drones perform the CPU and I/O work.

For scripts or repositories, call stage_bundle once per destination drone, then have the batch commands reference /var/lib/mcp-drone/staging/<bundle-id>. Jobs submitted for agent_id=alpha run as the provisioned drone-alpha Unix account, so the manifest must declare that agent on each target runner.

Current boundaries

This is a small v1 runner, not a general cluster scheduler. It does not flash media automatically, provide browser login, or implement remote HTTP MCP. Persistence promotion still expects an already-mounted data volume; the separate persistent installer handles whole-disk installation and LUKS2 root encryption. SMB, Kubernetes, and remote MCP transport remain optional future work.

USB writes are destructive and are never part of the build step. The image enrollment hook is intentionally explicit: it only registers when given a coordinator URL and one-time token.

Security defaults

The image profile disables root/password SSH, accepts only declared public keys, binds services to the private interface, and creates one Unix identity and share root per agent. The host builder never guesses a target disk.

Repository layout

schemas/manifest.schema.json   manifest contract
src/mcp_drone_os/              coordinator, MCP gateway, and state
runner/profile/                bootable ArchISO profile
runner/systemd/                service templates
tools/                         image build and QEMU test entrypoints
tests/                         unit and contract tests

Available Tools

13 tools
cancel_taskB

Cancel a queued or running task.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and delivers little: it doesn't say whether cancellation is reversible, whether a running task is killed gracefully or abruptly, whether the call is idempotent, or what permissions are needed. Only the applicable task states are disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One tight sentence with the verb front-loaded and zero filler. It is efficient, though the terseness is partly due to under-specification rather than maximal information density.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A destructive mutation tool with no annotations and no output schema needs more than a single sentence. Missing are reversibility, effect on a running task, and any indication of error behavior, leaving real gaps for an agent deciding whether to invoke it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter (task_id) is fully covered by the schema (100% coverage), so the schema does the heavy lifting. The description adds no meaning about the identifier's format or source, which is the baseline-3 outcome for one documented parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clear verb+resource: 'Cancel a ... task', with state scope ('queued or running') that hints at when it applies. It doesn't name or contrast any sibling (e.g. task_status), so it falls short of the 5 bar for explicit sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Queued or running' implies the valid task states, giving implied usage. However, there is no guidance on when to cancel versus checking task_status first, no prerequisites, and no mention of alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dispatch_pendingB

Dispatch queued tasks to connected drones.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It does not say whether dispatch is synchronous or async, whether it is idempotent or safe to retry, what happens to tasks with no available drone, or whether it mutates queue state—all material for a bulk dispatch operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler or repetition. It is efficient, though the brevity is partly achieved by omitting information an agent would need.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a state-changing fleet operation with no annotations, no output schema, and no parameters, the description is too thin. An agent cannot tell what the dispatch returns (task IDs? count?), whether it must be followed by task_status polling, or how it relates to the other eleven sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline for a parameterless tool applies. No misleading parameter language is present.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb ('Dispatch') plus resource ('queued tasks') plus destination scope ('connected drones'), which is clear on its own. However, it does not distinguish itself from the sibling 'run_jobs', which plausibly overlaps in meaning, so an agent cannot confidently choose between the two from the text alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No statement of when to use this versus 'run_jobs', 'stage_bundle', or 'reconcile_running', and no prerequisites (e.g., whether drones must be registered first). Usage must be inferred entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fetch_artifactC

Fetch one verified artifact to the coordinator host.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it discloses almost nothing beyond the one-line action. It does not say what 'verified' means, whether an existing file at the destination is overwritten, what permissions are required, or what happens when verification fails.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single, front-loaded sentence with no filler, which is appropriate sizing for a narrow tool. The terseness borders on under-specification, but that is penalized elsewhere rather than here.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and three required parameters, the description leaves key gaps: the meaning of 'verified', the relation to siblings like task_artifacts, and any failure or overwrite semantics. An agent can guess the happy path but cannot call this confidently in edge cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is reported at 100%, so the structured schema already documents the parameters and the baseline of 3 applies. The description adds only the loose notion that the fetched artifact lands on the coordinator host, which partially maps to 'destination' but adds no syntax or format detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (fetch), resource (one verified artifact) and destination (the coordinator host), so the action is unambiguous. It does not, however, distinguish itself from related siblings such as task_artifacts or stage_bundle, which an agent must infer from names alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no named alternative for retrieving artifacts (e.g. task_artifacts for listing). The word 'verified' hints at a gating condition but the description never explains it or says when this tool should be preferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fleet_snapshotC

Get drones, capacity, and task state.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden but discloses almost nothing. 'Get' implies a read, yet there is no statement about point-in-time semantics, freshness, permissions, or scope (whole fleet vs. one). Those gaps are exactly what an agent needs for a fleet-wide tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence, front-loaded with the verb and the enumerated resources. It is efficient with no filler, though its brevity reflects under-specification rather than tight editing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No parameters, no annotations, and no output schema, so the description must explain the return surface and usage context, and it does not. An agent does not know the snapshot's shape, granularity, or freshness, which matters for a fleet-state tool with this many siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema baseline of 4 applies. There is nothing for the description to clarify regarding inputs, and it does not misrepresent the no-arg signature.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a verb (Get) and three resources (drones, capacity, task state), so the agent knows the domain. However, 'capacity' and 'task state' are vague, and nothing distinguishes this fleet-wide snapshot from siblings like task_status or task_logs, which also surface task state. Purpose is readable but not sharply differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance and no mention of alternatives. An agent cannot tell from the description whether fleet_snapshot should be preferred over task_status or dispatch_pending when it needs drone/task information. Usage is only weakly implied by the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

promote_persistentC

Move a drone onto an already-mounted persistent data volume.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden. It states one precondition (the persistent volume must already be mounted), which is useful, but says nothing about whether existing drone data is preserved or overwritten, whether the operation is idempotent, what permissions are needed, or what state the drone ends up in.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the action front-loaded and no filler. The terseness does border on under-specification, but there is no wasted text to trim.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a state-changing operation with no annotations and no output schema, one sentence is not enough. An agent cannot tell what the call destroys, whether it is reversible, or what prerequisites beyond the mounted volume apply, and nothing distinguishes it from the surrounding lifecycle tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema declares drone_id as required but supplies no property documentation, and the description never mentions how the target drone is identified. With schema coverage reported at 100% (trivially, since there are no documented properties) and only one required parameter, this lands at the baseline rather than below it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a concrete verb and object ('Move a drone onto an already-mounted persistent data volume'), which tells an agent the operation is a relocation of an existing drone rather than a registration or dispatch. It does not differentiate itself from siblings like register_drone or stage_bundle, and the intent behind 'promote' (what changes for the drone's persistence) is not spelled out.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use, no when-not-to-use, and no named alternative among the many fleet/task siblings. The precondition 'already-mounted' is the only usage signal, and it is buried rather than framed as a selection criterion.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reconcile_runningB

Reconcile running task state with connected drones.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. 'Reconcile' implies state mutation and possible side effects on drones, but the description never states whether it writes changes, is idempotent, requires authorization, or what happens to mismatched tasks.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the action and scope front-loaded and no filler. Nothing is padded or redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a likely state-mutating operation with no annotations, no output schema, and no parameters, the description should explain what reconciliation actually does (correct drift, restart, cancel orphaned tasks) and its side effects. It leaves the agent guessing at the core behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to clarify; the baseline of 4 applies. No parameter detail is missing or needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (reconcile) and resources (running task state, connected drones), which is more than a restatement of the name. It does not, however, distinguish itself from siblings like task_status or dispatch_pending, which also touch task/fleet state.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of the sibling tools that touch similar state (task_status, dispatch_pending, cancel_task). The agent must infer the trigger condition entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

register_droneC

Register or heartbeat a drone. ssh_user may name the SSH account or config alias user.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full disclosure burden, and it only implies upsert-like behavior via "register or heartbeat". It omits auth requirements, what happens on duplicate drone_id, whether an existing record is overwritten, and any rate/retry semantics — all material for a state-mutating registration call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core action and no filler. The second sentence is terse to the point of being cryptic, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A mutation-style tool with three required params, no annotations, and no output schema needs more: registration/update semantics, duplicate handling, and expected response are all absent. The description is too thin to safely invoke without guessing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema documents drone_id, hostname, and address and the baseline is 3. The description adds only one useful clarification (ssh_user ambiguity), and that parameter is not visible in the listed schema, so its applicability is unclear.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Register or heartbeat a drone") and the dual mode is clear. Siblings like fleet_snapshot, run_jobs, and task_status are unrelated operations, so the agent can place this as the fleet-membership/registration entry point, though the description never explicitly contrasts itself with any sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance: nothing says when to register vs. when to heartbeat, whether it is idempotent, or whether this should be called before run_jobs/dispatch_pending. The single hint is the word "or", which leaves the choice to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_jobsC

Submit and, by default, dispatch one-shot or long-lived jobs as a batch.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry full behavioral burden. It only states submit/dispatch and mentions a default, but omits permissions, side effects, long-lived job semantics, cancellation behavior, and batch error handling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. It is efficient, though arguably too sparse for a batch dispatcher.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a batch job dispatch tool with no annotations, no output schema, and many related siblings, one sentence is not enough. It lacks prerequisites, return expectations, and side-effect disclosure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is reported at 100%, so the baseline is 3. The description gestures at 'jobs' and 'batch' but does not clarify agent_id, job structure, or dispatch control beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: submit and dispatch jobs as a batch. The default dispatch behavior is useful, but it does not distinguish itself from siblings like dispatch_pending or promote_persistent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use guidance or alternatives are given. The phrase 'by default' hints at a dispatch/not-dispatch option, but there is no routing advice against sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stage_bundleB

Copy a local script or directory to one drone once for reuse by many jobs.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden. It conveys that this writes content onto a drone, but says nothing about overwrite behavior for an existing bundle_id, permissions/auth needed, idempotency, or whether the copy is blocking. For a state-mutating tool this leaves significant gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence that front-loads the action and ends with the reuse rationale. Every clause earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-required-parameter mutation tool with no annotations and no output schema, the description is adequate on intent but thin on behavior — bundle_id uniqueness, overwrite semantics, and failure modes are unaddressed. The fully covered schema offsets some of this.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description loosely maps to source ('local script or directory') and hints at drone/bundle reuse, but adds no format, path, or uniqueness semantics beyond what the schema already documents.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource ('Copy a local script or directory') plus scope ('to one drone once for reuse by many jobs'), which clearly separates it from siblings like run_jobs or fetch_artifact. It stops short of naming an alternative explicitly, but the intent is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'once for reuse by many jobs' implies the staging-before-execution workflow, but there is no explicit when-to-use statement, no prerequisites, and no named alternative (e.g. when to stage versus passing source inline to run_jobs). Usage must be inferred from the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_artifactsC

List task artifact metadata.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full behavioral burden, and it discloses almost nothing. The word "metadata" implies contents are not returned and that fetch_artifact handles the actual bytes, but there is no statement about permissions, ordering, pagination, or what happens when a task has no artifacts.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single four-word sentence with the verb and resource front-loaded and no filler. It is efficient, though the extreme brevity borders on under-specification rather than disciplined conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple list tool with no output schema and no annotations, the description is minimally adequate: it tells the agent what is listed and for which task. It omits return shape hints, ordering, and its relationship to fetch_artifact, which are the main gaps an agent would face.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is reported at 100%, so the schema is expected to carry parameter documentation and the baseline is 3. The description adds nothing about task_id (e.g., whether it is required or scopes results), so it does not exceed the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: "List task artifact metadata." An agent can tell it produces a listing of metadata for a task's artifacts rather than artifact contents or task status. It does not explicitly distinguish itself from the sibling fetch_artifact, which is the tool most likely to be confused with it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is provided. With siblings like fetch_artifact, task_status, and task_logs present, the description never states which one to pick or under what condition, leaving the agent to infer that this is the metadata-listing variant.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_logsC

Read bounded task logs.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It implies a read-only operation but does not explain what 'bounded' means, whether logs are paginated or size-limited, what permissions are required, or what the response looks like.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

At four words, it is front-loaded and wastes nothing. It is appropriately terse for a simple read tool, though the unexplained 'bounded' qualifier keeps it from being fully clear.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations and no output schema, the description is incomplete. It leaves open which task's logs are returned, what 'bounded' means, how logs are formatted, and when to use this tool instead of its many siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents the required task_id parameter. The description adds no parameter meaning beyond the schema, which is the expected baseline when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb ('Read') and resource ('task logs'), so an agent can tell the fundamental operation. However, it does not differentiate this tool from siblings like task_artifacts or task_status, and the modifier 'bounded' is left unexplained.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is provided. It does not say when to choose task_logs over task_status, task_artifacts, or other sibling tools, nor does it imply any prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_statusC

Read one task.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. "Read" implies a non-mutating operation, but the description says nothing about permissions, error behavior when the task_id is unknown, or what state information is returned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is front-loaded and has zero filler, but it is under-specified rather than truly concise — brevity here reflects a lack of content, not economical phrasing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read tool with no annotations and no output schema, the description is too thin. It does not describe what a task's status comprises, leaving an agent without enough context to know when this tool is the right choice.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is reported at 100%, so the structured schema already documents the required task_id. The description adds no syntax, format, or semantics beyond what the schema provides, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Read one task" provides a verb (read) and a resource (task) but is generic and does not reflect the tool's actual name/scope (task_status). It does not distinguish it from siblings like task_logs or task_artifacts, which also read data for a task.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no conditions, and no mention of alternatives. An agent must infer that this is the status-inspection tool versus task_logs/task_artifacts purely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

user_eventC

Record a dashboard/user action for an agent.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. 'Record' implies a write, but the description says nothing about side effects, idempotency, required permissions, whether the event is persisted, or what the call returns. For a mutation with zero annotation coverage this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words. It is concise, but the brevity reflects under-specification rather than disciplined economy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a write tool with no annotations, no output schema, and two undocumented required parameters, the description should carry far more. It omits payload expectations, side effects, and any usage context, leaving the agent without enough to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema lists two required parameters (agent_id, action) but provides no property descriptions, so the description must clarify them. It only obliquely gestures at 'an agent' and never explains what an 'action' value should look like. The 100% coverage figure is misleading since there are essentially no documented properties to cover.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a verb ('Record') and a resource ('dashboard/user action for an agent'), so the basic intent is inferable. However, 'dashboard/user action' is vague and the description does nothing to distinguish this tool from the many task/agent-oriented siblings like task_status or register_drone. It is adequate but not specific enough to uniquely identify the operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool, when not to, or which sibling to prefer for related needs. The agent gets no routing signal at all beyond the vague noun phrase.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 13 tool updatesv0.1.0
    • First observedcancel_task
    • First observeddispatch_pending
    • First observedfetch_artifact
    • First observedfleet_snapshot
    • First observedpromote_persistent
    • First observedreconcile_running
    • First observedregister_drone
    • First observedrun_jobs
    • First observedstage_bundle
    • First observedtask_artifacts
    • First observedtask_logs
    • First observedtask_status
    • First observeduser_event

TDQS

B3.2/5.0

Scored across 13 tools

Disambiguation4/5

Most tools target clearly distinct resources or actions (task_status vs task_logs vs task_artifacts, fetch_artifact vs task_artifacts). The one soft overlap is run_jobs, which 'submits and by default dispatches' jobs, versus dispatch_pending and reconcile_running, which also manage task dispatch state, creating some ambiguity about when to use each. Otherwise boundaries are clean.

Naming Consistency4/5

All names use snake_case consistently. Action tools follow verb_noun (register_drone, cancel_task, dispatch_pending, fetch_artifact) while read tools use noun_noun (task_status, task_logs, task_artifacts, fleet_snapshot), which is a readable convention but not a single uniform pattern.

Tool Count5/5

13 tools sits well within the ideal range and each maps to a real facet of drone orchestration (fleet state, registration, job submission, staging, task inspection, artifact handling, dispatch/reconcile). No obvious filler or redundancy in the count.

Completeness4/5

The surface covers a full task lifecycle: register, stage, submit, dispatch, monitor, fetch artifacts, cancel, reconcile, and promote persistence. Minor gaps exist, such as no explicit drone deregistration/removal or bulk task listing, but core workflows have no dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    Durable MCP server for managing long-running jobs locally, over SSH, or on Slurm clusters. Jobs survive client disconnects and return exit codes, bounded logs, and JSON artifacts.
    11
    41 PyPI
    1
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to safely run bounded, sandboxed tasks on remote machines with persistent state and reviewable artifacts.
    Apache 2.0
  • A
    license
    A
    quality
    A
    maintenance
    Runs commands on machines you cannot SSH into by publishing steps to a transport the far side already reaches (git, file share, S3, relay), then reads back a log with every line timestamped in UTC. Seven tools cover send, status, logs, stall detection and a pre-flight check.
    7
    Apache 2.0
  • A
    license
    C
    quality
    B
    maintenance
    Exposes the SSH + tmux client over stdio so agents can drive hosts, tmux sessions, commands, broadcasts, tunnels and audit logs programmatically. It offers 47 grouped tools covering host and session management, connectivity checks, and background port forwarding.
    47
    MIT