Droid MCP Controller
Connects Factory Droid runs to Amp's MCP OAuth HTTP endpoint, attaching the amp-puck MCP server on task create and resume so Droid tasks can report into Amp/Puck conversations through the puck tool (with all other Amp tools denied). Includes workspace/reasoning/autonomy tooling for starting, continuing, inspecting, cancelling and listing Droid runs, and discovery of the live authenticated Factory model catalog.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Droid MCP ControllerStart a Droid task to review my latest changes without editing files."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Droid MCP controller
A small, same-host MCP server for Puck to start, inspect, cancel, and continue Factory Droid tasks. No UI, database service, Factory daemon, account API, or completion-notification dependency. Installation and secure connection steps are in HANDOFF.md.
Requires Node.js 22+, Linux or macOS, and an authenticated Droid CLI. The target
is Droid 0.233.0, normally ~/.local/bin/droid. The controller inherits the
service user's login environment; it does not collect credentials. Dependencies
are locked. uuid is overridden to 11.1.1 to fix GHSA-w5hq-g745-h8pq; the Factory
SDK's v4 usage is covered by the protocol E2E suite.
npm ci --ignore-scripts
npm run check
npm test
node src/server.mjs --config /absolute/path/to/config.jsonPuck calls seven tools
Tool | Required arguments | Behavior |
|
| Returns immediately with a new controller |
|
| Loads the prior run's Droid UUID and workspace; returns a new runId. |
|
| Durable state, UUID, timestamps, recent events, partial text, separate stderr. |
|
| Final text/outcome if available, otherwise explicitly partial progress. |
| none | Discover durable runs, newest first. |
|
| Interrupt; poll until terminal. Never undoes edits. |
| none | Current authenticated Factory model catalog, cached briefly; never a historical allowlist. |
Start/continue also accept autonomy: "off" | "low" | "medium" | "high" and
optional model and reasoningEffort. Every new turn uses the host's
defaultAutonomy (high unless explicitly configured), including continuation.
The host's maxAutonomy ceiling can only be changed in its local configuration.
off sets Spec mode plus autonomy off; the others set Auto mode plus the
requested level through JSON-RPC. Both source defaults and the example config
use defaultAutonomy: "high" and maxAutonomy: "high" for owner-authorized full
Auto execution. This means supported service-user access, not OS/root privileges
or isolation. Explicitly request autonomy: "off" for read-only work. To lower
the host ceiling, also lower its default so it does not exceed the ceiling.
Tool discovery includes the host's approved workspace roots, autonomy ceiling,
and reasoning setting. Choose a workspace from those roots rather than guessing
a repository path. Call droid_models({}) when choosing a model, then pass a
returned ID explicitly on every start/continue to avoid inheriting a costly
default. Do not guess or reuse IDs from old documentation or saved sessions.
Reasoning is independent of autonomy. Per-turn reasoningEffort overrides the
host setting on both start and resume. The effective explicit value is saved
and returned by status/result/list. Older records may lack that field; they
are not rewritten. Without a turn or host value, Droid uses default/saved reasoning.
Use only reasoning levels supported by the selected Factory model.
droid_result accepts offset (default 0) and limit (default 12000, max 16000)
in JavaScript string characters. Follow nextOffset until null. It returns the
SDK's final answer, not a concatenation of progress messages. droid_list also
supports offset/limit (default 25, max 100).
Typical Puck sequence (tool arguments, not shell commands):
{"requestKey":"review-42","workspace":"/approved/project","model":"YOUR_ENABLED_MODEL_ID","reasoningEffort":"high","autonomy":"off","prompt":"Review the changes; do not edit files."}Call
droid_models, choose a current ID, then calldroid_startwith those arguments. RetainrunIdandrequestKey.Poll
droid_status({runId})at a reasonable interval, e.g. 2–5 seconds. Recent complete assistant/tool/error events show progress; a quiet stream does not prove the task has finished.Once
terminalis true, inspectstate, then retrievedroid_result.resultAvailablemeans a terminal Droid result was persisted; it is not a success flag. SDK execution success requiresstate: "succeeded"and a successful outcome; task completion also requires checking the requested artifacts and unanswered questions. An empty answer is not completion evidence.Continue only from the newest controller run in the session, with
{runId, requestKey:"review-42-followup", model:"YOUR_ENABLED_MODEL_ID", autonomy:"off", prompt:"Explain the first finding."}. Retain the returned NEW runId. The Droid UUID stays the same.For immediate steering: cancel the current run, await terminal status, inspect its outcome, then continue. Busy sessions reject follow-ups rather than queue them ambiguously. Never operate the same UUID from another Droid process.
Keys are global to this controller state directory. Retrying exactly the same start/continue arguments returns the existing run, even after restart. Reusing a key for different arguments fails. Use a new key for intentional new work. Omitted autonomy/reasoning on a replay use the originally accepted values, not changed host defaults. Explicit changed reasoning conflicts with the key. An accepted intent is synced to disk before spawning the worker. This provides at-most-one controller launch per key, not exactly-once Droid side effects.
A controller run ID represents one turn. A Droid session is linear: only its
authoritative head (newest accepted run) may be continued. A stale ancestor
returns error, code: "not_session_head", and headRunId. A failed, interrupted,
cancelled or timed-out successor remains head; unknown outcomes block the family.
An exact replay of an already accepted key still returns that run rather than
launching another turn, even if its source is no longer head.
Read-only (autonomy=off) runs may share canonical workspace trees. Any non-off
run requires exclusive access to its canonical tree, including ancestors and
descendants. Symlink spellings resolve to the same lock. Disjoint siblings may
run concurrently, still subject to maxConcurrentRuns. Conflicts reject with
code: "workspace_busy", conflictingRunId, and the conflicting canonical
workspace; there is no queue. Locks last through worker/process cleanup, not
merely until a terminal result is received. This scheduling rule is not a
filesystem sandbox or a cross-controller lock against unrelated Droid processes.
Related MCP server: windlass
Models come from the live authenticated runtime
droid_models initializes a short-lived Droid connection using the same CLI,
service user, HOME and FACTORY_HOME_OVERRIDE as normal runs. It captures live
availableModels/available_models, closes the transient session and transport,
and never submits a prompt or approves a tool. The probe uses controller state
cwd rather than a target project. Factory may retain an empty initialization
session in its own storage; the controller creates no run for discovery.
No additional Factory API key is required.
The response includes fetchedAt and models, ordered by displayName then id.
IDs/display names are required; optional provider/reasoning/image/PDF metadata is
returned only when supplied by Factory. Disabled entries are omitted. The
controller maintains no hard-coded compatibility list or historical model merge.
Successful discovery replaces the entire in-memory cache. modelCacheTtlMs
defaults to 60000 (allowed 5000–600000); concurrent requests share one refresh.
After expiry, failed refresh returns code: "model_discovery_failed", never the
expired catalog as current. Droid remains the final execution/authorization
authority, even for an ID returned by discovery. Pricing is not fabricated.
For an Amp connection named Droid Grokbot, discover droid-grokbot with
tool_search, then call the seven functions through code_exec:
import { droid_start } from "droid-grokbot"
text(await droid_start({
requestKey: "review-42",
workspace: "/approved/project",
model: "YOUR_ENABLED_MODEL_ID",
autonomy: "off",
prompt: "Review the changes; do not edit files."
}))The module name depends on the connection name. code_exec is a tool dispatcher,
not a host shell; it cannot edit the host configuration or run timer-based polling.
Poll in later calls. A separate runner is needed only for host administration,
not for ordinary Droid execution. There is no controller completion callback or
live prompt injection: follow-ups are serial after the prior turn terminates.
Droid reports and questions use the actual Amp MCP
Optional private host configuration connects Droid to Amp's OAuth HTTP endpoint:
{"puck":{"conversationId":"T-11111111-1111-4111-8111-111111111111","url":"https://ampcode.com/mcp?profile=external-agent"}}Replace the example ID with the intended Puck conversation privately. Complete
Factory's supported MCP OAuth once as the service user; keep the same server URL
used for that authorization (an existing optional threadID query is accepted).
The controller never accepts OAuth tokens/headers and never initiates consent
during task execution. The URL alone does not route messages: the actual puck
tool's params.conversationID does.
The controller attaches amp-puck using SDK mcpServers on create and resume,
independent of task cwd or its .factory files. Before each turn it discovers
tools, unions every amp-puck___* ID except exactly amp-puck___puck with existing
disabledToolIds, updates settings and re-lists. Restored disables are preserved;
native tools are not restricted. The actual registered amp-puck___puck must be
allowed and every other Amp tool denied, or no prompt is submitted. Discovery or
settings errors and cancellation also prevent submission. An
SDK-injected server can be absent from listMcpServers despite successful tool
discovery, so the registered tool catalog determines readiness.
Start/continue optionally accept puckConversationId. New sessions default to
the host recipient; continuation retains its prior recipient unless explicitly
overridden. Replays retain their accepted routing, and changing it conflicts
with the same request key. Every turn receives coordination instructions plus
its recipient, controller run ID and Droid UUID, so routing guidance also applies
outside this repository. AGENTS.md records the same protocol for repository work.
Droid calls amp-puck___puck with {action:"send",params:{conversationID,message}}.
Ordinary coarse progress reports are sent once, fire-and-forget; continue
authorized independent work without waiting. For an actual question, a material
blocker, or before an expensive/irreversible step when steering is wanted, send
one CHECKPOINT: explicit recipient, run/session handles, task marker, completed
evidence, proposed action and exact decision needed. Consume a correlated completed
reply inline; otherwise queued/working replies use only that send's exact
replyHandle with read_reply, at most 6 reads in the natural tool loop.
Never resend (including timeout or ambiguous failure), wait solely for messaging,
or use latest-active/empty-params fallback. Without an answer within the bound,
end as BLOCKED with the marker, decision and handle; leave the dependent action
untouched. Call Puck directly even in Spec mode, never exit Spec just to message
or guess AskUser answers. Acceptance/queued/working/silence is not approval; a
reply steers only that checkpoint within existing task authorization. Prompt
wording cannot override off-mode permission cancellation. Puck can continue the
latest session head. This is requested in-turn coordination, not unsolicited
active-turn steering; wording tests do not prove model obedience.
Agent reports are not automatic controller completion notifications. Crashes, cancellation, declined native AskUser or MCP failures can prevent a report; Puck must still poll status/result. Message acceptance is not proof of receipt: live acceptance requires the unique agent-origin marker to arrive in the intended Puck conversation and a correlated reply to reach Droid.
OAuth is account authorization, not a demonstrated thread/tool-scoped grant.
The external-agent profile may expose additional Amp tools now or in the future.
Before task submission, the controller dynamically disables every discovered
amp-puck___* tool except exactly amp-puck___puck, then re-lists and fails
closed if any other Amp tool remains allowed. This client-side denial reduces
model exposure, not credential rights or OS access; an agent with High
service-user access is trusted with that user's environment. No credentials,
personal recipient IDs or host configuration belong in public source/history.
See Factory MCP and
Amp MCP.
Outcomes are explicit
State | Meaning |
| Active; keep polling. |
| Matching terminal SDK result has subtype |
| Matching terminal SDK result has subtype |
| Terminal agent failure, setup/transport failure, unexpected exit, or invalid result. |
| Controller cancelled without receiving a terminal result, usually during setup or after forced process termination. |
| Configured wall-clock deadline expired; never inferred from silence. |
| Controller died before a durable terminal result; never replayed or inferred as success. |
All but the first row have terminal: true. An unknown outcome is terminal for
the controller's polling loop but not a known Droid outcome. Continuation
of that UUID is blocked, even using an older run handle. Inspect/reconcile Droid
history and the workspace locally. Start a fresh controller task with a new key
after reconciliation, or use the saved UUID manually outside the controller.
Do not edit state to manufacture success.
High-autonomy turns record and approve offered single-use permissions with
ProceedOnce, never persistent ProceedAlways rules. Other levels decline
pending permission requests with Cancel; high also declines if no single-use
option is offered. AskUser questions are recorded and declined, not guessed.
SDK 0.9.1 maps a completed turn to success even after declined AskUser and with
empty final text; cancelled maps to interruption. The controller preserves that
SDK outcome, not a judgment that the user's task finished. Inspect recorded
ask_user_declined events and task artifacts before accepting completion. Puck
can answer the blocked question in a new prompt after terminal status. There is
no interactive permission queue or arbitrary session import; saved outcomes are
not rewritten to manufacture a different result.
Same-host durability and cleanup
stateDirectory contains private state.json, per-run *.result.json, and an
owner record. Writes use file+directory fsync and atomic replacement. A single
process owns the state; startup refuses a live owner, another host/HOME/profile,
or malformed state. A stale dead-process owner is recovered under a startup
lock. PID reuse conservatively blocks startup rather than risking two owners.
A controller crash stops its workers through IPC disconnect; each worker owns a separate process group containing Droid. Cancellation first sends protocol interrupt, then force-kills that group after its grace period. The session slot is not released until worker cleanup. Terminal results are persisted before cleanup, so restart can retain a known outcome even if cleanup was interrupted. Processes deliberately detached by Droid tools may escape that group: do not authorize background-service work unless separately supervised.
Keep the same user, HOME, FACTORY_HOME_OVERRIDE (if used), Droid session storage,
and workspace. Controller records do not replace Factory's own saved history.
Back up both. Progress is intentionally bounded: last 20 events, 4000 characters
per event field, and 16000-character text/stderr tails. Final SDK results are
stored with submitted user-message copies omitted. There is no automatic deletion
of completed records; archive them locally with the controller stopped, preserving
keys if deduplication is needed.
Timestamps are UTC ISO strings; Puck may display them in Asia/Bangkok.
Original submitted prompts are not retained in controller state. Prompt content
participates in the SHA256 intent fingerprint and is passed to the worker only
over IPC for the accepted execution. A crash before submission becomes unknown,
not a replay. Assistant echoes, permission context, stderr and final results can
still contain sensitive text; Factory's own session history is separate and
may retain prompts. Fingerprints are not encryption.
Version 1 state migrates atomically to version 2 at startup. Existing IDs,
fingerprints and results stay intact; raw prompt fields are removed. For each
known Droid UUID, the latest createdAt selects the initial sessionHeads entry
(equal timestamps use durable insertion order). Version 2 heads are authoritative,
not reconstructed from parent links. Invalid timestamps or missing/wrong-family
heads fail startup without resetting state. Existing unknown families stay blocked.
Back up stopped private state before upgrading. Downgrade requires restoring the
matched v1 backup and source, not running old code against v2 or editing history.
Migration cannot erase prompts from old backups or Factory history.
Existing v1 result files remain unchanged; new results omit SDK user-message
copies of the original input while retaining assistant/tool output and outcomes.
Configuration and security
Copy config.example.json and replace every path with an absolute path.
JSON paths do not expand ~. Only real directories at/below approved roots
can be launched; traversal, sibling-prefix and symlink escapes are rejected.
Resume must resolve to exactly the recorded workspace. State must be private
(700 directory, 600 files).
Optional settings:
Setting | Default |
|
|
|
|
|
|
| Unset; otherwise |
| 4 across distinct sessions |
| 3600000 (one hour wall clock, including setup) |
| 5000, also bounds terminal cleanup |
| 60000, bounded to 5000–600000 |
| Unset; optional private |
| 8787, HTTP only; binds 127.0.0.1 only |
| Required for HTTP; private file, ≥32 URL-safe random characters |
| Optional approved HTTPS proxy URL, ending in |
The listener implements stateless Streamable HTTP MCP POST /mcp, with constant-
time Bearer authentication on every request, approved Host/Origin checks, a 1 MiB
body cap and request timeouts. No unauthenticated HTTP option exists. No OAuth
provider, SSE notification channel, or automatic push back to Puck is provided.
Status/results work independently of a connected client.
Malformed JSON/MCP envelopes return HTTP 400. Unexpected server/transport failures
return generic HTTP 500 Internal MCP server error; details stay in local stderr.
Lifecycle/tool rejections remain normal MCP isError responses, not HTTP 500.
Workspace approval is a launch restriction, not an OS sandbox. Droid runs
as the service user; enabled tools, inherited MCP servers, project hooks, and
commands can access what that user can. Inspect trusted project/Factory config.
For hard file/network isolation use a dedicated restricted account/container
with the approved directories mounted. Read-only/Spec mode is Droid policy,
not a filesystem mount guarantee. Low/medium/high authorize real actions; high
is powerful. The controller never uses --skip-permissions-unsafe.
Treat the bearer token as remote access to this account within the selected ceiling. Anyone with it can read stored results and start tasks. Store it only in local/private credential settings, use TLS for every remote hop, and restrict the HTTPS proxy/tunnel to this controller. Never expose an orb preview as the user's installed controller. Rotate the token by replacing its file and restarting the server; update Amp's stored credential separately.
Workspace approval, outages, and safe upgrades
Workspace rejection: a correct authorization failure, not a reason to remove the allowlist. Ask the operator to add only the explicitly approved project root. Existing directories, real paths, sibling/traversal protection, and resume-cwd checks still apply. Configuration is loaded at process startup; coordinate a controller-only restart with every caller to activate changes.
HTTP 530: a failure at the HTTPS/Cloudflare boundary; it does not establish that Droid failed or stopped. Check the actual public route and Cloudflare's tunnel connection state, DNS, and bounded logs. Loopback readiness/metrics can appear healthy while public routing is down. Synthetic host DNS can also make a local HTTPS probe misleading; cloud-client acceptance is required. Reserve connector-only recovery when public routing fails; preserve active Droid runs. A controller restart cannot repair an unavailable tunnel. Retry read-only status/list/result after a short delay; replay start/continue only with the same key and identical arguments when acceptance is uncertain. Never create a second intent merely because HTTP failed.
Connector DNS: some managed hosts return synthetic addresses that cannot carry tunnel traffic. Fix only the connector's resolution/routing, not shared machine DNS. A pinned-edge override is a host workaround, not a portability guarantee; retain edge TLS verification and revisit it if edge addresses change.
Upgrade/restart: reserve the execution slot, check that all runs are terminal, back up private state, and restart only the changed service. Keep the healthy tunnel, bearer credential, HTTPS URL, user/HOME, and state store unchanged. Recheck remote discovery, prior results, list, and idempotency before resuming. Never rewrite installation state to make a failed upgrade appear successful.
Persistence: process supervision and durable records do not imply machine- boot startup. Use the host's supported startup mechanism when available; a managed container without systemd/launchd needs platform-specific provisioning. Do not claim boot persistence from a successful process restart alone.
Protocol evidence and tests
Uses published @factory/droid-sdk/node ProcessTransport with
createSession({transport}) / resumeSession(id,{transport}). This bypasses
the SDK helper's API-key requirement, not CLI authentication. The supported CLI
invocation is exec --input-format stream-jsonrpc --output-format stream-jsonrpc.
Settings use initialize/update requests, not CLI model/autonomy flags.
Terminal matching is agent_turn_completed.turnId == add_user_message.messageId;
idle, assistant text, silence, unrelated turn IDs and process exit cannot succeed.
npm test drives the real MCP SDK client, controller, worker, Factory SDK and
mock JSON-RPC peer. It covers exact final output, settings, UUID retention,
idempotency, serial steering, permissions/questions, forced/setup cancellation,
agent/process failures, crash/restart durability, process-group cleanup,
timeouts, malformed/unrelated events, directory escapes, authentication and
HTTP checks, STDIO, ownership and corrupt-state rejection. See test/README.md
for failure cases written before implementation. npm run smoke additionally
exercises a real installed CLI through MCP and writes sanitized repeatable
evidence; authenticated success is required for production handoff.
Sources: Factory SDK documentation, Droid exec, published SDK, reference bridge. The reference bridge was cloned/read, not changed or incorporated: its OpenAI-transcript/tool isolation contract differs from this session controller.
License
MIT; see LICENSE.
This server cannot be deployed
Maintenance
Related MCP Connectors
Local-first task manager: create, edit, and complete tasks, projects, and checklists via MCP.
Hosted MCP server for task-first delegation to remote workstations and workers.
Private-by-default, local-first memory/context/task orchestrator for MCP apps and agents.
Personal assistant MCP server with search, execute, packages, jobs, secrets, and integrations.
Related MCP Servers
- AlicenseCqualityDmaintenanceEnables local task management through the Model Context Protocol with tools for creating, updating, and deleting tasks. It also provides read-only resources for viewing task summaries and full task lists.412 npmISC
- AlicenseNot gradedqualityDmaintenanceEnables running and managing automated tasks with retry loops and machine-checkable success criteria via MCP tools.1 npmMIT
- AlicenseNot gradedqualityCmaintenanceEnables MCP clients to run bounded robotics tasks, retrieve independently verified results, and reuse checked data or artifacts from the local WorldFuzz platform.MIT
- AlicenseAqualityCmaintenanceEnables MCP clients to submit tasks to a local file-backed queue and later read or cancel them by ID, with results produced by a separately run worker.3MIT