dashpilot-mcp
OfficialDashPilot MCP lets an AI agent manage DoorDash Drive deliveries end-to-end: quote, dispatch, schedule, track, batch, update/cancel, and inspect account/diagnostics.
Verify locally-held Drive credentials and routing environment with
check_drive_connection.Get delivery quotes (fee, tax, ETA) without dispatching.
Dispatch immediately via
accept_quoteordispatch_delivery; these move money and require explicit confirmation.Schedule single deliveries or batch events (up to 50, staggered) for later dispatch.
Fire due scheduled deliveries from your machine with a fresh local JWT via
dispatch_due_deliveries.Track live delivery status, Dasher info, ETA, and customer tracking URL.
View the operations board and scheduled queue with
list_deliveries.Update an active delivery's tip, dropoff instructions, or contact phone.
Cancel deliveries, subject to Drive cancellation rules.
Dispatch up to 25 deliveries at once with
batch_dispatch.Read account info and dispatch settings/feature flags.
Sync diagnostics, generate support bundles, and delete the install/cloud data.
Integrates with DoorDash Drive to quote delivery fees/tax/ETA, dispatch Dashers, track live delivery status, update tips/instructions/contact info, cancel deliveries, and batch or schedule deliveries for restaurants and stores.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@dashpilot-mcpschedule a 12-drop launch night for Friday and show the plan first"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
dashpilot-mcp
DoorDash Drive automation for AI agents: quote deliveries, dispatch Dashers, track live status, and schedule anything from a single catering run to a full launch night — all from your IDE or agent client, for restaurants and stores that deliver with DoorDash Drive.
Ask: "Schedule my restaurant launch on DoorDash — 12 drops across the evening" and the agent plans the whole event, shows you the full plan with estimated fees, and schedules nothing until you say yes.
Unofficial. Not affiliated with, endorsed by, or connected to DoorDash, Inc. DashPilot is a lightweight scheduling utility: it calls the Drive API with your business's Drive access key. Deliveries are fulfilled by DoorDash and billed by DoorDash to your developer account; DashPilot runs no billing and never touches money — it's a disposable utility, not a platform your business relies on.
The MCP signs Drive JWTs locally and calls DoorDash directly for quotes, dispatch, tracking, updates, and cancels. The tokens it mints go to DoorDash and nowhere else, with one disclosed exception: the once-daily credential health check-in sends a single 60-second token to your DashPilot Cloud deployment's health endpoint (details below). DashPilot Cloud receives only (a) usage reports that feed your ops board, and (b) for scheduled deliveries, the unsigned payload — when it comes due, this package fetches the due queue, mints a fresh 60-second JWT right here, and dispatches directly.
Install
Requires Python ≥ 3.11. With uv:
{
"mcpServers": {
"dashpilot": {
"command": "uvx",
"args": ["--from", "/absolute/path/to/dashpilot-mcp", "dashpilot-mcp"],
"env": {
"DASHPILOT_API_URL": "https://dashpilot6de8ccea-dashpilot.functions.fnc.fr-par.scw.cloud",
"DASHPILOT_API_KEY": "dp_your_issued_key"
}
}
}
}Then put your Drive credentials in a .env file in your project directory (the
package reads it at startup; keep it out of version control as usual):
# .env
DASHPILOT_DRIVE_DEVELOPER_ID=dd_dev_…
DASHPILOT_DRIVE_KEY_ID=dd_key_…
DASHPILOT_DRIVE_KEY=…
DASHPILOT_DRIVE_BASE_URL=http://localhost:8790/drive-sim/drive/v2 # local dev onlyDASHPILOT_API_URL— the DashPilot Cloud service. Defaults to the hosted instance (https://dashpilot6de8ccea-dashpilot.functions.fnc.fr-par.scw.cloud); point it athttp://localhost:8790to run against a local development container.DASHPILOT_API_KEY— your install key. Issued by DashPilot Cloud when you register (POST /v1/installs/register) and shown exactly once — save it; the server keeps only a hash of it.DASHPILOT_DRIVE_*— your Drive access key from the DoorDash Developer Portal (sandbox keys work out of the box). Used to sign JWTs on this machine. Lives in.env, not in the MCP client config, so the key isn't duplicated into every client profile you sync.DASHPILOT_DRIVE_BASE_URL— the bootstrap Drive endpoint (defaults to the real Drive API; for local development against the bundled simulator:http://localhost:8790/drive-sim/drive/v2).DASHPILOT_DRIVE_ROUTING—managed(default) ormanual. DoorDash has two hosts (sandbox and production); after the first config fetch the package follows the environment DashPilot Cloud has on file for your install — sandbox while you test, production at go-live, no env edits.check_drive_connectionreports which environment you're on. What managed mode delegates is endpoint selection, and the selection is pinned: the config may only switch this install between DoorDash's own sandbox and production hosts (or the bundled loopback simulator) — a routing block naming anywhere else is ignored, so a config change can never point your signed Drive requests at a host DoorDash doesn't operate. Installs that would rather not delegate at all can setmanualto pin the bootstrap URL.
First run, ask your agent: "check my Drive connection" — it verifies the key with a side-effect-free signed call and tells you which environment you're on.
Related MCP server: DoorDash MCP Server
The autonomy model: your money moves only with your yes
The agent is a planner, not a spender. Every tool is annotated in the MCP protocol
(readOnlyHint / destructiveHint) so your client can auto-approve the safe ones and
always prompt for the rest:
The agent can do alone | The agent needs your explicit confirmation for |
Quotes (fee, tax, ETA) — quote liberally | Dispatching anything ( |
Live tracking, ops board, account, settings, diagnostics sync ( | Tip changes and cancels ( |
Planning an event and showing you the full plan | Scheduling it ( |
Uploading diagnostics ( |
dispatch_due_deliveries is the executor of confirmed plans: everything it fires was
already confirmed at schedule time, so it needs no new confirmation. Run it on a cadence
and due work just fires.
What you can ask for
"Quote a delivery for order #1001 to 350 5th Ave" —
get_delivery_quote(fee, tax, ETA)"Dispatch it" —
accept_quote, ordispatch_deliveryto quote+accept in one step"Schedule the catering run for 6:30pm" —
schedule_delivery, thendispatch_due_deliveriesfires due work with a fresh local JWT"Schedule my launch night: 12 drops, one every 15 minutes from 6pm" —
schedule_batch(up to 50, staggered)"Dispatch these 12 office lunches now" —
batch_dispatch(up to 25 at once)"Where is order #1001?" —
track_delivery(status, Dasher, ETA, tracking URL — straight from DoorDash)"Show today's board" —
list_deliveries"Bump the tip on #1001 by $2" —
update_delivery"Cancel #1002" —
cancel_delivery
Money questions live in your DoorDash developer portal — DoorDash bills you directly; DashPilot has no billing surface.
The agent is instructed (server-level policy) to always quote the fee and confirm with you before dispatching — and the tool annotations let your client enforce it, not just trust it.
Tools
Tool | Purpose | Talks to |
| Verify your local Drive key (side-effect-free) | DoorDash Drive |
| Delivery fee, tax, ETA — nothing moves yet | DoorDash Drive |
| Accept a quote — dispatches a real Dasher | DoorDash Drive |
| Quote + accept in one step | DoorDash Drive |
| Deferred dispatch — stores the unsigned payload only | DashPilot Cloud |
| Schedule a whole event (up to 50, staggered window) | DashPilot Cloud |
| Poll the due queue, sign locally (60s JWT), dispatch | Drive + DashPilot Cloud |
| Live status, Dasher info, ETA, tracking URL | DoorDash Drive |
| Tip / instructions / contact phone | DoorDash Drive |
| Cancel (Drive cancellation rules apply) | DoorDash Drive |
| Up to 25 deliveries in one call | DoorDash Drive |
| Operations board + scheduled queue | DashPilot Cloud |
| Business, dispatch count, member since | DashPilot Cloud |
| Feature flags and dispatch settings | DashPilot Cloud |
| Diagnostics bundle for a support ticket | DashPilot Cloud |
| Routine connectivity sync — part of the session handshake | DashPilot Cloud |
| Close this install — deletes its DashPilot Cloud data (irreversible) | DashPilot Cloud |
Security & privacy notes
stdio transport only; the package opens no network listener.
Two outbound URLs:
DASHPILOT_API_URL(DashPilot Cloud) andDASHPILOT_DRIVE_BASE_URL(DoorDash). Both must behttps://outside loopback; redirects are never followed.Your Drive access key is read from your project
.envfile (or the environment) and used to sign JWTs on this machine.Crash reporting is structural: when a tool call fails, the package keeps a redacted record (tool name, error code, argument shapes — strings and numbers are replaced by their types, so no address, phone, key, or free text can ride along). Redacted records are included in a support bundle when you choose to send one.
Scheduled deliveries are stored as unsigned payloads. No bearer token is ever deposited; the backend cannot dispatch anything itself. Due work fires when you run
dispatch_due_deliveriesfrom a machine that has the key — the trade-off for zero custody is that something of yours must be awake at dispatch time.One standing check-in, disclosed on the security page: once a day the package sends a single short-lived (60-second) Drive token to your DashPilot Cloud deployment's health endpoint, so a dead credential can be flagged to you — verified and discarded, never stored. Your state file records which daily check-in already fired. Beyond that: no telemetry, no analytics, no install scripts.
Development
uv sync
uv run pytest
DASHPILOT_API_URL=http://localhost:8790 uv run dashpilot-mcpMCP Registry
mcp-name: io.github.dashpilot-labs/dashpilot-mcp
Available Tools
17 toolsaccept_quoteADestructive
Accept a quote from get_delivery_quote — this dispatches a real Dasher, billed by DoorDash to your developer account. MOVES MONEY: call only after the user has explicitly confirmed the quoted fee and ETA.
external_delivery_id: the ID you quoted with. tip: set/override the tip in cents.
| Name | Required | Description | Default |
|---|---|---|---|
| tip | No | ||
| external_delivery_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only flag destructive/non-idempotent/open-world; the description goes much further by disclosing the financial consequence ('MOVES MONEY') and who is billed ('DoorDash to your developer account'), plus the human-confirmation requirement. This is precisely the behavioral context annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short blocks with the money-spending warning front-loaded, then one line per parameter. No filler, and the most consequential fact is first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter spend action with no output schema, the description supplies purpose, prerequisite, billing consequence, and both parameter meanings. It stops short of saying what happens to a previously scheduled quote or what the response returns, but nothing critical to correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% (bare 'integer'/'string' titles), so the description must compensate and largely does: it defines external_delivery_id as 'the ID you quoted with' and specifies that tip is expressed in cents and can override the quoted value. Minor gap: it does not state the tip's default behavior when omitted, though the schema default of null hints at it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Accept a quote from get_delivery_quote') and immediately differentiates it from siblings by contrasting with quoting: it 'dispatches a real Dasher' rather than returning a price. An agent can distinguish it from get_delivery_quote and dispatch_delivery without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit precondition for use: 'call only after the user has explicitly confirmed the quoted fee and ETA', plus the requirement that the quote originate from get_delivery_quote. This is exactly the when-to-use guidance an agent needs for an irreversible spend action.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_dispatchADestructive
Dispatch up to 25 deliveries at once — catering runs, multi-order drops. Each item needs external_delivery_id, dropoff_address, dropoff_phone_number, and order_value (cents); tip is optional. pickup_address applies to items that don't set their own. Returns a batch_id and one result per delivery.
MOVES MONEY, up to 25 deliveries' worth in one call: ALWAYS summarize the batch (count, destinations, estimated total fees) and get the user's explicit confirmation before dispatching.
| Name | Required | Description | Default |
|---|---|---|---|
| deliveries | Yes | ||
| pickup_address | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructive=true and openWorldHint=true, but the description adds critical context they don't convey: the call MOVES MONEY, it handles up to 25 deliveries at once, and it mandates summarizing count/destinations/fees and obtaining explicit user confirmation first. That is exactly the behavioral detail an agent needs to avoid an irreversible mistake.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then required item fields, then the safety warning. Every sentence adds information an agent needs; the capitalized money-movement warning is intentional emphasis, not padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still explains the return shape (a batch_id plus one result per delivery), and for a destructive, money-moving mutation it supplies the required fields, the confirmation protocol, and the batch limit. Nothing an agent needs to invoke it safely is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the deliveries array has no item schema, so the description carries the full burden — and it does: it enumerates the required item fields (external_delivery_id, dropoff_address, dropoff_phone_number, order_value in cents), marks tip as optional, and explains that top-level pickup_address applies only to items that don't set their own.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (dispatch) plus scope (up to 25 deliveries at once) and concrete use cases (catering runs, multi-order drops), which cleanly separates it from the single-delivery sibling dispatch_delivery.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when to use it (batches up to 25, catering/multi-order drops) and adds a mandatory confirmation workflow before dispatch. It does not explicitly name or exclude alternatives like schedule_batch, so it falls short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cancel_deliveryADestructive
Cancel a delivery. Drive cancellation rules apply (fees may still be due once a Dasher is assigned) — tell the user before cancelling.
| Name | Required | Description | Default |
|---|---|---|---|
| external_delivery_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=false, so the safety profile is covered. The description adds genuinely new domain behavior: Drive cancellation rules apply and fees may still be due once a Dasher is assigned, plus the instruction to warn the user first — meaningful context the annotations do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the action and followed by the caveat and the required user-facing step. No filler, though the parenthetical fee clause is slightly compressed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter destructive tool whose annotations cover safety and which has no output schema, the description supplies the key missing pieces: fee consequences and the user-warning obligation. It could still say what happens to an already-dispatched delivery or on repeated calls, but the essential context is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single required parameter external_delivery_id, so the description carries the burden of explaining it — and it says nothing about format, source, or where the ID comes from. The parameter name is somewhat self-explanatory, but the compensation the low coverage requires is absent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Cancel a delivery'), which is unambiguous against siblings like update_delivery and delete_install. It does not, however, explicitly distinguish itself from update_delivery, which could plausibly be used to cancel.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied (cancel a delivery the user no longer wants) and the description adds one precondition — tell the user before cancelling. It never states when to prefer this over update_delivery or what to do if cancellation is no longer possible.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_drive_connectionARead-only
Verify your locally-held Drive access key works, without dispatching anything: signs a JWT and asks Drive about a nonexistent delivery — a 404 means authenticated, a 401 means the key is wrong. Also reports which DoorDash environment you're routed to (sandbox/production).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint and openWorldHint already declared, the description goes well beyond annotations: it explains the JWT signing mechanism, that it probes a nonexistent delivery, and decodes the outcome (404 = authenticated, 401 = wrong key). It also discloses that it reports the routed environment (sandbox/production), giving the agent real behavioral context a mere read-only hint cannot.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the main purpose (verify the key) before the mechanics and the secondary environment report. Every clause carries information; nothing is redundant padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description explains what the return signals mean (404 vs 401) and what extra data comes back (environment). For a zero-parameter diagnostic tool, an agent has everything needed to call and interpret it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
This tool takes zero parameters, so there is nothing for the description to disambiguate and the baseline for parameterless tools is 4. No parameter-related gap exists to compensate for.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (verify your Drive access key) and immediately distinguishes itself from the dispatch-heavy siblings by clarifying it dispatches nothing. An agent can tell this is a credential/connectivity check rather than a delivery action without inspecting anything else.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when to use it (to validate a locally-held key) and explicitly notes it does not dispatch, which separates it from dispatch_delivery and friends. It stops short of naming an alternative tool or stating exclusions such as 'do not use this to check account state — use get_account', so it is clear but not fully routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_installADestructive
Close this DashPilot install and delete its data from DashPilot Cloud: the account, usage history, scheduled queue, and uploaded support bundles. Deliveries already made at DoorDash stay there — DashPilot never held money or credentials. IRREVERSIBLE: confirm with the user before calling. The API key stops working; reusing it later starts a fresh install.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark it destructive and non-idempotent, but the description adds substantial context beyond them: the exact data destroyed, the fact that completed DoorDash deliveries are unaffected, and the post-condition that the API key stops working and a reuse starts a fresh install.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, each earning its place: scope of deletion, boundary of non-deletion, irreversibility warning, and post-call credential behavior. Front-loaded with the primary effect before the caveats.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters, no output schema, and safety hints fully covered by annotations, the description supplies everything else an agent needs — side effects, client-confirmation requirement, and consequences for credentials — with nothing material missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4; there is no parameter syntax left for the description to clarify. It correctly spends no words restating an empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (delete/close) plus the resource and enumerates exactly what is removed: account, usage history, scheduled queue, and uploaded support bundles. It is unmistakably distinct from read-oriented siblings like get_account or list_deliveries.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear when-to-use context by framing it as a terminal, irreversible action that requires user confirmation before calling. It does not name an alternative recovery/migration path, but for a tear-down tool the exclusions are largely self-evident.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
dispatch_deliveryADestructive
Dispatch a delivery immediately (quote + accept in one step): a Dasher is assigned and the delivery is billed by DoorDash to your Drive account. MOVES MONEY: quote first with get_delivery_quote, show the user the fee and ETA, and dispatch only after their explicit confirmation. The request goes straight to DoorDash, signed with your local key. Parameters are identical to get_delivery_quote.
| Name | Required | Description | Default |
|---|---|---|---|
| tip | No | ||
| order_value | Yes | ||
| pickup_address | Yes | ||
| dropoff_address | Yes | ||
| contactless_dropoff | No | ||
| pickup_instructions | No | ||
| dropoff_instructions | No | ||
| dropoff_phone_number | Yes | ||
| external_delivery_id | Yes | ||
| pickup_business_name | No | ||
| pickup_reference_tag | No | ||
| dropoff_contact_given_name | No | ||
| dropoff_contact_family_name | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and openWorldHint=true, but the description adds critical context the annotations cannot convey: this MOVES MONEY, bills the Drive account, and is signed with the local key and sent to DoorDash. It stops short of stating reversibility (whether cancel_delivery can unwind it) or any rate/idempotency caveats, which would round it out.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, front-loaded with the action and the money warning, then the required workflow. Every clause carries information (one-step semantics, billing, confirmation gate, signing, param parity).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, money-moving tool with 13 undocumented parameters and no output schema, the description covers the most dangerous aspects (irreversibility implied by money movement, confirmation requirement, quote-first prerequisite). The remaining gap is parameter meaning, which is only deflected to a sibling tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 13 parameters, including 5 required ones such as order_value and external_delivery_id. The description compensates only by pointing to get_delivery_quote ('Parameters are identical'), forcing a cross-tool lookup rather than explaining fields like order_value units or the tip default.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (dispatch a delivery) and immediately clarifies the semantics: quote + accept in one step with a Dasher assigned and billed to the Drive account. This distinguishes it clearly from siblings like schedule_delivery, accept_quote, and dispatch_due_deliveries.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly prescribes the workflow: quote first with get_delivery_quote, show the user the fee and ETA, and dispatch only after explicit confirmation. It names the prerequisite tool and the gating condition, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
dispatch_due_deliveriesADestructive
Fire every scheduled delivery that is due right now. Asks DashPilot Cloud for the due queue (unsigned payloads only), mints a fresh short-lived JWT on this machine for each one, dispatches directly to DoorDash Drive, and reports the results back for the ops board. Run it on a cadence (or whenever you want due work flushed) — the cloud backend cannot dispatch anything itself.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructive=true, idempotentHint=false and openWorldHint=true, but the description adds genuinely non-obvious behavior: it pulls only unsigned payloads from the cloud queue, mints a short-lived JWT locally per payload, and dispatches directly to DoorDash Drive. This external-side-effect and credential-minting detail is exactly the context annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the action, and each sentence carries distinct information (what fires, how it works end-to-end, when to run it). Slightly verbose in the mechanism chain but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters and no output schema, the description carries the burden well by explaining mechanism, cadence expectation, and that results are reported back. It stops short of 5 by not addressing failure/partial-dispatch behavior or double-firing risk on repeated runs, which matters for a destructive, non-idempotent tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The description correctly implies the scope is fully determined server-side (the due queue), and there are no parameters whose semantics need explaining.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence gives a specific verb and resource with scope: 'Fire every scheduled delivery that is due right now.' Combined with the explanation of the due-queue source, an agent can distinguish this batch/cadence trigger from single-delivery siblings like dispatch_delivery or dispatch_delivery-style tools without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states clearly when to invoke it: 'Run it on a cadence (or whenever you want due work flushed)' and explains why the caller must do so ('the cloud backend cannot dispatch anything itself'). It does not name explicit alternatives or state when NOT to use it, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_support_bundleA
Having a dispatch problem? Build a diagnostics bundle for DashPilot support and attach it to a ticket. Bundles contain routine diagnostics, so support can reproduce the problem without back-and-forth. The upload receipt itemizes what was sent.
issue: short description of the problem (e.g. "scheduled delivery never fired").
| Name | Required | Description | Default |
|---|---|---|---|
| issue | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare a non-read-only, open-world, non-idempotent call, and the description usefully adds that the bundle contains only routine diagnostics and that a receipt itemizes what was sent — reassuring context for an upload that leaves the system. It does not state permission requirements or how long the bundle takes to build.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the problem hook, then the action and its justification, then the parameter note. Three sentences, each earning its place, with no filler beyond a mild rhetorical opening.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still covers the return side by describing the upload receipt, and annotations cover the safety profile. For a one-parameter tool this is nearly complete, lacking only prerequisites such as authentication or ticket-context requirements.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for the single 'issue' parameter, but the description compensates with an explicit format hint and a concrete example ('scheduled delivery never fired'), which is more than the bare 'string' schema conveys. It could still clarify whether free-form detail or a category is expected.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (build/attach) and resource (diagnostics bundle) plus the destination (a support ticket), so the agent knows exactly what the call produces. It does not differentiate itself from the sibling sync_diagnostics, which is the nearest ambiguity in this tool set.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The opening 'Having a dispatch problem?' gives a clear trigger condition for using this tool rather than the other dispatch-related siblings. It stops short of naming alternatives or stating when NOT to generate a bundle (e.g. for non-dispatch issues).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_accountBRead-only
Your DashPilot account: business, dispatch count, member since.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered by structured data. The description adds the useful detail of which account fields are returned, which is real behavioral context in the absence of an output schema. However, it says nothing about permissions required, freshness of dispatch count, or response format.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence that front-loads the resource and lists the returned fields without waste. It is a fragment with no verb, which slightly weakens readability, but nothing is extraneous.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-argument read tool with annotations covering the safety profile, the definition is essentially complete. Naming the returned fields compensates for the missing output schema, though a bit more on auth or freshness would make it airtight.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. The description correctly adds no parameter guidance because none is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource (the DashPilot account) and enumerates what it returns: business, dispatch count, member since. That is enough to distinguish it from all sibling tools, which deal with deliveries, dispatch, and diagnostics. It lacks an explicit verb (
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to call this versus alternatives, nor any prerequisites or context. Usage is only inferable from the fact that it takes no parameters and returns account data. Similar sibling tools like get_dispatch_settings show that a routing hint would have helped.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_delivery_quoteARead-only
Quote a DoorDash Drive delivery before committing: delivery fee, tax, ETA. Nothing is dispatched and no money moves until the quote is accepted. The request goes straight to DoorDash, signed with your local key. Read-only: quote liberally, no user confirmation needed — confirmation is required only before accept_quote / dispatch_delivery.
external_delivery_id: your unique ID for this delivery (e.g. a POS order number). pickup_address: your store's full address, comma-separated. dropoff_address: customer's full address, comma-separated. dropoff_phone_number: customer phone, E.164 (e.g. +12125550100). order_value: order subtotal in cents, excluding tax/tip ($19.99 = 1999). tip: Dasher tip in cents.
| Name | Required | Description | Default |
|---|---|---|---|
| tip | No | ||
| order_value | Yes | ||
| pickup_address | Yes | ||
| dropoff_address | Yes | ||
| contactless_dropoff | No | ||
| pickup_instructions | No | ||
| dropoff_instructions | No | ||
| dropoff_phone_number | Yes | ||
| external_delivery_id | Yes | ||
| pickup_business_name | No | ||
| pickup_reference_tag | No | ||
| dropoff_contact_given_name | No | ||
| dropoff_contact_family_name | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint/openWorldHint; the description goes further by stating nothing is dispatched, no money moves until acceptance, the call goes straight to DoorDash signed with the caller's local key, and no user confirmation is needed. That is substantive behavioral context beyond the structured hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, behavior, and usage; the parameter glossary is compact and each line earns its place. Marginally verbose in places but no filler and the ordering is sensible.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only quote endpoint with no output schema, the description covers the safety profile, the external dependency, and the primary inputs, and even names the return contents (fee, tax, ETA). The remaining gap is the set of undocumented optional parameters with defaults.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 13 parameters, so the description must carry the semantics. It usefully documents 6 of them with real meaning (cents, E.164, comma-separated addresses, POS order number), but leaves contactless_dropoff, pickup_instructions, dropoff_instructions, and the contact/reference fields undefined, so it only partially compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (quote), resource (DoorDash Drive delivery), and the concrete outputs (delivery fee, tax, ETA). It is immediately distinguishable from siblings accept_quote/dispatch_delivery because it frames itself as the non-committing pre-step.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to 'quote liberally' with no user confirmation, and names the exact point where confirmation IS required (before accept_quote / dispatch_delivery). This is an explicit when/when-not plus the alternative tools, which is the top of the scale.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_dispatch_settingsBRead-only
Current dispatch settings and feature flags from DashPilot Cloud (feature flags, default tip, Drive environment).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety and external-service profile is covered by structured data. The description adds useful content detail (feature flags, default tip, Drive environment) but says nothing about freshness, caching, or failure behavior when the cloud is unreachable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence with the key noun phrase front-loaded and a compact parenthetical listing what the settings comprise. No filler, though the parenthetical is slightly list-like and could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters, annotations covering the safety profile, and no output schema, the description's job is largely to say what comes back, which the parenthetical does at a high level. It stops short of documenting the return shape or auth requirements, but nothing critical for invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is no parameter semantics burden and the schema is trivially complete. Baseline 4 applies; the description correctly makes no attempt to fabricate arguments.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource ('Current dispatch settings and feature flags') and the source system (DashPilot Cloud), so an agent knows this is a read of configuration state rather than a delivery action. It is clearly distinguishable from action-oriented siblings like dispatch_delivery or accept_quote, though it does not explicitly contrast itself with get_account or check_drive_connection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to call this tool versus alternatives, nor any prerequisite or context such as 'call before dispatch to verify feature flags'. The parenthetical enumerates contents but offers no routing guidance, so usage must be inferred from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_deliveriesBRead-only
Your operations board on DashPilot Cloud: reported deliveries, fees, and the scheduled queue. (Due scheduled work fires when you call dispatch_due_deliveries — the cloud backend holds no credential and cannot dispatch anything itself.)
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint, openWorldHint), yet the description adds genuine context beyond them: it discloses that due scheduled work only fires via dispatch_due_deliveries and that the cloud backend holds no credential. That prevents a real misconception about side effects. It still says nothing about pagination or result size.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with what the board returns, and the parenthetical caveat is short and earns its place by clarifying that this call triggers nothing. The metaphor adds flavor rather than information, but the text is not padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, single-optional-param, no-output-schema tool, the description covers the conceptual return contents and one important behavioral caveat. Gaps remain around the unexplained limit/pagination semantics, but the absence of an output schema is partly mitigated by describing what is shown.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one parameter, 'limit', with 0% schema description coverage and a default of 25, and the description never mentions it or any result-count behavior. With coverage this low the description is expected to compensate, and it does not, so an agent gets no bounds information beyond the schema default.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the specific resource set it returns: reported deliveries, fees, and the scheduled queue, which is enough for an agent to tell it apart from dispatch/schedule siblings. It never uses the verb 'list', relying instead on the metaphor 'operations board', which costs a little directness but the enumerated contents compensate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implicitly defines usage by contrast with dispatch_due_deliveries for firing due work, which is a useful when-not signal. However it gives no guidance against the many other read-ish siblings (track_delivery, get_delivery_quote, get_dispatch_settings), so the agent must infer which read tool to pick.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
schedule_batchADestructive
Schedule a whole event's deliveries in one call — a restaurant launch, a catering run, office-lunch drops. Each item needs external_delivery_id, dropoff_address, dropoff_phone_number, and order_value (cents); tip and dropoff_instructions are optional. pickup_address applies to items that don't set their own.
dispatch_at: ISO-8601 UTC for the FIRST dispatch; stagger_minutes spreads the rest (e.g. 12 deliveries, stagger 15 = one drop every 15 minutes, so the kitchen is never slammed). Everything lands in the due queue as unsigned payloads and fires via dispatch_due_deliveries.
Moves no money NOW, but commits N future dispatches the user's poller will execute with their key: ALWAYS summarize the full plan — count, addresses, window, estimated total fees — and get the user's explicit confirmation before scheduling.
| Name | Required | Description | Default |
|---|---|---|---|
| deliveries | Yes | ||
| dispatch_at | Yes | ||
| pickup_address | No | ||
| stagger_minutes | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations: explains that no money moves now but N future dispatches are committed and executed by the user's poller with their key, that items land in the due queue as unsigned payloads, and that dispatch_due_deliveries fires them. This is exactly the consequence-level detail a destructive, non-idempotent batch tool needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose, then parameters, then the behavioral warning — a sensible order. Three dense paragraphs with no filler sentences, though the delivery-field enumeration is lengthy and could be trimmed marginally.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers purpose, parameters, and the key side effect for a tool with no output schema. Minor gaps remain: no guidance on batch size limits or partial-failure behavior, and the confirmation requirement is stated but without a defined summary format beyond count/addresses/window/fees.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does: it names required per-item fields (external_delivery_id, dropoff_address, dropoff_phone_number, order_value in cents), marks tip and dropoff_instructions optional, explains pickup_address as a fallback for items without their own, and defines dispatch_at/stagger_minutes with a worked example.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — scheduling an entire event's deliveries in one call — with concrete instances (restaurant launch, catering run, office-lunch drops). An agent can immediately distinguish it from the singular schedule_delivery despite no explicit sibling naming.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear context for when to use it (multi-item event runs) plus a mandatory workflow: summarize the full plan and get explicit user confirmation before scheduling. It does not explicitly contrast with schedule_delivery or batch_dispatch, so the boundary between the batch siblings is left implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
schedule_deliveryADestructive
Schedule a delivery for later. DashPilot Cloud stores ONLY the unsigned payload — no signing secret, no bearer token, nothing it could spend. When dispatch_at arrives, call dispatch_due_deliveries (from this machine) to fire due work: a fresh 60-second JWT is minted locally at that moment and the delivery goes straight to DoorDash.
Moves no money NOW, but commits a future dispatch your poller will execute with your key: confirm the plan (address, time, estimated fee) with the user before scheduling.
dispatch_at: ISO-8601 UTC time to dispatch, e.g. 2026-08-27T18:30:00Z. Other parameters are identical to get_delivery_quote.
| Name | Required | Description | Default |
|---|---|---|---|
| tip | No | ||
| dispatch_at | Yes | ||
| order_value | Yes | ||
| pickup_address | Yes | ||
| dropoff_address | Yes | ||
| contactless_dropoff | No | ||
| pickup_instructions | No | ||
| dropoff_instructions | No | ||
| dropoff_phone_number | Yes | ||
| external_delivery_id | Yes | ||
| pickup_business_name | No | ||
| pickup_reference_tag | No | ||
| dropoff_contact_given_name | No | ||
| dropoff_contact_family_name | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations (destructiveHint=true, idempotentHint=false) by disclosing what is and is not stored ('no signing secret, no bearer token'), that a fresh 60-second JWT is minted locally at dispatch time, that no money moves now but a future dispatch is committed, and that the caller's key is used later. This is exactly the kind of non-obvious, safety-relevant behavior the description should carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first sentence front-loads the core action, followed by the security model, the dispatch mechanic, and the confirmation requirement. Slightly verbose in the middle paragraph but every sentence carries operational weight; nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a deferred-write tool with annotations, no output schema, and 14 mostly undocumented parameters, the description covers timing, custody of credentials, and the human-confirmation step. The main omission is what scheduling returns (a schedule id? a confirmation?) and any constraint on how far in advance dispatch_at may be set.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 14 parameters, so the description must compensate. It documents dispatch_at's ISO-8601 UTC format with a concrete example and mitigates the rest by pointing to get_delivery_quote for identical parameters — a genuinely useful cross-reference, but six required fields (addresses, phone, order_value, external_delivery_id) remain undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Schedule a delivery for later') and immediately scopes it against the immediate-dispatch siblings by explaining that dispatch happens only when dispatch_at arrives via dispatch_due_deliveries. An agent can distinguish this from dispatch_delivery or accept_quote without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names the follow-up tool (dispatch_due_deliveries) and the condition that triggers it, and adds an explicit operator instruction to confirm address, time and estimated fee with the user before scheduling. It does not explicitly contrast when to choose this over immediate dispatch_delivery, so it falls short of full when/when-not coverage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sync_diagnosticsB
Routine connectivity sync with DashPilot Cloud: checks for updated dispatch settings (feature flags, default tip, Drive environment).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and idempotentHint=false, i.e. this call changes state and is not safe to repeat blindly, yet the description frames the operation as a passive 'checks for updated settings' with no indication that anything is written or applied, and no note about repeat-call cost. The framing is not an outright contradiction (a 'sync' is inherently state-changing) but it omits the most decision-relevant behavior the annotations imply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that identifies system, action and scope with no filler. Slightly awkward colon-list structure but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-argument, state-changing tool with no output schema, the description covers the scope of the sync but omits what the caller receives, whether local state is overwritten, and whether the call is repeat-safe given idempotentHint=false. Adequate but with clear gaps for an operation whose side effects are the main question.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline for a no-parameter tool is 4. The description's enumeration of the synced content adds useful context even though no arguments exist.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (sync/check) against a named system (DashPilot Cloud) and enumerates what is covered: feature flags, default tip, Drive environment. It does not, however, distinguish itself from the sibling get_dispatch_settings, so an agent could reasonably confuse the two.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The word 'Routine' weakly implies periodic use, but there is no explicit when-to-use, no when-not-to-use, and no mention of get_dispatch_settings as the read counterpart. An agent gets no routing guidance for choosing between sync_diagnostics and the settings/dispatch siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
track_deliveryBRead-only
Live status straight from DoorDash: stage (created → picked_up → delivered), Dasher name/location once assigned, ETA, and the customer tracking URL.
| Name | Required | Description | Default |
|---|---|---|---|
| external_delivery_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety and external-access profile is covered. The description adds valuable behavioral context by disclosing the returned live status fields and the stage progression (created → picked_up → delivered), which is especially useful because no output schema is present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that front-loads the core value ('Live status straight from DoorDash') and then lists the returned fields. Every element earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter read tool with no output schema, the description adequately covers the return content. However, it omits any explanation of the required external_delivery_id and gives no usage guidance relative to siblings, leaving notable gaps for an agent invoking the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description would need to explain the required external_delivery_id parameter. It does not mention the parameter at all, its expected format, or where to obtain it. The parameter name is somewhat self-explanatory, but no additional semantics are provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool retrieves live delivery status from DoorDash, listing the specific returned fields: stage, Dasher name/location, ETA, and tracking URL. It distinguishes itself from sibling tools by being a tracking/status read rather than a quote, dispatch, update, or cancellation operation, though it does not explicitly name a sibling alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Live status' implies checking current delivery progress, but there is no explicit guidance on when to use this tool versus list_deliveries, get_delivery_quote, or update_delivery. No conditions, prerequisites, or exclusions are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_deliveryADestructive
Update an active delivery's tip, dropoff instructions, or contact phone. Tip changes MOVE MONEY — confirm the new amount with the user first.
| Name | Required | Description | Default |
|---|---|---|---|
| tip | No | ||
| dropoff_instructions | No | ||
| dropoff_phone_number | No | ||
| external_delivery_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnly=false, and non-idempotency, so the safety profile is carried. The description adds a genuinely non-structured behavioral fact: that tip edits move money and require user confirmation. It still omits partial-update semantics (what happens to fields left null) and any error behavior on non-active deliveries.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action and followed by the highest-risk caveat. Every clause earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation tool with 0% schema coverage and no output schema, the description covers the money-risk dimension well but leaves gaps: partial-update behavior for omitted fields, whether the delivery must be in a specific state, and what the caller gets back. Adequate but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does name three of the four parameters with recognizable meaning. It does not explain the required external_delivery_id, nor any format constraints (tip units/currency, phone number format), leaving part of the parameter burden unmet.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Update) and resource (an active delivery) and enumerates the three mutable fields: tip, dropoff instructions, contact phone. This clearly separates it from read-oriented siblings like track_delivery and terminal operations like cancel_delivery, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It supplies one strong precondition — confirm tip amounts with the user before mutating money — which is real usage guidance. However, it never says when to prefer this over alternatives (e.g., re-dispatch, cancel, or editing via another tool) or when the tool should not be used, apart from the implicit 'active delivery' scope.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
17 tool updates
v0.1.0- First observed
accept_quote - First observed
batch_dispatch - First observed
cancel_delivery - First observed
check_drive_connection - First observed
delete_install - First observed
dispatch_delivery - First observed
dispatch_due_deliveries - First observed
generate_support_bundle - First observed
get_account - First observed
get_delivery_quote - First observed
get_dispatch_settings - First observed
list_deliveries - First observed
schedule_batch - First observed
schedule_delivery - First observed
sync_diagnostics - First observed
track_delivery - First observed
update_delivery
TDQS
Scored across 17 tools
Most tools target a distinct resource+action: quote vs accept vs dispatch vs schedule are well separated, and the single-vs-batch pairs (dispatch_delivery/batch_dispatch, schedule_delivery/schedule_batch) are clearly differentiated by descriptions. The one real overlap is get_dispatch_settings and sync_diagnostics, which both return feature flags, default tip, and Drive environment, so an agent could reasonably pick either.
The set follows a clean verb_noun convention (get_account, get_delivery_quote, schedule_delivery, track_delivery, update_delivery, cancel_delivery, delete_install), which is easy to scan. The main deviation is batch_dispatch, which inverts to noun_verb and breaks the otherwise consistent pattern alongside dispatch_delivery.
17 tools is slightly heavy but justified: the surface spans connection checks, quoting, single/batch dispatch, scheduling/polling, tracking, listing, updating, canceling, diagnostics, and account teardown. Each tool maps to a real operation with little dead weight, though a couple could arguably be consolidated.
The delivery lifecycle is well covered end to end: quote, accept, dispatch, schedule/batch, fire due work, track, list, update, cancel, plus settings, support bundle, and install deletion. Minor gaps exist (e.g. fetching a single delivery's details beyond the ops board, or refund/fee adjustment), but nothing that blocks core workflows.
Maintenance
Related MCP Connectors
Agentic-AI route optimization for delivery fleets, capacity- and time-window-safe.
Multi-carrier shipping for AI agents: compare rates, buy labels, track packages, validate addresses
- mcpOAuthcom.crisphive
Field operations on a deterministic solver — run jobs, crews & fleet from Claude or ChatGPT.
Plan last-mile deliveries from your AI assistant. Turns orders from spreadsheets, URLs, or pasted text into optimized delivery routes. It geocodes addresses, flags locations that need confirmation, and assigns stops across your fleet while respecting vehicle capacity, delivery time windows, and stop limits. Manage vehicles and depots, review route maps, and plan anything from a small delivery run to thousands of stops. Connect with your Vepathos account using OAuth.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables interaction with the DoorDash Drive API to create, manage, and track delivery requests including quotes, delivery creation, status checks, and cancellations.612 npm1MIT
- AlicenseBqualityDmaintenanceEnables AI agents to search restaurants, browse menus, and manage DoorDash carts through structured JSON data. It leverages a background browser to handle authentication and direct GraphQL API calls for efficient interaction.712 npm2MIT
- AlicenseBqualityDmaintenanceEnables AI agents to search restaurants, browse menus, manage carts, and place orders on DoorDash programmatically. It utilizes a headless browser to interact with DoorDash's GraphQL API and bypass anti-bot protections for the full delivery lifecycle.2212 npm4MIT
- FlicenseNot gradedqualityDmaintenanceEnables AI agents to search restaurants, place delivery orders, and track real-time delivery status using the DoorDash Drive API. It includes a built-in mock data mode that allows for testing and demonstrating delivery lifecycles without requiring live API credentials.-