Skip to main content
Glama

Server Details

Manage your Cycle infrastructure from any AI assistant. Deploy and reconfigure containers, VMs, and full stacks; create environments; provision servers and manage clusters; handle DNS records, scoped variables, image sources, and volumes. Troubleshoot with logs, metrics, telemetry, and instance console commands. Authenticates via OAuth with read, write, and exec scopes. Tools outside your granted scopes are hidden, and every change is previewed before you confirm it.

Ownership verified
Status
Healthy
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL

TDQS

A4/5.0

Scored across 44 tools

Disambiguation4/5

Most tools target a distinct resource+action (delete_container vs delete_environment vs delete_stack, cycle_control_container vs reconfigure_container), and the long descriptions explicitly contrast overlapping ones. The main overlap is observational: get_telemetry, query_metrics, capture_stream and get_logs all serve diagnostics, but each description delineates snapshot vs history vs live vs search well enough that misselection is unlikely.

Naming Consistency4/5

Almost all names are snake_case verb_noun (list_servers, deploy_application, reconfigure_container, migrate_instances), which is consistent and readable. Deviations are minor: the control family is prefixed resource-style (cycle_control_container/environment/virtual_machine) instead of a plain verb, and decommission_server sits alongside delete_* for what is effectively deletion.

Tool Count2/5

At 44 tools the surface is very large and heavy for an agent to navigate, spanning containers, VMs, servers, clusters, DNS, images, stacks, volumes, variables, telemetry, jobs and migrations. The breadth is arguably real, but the count exceeds the well-scoped range and several tools could be consolidated (e.g. the control_* triple).

Completeness4/5

Coverage is strong: full lifecycle for containers, VMs, environments, stacks, DNS records, image sources, volumes and clusters, with explicit preview/confirm workflows and even a get_more_tools escape hatch. Minor gaps exist and are self-acknowledged (no zone deletion, no scoped-variable delete, no VM volume tool), which agents can mostly work around.

Available Tools

44 tools
capture_streamCapture a Live StreamA
Read-onlyIdempotent
Inspect

Watch a live Cycle stream for a few seconds and return what arrived, plus the temporary credentials so the user can keep watching with their own websocket client. Targets:

  • instance_console (environment + container [+ instance]): the container's live console output — crash messages and startup errors as they happen.

  • server_console (server): the host's agent console output.

  • instance_telemetry (environment + container [+ instance]): live per-second JSON resource snapshots for one instance.

  • server_telemetry (server): live JSON stats for a server.

Use this when snapshot data is not enough: logs come back empty, telemetry lags (snapshots can trail by ~10 minutes), or the problem is happening right now. An empty capture is a valid observation — it means the stream was quiet for the window.

Streams replay recently buffered output when you connect, so a capture is NOT all fresh: frames arriving in roughly the first second are counted as replayed BACKLOG (backlog_bytes) and may predate the capture by a long time; only what arrives after that counts as live output (live_bytes). Console captures insert a marker line between the two; telemetry frames carry [T+offset] arrival prefixes. backlog_bytes>0 with live_bytes=0 means nothing new happened while watching — never treat replayed backlog as proof of what the process is doing right now. A large capture is trimmed to the most recent ~16KB and flagged truncated (with total_bytes = the untrimmed size).

The returned token is short-lived, observation-only, and expires on its own; there is no revoke step. Read-only: nothing can be sent to the stream through this tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
serverNoServer hostname, nickname, or ID. Required for server targets.
targetYesWhich live stream to capture.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
instanceNoInstance id or hostname. Optional when the container has a single instance.
containerNoContainer. Required for instance targets.
environmentNoEnvironment. Helps resolve the container for instance targets.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.
duration_secondsNoHow long to capture, 1-30 seconds.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnly/idempotent/destructive/openWorld, and the description adds substantial non-obvious behavior: backlog replay on connect, the backlog_bytes vs live_bytes distinction, marker lines and [T+offset] prefixes, ~16KB trimming with a truncated flag and total_bytes, and the short-lived observation-only token with no revoke step. The read-only claim is consistent with readOnlyHint=true.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core action is front-loaded, then targets are bulleted, then behavioral caveats — a sensible structure. It is dense but nearly every sentence carries a distinct behavioral fact; it is on the long side for a single tool, which keeps it just under the top score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden and does so thoroughly: it explains backlog_bytes/live_bytes, truncated, total_bytes, and the returned token. Combined with the target semantics and replay warnings, an agent has everything needed to call and interpret this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds real meaning by mapping each target enum value to the parameters it requires (container/environment/instance for instance targets, server for server targets). That pairing is not derivable from the terse schema descriptions alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource+scope ('Watch a live Cycle stream for a few seconds and return what arrived, plus the temporary credentials') and enumerates all four targets with their required scope (environment + container [+ instance] vs server). This lets an agent distinguish it from siblings like get_logs, get_telemetry, and query_metrics without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states exactly when to reach for this tool ('snapshot data is not enough: logs come back empty, telemetry lags (snapshots can trail by ~10 minutes), or the problem is happening right now'), effectively contrasting it with the snapshot-based siblings. It also clarifies that an empty capture is a valid observation, removing the main ambiguity an agent would face when choosing between this and get_logs/get_telemetry.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_dns_propagationCheck DNS Resolution and PropagationA
Read-onlyIdempotent
Inspect

Answer "why doesn't this domain work yet" by asking DNS directly, from two sides at once: the zone's OWN authoritative nameservers (read from its live NS delegation, so whatever actually answers — one of Cycle's na/eu pools, or an outside provider) and the public recursive resolvers (what a real client sees). Comparing the two separates failures that look identical from a browser: resolved, propagating (nameservers have it, public resolvers not yet — WAIT for the TTL, change nothing), no-address (name exists, no A/AAAA — on Cycle, a LINKED record whose target has not published addresses), nxdomain (never created or deleted), or unknown (nothing reachable — proves nothing either way). The verdict's detail explains each case in context.

It also reports whether Cycle itself holds a record for this name (zone, type, and for LINKED records the container/VM/deployment it targets), which distinguishes "never created in Cycle" from "exists and still spreading".

Pass wait_seconds to keep re-probing until the name resolves everywhere. Read-only: only asks questions of DNS servers. To create or repoint a record, use manage_dns_record.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesFully qualified domain name to probe, e.g. 'app.example.com'. Pass the hostname a user would actually visit, not the zone.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
record_typesNoRecord types to query. Defaults to a and aaaa, which is what matters for reaching a container or VM.
wait_secondsNoKeep re-probing until the name resolves consistently, up to this many seconds. 0 (default) returns a single snapshot. Propagation usually outlasts any single call — a 'propagating' verdict after the wait is not a failure, just call again.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.
nameserver_groupNoForce which Cycle nameserver pool is treated as authoritative ('all' queries every pool). Omit to follow the zone's own delegation, which is what actually answers for it.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only/idempotent/open-world, so the bar is lower, yet the description adds substantial context: a full verdict taxonomy (resolved, propagating, no-address, nxdomain, unknown), the meaning of each, the caveat that 'unknown' proves nothing, and that it reports whether Cycle holds a record (zone/type/target for LINKED). This goes well beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and the two-sided mechanism are front-loaded, and sentences are information-dense. It is long for a single tool, with minor redundancy (the read-only claim restates readOnlyHint), but nearly every sentence carries semantic weight, so it stays structured rather than bloated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of explaining returns, and it does so via the verdict taxonomy and the Cycle-record distinction. An agent has enough to call it correctly and interpret results, including the propagation-over-time caveat.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds meaning: it explains what passing wait_seconds does ('keep re-probing until the name resolves everywhere') and how to interpret a propagating verdict after a wait ('not a failure, just call again'). It does less for the record_types/nameserver_group nuance, which stays in the schema, keeping this at 4 rather than 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific action (check DNS resolution/propagation) and a specific resource (DNS), and immediately frames it with an observable verb+scope: 'asking DNS directly, from two sides at once.' It clearly distinguishes itself from manage_dns_record and delete_dns_record by scoping to read-only diagnosis.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States the exact triggering situation ('why doesn't this domain work yet'), enumerates the verdicts that guide follow-up action (propagating -> WAIT, change nothing), and explicitly names the alternative for mutations: 'To create or repoint a record, use manage_dns_record.' When/when-not and the alternative are both present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_server_provisioningCheck Server Provisioning ProgressA
Read-onlyIdempotent
Inspect

Track servers coming online after deploy_servers: reports the deployment job's state and every server in the cluster with its current state (new → provisioning → configuring → live). A server is ready once it reaches 'live', meaning CycleOS booted and the server is checking in. Pass wait_seconds to keep polling up to that bound — progress updates stream to the client while it waits — or 0 (the default) for a single snapshot. Provisioning almost always outlasts a single call: bare metal takes many minutes, so the reliable pattern is repeated snapshot calls until ready is true, not one long wait. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
jobNoThe deployment job ID returned by deploy_servers, to track alongside the servers.
clusterYesThe cluster whose servers to check, as passed to deploy_servers.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
wait_secondsNoMax seconds to keep polling until the job completes and every server in the cluster is live. 0 returns a single snapshot. While waiting, progress updates are streamed to the client. Provisioning takes far longer than any single call can wait, so repeated 0-wait snapshots are the normal way to track it.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well past the annotations: it discloses the state machine (new → provisioning → configuring → live), the readiness condition ('live' means CycleOS booted and the server is checking in), the streaming behavior while waiting, and realistic timing expectations for bare metal. 'Read-only' is redundant with readOnlyHint but consistent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose is front-loaded in the first clause, and each subsequent sentence carries distinct information (state machine, readiness, polling strategy). It is on the dense side and slightly repetitive about provisioning taking long, which keeps it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of explaining returns, and it does: job state plus per-server state and the meaning of 'ready'. Combined with the polling guidance, an agent has everything needed to call and interpret this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds real meaning for wait_seconds: it defines 0 as a single snapshot, describes progress streaming while waiting, and warns that provisioning outlasts any single wait. The job and cluster parameters are only implied, not explained further.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Track servers coming online after deploy_servers') and scopes the return ('the deployment job's state and every server in the cluster with its current state'). It clearly distinguishes itself from deploy_servers (the predecessor) and from plain list_servers by framing itself as post-deploy progress tracking.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly positions the tool in a workflow ('after deploy_servers') and prescribes the operating pattern: repeated snapshot calls until ready is true rather than one long wait. It doesn't name a competing sibling (e.g. get_deployment_status) or state when this is the wrong tool, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_environmentCreate an EnvironmentAInspect

Create a new environment on a Cycle cluster and, by default, start it so discovery DNS and the scheduler are running and it is ready to deploy into. The load balancer is only created when the environment's services start with a public container present; deploy_application handles that after deploying public containers.

An environment is a group of containers with a private network between them. The private network is IPv6-only; containers resolve each other by hostname via the discovery service.

legacy_networking enables private IPv4 in the environment. This choice is PERMANENT — it cannot be changed after creation — and should be avoided unless an application truly cannot support IPv6. Prefer enabling the app's IPv6 support instead (e.g. mongod --ipv6).

The cluster is the set of servers the environment schedules onto; find cluster identifiers with list_servers. Creating an environment is a mutation: confirm name, cluster, and legacy_networking with the user before calling.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesHuman-readable name for the environment.
startNoStart the environment after creating it so its services are running. Default true.
clusterYesCluster identifier to create the environment on. Use list_servers to see available clusters.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
identifierNoOptional identifier slug; generated from the name when omitted.
descriptionNoOptional description of the environment's purpose.
wait_secondsNoMax seconds to wait for the start job. 0 returns immediately after the job is accepted.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.
legacy_networkingNoEnable private IPv4 (legacy mode). PERMANENT — cannot be changed after creation. Avoid unless the app cannot do IPv6.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses several traits an agent needs: the environment is started by default so DNS and the scheduler are running, the load balancer is created only later by deploy_application, and legacy_networking is PERMANENT and irreversible. It also flags the call as a mutation requiring user confirmation, which goes well past what readOnlyHint/destructiveHint already state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The definition is front-loaded with the core action and default behavior, then layers in network semantics and the legacy_networking warning. It is longer than typical but each paragraph carries distinct information; only the environment-definition paragraph is slightly incidental.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter mutation tool with no output schema, the description covers everything an agent needs: what gets created, defaults, irreversibility of legacy_networking, prerequisite lookup via list_servers, and the need for user confirmation. Annotations cover the safety profile and the schema covers all parameters, so no meaningful gap remains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description meaningfully enriches key parameters: legacy_networking's permanence and IPv6 preference (with a concrete example, mongod --ipv6), the default-start behavior, and the cluster parameter's link to list_servers. This adds real semantics beyond the schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Create a new environment on a Cycle cluster') and immediately defines what an environment is (a group of containers with a private IPv6 network). It is clearly distinguishable from siblings like delete_environment, list_environments, and cycle_control_environment, and it even clarifies scope relative to deploy_application's role with the load balancer.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives strong context: find cluster identifiers with list_servers, and confirm name/cluster/legacy_networking with the user before calling. It also routes the load-balancer concern to deploy_application. It stops short of explicitly stating when NOT to use it (e.g. versus cycle_control_environment), so it is clear context rather than full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cycle_control_containerControl ContainerA
Destructive
Inspect

Start, stop, or restart a container on Cycle. Waits up to wait_seconds for the job to finish and returns the container's final state. Use list_containers first only if unsure what exists.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionNoLifecycle action. stop and restart interrupt running workloads.start
contextNoWhy are you calling this tool? Briefly describe the user's goal.
containerYesContainer to act on.
environmentYesEnvironment the container lives in. Required; it scopes the lookup.
wait_secondsNoMax seconds to wait for the job to complete. 0 returns immediately after the job is accepted.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (destructiveHint=true, idempotentHint=false), so the bar is lower; the description still adds blocking semantics ('waits up to wait_seconds for the job to finish') and return content ('container's final state'), neither of which is derivable from annotations alone. It omits what happens if the wait times out or the job fails.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, action scope front-loaded, followed immediately by the timing behavior and the one routing caveat. No filler or repetition of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With six parameters, no output schema, and a destructive mutation, the description usefully tells the agent that the call blocks and yields the final container state. Remaining gaps are minor (failure/timeout behavior), and the destructive profile is carried by annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, including enum values, defaults, and the wait_seconds=0 immediate-return case, so the baseline is 3. The description's wait_seconds mention merely restates the schema and adds no syntax or format detail beyond it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States specific verbs (start, stop, restart) plus the resource (container) and platform (Cycle), which cleanly separates it from siblings such as reconfigure_container, delete_container, and cycle_control_environment. An agent can pick this tool without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit conditional for the alternative: 'Use list_containers first only if unsure what exists.' That is real routing guidance, but it doesn't address when to choose this over reconfigure_container or the environment-level control tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cycle_control_environmentControl EnvironmentA
DestructiveIdempotent
Inspect

Start or stop a Cycle environment. Starting brings up the environment's service containers: discovery (the environment's private DNS), scheduler, and — only once the environment has a public container — the load balancer. Containers in an environment cannot resolve each other by hostname until discovery is running — start the environment BEFORE deploying or setting up multi-container applications in it. Starting an already-live environment is safe and restarts any stopped services. Each service is itself a container and can be started or restarted individually with cycle_control_container using its container id.

Stopping an environment stops its services AND every container in it — treat it as highly disruptive.

Both actions are mutations: confirm with the user before calling. Waits up to wait_seconds for the job and returns the environment's state.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionNoLifecycle action. stop halts the environment's services and every container in it.start
contextNoWhy are you calling this tool? Briefly describe the user's goal.
environmentYesEnvironment to act on.
wait_secondsNoMax seconds to wait for the job to complete. 0 returns immediately after the job is accepted.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: it enumerates exactly what starting provisions (discovery/DNS, scheduler, conditional load balancer), what stopping destroys (services AND every container), that re-starting a live environment is safe and restarts stopped services, and that the call blocks up to wait_seconds and returns the environment state. Annotations only give the generic destructive/idempotent/openWorld flags.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well front-loaded: what it does, then ordering guidance, then sibling routing, then the destructiveness warning. Slightly redundant, since 'stopping stops services and every container' largely restates the schema's action description, costing a little efficiency.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-action lifecycle mutation with no output schema, it covers provisioning effects, ordering constraints, destructiveness, sibling alternative, and the confirmation requirement, and it states what is returned. Nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds modest value by clarifying the blast radius of the stop action and that the call waits up to wait_seconds and returns state, but it does not add syntax or format detail beyond what the schema already documents.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb pair (start/stop) and the resource (Cycle environment), and distinguishes itself from the same-granularity sibling cycle_control_container by explicitly saying per-service control lives there. An agent can tell what this does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete precondition ('start the environment BEFORE deploying or setting up multi-container applications') and a clear alternative at a different granularity ('Each service... can be started or restarted individually with cycle_control_container using its container id'). It also warns the stop action is highly disruptive, which routes the agent away from casual use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cycle_control_virtual_machineControl Virtual MachineA
Destructive
Inspect

Start, stop, or restart a virtual machine on Cycle. Waits up to wait_seconds for the job to finish and returns the VM's final state. Starting requires a hypervisor-capable server in the cluster (list_servers reports 'virtualization'); the first boot downloads the disk image and can exceed the wait — the job keeps running on Cycle. Log in with the VM's SSH keys or the root password Cycle generated at creation (retrievable for ~10 minutes afterward). stop and restart interrupt the running guest OS.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionNoLifecycle action. stop and restart interrupt the running guest OS.start
contextNoWhy are you calling this tool? Briefly describe the user's goal.
environmentYesEnvironment the VM lives in. Required; it scopes the lookup.
wait_secondsNoMax seconds to wait for the job to complete. 0 returns immediately after the job is accepted.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.
virtual_machineYesVirtual machine to act on.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial behavior beyond the annotations (destructiveHint=true, idempotentHint=false): it waits up to wait_seconds then returns the VM's final state, the first boot may exceed the wait while the job continues asynchronously on Cycle, and credentials (SSH keys or generated root password) are retrievable for only ~10 minutes. This is exactly the operational nuance annotations cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and timeout behavior, then layers prerequisites and caveats — no padding in the operational sentences. The SSH/root-password login detail is useful context but slightly tangential to invoking the tool, keeping it from a perfect score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description still explains the return ('the VM's final state'), the async semantics when the wait elapses, prerequisite cluster capability, and the destructive effect of stop/restart. For a 6-parameter mutation tool with full annotation coverage, nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning about how wait_seconds interacts with startup: the wait may time out on first boot while the job keeps running, and 0 returns immediately per the schema. It also reinforces that action is restricted to lifecycle verbs. It does not clarify context/environment/conversation_id beyond the schema, so not a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States specific verbs (start/stop/restart) and the resource (virtual machine on Cycle), which cleanly distinguishes it from siblings like cycle_control_container, cycle_control_environment, and delete_virtual_machine. An agent can identify the operation without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives real prerequisites and context: starting requires a hypervisor-capable server (verifiable via list_servers reporting 'virtualization'), and stop/restart interrupt the running guest OS. It does not explicitly name when to prefer a sibling tool (e.g., run_vm_command, reconfigure_virtual_machine), so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

decommission_serverDecommission ServerA
Destructive
Inspect

PERMANENTLY decommission (delete) an infrastructure server from the Cycle hub. This is IRREVERSIBLE through this tool — the server is removed from the cluster and, for provider-managed servers, typically destroyed at the provider. Cycle refuses to decommission a server that still has running instances, and this tool never forces past that: the environments or containers using the server must be deleted (or their instances stopped/migrated) first.

This tool is two-step by design and NOTHING is decommissioned on the first call:

  1. Call without confirm. The server is resolved and the response describes it — hostname, nickname (if set), ID, cluster, provider/model/location, when it was created, its current uptime, and every instance still placed on it with its state and owning environment. Present ALL of those details to the user verbatim.

  2. Only after the user explicitly approves decommissioning that specific server, call again with confirm:true and server set to the exact 24-char hex ID from step 1 (hostnames are not accepted with confirm — the ID binds the action to what the user approved).

Never set confirm:true unless the user has just approved this exact decommission; a general instruction like "scale down the cluster" is not sufficient.

ParametersJSON Schema
NameRequiredDescriptionDefault
serverYesServer to decommission. Without confirm: hostname, nickname, or 24-char hex ID. With confirm:true: MUST be the exact 24-char hex ID returned by the preview call.
confirmNoSet true ONLY after the user has explicitly approved decommissioning this exact server (shown its hostname, nickname, ID, uptime, and any instances on it). Requires server to be the exact hex ID from the preview call.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
wait_secondsNoconfirm only: max seconds to wait for the decommission job to complete (default 60). 0 returns immediately after the job is accepted.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true and idempotentHint=false, but the description adds critical behavior beyond them: irreversibility, provider-side destruction for managed servers, the refusal condition on running instances, the two-step design where nothing happens on the first call, and the ID-binding requirement for confirm.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the permanent/irreversible warning and structured with numbered steps, so it is easy to scan. It is somewhat verbose, repeating the 24-char hex ID and 'explicitly approved' constraints several times, which costs a point.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive two-step tool with no output schema, the description covers the prerequisites, the preview response contents, the confirmation contract, and the safety guardrails. An agent has everything needed to call it correctly and safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real semantics: hostname/nickname accepted only without confirm, confirm requires the exact 24-char hex ID from the preview, and confirm:true is forbidden absent explicit user approval. It does not elaborate on wait_seconds beyond what the schema already says.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (decommission/delete) plus resource (infrastructure server) and scope (from the Cycle hub, permanent). It is clearly distinguishable from siblings such as delete_virtual_machine, delete_environment, and delete_container by naming the server-level resource explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly covers when to use (permanent server removal), when not (refuses if running instances; instances must be stopped/migrated or environments/containers deleted first), and the mandatory two-step preview-then-confirm workflow with the guardrail that confirm requires just-in-time user approval of this exact server.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_containerDelete ContainerA
Destructive
Inspect

PERMANENTLY delete a container on Cycle. This is IRREVERSIBLE: it destroys the container, every one of its instances, and any data on stateful local instances.

This tool is two-step by design and NOTHING is deleted on the first call:

  1. Call without confirm. The container is resolved and the response describes exactly what would be destroyed — name, ID, environment, when it was created, and how many instances exist, their state, and which servers they run on. Present ALL of those details to the user verbatim.

  2. Only after the user explicitly approves deleting that specific container, call again with confirm:true and container set to the exact 24-char hex ID from step 1 (names are not accepted with confirm — the ID binds the deletion to what the user approved).

Never set confirm:true unless the user has just approved this exact deletion; a general instruction like "clean up the environment" is not sufficient. Containers with deletion protection (lock) cannot be deleted until unlocked; reconfigure_container with lock:false lifts it, which is itself a change the user must approve.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmNoSet true ONLY after the user has explicitly approved deleting this exact container (shown its name, ID, creation date, and instances). Requires container to be the exact hex ID from the preview call.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
containerYesContainer to delete. With confirm:true: MUST be the exact 24-char hex ID returned by the preview call.
environmentNoEnvironment the container lives in. Required when container is a name or identifier; ignored when an exact hex ID is given.
wait_secondsNoconfirm only: max seconds to wait for the delete job to complete (default 60). 0 returns immediately after the job is accepted.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only say destructiveHint=true and idempotentHint=false; the description goes far beyond by disclosing irreversibility, exactly what is destroyed, the two-step preview-then-confirm safety contract, that the ID (not the name) binds the deletion to what the user approved, and that deletion protection blocks the call until unlocked.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the irreversible consequence in the first sentence, then numbers the two-step flow so the ordering is unmistakable. Despite its length, every sentence carries distinct, load-bearing safety information rather than restating the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation tool with no output schema, it compensates by describing the contents of the step-1 preview response (name, ID, environment, creation date, instance count/state, servers), which is the return value the agent most needs to relay. Nothing required to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds real cross-parameter meaning: confirm:true requires the exact 24-char hex ID from the preview and rejects names, which clarifies how container, environment, and confirm interact. It doesn't cover wait_seconds or conversation_id semantics in prose, keeping it just short of a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource with explicit scope and severity ('PERMANENTLY delete a container'), and immediately enumerates the blast radius (container, all instances, stateful local instance data). An agent can distinguish this from delete_environment, delete_stack, or delete_virtual_machine purely by the named resource.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit two-step protocol: call without confirm first, present the preview verbatim, then only call again with confirm:true after explicit user approval. It also states the when-not condition ('a general instruction like clean up the environment is not sufficient') and names the sibling to use for the locked case (reconfigure_container with lock:false).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_dns_recordDelete DNS RecordA
Destructive
Inspect

PERMANENTLY delete a single DNS record from a Cycle zone. This is IRREVERSIBLE: the record stops resolving, and for a linked record Cycle also tears down its load-balancer routing and stops maintaining its TLS certificate.

This tool is two-step by design and NOTHING is deleted on the first call:

  1. Call without confirm. The record is resolved and the response describes exactly what would be removed — its name, ID, type, full payload, resolved domain, and zone. Present ALL of those details to the user verbatim.

  2. Only after the user explicitly approves deleting that specific record, call again with confirm:true, zone set to the zone from step 1, and record_id set to the exact 24-char hex ID from step 1 (names are not accepted with confirm — the ID binds the deletion to what the user approved).

Addressing for the preview: pass a full 'domain' (e.g. 'app.example.com' — the covering zone and record name are derived; apex = '@'), or 'zone' plus 'name' (or 'record_id'). When several records share the name, 'type' or 'record_id' disambiguates.

Never set confirm:true unless the user has just approved this exact deletion; a general instruction like "clean up the zone" is not sufficient. To create or repoint records use manage_dns_record; to delete a whole zone, surface that as a gap — this tool only deletes single records.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoRecord name within the zone ('@' for apex, '*' for wildcard).
typeNoRecord type; disambiguates when several records share a name.
zoneNoZone origin (e.g. 'example.com') or 24-char hex ID. Required with confirm.
domainNoFull domain, e.g. 'app.example.com'; the covering zone and record name are derived. Alternative to zone+name.
confirmNoSet true ONLY after the user has explicitly approved deleting this exact record (shown its name, ID, type, payload, and zone). Requires zone plus record_id as the exact hex ID from the preview call.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
record_idNoExact record ID (with 'zone'); disambiguates when several records share a name. With confirm:true this MUST be the exact 24-char hex ID returned by the preview call.
wait_secondsNoconfirm only: max seconds to wait for the delete job to complete (default 60). 0 returns immediately after the job is accepted.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true and readOnlyHint=false, but the description adds what an agent cannot infer: irreversibility, that the record stops resolving, and that linked records lose LB routing and TLS maintenance. It also discloses the two-phase safety gate and the name-vs-ID binding rule.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Long, but front-loaded with the irreversibility warning and structured as a numbered flow, so it scans well. For an irreversible destructive tool the length is largely earned, though a few sentences could be tightened.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema and 9 parameters with none required, the description compensates by describing what the preview call returns (name, ID, type, payload, domain, zone) and the wait_seconds/conversation_id lifecycle. An agent has everything needed to call it correctly and safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds cross-parameter interaction rules the schema lacks: domain derivation vs zone+name addressing, type/record_id disambiguation, and that names are rejected with confirm in favor of the exact 24-char hex ID from the preview.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('PERMANENTLY delete a single DNS record from a Cycle zone') and explicitly differentiates from siblings: create/repoint via manage_dns_record, and zone deletion is out of scope. An agent can route correctly without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides an explicit two-step protocol, states the precondition for confirm:true ('only after the user explicitly approves'), warns that vague instructions like 'clean up the zone' are insufficient, and names alternatives. Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_environmentDelete EnvironmentA
Destructive
Inspect

PERMANENTLY delete an environment on Cycle and EVERYTHING in it: all containers and their instances (including data on stateful instances), all virtual machines and their disks, all scoped variables, and the environment's service containers (discovery, VPN, load balancer). This is IRREVERSIBLE.

This tool is two-step by design and NOTHING is deleted on the first call:

  1. Call without confirm. The environment is resolved and the response inventories exactly what would be destroyed — the environment's name, ID, and creation date, every container with its instance count, instance totals and the servers they run on, every VM, and every scoped variable. Present ALL of those details to the user verbatim.

  2. Only after the user explicitly approves deleting that specific environment and its full contents, call again with confirm:true and environment set to the exact 24-char hex ID from step 1 (names are not accepted with confirm — the ID binds the deletion to what the user approved).

Never set confirm:true unless the user has just approved this exact deletion; a general instruction like "clean things up" is not sufficient. To delete a single container instead, use delete_container.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmNoSet true ONLY after the user has explicitly approved deleting this exact environment (shown its name, ID, creation date, and the full inventory of containers, instances, VMs, and scoped variables). Requires environment to be the exact hex ID from the preview call.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
environmentYesEnvironment to delete. With confirm:true: MUST be the exact 24-char hex ID returned by the preview call.
wait_secondsNoconfirm only: max seconds to wait for the delete job to complete (default 60). Environments with many workloads take longer. 0 returns immediately after the job is accepted.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=false, but the description goes well beyond them: it specifies that the first call is non-destructive, that the preview returns an inventory that must be shown verbatim, that the ID (not the name) binds the deletion, and that the operation is irreversible. This is exactly the extra context annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The irreversible warning and full destruction list are front-loaded in the first two sentences, then the two-step flow is numbered so the agent can follow it sequentially. Every sentence carries an actionable constraint despite the length, which is proportionate to the risk.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by describing the preview response contents in detail (name, ID, creation date, container/instance/server counts, VMs, scoped variables). Combined with the destructive annotations, an agent has everything needed to call this safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so the baseline is 3, but the description adds real semantic weight: names are rejected when confirm is set, the ID must be the exact 24-char hex value from the preview, and confirm must be tied to the specific approval. It adds little on wait_seconds/conversation_id, which the schema already documents.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ('PERMANENTLY delete an environment') and immediately enumerates the exact blast radius (containers/instances, VMs/disks, scoped variables, service containers). It also explicitly distinguishes itself from the nearest sibling: 'To delete a single container instead, use delete_container.'

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit two-step protocol with the condition that gates step 2 ('Only after the user explicitly approves...'), a negative condition ('Never set confirm:true unless...'), and a concrete example of insufficient consent ('clean things up'). Alternatives are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_image_sourceDelete an Image Source or Its ImagesA
Destructive
Inspect

PERMANENTLY delete an image source together with every image built from it, or — when image_ids are given — only those images while keeping the source. IRREVERSIBLE.

An image a container is currently running on cannot be deleted; the preview lists such containers as a blocker. Stack builds do not use image sources, so this tool never touches a stack's images — use delete_stack for those.

Two-step, nothing is deleted on the first call:

  1. Call without confirm. The response inventories the source, the images that would go, and any containers on them. Present it to the user verbatim.

  2. Only after the user explicitly approves that specific deletion, call again with confirm:true, the exact 24-char hex source ID from step 1, and the same image_ids (if any).

ParametersJSON Schema
NameRequiredDescriptionDefault
sourceYesImage source: name, identifier, resource path ('image-source:...'), or 24-char hex ID. With confirm:true it MUST be the hex ID from the preview.
confirmNoSet true ONLY after the user has explicitly approved this exact deletion from the preview.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
image_idsNoDelete only these images (24-char hex IDs from list_image_sources) and keep the source. Omit to delete the source and all its images.
wait_secondsNoconfirm only: max seconds to wait for the delete job(s) (default 60). 0 returns after they are accepted.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark it destructiveHint=true/readOnlyHint=false, but the description goes well beyond them: IRREVERSIBLE consequences, the fact that nothing is deleted on the first call, the container-on-image blocker, and the exact confirmation contract. This is rich behavioral context beyond the structured hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the IRREVERSIBLE warning and the safety-relevant constraints before the numbered procedure. It is somewhat long, but each sentence carries distinct information for a destructive two-step tool, so little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Complete for a destructive tool: it covers purpose, blockers, sibling routing, the two-step approval workflow, and what the preview response contains. With no output schema, the description compensates by describing the preview inventory contents.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds workflow meaning: it explains that source must be the hex ID from the preview when confirming, that image_ids selects image-only deletion vs. source-plus-images, and that context/wait_seconds/conversation_id follow specific call patterns. Some of this duplicates the schema, but the contextualization is real added value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (delete) and resource (image source and/or its images), names the two modes (source+all images vs. only listed images), and explicitly distinguishes itself from delete_stack. An agent can identify it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use and when-not-to-use guidance: stack builds are routed to delete_stack, running containers block deletion (surfaced as a preview blocker), and the mandatory two-step approval flow is spelled out step by step. Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_stackDelete StackA
Destructive
Inspect

PERMANENTLY delete a stack on Cycle together with ALL of its builds. This is IRREVERSIBLE.

Cycle refuses to delete a stack while any container deployed from it exists, so the preview reports such containers as a blocker. Teardown order: delete the containers first (delete_environment removes every container in an environment; delete_container removes one), then the stack. Linked DNS records do not block anything and can be deleted at any point, in parallel.

This tool is two-step by design and NOTHING is deleted on the first call:

  1. Call without confirm. The stack is resolved and the response inventories what would be destroyed — the stack's name, ID, identifier, and creation date, its builds — and lists any containers deployed from it that block the deletion. Present ALL of those details to the user verbatim.

  2. Only after the user explicitly approves deleting that specific stack, call again with confirm:true and stack set to the exact 24-char hex ID from step 1 (names are not accepted with confirm — the ID binds the deletion to what the user approved).

Never set confirm:true unless the user has just approved this exact deletion; a general instruction like "clean things up" is not sufficient.

ParametersJSON Schema
NameRequiredDescriptionDefault
stackYesStack to delete. Without confirm: name, identifier, resource path ('stack:...'), or 24-char hex ID. With confirm:true: MUST be the exact 24-char hex ID returned by the preview call.
confirmNoSet true ONLY after the user has explicitly approved deleting this exact stack (shown its name, ID, builds, and the containers deployed from it). Requires stack to be the exact hex ID from the preview call.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
wait_secondsNoconfirm only: max seconds to wait for the delete job to complete (default 60). 0 returns immediately after the job is accepted.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the destructiveHint annotation by disclosing irreversibility, the platform's refusal to delete a stack with live containers and how that appears in the preview, the two-step no-op-on-first-call protocol, and the ID-binding rule. These are non-obvious behaviors an agent could not infer from annotations or schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the irreversible warning, then a numbered two-step protocol — every sentence carries operational weight. It is long, but the length is proportionate to the destructive two-phase workflow; only the verbatim-presentation instruction is mildly over-prescriptive.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, yet the description specifies what the preview returns (name, ID, identifier, creation date, builds, blocking containers) and what confirm does, so an agent can drive the whole workflow correctly without further information.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real semantic value: it explains why confirm:true requires the exact 24-char hex ID ('the ID binds the deletion to what the user approved') and that names are rejected with confirm, plus the preview/confirm role split for the stack parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource with scope: 'PERMANENTLY delete a stack on Cycle together with ALL of its builds.' It is immediately distinguishable from sibling delete tools (delete_container, delete_environment, delete_dns_record) because it names the teardown ordering relative to them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use, when-not-to-use, and alternatives: containers must be deleted first via delete_environment or delete_container, DNS records need not block, and confirm:true is forbidden unless the user just approved this exact deletion ('clean things up' is called out as insufficient).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_virtual_machineDelete Virtual MachineA
Destructive
Inspect

PERMANENTLY delete a virtual machine on Cycle. This is IRREVERSIBLE: it destroys the VM and its local volumes, including the base (boot) volume and everything the guest OS wrote to them. A running VM is powered off, not shut down gracefully.

This tool is two-step by design and NOTHING is deleted on the first call:

  1. Call without confirm. The VM is resolved and the response describes exactly what would be destroyed — name, ID, identifier, state, image, when it was created, and its volumes. Present ALL of those details to the user verbatim.

  2. Only after the user explicitly approves deleting that specific VM, call again with confirm:true and virtual_machine set to the exact 24-char hex ID from step 1 (names are not accepted with confirm — the ID binds the deletion to what the user approved).

Never set confirm:true unless the user has just approved this exact deletion; a general instruction like "clean up the environment" is not sufficient. VMs with deletion protection (lock) cannot be deleted until unlocked; reconfigure_virtual_machine with lock:false lifts it, which is itself a change the user must approve. Attached EXTERNAL (SAN) volumes are detached rather than destroyed, but local volumes are gone for good — if the guest holds data worth keeping, copy it out first (run_vm_command).

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmNoSet true ONLY after the user has explicitly approved deleting this exact VM (shown its name, ID, state, image, creation date, and volumes). Requires virtual_machine to be the exact hex ID from the preview call.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
environmentNoEnvironment the VM lives in. Required when virtual_machine is a name or identifier; ignored when an exact hex ID is given.
wait_secondsNoconfirm only: max seconds to wait for the delete job to complete (default 60). 0 returns immediately after the job is accepted.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.
virtual_machineYesVM to delete. With confirm:true: MUST be the exact 24-char hex ID returned by the preview call.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, but the description adds substantial beyond-annotation context: exactly what is destroyed (VM, local/boot volumes, guest OS writes), the non-graceful power-off behavior, the two-step no-op-on-first-call contract, the lock/deletion-protection block, and external SAN volumes being detached rather than destroyed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the critical irreversible warning and structures the two-step flow as a clear numbered list. It is fairly long, but nearly every sentence carries safety-relevant information; minor tightening is possible but nothing is gratuitous.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, non-idempotent tool with no output schema and one required parameter, the description covers everything an agent needs: prerequisites (unlock), what to present to the user, what survives vs is destroyed, and cross-tool alternatives. Nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so baseline is 3, but the description meaningfully reinforces the most safety-critical parameter constraints: confirm must bind to the exact 24-char hex ID from the preview call, names are not accepted with confirm, and what step-1 returns. This adds operational nuance beyond the schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (PERMANENTLY delete) and resource (virtual machine on Cycle) and immediately scopes it against sibling delete tools by naming exactly what gets destroyed (VM plus local/boot volumes). The irreversible power-off vs graceful shutdown distinction further disambiguates it from reconfigure or power tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly defines the two-step flow and when to proceed, including the exact trigger for confirm:true ('user has just approved this exact deletion') and an explicit exclusion ('a general instruction like clean up the environment is not sufficient'). It also routes to alternatives (reconfigure_virtual_machine for lock:false, run_vm_command to copy data out) with conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

deploy_applicationDeploy an ApplicationA
Destructive
Inspect

Deploy a multi-container application onto Cycle. You supply the app knowledge (image, command, env, ports, volumes, how many members and how they address each other); the tool turns it into a Cycle stack spec, creates and builds the stack, and deploys it into an environment. For a workload that truly needs a full virtual machine use deploy_virtual_machine.

Modes: by default everything deploys through a stack (containers that version and deploy together). One application is ONE call and one stack: every container it needs — data nodes, UI, sidecars — goes in the same 'containers' list. Containers are created stopped, so "incremental" means starting and checking them one at a time afterwards, not deploying them in separate calls; a stack cannot be extended with more containers later. Separate stacks are for independent applications that release on their own. stack:false creates one-off containers directly in the environment with no stack objects. Every rule below applies to both.

Existing stacks: a call with 'containers' ALWAYS creates a new stack. When the user names a stack that already exists — built by this tool, the dashboard, or from a git repo — pass stack_id instead and leave 'containers' out (or pass the original list only to link domains). It deploys the stack's latest usable build into the environment; build_id picks a specific build, rebuild:true generates a fresh build from the stack's source first. Containers the stack already has in that environment are left untouched unless 'redeploy' says to update them — rebuild:true plus redeploy:{reimage:true} is how a new image tag or commit is rolled out to a running deployment. preview:true reports the stack's recent builds, the containers the chosen build creates, and which existing ones would be updated.

MUST GET RIGHT — Cycle accepts these and the application is then silently unreachable or insecure. Preview checks them and an error finding blocks the deploy:

  • Listen on ::. The private network is IPv6-ONLY (discovery serves AAAA records) and the load balancer reaches backends over it, so a process on 0.0.0.0 or localhost is unreachable from siblings and from the LB. IPv4-default apps need IPv6 enabled explicitly (mongod --ipv6, HOST=::), and so do client libraries (Node's ioredis needs ?family=6 on the URL).

  • TLS through the LB needs BOTH "443:" and "80:80" in ports. "443:80" is Cycle notation: LB ingress 443 routes to backend 80 with TLS terminated at the LB.

  • Secrets never go in env or args — both are stored in plain text in the stack. Sequence: deploy (containers are created stopped) → create scoped variables with manage_scoped_variable and source.secret:true → start with cycle_control_container. First boot reads them (postgres initdb reads POSTGRES_PASSWORD once; redis needs its config file present).

  • Stateful containers deploy as a SINGLE instance with their own volume. Model a clustered app as N separate single-instance stateful containers (mongo-0, mongo-1, ...), each with its own volume and unique hostname. Containers reach each other by hostname; the environment must be live for discovery to resolve them.

  • Placement is unconstrained by default — a container may land on ANY server whose pool allows it, so cluster members can spread across providers. Pin hardware with constraints.node.tags.all (every tag present) or .any, and confirm the tag exists with list_servers BEFORE deploying; a tag matching no server leaves the container with no deployment target.

  • public defaults to "disable". Set it only to expose a container.

Images — each container sets exactly ONE of: image (a Docker Hub target like 'mongo:7'; an existing image source with the same origin is reused), image_source (already on Cycle — prefer it whenever the user has one, it carries their registry credentials and build config), or dockerfile (repo or targz_url for Cycle to build). Registry and git credentials are never accepted directly.

Backups — 'backups' (stateful containers only) has Cycle run 'command' on the cron 'schedule' inside the container and ship its STDOUT to a backup-capable hub integration ('destination', e.g. Backblaze B2); 'restore_command' reads a backup from STDIN. All four are required; the destination must already be enabled on the hub. Commands may reference env vars (mongodump -u $USER) but the tools they call must exist in the image.

Domains — 'domain' exposes a container through the environment load balancer. A DNS zone covering it must already exist; the tool creates a LINKED record pointing at the container and never overwrites an existing one (conflicts are errors to resolve with the user). Forces public:"enable" and requires ports; TLS on the record follows from a '443:' mapping. Cycle creates the environment load balancer only when the environment's services start with a public container present, so after deploying public containers into an environment without one the tool starts the environment's services itself; if that fails the response carries 'load_balancer_missing'.

'application' hint — pass an app name with NO containers to get recommended per-container config, topology, and caveats, then compose the container list from it. Unknown names return the known names instead of an error.

Workflow: 1) call with preview:true — returns the exact spec that would deploy plus checklist findings, and makes NO changes. 2) Confirm with the user, then call again without preview. Never deploy without explicit confirmation.

Asynchronous and RESUMABLE: containers appear only at the last step, so an early empty list_containers means "still working". Every response carries stack/build/job ids, 'phase', and 'recommended_action'. After a transport failure, repeat the call with the same name/deployment_id and follow 'recommended_action' ('wait', 'resume' with stack_id, or 'nothing_to_do'). A stack_id call is idempotent at every phase.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoStack name; also the default container prefix.
stackNoDefault true: deploy via a Cycle stack. false: create or reuse the image source, import the image, and create the container(s) directly in the environment with no stack or build. Preview and domains work the same in both modes.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
previewNoReturn the spec that would deploy plus checklist findings, making NO changes. Always run this first to confirm intent with the user.
rebuildNoWith stack_id: create and generate a fresh build from the stack's source before deploying, even when a live build exists (picks up new image tags or repo commits). Default false reuses the latest build; only a failed or deleted latest build is replaced.
build_idNoWith stack_id: deploy this specific build instead of the latest. A failed or deleted build is rejected.
redeployNoWith stack_id, when the stack's containers already exist in the environment: update them to the chosen build. At least one of reimage/reconfigure must be true. Containers in the build that are missing from the environment are created either way.
stack_idNoDeploy an EXISTING stack (ID, identifier, resource path, or name) into 'environment' instead of creating one — to reuse a stack the user already has, or to continue an interrupted deployment. Picks up from the chosen build's current state (generate, deploy, containers, DNS) without redoing completed steps. 'containers' is optional here and only used to link domains. Not combinable with stack:false.
containersNoThe containers to deploy. Compose these yourself; use the 'application' hint for guidance.
applicationNoOptional app-hint name, free-form. Curated hints: elasticsearch, mongodb, postgres, redis; other names return no hint (not an error). With no 'containers', returns the hint instead of deploying.
environmentYesTarget environment (must already exist).
wait_secondsNoMax seconds this call blocks on the deploy pipeline (default 60). 0 submits the next step and returns immediately. Every response carries stack/build/job ids and 'phase', so a later call with stack_id continues from wherever this one stopped.
checklist_ackNoChecklist error codes from a preview (e.g. 'listener.ipv4-bind') to deploy in spite of. Only after telling the user exactly what the finding says will break; an unacknowledged error blocks the deploy.
deployment_idNoIdempotency key (lowercase slug) used as the stack identifier; defaults to a slug of name. A repeat call with the same key refuses to create a duplicate stack and reports the existing one with the ids needed to resume. Pass a fresh value to deliberately create a separate stack.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only flag destructive/openWorld/non-idempotent; the description goes far beyond by disclosing that containers are created stopped, that secrets in env/args are stored in plain text, that stateful containers are forced single-instance, that IPv6-only networking and '443:<port>' plus '80:80' are required, and that failures are resumable via stack_id with 'phase'/'recommended_action'. This is exactly the context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded and organized under clear headers (MUST GET RIGHT, Images, Backups, Domains, Workflow), so the agent can scan to what it needs. It is nonetheless very long and overlaps substantially with the already-detailed schema, which costs it the top score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 15-parameter, nested, destructive, asynchronous tool with no output schema, the description covers everything an agent needs: the preview-then-confirm workflow, resume semantics after transport failure, idempotency key behavior, and named response fields (phase, recommended_action, load_balancer_missing). Nothing material is left to inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real cross-parameter meaning the schema does not: when to pass stack_id versus containers, that 'containers' always creates a new stack, how rebuild+redeploy:{reimage:true} rolls out a new image, and that domain forces public:'enable'. The schema descriptions are themselves verbose, so the marginal gain is real but not maximal.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Deploy a multi-container application onto Cycle') and immediately names what it is not: 'For a workload that truly needs a full virtual machine use deploy_virtual_machine.' The container-vs-VM boundary is drawn explicitly, so an agent can route without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when/when-not rules across every branch: new stack vs existing stack_id, stack:false for one-off containers, preview before every real deploy, manage_scoped_variable for secrets, cycle_control_container for starting, list_servers to verify constraint tags. Alternatives are named with the condition that selects them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

deploy_serversDeploy Servers from an Infrastructure ProviderAInspect

Provision new servers (bare metal or VMs running CycleOS) from a connected infrastructure provider into a cluster. THIS COSTS THE USER REAL MONEY, billed by the provider directly — never call it with confirm=true unless the user has explicitly approved the exact plan and its estimated cost in this conversation. Without confirm, the call is a safe dry run: it validates the model, location, and zone against live provider data and returns the plan with the estimated monthly cost increase at the provider for the user to approve. At most 10 servers per call. Use get_deployable_server_models first to find valid models and locations. Deploying returns a Cycle job; track it and the servers coming online with check_server_provisioning.

ParametersJSON Schema
NameRequiredDescriptionDefault
clusterYesCluster to provision the servers into. Clusters are isolation boundaries: servers in different clusters share no network and cannot host the same environments, which makes separate clusters the right tool for dev/staging/prod separation. A cluster the credential cannot reach — never created, or created but out of its ACL scope, which Cycle reports identically — fails the deploy with a permission error; create it first with manage_cluster.
confirmNoMust be true to actually deploy. When false or omitted, the tool only validates the request and returns the deployment plan with its estimated cost so the user can approve it. Set true ONLY after the user has explicitly approved this exact plan and its estimated cost.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
serversYesServers to provision. At most 10 total (including count) per call.
providerYesThe infrastructure-provider integration to deploy through: integration ID, identifier, or vendor. All servers in one call deploy through this provider.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnly=false, destructive=false, idempotent=false, openWorld=true. The description adds the facts that matter most: real money is billed by the provider, confirm=false is a validating dry run that returns a cost estimate, a hard 10-server-per-call cap, permission failures for unreachable clusters, and that the call returns an asynchronous Cycle job. That is materially more than the annotations convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the cost warning, which is exactly what an agent must see first, and each sentence carries information. It is somewhat long, and the confirm approval rule is stated twice (description and confirm parameter), which is deliberate emphasis but mild redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter mutation tool with no output schema and no destructive annotation, the description covers the safety profile, the async return path (Cycle job tracked via check_server_provisioning), the discovery prerequisite, and the cost consequence. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3; the description still earns credit by explaining the confirm dry-run/approve semantics and the 10-server aggregate limit that the schema alone does not spell out. It does not, however, add much on cluster/provider resolution or the nested server fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Provision new servers (bare metal or VMs running CycleOS) from a connected infrastructure provider into a cluster'), and immediately distinguishes the tool's cost-bearing nature. An agent can tell this is infrastructure provisioning rather than an application/VM deploy without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use guidance: dry run without confirm, real deployment only after explicit user approval of the exact plan and cost, plus named prerequisites (get_deployable_server_models) and follow-up (check_server_provisioning). It does not, however, differentiate itself from near-name siblings like deploy_virtual_machine or deploy_application, which an agent could plausibly confuse with this.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

deploy_virtual_machineDeploy a Virtual MachineA
Destructive
Inspect

Deploy a virtual machine onto Cycle. Prefer containers (deploy_application) for ordinary workloads; choose a VM only for hard isolation requirements, a custom OS or kernel (custom modules, non-Linux), or legacy software that cannot be containerized.

Platform rules this tool applies or checks:

  • The environment's cluster must contain a hypervisor-capable server. list_servers reports 'virtualization' per server; the response notes when no live server in the cluster confirms it. When hardware has to be provisioned for VMs, get_deployable_server_models reports 'hypervisor' per model (and takes hypervisor_only).

  • Image: exactly one of base_image or image_url. Call with NEITHER to fetch the live base-image catalogue (version identifiers, supported/UEFI flags) without deploying — do that first unless the user named an exact image; prefer versions marked supported. iPXE and external-volume image sources are not exposed here (a 'base' volume backed by a SAN volume is a different thing and IS supported).

  • Resources are explicit: ram is required, plus exactly one of cores or cpu_pin.

  • Storage: volumes MUST include the boot volume, identifier 'base'; creates without one are refused. Other volumes attach as RAW BLOCK DEVICES the guest must partition, format, and mount — remind the user.

  • Access: Cycle generates a root password at create. It is returned here but retrievable for only ~10 minutes, so relay it to the user promptly (afterwards reconfigure_virtual_machine sets a new one). SSH keys attach at provision time: ssh_keys references existing environment-scoped keys, new_ssh_keys creates and attaches them. Serial-over-SSH console access needs no VM networking: get_vm_console_access mints credentials for the user's own interactive session, and run_vm_command runs a single command and returns its output (in-guest setup like partitioning a volume goes through it).

  • Networking follows the container conventions: IPv6-ONLY private network, hostname defaults to the identifier, public defaults to 'disable', ports map like '443:443'. Point a domain at the VM afterwards with manage_dns_record (records can link to VMs).

  • Placement: constraints.node.tags.all/.any restricts which tagged servers may host the VM — same tag model as deploy_application; useful when only some servers are hypervisor-capable.

Workflow:

  1. If the user hasn't picked an image, call with no base_image/image_url to list the base images.

  2. Call with preview:true — returns the exact create request (read-only lookups only; nothing is created) to confirm with the user. Never create without explicit confirmation.

  3. Call again without preview. The VM is created and, unless start:false, started. The first boot downloads the disk image and can outlast wait_seconds; the job keeps running on Cycle (check list_virtual_machines or get_jobs).

Retries are safe: identifier (default: slug of name) is the idempotency key — a repeat call refuses to create a duplicate and reports the existing VM. Pass a fresh identifier to deliberately create another alongside it.

ParametersJSON Schema
NameRequiredDescriptionDefault
ramNoRAM limit, e.g. '2G'. At least 512M, less than 65G.
nameNoName for the virtual machine. Required except when listing base images.
coresNoNumber of vCPU cores (1-32).
portsNoPort mappings like '443:443'.
startNoStart the VM right after creation (default). false leaves it stopped for cycle_control_virtual_machine.
publicNoPublic network access. Defaults to 'disable'.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
cpu_pinNoPin the VM to specific host cores/ranges, e.g. '0-3' ('x' = the host's max core).
previewNoReturn the exact create request and make NO changes. Always run this first and confirm with the user.
volumesNoVolumes. Must include the boot volume, e.g. {identifier: 'base', size: '10G'}.
hostnameNoPrivate network hostname. Defaults to the identifier.
ssh_keysNoExisting VM SSH keys to attach, by name or 24-char hex ID. Keys are environment-scoped.
image_urlNoURL of a custom disk image to boot from.
os_flavorNoGuest OS flavor for platform preconfiguration. 'windows' adds virtio-win drivers and mounts a drive of provisioning scripts (github.com/cycleplatform/windows-vm-utils) because Windows lacks cloud-init — tell the user to run them for network setup inside the guest.
base_imageNoCycle base-image VERSION identifier from the catalogue this tool returns when called with no image.
identifierNoIdentifier slug and idempotency key; defaults to a slug of name.
constraintsNoRestrict which tagged servers may host this VM (tags live on servers). Confirm tags with list_servers first.
environmentNoTarget environment (must already exist and be live). Required except when listing base images.
allocate_ramNoPreallocate the RAM instead of growing on demand.
new_ssh_keysNoSSH keys to create in the environment from user-supplied public keys, then attach.
wait_secondsNoMax seconds to wait on the start job (default 60). 0 submits and returns immediately.
allocate_coresNoReserve the cores exclusively for this VM. Only with cores.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Far exceeds what the annotations (destructiveHint=true, openWorldHint=true, readOnlyHint=false) convey: the ~10-minute root-password retrieval window, the image-download-on-first-boot that can outlast wait_seconds, the fact the job continues on Cycle after return, preview being read-only, and the raw-block-device caveat for non-boot volumes. The only nuance is that the 'retries are safe / identifier is the idempotency key' framing sits in mild tension with idempotentHint=false, but the description explains the mechanism precisely rather than claiming a false behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Given 23 parameters and a multi-step deployment flow, the length is justified and the bullets are front-loaded with the routing decision first. There is minor redundancy — the preview guidance repeats the schema's own preview description, and the workflow restates rules already stated above.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description still covers what comes back (root password, preview request, job id semantics) and where to check progress (list_virtual_machines, get_jobs). Prerequisites (live environment, hypervisor-capable server, existing SSH keys) and follow-up tools (reconfigure_virtual_machine, manage_dns_record, run_vm_command) are all named.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema coverage is 100%, the description adds cross-field rules the schema cannot express: exactly one of base_image or image_url, calling with neither to fetch the image catalogue, ram required plus exactly one of cores or cpu_pin, and the mandatory 'base' boot volume. These constraints materially change how the agent assembles arguments.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Deploy a virtual machine onto Cycle') and immediately differentiates from the sibling deploy_application by naming the conditions that select the VM path instead. An agent can route between the two tools without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when NOT to use it (ordinary workloads → deploy_application) and enumerates the three legitimate use cases (hard isolation, custom OS/kernel, un-containerizable legacy software). The numbered workflow, the 'never create without explicit confirmation' rule, and the pointer to list_servers/get_deployable_server_models for hypervisor capability leave nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

diagnoseTriage a ProblemA
Read-onlyIdempotent
Inspect

Triage a "something is wrong" report on Cycle. Give it the environment, container, and/or server the user suspects; it fans out the platform's high-signal diagnostic reads concurrently and returns findings ranked critical > warning > info, each naming the follow-up tool (with arguments) to run next. Always start here when the problem is vague, then drill down with the suggested tools.

Problem-class guide — pass as focus, or omit to run every check the scope allows:

  • Container or VM unreachable from its URL → focus "unreachable" with environment (+ container). Covers virtual machines as fully as containers — a LINKED record can point at either. Cross-checks three independent layers so a pass at one never masks a failure at another: (1) ingress CONFIG — LINKED records vs the target's public network and port mappings against the LB's actual controller config, DMZ records included; (2) DNS RESOLUTION — each linked domain against its zone's authoritative nameservers AND public resolvers, which is the only way to see an unpropagated or address-less record; (3) an END-TO-END synthetic HTTP GET at each linked domain (max 5, from the MCP server's network) — probe.http.ok means VERIFIED SERVING, and 502-504 usually means the app is bound to localhost or IPv4-only instead of :: . Also covers LB/gateway/discovery service health, instance readiness, and LB destination errors. A target can be fully healthy yet unreachable — config mismatches are the most common cause; every finding carries a machine-readable code and explains itself.

  • Container stopped, crashing, or restarting → focus "crashloop" with container. Checks state drift, broken instances, restart/healthcheck events, and error-pattern logs.

  • Out of disk / containers can't write → focus "storage" with server (or container — a "no space left" log traces to its host). Checks storage pool and mount utilization, storage-full events.

  • CPU/RAM running low → focus "resources" with server or container. Checks load vs cores, RAM headroom, allocation pressure, instance OOM/throttling.

  • Containers not talking to each other → focus "networking" with environment (+ server for mesh problems). Checks discovery/VPN services, mesh and neighbor events; suggests the neighbor_latency metrics preset.

  • Stack build stuck or failed, deploy never produced containers → focus "build" with stack (+ build id), or with a container deployed from the stack. Reads the build's state and error and every image build it contains, attaching the tail of each failed build log as evidence.

  • LB weirdness / 502s → focus "load_balancer" with environment. Checks LB service state, DNS-record-to-container port routing, disconnect reasons, per-destination response codes, plus the synthetic HTTP probe above. LB telemetry lags several minutes and 404s when absent — that never means the LB itself is gone.

A finding of category "platform" means a diagnostic read itself errored or could not be attempted — run the suggested tool manually. healthy=true means no warning-or-worse findings in the window (default: the last hour). Anything that could NOT be verified (an unprobed ingress path, unreachable DNS) is reported as a warning so healthy is never true by omission. Read-only: this never changes anything. For in-container investigation afterwards, use run_instance_command; for a VM guest, run_vm_command; for live output, capture_stream.

ParametersJSON Schema
NameRequiredDescriptionDefault
buildNoSpecific stack build ID to inspect instead of the latest. Requires stack.
focusNoProblem class to prioritize; omit to run every check applicable to the scope.
stackNoStack to diagnose: name, identifier, or ID. Checks its latest build and that build's images; a container deployed from a stack implies this.
serverNoServer to diagnose: hostname, nickname, or ID. Use for storage-full, resource exhaustion, or host-down suspicions.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
containerNoContainer to diagnose. Best for crash loops and won't-start problems; its environment is diagnosed too.
environmentNoEnvironment to diagnose. Use for app-level problems: unreachable URLs, 502s, containers not talking.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.
lookback_minutesNoHow far back to scan events, logs, and telemetry, 5-1440 minutes.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive), but the description adds substantial behavior beyond them: findings are ranked, each names its follow-up tool with arguments, 'platform' category means a diagnostic read itself failed, healthy=true only means no warning-or-worse in the window, unverified paths are downgraded to warnings so healthy is never true by omission, and the default lookback is the last hour.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The critical guidance is front-loaded in the first two sentences, and the bulleted problem-class guide is structured and scannable. It is long and dense, but for a 9-parameter tool with seven focus classes almost every clause carries decision-relevant content, so the length is largely earned rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description explains the return shape (ranked findings with machine-readable codes, a healthy flag, a conversation_id line to thread back), the read-only guarantee, and how to interpret failure/omission cases. For a broad triage tool with 9 optional params, nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% (baseline 3), but the description adds real meaning: which scope parameter to pair with each focus class, that build requires stack, that a container implies its stack, when to use server vs container vs environment, and that telemetry probes are capped at 5 domains. This is genuinely beyond the schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a precise verb+resource ('Triage a "something is wrong" report on Cycle') and immediately describes the mechanism: concurrent high-signal diagnostic reads with findings ranked critical > warning > info. It also names the sibling tools it defers to for drill-down (run_instance_command, run_vm_command, capture_stream), so an agent can place it against alternatives without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use ('Always start here when the problem is vague, then drill down with the suggested tools') plus a full problem-class guide that maps each focus value to the triggering symptom and the required scope parameter. It even notes the omission case ('omit to run every check the scope allows'), leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_deployable_server_modelsGet Deployable Server ModelsA
Read-onlyIdempotent
Inspect

Discover what infrastructure can be provisioned on this hub. Called with no arguments, it lists the enabled infrastructure-provider integrations. Pass 'providers' to also list the server models each offers — specs, datacenter locations, and estimated pricing — limited to one or two providers per call since some offer hundreds of models; narrow further with category, location, max_monthly_price_usd, and hypervisor_only. category matches each PROVIDER's own naming (values like 'bare-metal', 'general', 'high-frequency', 'high-performance', 'optimized-dedicated', 'vx1' — provider-specific, not a Cycle taxonomy), so list without it first and read the real categories off the results; a filter that matches nothing reports the categories that were actually available. To find hardware that can host virtual machines use hypervisor_only, NOT a category guess. Use it to answer questions like 'what infrastructure providers are enabled on my hub', 'what bare metal servers does offer', 'what servers are available under $50/mo', or 'which models can host virtual machines', and to verify a specific model and location exist before any server deployment is considered. All prices are estimates reported by the provider and are billed by that provider directly to the user's account there, not by Cycle. Read-only: this tool never provisions anything.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum models returned per provider, after sorting by estimated monthly price ascending. Default 50, max 200.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
categoryNoOnly include models whose provider category (or class) matches. Case-insensitive substring match. Values are the PROVIDER's own category names and differ per provider (e.g. 'bare-metal', 'general', 'high-frequency', 'high-performance', 'optimized-dedicated', 'vx1') — they do not describe what a model can run, so do NOT filter on a guessed value like 'virtual-machine'. To find VM-capable hardware use hypervisor_only instead. Call without this filter first and read the categories back off the returned models.
locationNoOnly include models available in a matching provider location. Matches location ID, abbreviation, name, provider code, or geographic city/region/country (case-insensitive substring).
providersNoProvider integrations to list server models for. Each entry matches an integration by ID, identifier, name, or vendor (case-insensitive substring). Omit to just see which infrastructure-provider integrations are enabled. Limit to one or two providers per call — some offer hundreds of models.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.
hypervisor_onlyNoOnly include models that support virtual machines (hardware virtualization). Models whose provider does not report the capability are excluded. Default false.
include_incompatibleNoAlso include models marked incompatible with the Cycle platform. Default false.
max_monthly_price_usdNoOnly include models whose ESTIMATED monthly price is at or below this many US dollars.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly/idempotent annotations, it discloses that prices are provider-reported estimates billed directly by the provider (not Cycle), that a non-matching filter reports the categories actually available, and that the tool never provisions anything. These are behavioral traits the annotations do not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and usage, and every sentence carries information. However it is delivered as one long dense paragraph with some redundancy (the category caveat repeats in both description and schema), so structure could be tighter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter, filter-heavy discovery tool with no output schema, the description covers the decision-making an agent needs: what each filter does, common pitfalls, return contents, and preconditions before deployment. Nothing essential to correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds real meaning: it explains why to limit to one or two providers (some offer hundreds of models), clarifies that category values are provider-specific and not a Cycle taxonomy, and distinguishes category from hypervisor_only. It reinforces rather than replaces the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ('Discover what infrastructure can be provisioned on this hub') and enumerates exactly what is returned — enabled integrations, then server models with specs, locations, and pricing. It is clearly a read-only discovery tool and distinguishable from deployment siblings like deploy_servers or deploy_virtual_machine.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: use hypervisor_only to find VM-capable hardware rather than guessing a category, call without a category filter first to read real categories off results, and use the tool to verify a model/location exists before deployment. It even supplies concrete example questions that map to invocation patterns.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_deployment_statusGet Deployment StatusA
Read-onlyIdempotent
Inspect

Report where deployments (Cycle stacks and their builds) stand — the reconciliation tool when a deploy_application call was interrupted, timed out, or its response was lost. Read-only.

Without arguments it lists recent stacks, newest first, each with its latest build. Pass 'stack' for one stack's detail including recent builds and their errors. Pass 'environment' to also see which of a stack's containers exist there — containers are created only by the final deploy step, so an empty container list on a live build means "generated but not yet deployed into this environment", not "failed".

Each stack carries a derived 'phase':

  • generating: the latest build is still compiling/importing (new, building, importing, verifying, saving) — the pipeline is in flight, nothing is wrong.

  • deploying: the build is being deployed into an environment right now.

  • generated: the build is live; with 'environment' given, no containers from it exist there yet — the deploy step still needs to run.

  • deployed: the build is live and its containers exist in the given environment.

  • failed: the latest build ended with an error (shown on the build).

To deploy a listed stack — reuse it in an environment, or continue an interrupted deployment — call deploy_application with stack_id set to the stack's id: it deploys the latest usable build, or a fresh one with rebuild:true, and picks up from the build's current state. Never resubmit a deployment from scratch after a timeout; check here first.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax stacks to return when listing (no 'stack' given), 1-20.
stackNoStack to inspect in detail: 24-char hex ID, identifier (= deployment_id from deploy_application), resource path ('stack:...'), or name. Omit to list recent stacks.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
environmentNoEnvironment. When given, each reported stack also lists its containers in that environment, distinguishing 'generated' from 'deployed'.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only/idempotent, but the description adds substantial behavioral context beyond them: the derived 'phase' state machine (generating/deploying/generated/deployed/failed) and the counterintuitive container semantics ('empty container list ... means generated but not yet deployed, not failed'). This is exactly the non-obvious behavior an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then the no-arg behavior, then parameter semantics, then a bulleted phase glossary. Long but nearly every sentence earns its place; the phase bullets are dense but necessary given no output schema. Slightly verbose in the closing deploy_application paragraph.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema and a non-trivial derived-state return shape, the description fully specifies what comes back (phases, builds, container lists) and how to interpret edge cases. An agent can call and interpret it correctly without further context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: 'environment' is explained as the switch that distinguishes 'generated' from 'deployed', and 'stack' is tied to the deploy_application deployment_id. It doesn't elaborate on limit/context/conversation_id, which the schema already covers.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Report where deployments ... stand') and immediately frames it as the reconciliation tool for interrupted/timed-out deploy_application calls. This clearly distinguishes it from the sibling deploy_application (write) and status siblings like get_jobs/get_logs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('when a deploy_application call was interrupted, timed out, or its response was lost'), how to escalate to the alternative (call deploy_application with stack_id), and an explicit prohibition ('Never resubmit a deployment from scratch after a timeout; check here first').

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_hubDescribe the Authenticated HubA
Read-only
Inspect

Describe the hub this session is authenticated to, plus the account's membership role and Cycle capabilities in it. Call it only when the user asks about their hub, account, role, or permissions. Do not call it to pre-check whether another tool is allowed: call that tool directly and report its error if Cycle denies it.

ParametersJSON Schema
NameRequiredDescriptionDefault
contextNoWhy are you calling this tool? Briefly describe the user's goal.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=true, so safety and idempotency are covered. The description adds non-obvious behavioral context: it is a reporting/description call, not an authorization gate, and errors from other tools should be surfaced rather than pre-empted. It does not discuss return format, but no output schema exists, so a small gap remains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the purpose, then the when-to-use, then the anti-pattern exclusion. No filler, no repetition of schema or annotation content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless read-only description tool with no output schema, the description delivers purpose, when-to-use, when-not-to-use, and the key anti-pattern. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both parameters (context, conversation_id) are already fully documented, and the description correctly does not restate them. The conversation_id semantics, including the omit-on-first-call rule, live in the schema and are not repeated, so no extra value or harm is added here. Baseline would be 3, but the usage guidance around not using this tool as a pre-check effectively scopes the call, nudging it to 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Describe) and resource (the hub), and enumerates exactly what is returned: the hub, the account's membership role, and Cycle capabilities in it. It is clearly distinguishable from siblings like list_servers or get_jobs, which return different resources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit when-to-use trigger ('only when the user asks about their hub, account, role, or permissions') and an explicit when-not-to-use rule ('Do not call it to pre-check whether another tool is allowed: call that tool directly'). That is a direct routing instruction against a plausible misuse.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_jobsGet Job StatusA
Read-onlyIdempotent
Inspect

Check the status and progress of one or more Cycle jobs by ID. Cycle runs mutating work asynchronously as jobs; tools that submit work (like migrate_instances or cycle_control_container) return job IDs you can track here. Pass wait_seconds to block until every job reaches a terminal state (completed, error, or expired) up to that bound, or 0 to get an immediate snapshot. all_done is true when every requested job has completed successfully.

ParametersJSON Schema
NameRequiredDescriptionDefault
jobsYesJob IDs to check (24-char hex). Pass several to track a multi-instance migration in one call.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
wait_secondsNoMax seconds to wait for the jobs to reach a terminal state. 0 returns the current state immediately.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds real behavioral context the annotations cannot: the terminal-state enumeration (completed, error, expired), the blocking semantics of wait_seconds, and the meaning of all_done.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three front-loaded sentences with no redundancy: purpose first, origin of job IDs second, the blocking/snapshot decision third. Every sentence carries information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-param read tool with no output schema, the description covers the async-work model and the wait behavior well. It could say slightly more about what the per-job result contains, but the core decision inputs are all present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds meaning beyond the schema by explaining that wait_seconds blocks until all jobs reach a terminal state and by spelling out what those terminal states are. It also clarifies the multi-ID use case for tracking a migration in one call.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Check the status and progress of ... Cycle jobs by ID') and distinguishes itself from list/control siblings by framing jobs as the async output of mutating tools. An agent can place it precisely in the lifecycle without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says the job IDs come from tools that submit work (migrate_instances, cycle_control_container) and explains the wait-vs-snapshot choice. It does not name a competing sibling to avoid or state exclusions, but the usage context is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_logsSearch Container LogsA
Read-onlyIdempotent
Inspect

Search aggregated logs from Cycle containers. First stop for crash loops, application errors, and confirming "out of memory" / "no space left" style failures. Use after diagnose points at a specific container or instance, or directly when the user names one.

Scope with environment for environment-wide logs, add container to narrow, add instance to narrow further — the most specific reference wins. Search is optional; without it the newest lines in the range are returned. Use search_type "regexp" with RE2 syntax for patterns like "(?i)error|panic". context_window returns surrounding lines after each match (lines sharing a context_window ID belong to the same match).

Defaults to the last hour. If nothing comes back, widen the range or loosen the search — a container that just started may also not have shipped logs yet. Read-only.

To wait for a line to APPEAR (readiness/migration-done markers like "database system is ready"), pass search plus wait_seconds: the call polls until the first match or the budget elapses, instead of you sleeping between calls. A no-match result after the wait is not failure — repeat the same call to keep waiting.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum log lines to return, 1-500.
searchNoText or pattern to match log messages against. Omit to return the newest lines in the range.
compactNoReduce structured JSON lines (ECS, logrus, zap...) to level, message, and error fields, dropping envelope fields. Non-JSON lines pass through unchanged. Use for chatty JSON loggers such as Elasticsearch.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
instanceNoInstance id or hostname. Narrowest scope; requires container.
containerNoContainer to pull logs from. Narrows the scope to one container.
range_endNoRFC3339 end of the log window. Defaults to now.
environmentNoEnvironment to pull logs from. Broadest scope.
range_startNoRFC3339 start of the log window. Defaults to one hour before range_end.
search_typeNoraw = exact substring match (default). regexp = RE2 pattern, e.g. (?i)error|panic.
wait_secondsNoMax seconds to block waiting for the first matching line, polling every few seconds (0 = return immediately). Requires search; not combinable with range_end. Values above 60 are clamped — longer blocking calls are unreliable through connectors; repeat the call to keep waiting.
context_windowNoLines of surrounding context to include after each match, 0-10.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the readOnly/idempotent annotations: default one-hour window, the 'most specific reference wins' scope precedence, wait_seconds polling semantics with a 60s clamp and the explicit note that a no-match after waiting is not failure, and the context_window grouping behavior. These are exactly the traits an agent needs and none are in the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and usage, then scope, defaults, failure handling, and the wait pattern in a logical order. Dense and information-rich, but four paragraphs covering a 13-parameter tool is at the upper edge of comfortable length; nothing is truly wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers purpose, routing, scope rules, defaults, edge cases (empty results, unshipped logs), and the blocking-wait workflow for a 13-parameter tool with no output schema. An agent has everything needed to call it correctly on the first try.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description still adds real semantics beyond the schema: scope precedence across environment/container/instance, search being optional, RE2 syntax for search_type, and the meaning of shared context_window IDs. Slightly above baseline because it disambiguates how parameters interact.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Search aggregated logs from Cycle containers') and immediately scopes it against siblings by naming the scenarios (crash loops, OOM, disk-full). An agent can distinguish it from get_telemetry, query_metrics, and get_jobs without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit routing: 'First stop for crash loops...', 'Use after diagnose points at a specific container or instance, or directly when the user names one,' plus a fallback ('If nothing comes back, widen the range or loosen the search'). It names the sibling (diagnose) and the condition that selects this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_more_toolsRequest a Missing ToolA
Read-onlyIdempotent
Inspect

Call this whenever the user's request cannot be satisfied by the available tools. Describe the missing tool or capability so it can be considered for a future release, then continue helping the user as best you can with the tools that exist.

ParametersJSON Schema
NameRequiredDescriptionDefault
contextNoWhy are you calling this tool? Briefly describe the user's goal.
capabilityYesWhat tool or capability you needed but could not find, and what the user was trying to accomplish.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world). The description adds genuinely useful context beyond them: the call goes to 'a future release,' so the agent learns it produces no immediate capability and should not block on it, and should keep working with existing tools.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, no filler, with the trigger condition front-loaded and the follow-up behavior second. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description adequately covers the trigger and the post-call instruction, and the schema handles the conversation_id threading. A brief note on what happens to the request (tracked vs. answered immediately) would close the remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already explains 'capability', 'context', and the conversation_id reuse rule. The description adds no parameter-level detail (e.g., how detailed 'capability' should be), so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action (request a missing tool/capability) on a specific resource, and is unmistakably distinct from all infrastructure siblings like deploy_application or list_servers. An agent can tell instantly this is the meta-tool for unsatisfiable requests.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger condition: 'whenever the user's request cannot be satisfied by the available tools.' It also instructs to continue helping afterwards, which implicitly defines the fallback behavior. What's missing is guidance on repeated/batched calls, but the when-to-use is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_telemetrySummarize Resource TelemetryA
Read-onlyIdempotent
Inspect

Summarized resource telemetry from Cycle, with anomaly flags. Pick a target:

  • target=instance (needs environment + container, plus instance when there are several): per-instance CPU usage and throttling, memory usage vs limit and OOM fail count, swap, network errors, process count. Use for one hot or crashing container.

  • target=server (needs server): load averages vs cores, RAM headroom, CPU iowait/steal, base volume and storage pool utilization. Use for "server out of resources / disk" suspicions.

  • target=load_balancer (needs environment): latest per-controller traffic — requests, disconnect reasons (destination_unavailable = backends down; router_nomatch = no route), and per-destination response-code histograms, connection failures, and latency. Use for 502s and unreachable-URL problems. Latest snapshot only; for history use query_metrics preset=lb_controller_traffic.

Returns min/avg/max/latest summaries rather than raw series. Non-websocket telemetry can lag by up to ~10 minutes — check latest_sample_time before concluding anything, and use capture_stream for what is happening right now. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
serverNoServer hostname, nickname, or ID. Required for target=server.
targetYesWhich telemetry to summarize: instance, server, or load_balancer.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
instanceNoInstance id or hostname for target=instance. Optional when the container has a single instance.
containerNoContainer. Required for target=instance.
range_endNoRFC3339 end of the window (instance and server targets). Defaults to now.
environmentNoEnvironment. Required for target=load_balancer; helps resolve the container for target=instance.
range_startNoRFC3339 start of the window (instance and server targets). Defaults to one hour before range_end.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/idempotent/destructive=false, so safety is covered; the description goes further by disclosing staleness (non-websocket telemetry can lag ~10 minutes), instructing to check latest_sample_time, distinguishing latest-snapshot-only from historical data, and noting it returns min/avg/max/latest summaries rather than raw series.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the summary statement, then a scannable bullet per target, then global caveats. Dense but every sentence carries routing or freshness information; the enumerated metric lists make it longer than strictly necessary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter, no-output-schema tool, it covers target selection, per-target requirements, staleness, and return granularity. Remaining gaps are minor, e.g. it does not restate the conversation_id chaining protocol, though that is documented in the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, giving a baseline of 3, but the description adds cross-parameter conditional logic the schema cannot express: which fields are required per target (instance needs environment + container, plus instance when several exist; server needs server; load_balancer needs environment). That materially improves correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (summarize) and resource (resource telemetry from Cycle) and immediately enumerates the three target modes with the exact metrics each returns. An agent can distinguish this from query_metrics and capture_stream without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit routing per target: 'Use for one hot or crashing container', 'Use for server out of resources / disk suspicions', 'Use for 502s and unreachable-URL problems'. It also names alternatives and their trigger conditions — query_metrics preset=lb_controller_traffic for history and capture_stream for real-time.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_vm_console_accessGet Serial Console Access to a VMA
Destructive
Inspect

Mint serial-over-SSH credentials so the USER can open their own interactive console session on a virtual machine, and return the exact ssh command to run. Use this when a person wants to drive the console themselves — poke around, watch a boot, run an interactive installer or a TUI. To run a specific command and get its output back, use run_vm_command instead; do NOT take these credentials and try to drive an interactive shell yourself.

What comes back: the ssh command, the connection address, the connection secret (used at ssh's own password prompt), and the token's expiry. The credentials reach the VM's serial console through the Cycle gateway — no VM networking required, so this works even with public networking disabled or the guest's network broken.

Logging IN to the guest is separate: past ssh, the console lands at the guest's own login prompt. The response includes the generated root password when it is still retrievable (~10 minutes after creation); after that the user needs their own guest credentials or an attached SSH key.

Relay the credentials to the user plainly so they can paste the command. They are short-lived and they grant root-capable console access — treat them as a secret, do not write them into files, commits, or issues.

action:"revoke" expires ALL serial-over-SSH tokens for the VM, immediately disconnecting every active console session on it (including a session the user is in the middle of, and anyone else's). It is a mutation: confirm with the user before calling it.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionNogrant (default) mints credentials for a new console session. revoke expires ALL of the VM's console tokens, disconnecting every active session — confirm with the user first.grant
contextNoWhy are you calling this tool? Briefly describe the user's goal.
environmentYesEnvironment the VM lives in.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.
virtual_machineYesVM to get console access to.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: it explains that credentials reach the console via the Cycle gateway (works with networking disabled), that returning root password is only retrievable ~10 minutes, and critically that action:'revoke' destroys ALL tokens for the VM and disconnects every active session including a mid-session user, with an instruction to confirm first. This is rich disclosure for a destructive, non-idempotent mutation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and sibling routing, then layered detail (what comes back, login nuance, secret handling, revoke semantics) in distinct, purposeful paragraphs. Despite its length, the security-critical nature of minting root-capable credentials justifies every section.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description enumerates the return payload (ssh command, address, secret, expiry, root password) and covers the two-mode behavior (grant vs revoke), credential handling, and login separation. An agent has everything needed to call and relay results correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the schema already defines 'action' enum semantics, so the baseline is 3. The description reinforces the meaning of action:'revoke' and clarifies what the grant returns, but adds no parameter syntax or format detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource combination ('Mint serial-over-SSH credentials' for a virtual machine) and the scope of use (interactive console session driven by the user). It explicitly distinguishes itself from the sibling run_vm_command, so an agent can select between the two without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the when ('a person wants to drive the console themselves — poke around, watch a boot, run an interactive installer or a TUI') and the alternative ('To run a specific command and get its output back, use run_vm_command instead'), plus a hard exclusion ('do NOT take these credentials and try to drive an interactive shell yourself').

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_containersList ContainersA
Read-onlyIdempotent
Inspect

List containers in the Cycle hub, optionally scoped to one environment. Use to discover what exists, check current state, or disambiguate when multiple containers share a name. Prefer passing environment to narrow results. If has_more is true, call again with page = next_page.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNoPage number; use next_page from a previous response.
limitNoResults per page.
stateNoOnly return containers currently in this state.
searchNoFree-text match against container names and identifiers, e.g. 'api'.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
environmentNoEnvironment to scope to: identifier slug or 24-char hex ID. Omit to list across the whole hub.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, open-world. The description adds meaningful behavioral detail beyond that: the pagination contract ('If has_more is true, call again with page = next_page'), referencing response fields not exposed in any output schema. Useful, but no mention of ordering, result caps, or default scope richness.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, front-loaded with purpose then usage then the operational pagination tip. Every sentence carries weight; there is no filler or repetition of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool whose annotations cover the safety profile and with no output schema, the description covers purpose, filtering, and pagination adequately. It could say more about return shape or default result caps, but nothing critical for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so every parameter (page, limit, state, search, context, environment, conversation_id) is already fully documented, including the 'omit to list across the whole hub' note. The description's environment preference and page=next_page hints add only marginal meaning over the schema. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (containers) plus the scope (in the Cycle hub, optionally scoped to one environment). An agent can distinguish this from list_environments, list_servers, and list_virtual_machines by the resource noun alone. No ambiguity about what the tool returns.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear usage contexts: discover what exists, check current state, or disambiguate same-named containers. Also advises preferring the environment parameter to narrow results. However it never names an explicit alternative tool or a when-not-to-use condition, so it stops short of the top tier.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_dns_zonesList DNS Zones and RecordsA
Read-only
Inspect

List the hub's DNS zones (origin, hosted, state), or the records of one zone when 'zone' is given. Hosted zones are served by Cycle directly; non-hosted zones hold records that configure the load balancer/TLS but require the user to manage the actual DNS externally. Read-only. Use manage_dns_record to create or repoint records.

ParametersJSON Schema
NameRequiredDescriptionDefault
zoneNoOptional zone to fetch records for: origin (e.g. 'example.com') or 24-char hex ID. Omit to list all zones.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered and the 'Read-only' line is reinforcement. The description adds real semantic context the annotations don't: hosted zones are served by Cycle directly, while non-hosted zones only configure LB/TLS and require external DNS management. That distinction materially affects interpretation of returned data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core operation and scope before the explanatory clause, and every sentence carries information. The hosted/non-hosted sentence is longer but earns its place by clarifying data meaning; overall tight, though slightly dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with no output schema and annotations covering the safety profile, the description supplies the needed operational and semantic context. Return-format details are absent but the absence of an output schema means some expectation-setting could still help; otherwise complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents 'zone', 'context', and 'conversation_id' fully, including the origin/24-char hex ID format. The description restates the zone dual-behavior but adds no syntax or format detail beyond the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (list) and resource (DNS zones/records), plus the scope split: all zones by default, or one zone's records when 'zone' is given. It also distinguishes itself from manage_dns_record, which handles writes, so an agent can route correctly without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly explains the two usage modes (omit 'zone' to list all, supply it to list one zone's records) and names the write alternative manage_dns_record for create/repoint. Lacks explicit when-not-to-use framing beyond the write redirect, so a 4 rather than 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_environmentsList EnvironmentsA
Read-onlyIdempotent
Inspect

List all environments in the authenticated Cycle hub, with their current state. An environment must be 'live' for its containers and services (discovery DNS, load balancer, scheduler) to run — deploying into a non-live environment yields containers with no working networking.

ParametersJSON Schema
NameRequiredDescriptionDefault
contextNoWhy are you calling this tool? Briefly describe the user's goal.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds non-obvious domain behavior: a non-live environment yields containers with no working networking, which is a real diagnostic constraint an agent would not get from the structured fields alone. It does not describe result volume or paging, but for a simple list call that is minor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action and scope; the second sentence is longer but delivers the one non-obvious fact about environment state. No filler or restatement of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by stating what is returned (each environment's current state). For a two-parameter, zero-required, read-only list tool with full annotation coverage, this is essentially complete; only the exact fields per environment are unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters (context, conversation_id) carry their own explanations, including the first-call/omit-then-echo convention for conversation_id. The description adds nothing about parameters, which is the expected baseline when the schema already does the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("List all environments in the authenticated Cycle hub") plus the data returned ("with their current state"). The scope (authenticated hub) helps separate it from mutating siblings like create_environment, delete_environment, or cycle_control_environment, though it never names an alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: an agent reading "list all environments ... with their current state" can infer this is the discovery/inspection call preceding environment mutations. However, there is no explicit when-to-use or when-not-to-use guidance, nor any pointer to alternatives such as get_hub or cycle_control_environment.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_image_sourcesList Image Sources and ImagesA
Read-only
Inspect

List the hub's image sources — where container images are pulled or built from (Docker Hub target, private registry, Dockerfile repo) — with per-state image counts and their newest images. Pass 'source' to inspect one source with its full image history. Read-only.

Stacks do NOT use image sources: a stack build produces its own images (source type 'stack-build'), which are owned by the stack and vanish with it. Image sources come from deploy_application with stack:false (one per Docker Hub target, reusing an existing match) or from the dashboard, so a source with no live image is usually left over from a deleted one-off container.

ParametersJSON Schema
NameRequiredDescriptionDefault
imagesNoInclude each source's newest images (default true). false returns sources and counts only.
searchNoFilter the listing by name or identifier substring. Ignored with 'source'.
sourceNoOne image source to inspect: name, identifier, resource path ('image-source:...'), or 24-char hex ID. Omit to list all.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered and the description's 'Read-only' is largely redundant. It still adds genuine behavioral context the annotations lack: sources originate from deploy_application with stack:false or the dashboard, and an orphaned source is usually a leftover from a deleted one-off container, which helps the agent interpret results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Content is front-loaded with the core purpose and the 'source' switch, and the stack/origin caveat earns its place by preventing misinterpretation of empty sources. The final sentence is dense and multi-clause, but every clause carries disambiguating information rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of describing returns and does so: per-state image counts, newest images, and full image history when 'source' is set. Combined with the origin/lifecycle explanation, an agent has enough to call and interpret this correctly, though pagination or result-size behavior is unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds cross-parameter semantics the schema does not spell out: 'source' switches the tool into full-image-history mode, and the images/search interplay is framed as mode selection rather than independent filters. This goes modestly beyond the per-parameter schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('List the hub's image sources') and immediately defines what an image source is, including the three concrete origins (Docker Hub target, private registry, Dockerfile repo). It explicitly carves out what this tool is NOT about ('Stacks do NOT use image sources'), which disambiguates it from the stack-related siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear conditional usage: pass 'source' to inspect a single source with full image history, omit to list all. It also tells the agent when results are not stack-related and why a source might be stale. It stops short of naming a sibling alternative (e.g. manage_image_source) for acting on a source, so it is strong but not fully prescriptive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_scoped_variablesList Scoped VariablesA
Read-onlyIdempotent
Inspect

List an environment's scoped variables — environment-level values/secrets injected into containers at runtime. Shows each variable's scope (global or specific containers), the containers it currently reaches, its delivery method(s) — environment variable, file, and/or internal API unix socket — and whether its value is raw or fetched from a URL at container start. Raw values are omitted unless include_values is true.

ParametersJSON Schema
NameRequiredDescriptionDefault
contextNoWhy are you calling this tool? Briefly describe the user's goal.
environmentYesenvironment to inspect
include_valuesNoinclude raw source values in the output; defaults to false so secret values stay out of context
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description goes beyond them usefully: raw values are withheld by default, delivery methods can be env var/file/unix socket, and values may be raw or URL-fetched at container start. It does not mention pagination or result size for what could be a long list.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense paragraph that front-loads what the tool is and then enumerates the returned fields; every clause carries information rather than filler. It is slightly long but no sentence is redundant with the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description compensates by enumerating the returned fields (scope, reached containers, delivery methods, raw vs URL-fetched), which is exactly what an agent needs to interpret results. Minor gaps remain around ordering, size limits, and the conversation_id/context tracking parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters are already documented, including include_values' default and its secret-exposure rationale. The description's restatement of the include_values behavior adds no syntax or format detail beyond the schema, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource ('List an environment's scoped variables') and immediately defines the domain jargon — environment-level values/secrets injected into containers at runtime. It is clearly distinguishable from the write counterpart manage_scoped_variable by the verb alone, but it never names that sibling explicitly, so it falls short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the tool name and the read-only framing, and the include_values default hints at a secret-safe inspection use case. However, there is no explicit when-to-use statement, no mention of when to reach for manage_scoped_variable instead, and no prerequisite guidance for the required environment parameter.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_serversList ServersA
Read-onlyIdempotent
Inspect

List the physical/virtual servers making up your Cycle infrastructure, grouped by cluster (Cycle uses clusters for physical isolation). Use to see what hardware exists, which cluster it belongs to, and its provider and state. hostname and nickname are matched as case-insensitive substrings, so a partial hostname is enough to find a specific server. cluster and state are exact filters applied by the Cycle API.

Each server also reports constraints: constraints.tags is the list of tags on that server, and constraints.allow.pool is true when the server accepts containers that specify no tags. Placement is driven by these SERVER tags — a container restricts itself to a subset of servers via its own config.deploy.constraints.node.tags (see deploy_application), and a Cycle cluster does NOT narrow placement on its own. Read tags here to confirm a tag exists before referencing it in a container constraint; note that if every server has allow.pool true, an unconstrained container can land on any of them.

Each server also reports 'virtualization': true means it can host virtual machines (see deploy_virtual_machine); absent means the capability is unknown (node stats unavailable).

EVERY cluster visible to the credential is listed, including clusters that hold no servers at all (count 0, empty servers list) — a cluster exists on Cycle independently of any hardware in it, so an empty cluster here is a real cluster awaiting servers, not a missing one. This is the authoritative list of what clusters exist.

ParametersJSON Schema
NameRequiredDescriptionDefault
stateNoOnly return servers currently in this state.
clusterNoExact cluster identifier to scope to. Clusters are Cycle's unit of physical isolation.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
hostnameNoCase-insensitive substring of a server hostname. A partial hostname is enough to locate a specific server.
nicknameNoCase-insensitive substring of a server's nickname (custom identification name).
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive), and the description adds substantial context beyond them: substring vs exact filter semantics enforced by the API, that empty clusters are still returned (count 0), that clusters do NOT narrow placement, and that an absent 'virtualization' field means unknown capability. That is real behavioral disclosure an agent could not infer from structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose and filter semantics, then moves to secondary details (constraints, virtualization, empty clusters) in separate paragraphs. The tag-placement paragraph is long but each sentence carries distinct operational meaning; the only mild cost is overall length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and six optional parameters, the description compensates well: it enumerates the returned fields (provider, state, constraints.tags, constraints.allow.pool, virtualization) and explains the edge case of empty clusters. An agent has everything needed to call and interpret this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds genuine meaning: hostname and nickname are case-insensitive substring matches (a partial hostname suffices) while cluster and state are exact filters applied by the API. It also clarifies that a Cycle cluster does not itself restrict placement, which changes how an agent should interpret the cluster parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource+scope: 'List the physical/virtual servers making up your Cycle infrastructure, grouped by cluster.' That distinguishes it cleanly from siblings like list_containers, list_virtual_machines, and list_environments, and the follow-up sentence states exactly what each entry reveals (hardware, cluster, provider, state).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear usage context ('Use to see what hardware exists, which cluster it belongs to, and its provider and state') and a concrete workflow trigger — read tags here to confirm a tag exists before referencing it in a container constraint. It cross-references deploy_application and deploy_virtual_machine but never states when NOT to use this tool versus a sibling lister.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_virtual_machinesList Virtual MachinesA
Read-onlyIdempotent
Inspect

List virtual machines on the hub with their current state, image, resources, and assigned public IPs. Filter by environment and/or state. Use this to check on a VM (e.g. whether a deploy or start finished — the disk image download on first boot can take a while) and to find a VM's ID before controlling it. VM names are not guaranteed unique; disambiguate with the identifier or ID. Created VMs live in an environment and share its IPv6-only private network with containers there.

ParametersJSON Schema
NameRequiredDescriptionDefault
stateNoOnly return VMs currently in this state.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
environmentNoOnly list VMs in this environment. Omit for all environments.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds genuinely useful behavior beyond them: VM names are not unique (disambiguate via identifier/ID), VMs share an IPv6-only private network with containers in the environment, and first-boot image download can make things appear stuck.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose then filters then usage guidance, and each sentence carries information. It is slightly dense with the VM/container networking aside, but nothing is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly summarizes the returned fields, and it covers filtering and the name-uniqueness caveat. It stops short of mentioning pagination or result size behavior for a list tool that could return many VMs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents state, environment, context, and conversation_id. The description restates the environment/state filtering but adds no syntax or format detail beyond the schema, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List virtual machines on the hub') and enumerates what is returned (current state, image, resources, assigned public IPs). This clearly distinguishes it from siblings like list_containers and list_servers, which cover different resource types.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete usage contexts: checking whether a deploy or start finished, and finding a VM's ID before controlling it (i.e., before cycle_control_virtual_machine or run_vm_command). It does not explicitly name alternative tools or state when-not to use it, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_clusterCreate or Update a ClusterA
Destructive
Inspect

Create or update a cluster — Cycle's unit of physical isolation, grouping servers that share a private network and can host the same environments. Servers in different clusters share no network, which makes separate clusters the right tool for dev/staging/prod separation. Use list_servers to see existing clusters and what they hold.

Creating a cluster is a hub-level change but provisions no infrastructure and costs nothing on its own; servers are added separately with deploy_servers.

Cycle deliberately returns the same 403 permission error for a cluster that does not exist and for one the credential has no ACL access to, so neither this tool nor any other can tell those apart. That means a create may fail because the cluster already exists out of view — the error will say so, and the fix is access to it, not a second create.

update only changes non_essential; Cycle exposes no rename, and monitoring tiers are configured elsewhere.

Called with no action, this tool just ensures the cluster exists: it creates a missing one, and reports an existing one unchanged (action "none") rather than trying to update it. Pass action explicitly when you specifically mean to create or to modify.

Always call with preview:true first — it validates and reports what would change, touching NOTHING. Get explicit user confirmation, then call again without preview.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionNo'create' makes a new cluster, 'update' modifies an existing one (non_essential is the only updatable field). Omit to just ensure the cluster exists: it creates when no cluster with that identifier is visible, and reports the existing one unchanged when there already is one. Pass 'create' explicitly to get an error instead when it already exists.
clusterYesCluster to create or update: identifier for create, identifier or 24-char hex ID for update.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
previewNoValidate and report what would change without touching anything. Always call with preview:true first, confirm with the user, then call again without it.
non_essentialNoMark the cluster non-essential, excluding its resources from certain metrics and summaries by default. The only field Cycle allows updating on an existing cluster.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A4.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=false, and openWorldHint=true, so the risk profile is covered. The description does add real value beyond them: the cost/scope note ('provisions no infrastructure and costs nothing on its own'), the deliberate 403 ambiguity between missing and inaccessible clusters, and the fact that only non_essential is updatable. However it never says what 'update' destroys or how the non_essential toggle affects existing metrics, so it does not fully exploit the room annotations leave open.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Five short paragraphs, each carrying distinct load: identity/scope, cost and isolation model, the 403 trap, update limits, default semantics, and the preview workflow. Purpose is front-loaded in the first clause, and nothing is padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, non-idempotent mutation with 6 parameters and no output schema, the description covers scope, cost, error semantics, the updateable field, the no-action default, and the mandatory preview/confirm cycle. There is no return-value explanation needed and none is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description goes beyond the schema by framing the omit-action default in practical terms ('just ensures the cluster exists... reports an existing one unchanged') and by connecting the 403 quirk to why a bare create can fail, which is the semantic an agent is most likely to get wrong when choosing action values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource pair ('create or update a cluster') and immediately defines what a cluster is ('Cycle's unit of physical isolation, grouping servers that share a private network'). It distinguishes itself from siblings by routing the agent to list_servers for inspection and deploy_servers for adding servers, so the agent can place this tool without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use context (separate clusters for dev/staging/prod), names the sibling tools to use instead for other steps (list_servers, deploy_servers), and rules things out ('Cycle exposes no rename, and monitoring tiers are configured elsewhere'). The mandatory preview-then-confirm workflow is stated as a hard rule, leaving no inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_dns_recordCreate or Update a DNS RecordA
Destructive
Inspect

Create or update (repoint) a single DNS record in a Cycle zone. Use list_dns_zones to inspect zones and records first. This tool cannot delete records — use delete_dns_record for that.

Addressing: pass a full 'domain' (e.g. 'app.example.com' — the covering zone and record name are derived; apex = '@'), or 'zone' plus 'name'. For update, 'record_id' (with 'zone') disambiguates when several records share a name.

Types: a/aaaa (value = IP), cname/alias/ns (value = target domain), mx (value = mail host, priority), srv (value = target, port, priority, weight), txt (value = text), caa (tag + value), linked — a Cycle-managed pointer at a container, tagged deployment, or virtual machine via the 'linked' object; Cycle wires up the IPs, load-balancer routing, and TLS certificates automatically (the environment load balancer must be running for traffic to flow).

Semantics: create NEVER overwrites — an existing record with the same name and type is an error; use action:update to repoint it. update REPLACES the record's whole type payload with what you pass, and records cannot be renamed. In a non-hosted zone, records still configure the load balancer/TLS but the user must manage the zone's public DNS externally.

Always call with preview:true first — it validates and returns a from/to diff, changing NOTHING. Get explicit user confirmation, then call again without preview.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagNocaa tag, e.g. 'issue'.
nameNoRecord name within the zone ('@' for apex, '*' for wildcard).
portNosrv port (required for srv).
typeYesRecord type. Also the payload written on update.
zoneNoZone origin (e.g. 'example.com') or 24-char hex ID.
valueNoRecord data: IP for a/aaaa, target domain for cname/alias/mx/ns/srv, text for txt, value for caa. Not used for linked.
actionYescreate adds a new record; update replaces an existing record's type payload (e.g. repointing a linked record).
domainNoFull domain, e.g. 'app.example.com'; the covering zone and record name are derived. Alternative to zone+name.
linkedNoTarget for type 'linked'. Exactly one of: container (optionally + deployment_tag to follow that tagged deployment), or vm.
weightNosrv weight.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
previewNoValidate and return the from/to diff, changing NOTHING. Always run this first; confirm with the user, then call again without preview.
priorityNomx (required) / srv priority.
record_idNoExact record ID for update; disambiguates when several records share a name.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the destructiveHint/readOnlyHint annotations, it discloses create-never-overwrites, update-replaces-whole-payload, records-cannot-be-renamed, the preview diff contract, non-hosted-zone DNS implications, and that linked records auto-wire IPs/LB/TLS with a load-balancer-running prerequisite. This is rich behavioral context the annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is long, but front-loaded with purpose and organized into labeled sections (Addressing, Types, Semantics, preview workflow) that a reader can scan. A few sentences restate schema content, so it is not maximally tight, but for a 15-parameter nested tool the length is defensible.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex mutation tool with 15 parameters, a nested linked object, and no output schema, the description covers addressing, per-type payloads, create/update semantics, deletion boundary, and the mandatory preview-confirm flow. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: the domain-vs-zone+name addressing derivation (apex='@'), record_id disambiguation, and a per-type value format table (mx priority, srv port/weight, caa tag). Some of this overlaps the schema's own value/linked descriptions, which keeps it from a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence names a specific verb pair (create/update), the resource (a single DNS record), and the scope (in a Cycle zone). It also explicitly distinguishes itself from siblings list_dns_zones and delete_dns_record, so an agent can route correctly without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states when to use each alternative: list_dns_zones to inspect first, delete_dns_record for removal, action:update when a record already exists, and the mandatory preview-then-confirm workflow. When-to-use, when-not-to-use, and the correct alternative are all present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_external_volumesManage External VolumesA
Destructive
Inspect

Manage external volumes: storage outside a server's local disks (SAN over iSCSI, Ceph RBD, AWS EBS) that a container mounts via deploy_application volumes.external. A volume belongs to one cluster and location and lists the servers allowed to mount it.

list and describe are reads. scan has an integration report the volumes it exposes to a cluster; they then appear in list with state new/ready. create registers one volume: the named servers fix its cluster and location, the integration authenticates to the storage, and source names the device (lun; image+pool; volume_id, or create_size to allocate a new EBS volume). set_servers replaces the mountable server list. delete is refused while a container uses the volume; delete_source_device also destroys the backing device.

Cycle's catalog decides which source types exist, which are creatable, and which attachment types and modes each supports; it is returned in 'sources' on scan and create. Modes count concurrent attachers (single-instance, single-node, multi-node) and access (writer, read-only); multi-node-writer needs a cluster-aware filesystem.

scan, create, set_servers, delete: preview:true first, confirm with the user, then call again without preview.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNocreate: default single-instance-writer.
nameNocreate: display name.
actionYes
sourceNocreate: device details for the source_type.
unusedNolist: only volumes no container or VM is using.
volumeNoExternal volume for describe, set_servers, delete: name, identifier, or ID.
clusterNoCluster identifier. Filters list; required for scan.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
previewNoResolve everything and report what would be submitted, making NO changes.
serversNoServers by hostname, nickname, or ID. Required for create and set_servers; optional scan scope.
attachmentNocreate: default filesystem; block presents a raw device.
identifierNocreate: slug; defaults from name.
create_sizeNocreate: allocate a new device of this size, e.g. '100G'.
descriptionNocreate: free-text note.
integrationNoStorage integration by name, identifier, vendor, or ID. Required for scan and create.
source_typeNoRequired for scan and create.
wait_secondsNoMax seconds to wait on the job. 0 returns once it is accepted.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.
delete_source_deviceNodelete: also destroy the device on the storage system. Irreversible.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only supply the coarse destructive/readOnly/idempotent flags; the description adds much more: delete is refused while a container uses the volume, delete_source_device irreversibly destroys the backing device, scan returns an integration report whose volumes appear in list with state new/ready, and the catalog returned in 'sources' governs creatable source types and attachment/mode support. The preview:true protocol is also spelled out. This is rich disclosure beyond the annotations and is consistent with destructiveHint=true.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but well-organized into three paragraphs covering the resource, the action semantics, and the catalog/mode rules, with the preview protocol called out last. Every sentence carries information and the resource definition is front-loaded, though the block is long enough that an agent must read carefully rather than skim.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 19-parameter tool with no output schema and nested objects, the description covers action semantics, deletion safety, the catalog returned on scan/create, and the confirmation workflow, which is close to complete. Minor gaps remain around job/wait behavior (wait_seconds) and result states beyond 'new/ready', but the essential calling information is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 95%, so the baseline is 3, but the description genuinely adds meaning the schema lacks: it explains that 'source' names the device by lun / image+pool / volume_id / create_size, that modes count concurrent attachers and access and that multi-node-writer needs a cluster-aware filesystem, and that set_servers replaces (not appends to) the mountable server list. These clarify semantics the bare enum descriptions do not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ('Manage external volumes') and then defines the resource precisely: storage outside a server's local disks (SAN over iSCSI, Ceph RBD, AWS EBS) mounted via deploy_application volumes.external. It enumerates every action (list, describe, scan, create, set_servers, delete) and what each does, so the tool is unmistakable against siblings like manage_image_source or deploy_application.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly separates reads (list, describe) from scan/create/set_servers/delete and routes container mounting to deploy_application volumes.external, giving real context for when each action applies. It also prescribes a preview-then-confirm-then-call workflow for the four mutating actions. It stops short of naming an explicit alternative tool to use instead of this one for any case, so not a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_image_sourceEdit an Image SourceA
DestructiveIdempotent
Inspect

Edit an image source's name, description, identifier, builder integration, or registry credentials. The origin itself (which registry or repo, which target) cannot be changed — create a new source for that.

Credentials pass through this conversation in plain text, so prefer a webhook (oci-registry) or leave auth to the dashboard when the user has a choice; never invent or reuse a token the user did not supply for this purpose. The response reports the auth mode, never the secret.

Workflow: call with preview:true to see the field-by-field changes, confirm with the user, then call again without preview. An update takes effect on the next import or build; running containers are untouched.

ParametersJSON Schema
NameRequiredDescriptionDefault
authNoReplace the registry credentials. docker-hub: username + token. docker-registry: username + token or password. oci-registry: username + token, or webhook_url. clear:true removes credentials.
nameNoNew display name.
sourceYesImage source to edit: name, identifier, resource path ('image-source:...'), or 24-char hex ID.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
previewNoShow the changes without applying them. Always run this first.
identifierNoNew slug identifier.
descriptionNoNew description; empty string clears it.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.
builder_integrationNoHub integration (ID or identifier) that performs builds for this source.

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, openWorldHint=true, and readOnlyHint=false. The description adds meaningful context beyond that — credentials travel in plain text, tokens must never be invented or reused, and the update only takes effect on the next import/build while running containers are untouched — but does not cover permissions/ownership requirements or error behavior, keeping it at a solid 3 rather than higher.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short paragraphs, well front-loaded with the mutable-field list followed by safety and workflow guidance. Every sentence earns its place, though the response-format and workflow notes could be tightened slightly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully tells the agent what the response contains (auth mode, not the secret) and how the change propagates (next import/build). For a 9-param destructive mutation with a nested auth object it is nearly complete, missing only permissions/error detail.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, including the nested auth object with per-field descriptions, so the schema carries the bulk of parameter meaning and the baseline is 3. The description reinforces the auth semantics (response reports auth mode, never the secret) and the preview convention, but adds little syntax or format detail beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Edit) and resource (image source), enumerates the mutable fields (name, description, identifier, builder integration, registry credentials), and explicitly bounds the operation by naming what cannot change (the origin/registry/target). It also implicitly distinguishes itself from list_image_sources and delete_image_source by naming the edit surface.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when-to-use workflow (call with preview:true, confirm with the user, then call again without preview) and a when-not-to alternative (create a new source to change the origin). It also advises preferring webhook auth or dashboard-managed auth when the user has a choice, which is actionable guidance rather than vague context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_scoped_variableCreate or Update a Scoped VariableA
Destructive
Inspect

Create or update a scoped variable — an environment-level value or secret injected into containers at runtime. Use list_scoped_variables first to see what exists. This tool cannot delete variables.

A variable has a scope (which containers receive it), access (how it is delivered: env var, file, and/or the in-container internal API — the most secure option), and a source (a raw stored value, or a URL fetched when the container starts, with optional auth for third-party secret services). Any sensitive raw value — a password, API key, token, private key, certificate, or connection string carrying credentials — MUST be created with source.secret:true so Cycle treats it as a secret and this server never returns it.

create never overwrites — an existing variable with the same identifier is an error. update addresses the variable by environment + identifier (or variable_id when identifiers collide), replaces each section you pass (scope, access, source) wholesale, leaves omitted sections unchanged, and renames via new_identifier. Env-variable and file delivery apply when a container (re)starts; running containers keep the old value until restarted.

Always call with preview:true first — it validates and returns the from/to diff plus the containers reached, changing NOTHING. Get explicit user confirmation, then call again without preview.

ParametersJSON Schema
NameRequiredDescriptionDefault
scopeNoWhich containers receive the variable. With neither global nor a container list it reaches none.
accessNoHow the value is delivered; set any combination.
actionYescreate adds a new variable; update modifies an existing one.
sourceNoThe variable's value. Required for create.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
previewNoValidate and return the from/to diff plus reached containers, changing NOTHING. Always run this first.
identifierNoVariable identifier (a-zA-Z0-9 and dashes). Required for create; addresses the variable on update.
environmentYesEnvironment holding the variable.
variable_idNoExact variable ID; disambiguates an update when several variables share an identifier.
new_identifierNoupdate only: rename the variable to this identifier.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only say destructive/non-idempotent/open-world; the description adds substantial context beyond that: create errors on an existing identifier, update replaces each passed section wholesale while leaving omitted sections unchanged, env-var and file delivery only apply on container (re)start so running containers keep the old value, secret values are never returned, and preview changes nothing. These are exactly the behavioral traits annotations cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Long but dense and well-organized into purpose, variable model, secret rule, action semantics, and preview workflow, with the purpose and the preview rule front-loaded. Almost every sentence carries operational weight; a small amount of the scope/access/source exposition overlaps with the schema, keeping it short of a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 11-parameter nested tool with no output schema, the description supplies the full picture an agent needs: the object model, create-vs-update semantics, restart timing, secret handling, and the mandatory preview-confirm-commit loop. Nothing required to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description earns above baseline by explaining the conceptual model behind the nested objects — scope (which containers receive it), access (env var/file/internal API), source (raw vs URL-fetched with optional auth) — plus the secret:true requirement and new_identifier rename semantics. It adds meaning beyond field-level descriptions, though it stops short of covering every parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb pair (create/update) and resource (scoped variable), then defines what that resource is — an environment-level value or secret injected into containers at runtime. It also explicitly carves itself out from siblings by noting it cannot delete and pointing to list_scoped_variables. An agent can distinguish this from list_scoped_variables and the delete_* tools without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit ordered workflow: call list_scoped_variables first, always call with preview:true first, get user confirmation, then call again without preview. It also names the one thing it cannot do (delete). When-to-use and the correct call sequence are both spelled out rather than left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

migrate_instancesMigrate Container InstancesA
Destructive
Inspect

Migrate a container's or virtual machine's instances between servers on Cycle, or revert a recent migration. Pass exactly one of container or virtual_machine. Cycle backs every VM with a container, so a VM migrates through that container's single instance (select it with all_instances:true or name it in targets); the response reports both the VM and its backing container.

Cycle migrates instances across any infrastructure it manages — between servers, data centers, cloud providers, and on-prem hardware. Not every server is a valid target: containers carry tag restrictions and other constraints, so this tool only accepts destinations Cycle reports as compatible for the container. A migration is reversible: the original instance is retained until Cycle's purge window elapses (roughly 3 hours for stateful instances), during which action:"revert" restores it on its source server. Only retained source instances are revertable; running destination copies are skipped. Load balancer instances cannot be migrated.

Workflow:

  1. Call with preview:true and your selection. It makes NO changes and returns the selected instances — each with its current server (id and name) and whether it is stateful — plus, for migrate, the compatible destination servers. The plan comes from read-only lookups; Cycle does not validate it.

  2. If the destination is unclear, present the compatible servers and let the user choose. With NO compatible servers migration is impossible — tell the user why (tag or infrastructure constraints).

  3. Confirm the specific move with the user, then call again without preview. Never migrate or revert without explicit confirmation.

Select instances with exactly one of: targets (specific instances, each optionally with its own destination so they can spread across servers), source_server ("move this container off nuc-bear"), or all_instances. A top-level destination_server is the default for selected instances without their own and is required with source_server or all_instances.

copy_volumes applies to stateful instances and defaults to true so data is never silently dropped. A VM's local volumes — including its boot disk — are what it moves, so copy_volumes:false lands the VM on empty local storage and is almost never wanted. External (SAN) volumes are attached, not copied, and are unaffected either way.

Asynchronous: submits one job per instance and returns without waiting (wait_seconds gives a short bounded wait for quick moves). Track the job_ids with get_jobs, then call again with preview:true to confirm each instance reports its destination server.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionNomigrate moves instances to a destination server; revert restores recently migrated instances on their source server, within Cycle's purge window.migrate
contextNoWhy are you calling this tool? Briefly describe the user's goal.
previewNoResolve and return the plan without making changes. Always run this first.
targetsNoSpecific instances to act on.
containerNoContainer whose instances to migrate. Exactly one of container or virtual_machine.
environmentYesEnvironment the container or VM lives in. Required; it scopes the lookup.
copy_volumesNoWhether stateful instances copy their volume contents to the destination (default true). Ignored for non-stateful instances and for revert.
wait_secondsNoBounded wait for the submitted jobs (default 0: return immediately and track with get_jobs). Stateful migrations often run far longer than any wait.
all_instancesNoSelect every instance of the container.
source_serverNoSelect every instance of the container currently on this server (hostname, nickname, or ID). Mutually exclusive with targets.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.
virtual_machineNoVM to migrate. Exactly one of container or virtual_machine.
destination_serverNoDefault destination for selected instances without their own; required with source_server or all_instances. Hostname, nickname, or 24-char ID of one of the container's compatible servers. Ignored for revert.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructive/openWorld/non-idempotent, and the description adds substantial extra context beyond them: a ~3-hour purge window making migrations revertable, that only retained source instances are revertable, that load balancer instances cannot be migrated, that preview makes NO changes, that copy_volumes:false lands a VM on empty local storage, and that jobs are submitted asynchronously.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is long (~350 words) but front-loaded with the core action and organized into labeled sections and a numbered workflow, so structure carries the length. A few clauses (e.g. the VM-backing-container aside) could be trimmed, but nearly every sentence earns its place for a 13-parameter destructive tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no output schema, the description explains the response shape (reports both the VM and its backing container, returns job_ids) and how to verify completion via preview:true plus get_jobs. Together with the workflow and edge-case coverage, an agent has everything needed to call this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, but the description adds real semantics the schema lacks: the exactly-one-of container/virtual_machine rule, the interaction between targets/source_server/all_instances, the defaulting behavior of a top-level destination_server, and the stateful-vs-external volume distinction for copy_volumes. It meaningfully extends rather than restates the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb (migrate/revert) and resource (container or VM instances between servers on Cycle), plus the exactly-one-of constraint for container vs virtual_machine. This clearly distinguishes it from siblings like reconfigure_container, deploy_virtual_machine, and cycle_control_* without needing to open the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit 3-step workflow (preview first, present compatible servers, confirm before mutating), a rule to never migrate or revert without explicit confirmation, and the conditions under which migration is impossible (tag or infrastructure constraints). It also routes async tracking to the sibling tool get_jobs by name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

query_metricsRun a Metrics AggregationA
Read-onlyIdempotent
Inspect

Run a canned aggregation over Cycle's metrics store. Presets:

  • container_instance_count (scope: container): Hourly instance count for a container; reveals flapping, autoscaling, and unexpected shrinkage.

  • discovery_resolutions (scope: environment): Hourly internal DNS resolution counts for an environment's discovery service (lookups, cache hits, not-founds). Non-zero not-founds are the smoking gun for containers unable to find each other.

  • lb_controller_traffic (scope: environment): Hourly load balancer controller metrics for an environment (requests, connections, disconnect reasons, per-destination counters). Historical complement to get_telemetry target=load_balancer.

  • neighbor_latency (scope: server): Latest server-to-server mesh latency per neighbor; negative latency_ms means the neighbor is unreachable (outage).

  • server_cpu (scope: server): Hourly CPU usage trend for a server, per CPU state metric.

  • server_ram (scope: server): Hourly RAM trend for a server (available/free/total KB metrics).

Pass the input matching the preset's scope (server, environment, or container). Metrics data can lag by up to ~10 minutes; each row carries the sample time. No free-form queries — if none of the presets fit, say so rather than improvising. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
presetYesWhich canned aggregation to run. See the tool description for what each returns.
serverNoServer hostname, nickname, or ID. Required for server-scoped presets.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
containerNoContainer. Required for container-scoped presets.
environmentNoEnvironment. Required for environment-scoped presets.
lookback_hoursNoHow many hours back to aggregate, 1-72.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive), so the bar is lower, and the description still adds real operational context: metrics can lag up to ~10 minutes, each row carries the sample time, and the tool is read-only. It stops short of describing result shape or row limits, but for a read-only aggregation this is solid added value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and the no-free-form rule are front-loaded, and the bulleted preset list is scannable. The parenthetical explanations inside each bullet are slightly verbose but each one describes a diagnostic signal an agent would otherwise have to guess at, so the length is mostly earned.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of explaining what comes back, and it does so per preset (counts, latency per neighbor, CPU/RAM trends) plus freshness caveats. Remaining gaps are minor: no statement of row counts, pagination, or the exact fields in each result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds the preset-to-scope mapping inline ('scope: container', 'scope: environment', 'scope: server'), letting the agent resolve which of server/container/environment to supply without cross-referencing per-field text. It adds little on lookback_hours or conversation_id beyond the schema, keeping it at a 4 rather than 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ('Run a canned aggregation over Cycle's metrics store') and then enumerates all six presets with the scope and the signal each one surfaces. It explicitly differentiates itself from the sibling get_telemetry ('Historical complement to get_telemetry target=load_balancer'), so an agent can choose between them without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States the selection rule directly: pick a preset, then pass the input matching that preset's scope (server/environment/container). It also gives an explicit exclusion — 'No free-form queries — if none of the presets fit, say so rather than improvising' — which is exactly the kind of when-not guidance that prevents misuse.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reconfigure_containerReconfigure a ContainerA
DestructiveIdempotent
Inspect

Change an existing Cycle container: its runtime config (start command/args, environment variables, ports, hostname, public network mode, CPU/RAM, automated backups), its instance count, its image, its volumes (extend, read-only), or its metadata (name, identifier, annotations, lock, deprecate). Adding or removing volumes is a redeploy, not a reconfigure. The virtual-machine equivalent is reconfigure_virtual_machine.

Each group reaches Cycle through a different mechanism and differs in disruption; every change in the preview diff is tagged with an 'impact':

  • 'restart' — runtime config. Cycle's reconfigure REPLACES the container's entire config, so this tool fetches the current config, applies only the fields you pass, and submits the whole result; everything you don't pass keeps its value (including deployment constraints, integrations, and devices). Applying RESTARTS every instance; there is no way to change runtime config without a restart. volumes.max_size and volumes.read_only rewrite the volume set through Cycle's volumes.reconfigure task, which also restarts every instance. 'backups' (stateful containers only) merges the passed fields over the current backups integration — destination, command, schedule and restore_command must all be present once merged; backups:{disable:true} removes it.

  • 'scale' — instance count. Instances cannot be changed by reconfiguring; the tool submits Cycle's scale task, which does not rewrite the config. More than 1 instance on a stateful container is allowed but NOT recommended (the tool applies it with a warning note) — prefer separate single-instance containers.

  • 'reimage' — moves the container onto a different image built from the image source it ALREADY uses, keeping its runtime config; instances restart on it. reimage:true imports a fresh image to pick up the latest. image_id reimages onto any live image from that same source, including an older one — that is how you ROLL BACK a bad image: preview with reimage:true, and the response's 'image_options' lists the source's live images newest first with build times and marks the running one; pass the chosen id as image_id (it takes precedence over reimage:true, so nothing is imported). A rollback keeps the CURRENT runtime config, so check the args, env and volume paths still suit the older image and reconfigure them in the same call if not. Switching image sources is a redeploy, not a reimage, and is refused. Stack-deployed containers are refused too — reimage them by building and redeploying the stack. If an import outlives the wait, the response carries the new image_id; call again with it to finish without importing a second copy.

  • 'volume' — volumes.extend grows a local volume's allocation in place with an instance task, no restart; sizes only increase and are capped by the volume's max_size. Passing volumes at all returns the current volumes, each with its max_size and per-instance used/allocated, so a preview with an empty volumes object is a read.

  • 'none' — metadata. Applies immediately: no job, no restart, instances untouched.

Always call with preview:true first — it fetches the current state and returns a field-by-field from/to diff with each change's impact, making NO changes. The diff is computed by this tool, not validated by Cycle; use it to confirm the change matches the user's intent. Get explicit confirmation, then call again without preview.

Changes from several groups are applied in order — metadata, runtime config, volumes, image, instances — and each step is reported in 'steps'. The run stops at the first failed step, so a partial result is possible; read 'steps' rather than assuming all-or-nothing.

ParametersJSON Schema
NameRequiredDescriptionDefault
envNoEnvironment variables to set, merged into the existing ones (a tool convenience, not a Cycle feature).
ramNoNew RAM limit, e.g. '1G'. Only when the user asks for one.
argsNoReplacement arguments for the start command, e.g. '--replSet rs0 --bind_ip_all --ipv6'.
lockNotrue blocks delete_container until unlocked; false lifts the lock.
nameNoNew display name.
portsNoComplete replacement port list, e.g. ['27017:27017'] — pass every port wanted.
publicNoNew public network access mode.
sysctlNoKernel sysctls applied inside the container, e.g. {"net.core.somaxconn": "4096"}. Only namespaced keys work (net.*, fs.mqueue.*, kernel IPC shm/msg/sem); vm.* keys such as vm.max_map_count are host-level and are rejected.
backupsNoAutomated backups: Cycle runs 'command' on 'schedule' and ships its STDOUT to 'destination'; 'restore_command' reads a backup from STDIN.
commandNoNew start command path, e.g. 'mongod'.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
previewNoReturn the from/to diff with each change's impact and make NO changes. Always run this first.
reimageNoImport a fresh image from the container's current image source and move onto it. With preview:true, also returns 'image_options' without importing.
rlimitsNoProcess rlimits keyed by setrlimit(2) name, e.g. {"RLIMIT_NOFILE": {"soft": 65535, "hard": 65535}}. Cannot be raised from inside the container. Elasticsearch/OpenSearch bootstrap checks need RLIMIT_NOFILE >= 65535.
volumesNoVolume changes keyed by mount path. An empty object just reports the current volumes and usage.
hostnameNoNew private-network hostname.
image_idNoExisting live image from the same source to reimage onto — a rollback target from 'image_options', or the image_id of an import that outlived the wait. Takes precedence over reimage.
containerYesContainer to reconfigure.
deprecateNoMark the container deprecated (true) or undo it (false).
instancesNoDesired instance count.
unset_envNoEnvironment variable names to remove.
cpu_sharesNoNew CPU limit in shares: 10 = one full thread/vCPU; Cycle's default is 2. Only when the user asks for a CPU limit.
identifierNoNew identifier slug. On a stack-deployed container this desynchronizes it from the stack file, which keys containers by identifier.
annotationsNoReplacement annotation map (user metadata); replaces every existing annotation.
environmentYesEnvironment the container lives in.
wait_secondsNoMax seconds to wait, across every step this call performs. 0 returns as soon as the jobs are accepted.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructive/openWorld/idempotent but say nothing about mechanism; the description supplies the rest: config is fully replaced so unpassed fields persist, every runtime change restarts instances, scale uses Cycle's scale task, volumes.extend is in-place with no restart, metadata applies with no job, and partial application is possible because the run stops at the first failed step.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Long, but the size is justified by a 27-parameter mutation tool and it is front-loaded with the purpose, then organized into impact-tagged bullets ('restart', 'scale', 'reimage', 'volume', 'none'). Some sentences restate schema content (e.g. volume/max_size behavior appears in both), which is the only slack.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden itself, describing the preview from/to diff, the 'image_options' list with build times and running-image marker, the 'steps' array, and the new image_id returned when an import outlives the wait. Nothing an agent needs to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds semantics the schema alone doesn't convey: which parameter group triggers which Cycle mechanism, that backups merges over the current integration while annotations replaces, that image_id takes precedence over reimage, and the impact tag each change carries. It stops short of documenting the two required identity params beyond what the schema states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Change an existing Cycle container') and enumerates the changeable groups (runtime config, instance count, image, volumes, metadata). It explicitly names the sibling equivalent (reconfigure_virtual_machine) and distinguishes reimage from redeploy, so an agent can separate it from delete_container, manage_external_volumes, and deploy_application.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use ('Always call with preview:true first... Get explicit confirmation, then call again without preview') and when-not ('Adding or removing volumes is a redeploy, not a reconfigure'; 'Switching image sources is a redeploy... and is refused'; stack-deployed containers refused). It also routes to the right alternative for VMs and for rollbacks.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reconfigure_virtual_machineReconfigure a Virtual MachineA
DestructiveIdempotent
Inspect

Change an existing Cycle virtual machine: its config (hostname, public network mode, ports, egress, CPU, RAM), its guest root password, its assigned SSH keys, or its metadata (name, identifier, annotations, lock, deprecate). Volumes are NOT handled here — VM volume changes exist on Cycle but have no MCP tool yet; surface that gap rather than improvising. The boot image is fixed at create; there is no VM equivalent of a container reimage. The container equivalent is reconfigure_container.

Each group reaches Cycle through a different mechanism and differs in disruption; every change in the preview diff is tagged with an 'impact':

  • 'restart' — config. Cycle's reconfigure REPLACES the VM's entire config, so this tool fetches the current config, applies only the fields you pass, and submits the whole result; everything you don't pass keeps its value (including deployment constraints, startup/shutdown policies, telemetry, and runtime attachments). Applying RESTARTS the VM — a hard interruption of a running guest OS.

  • 'root-password' — the guest's root password, changed by a Cycle job inside the VM. Neither the preview nor the result echoes it; whoever sets it must record it themselves, since Cycle only makes a VM's root password readable for ~10 minutes after creation.

  • 'none' — metadata and SSH key assignment. Applies immediately: no job, no restart. Cycle stores one flat list of assigned keys; add_ssh_keys/remove_ssh_keys read it, apply your changes, and write it back. Keys are written into the guest at provision time, so a running guest may not honor an added key until it restarts, and removing a key does NOT reliably revoke access for a running guest — to be certain, restart the VM (cycle_control_virtual_machine) and verify from inside the guest.

Always call with preview:true first — it fetches the current state and returns a field-by-field from/to diff with each change's impact, making NO changes. The diff is computed by this tool, not validated by Cycle; use it to confirm the change matches the user's intent. Get explicit confirmation, then call again without preview.

Changes from several groups are applied in order — metadata and SSH keys, then config, then root password — and each step is reported in 'steps'. The run stops at the first failed step, so a partial result is possible; read 'steps' rather than assuming all-or-nothing.

ParametersJSON Schema
NameRequiredDescriptionDefault
ramNoNew RAM limit, e.g. '4G'. At least 512M, less than 65G.
lockNotrue blocks delete_virtual_machine until unlocked; false lifts the lock.
nameNoNew display name.
coresNoNew core count (1-32), placed by Cycle. Mutually exclusive with cpu_pin; clears any pinning.
portsNoComplete replacement port list, e.g. ['443:443'] — pass every port wanted.
publicNoNew public network access mode.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
cpu_pinNoPin to specific host cores, e.g. '0-3' ('x' = the host's max core). Mutually exclusive with cores; clears the core count.
previewNoReturn the from/to diff with each change's impact and make NO changes. Always run this first.
hostnameNoNew private-network hostname.
deprecateNoMark the VM deprecated (true) or undo it (false).
identifierNoNew identifier slug.
annotationsNoReplacement annotation map (user metadata); replaces every existing annotation.
environmentYesEnvironment the VM lives in.
add_ssh_keysNoExisting environment-scoped SSH keys to assign, by name or 24-char ID.
allocate_ramNoPreallocate the RAM limit rather than letting the guest grow into it.
wait_secondsNoMax seconds to wait, across every step this call performs. 0 returns as soon as the jobs are accepted.
root_passwordNoNew guest root password, at least 10 characters. Confirm with the user and have them record it before applying.
allocate_coresNoReserve the cores exclusively so no other VM can use them.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.
remove_ssh_keysNoAssigned SSH keys to unassign, by name or 24-char ID.
virtual_machineYesVM to reconfigure.
egress_via_gatewayNoRoute outbound traffic through the Cycle gateway instead of the host.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations (destructive=true, idempotent=true, openWorld=true) only set the safety profile; the description adds far more: per-group 'impact' tags, the fact that config changes restart the guest, that root password is never echoed and only readable ~10 min post-create, that SSH key removal does not reliably revoke access, and that multi-group runs are ordered and can partially fail. This is exactly the behavioral context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Long but well organized with grouped bullets and impact labels; each sentence carries distinct operational information. Slightly dense and could trim a few clauses, but nothing is filler given 23 parameters and multi-step behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates fully by explaining the preview diff, the 'steps' reporting, and the partial-failure model. An agent has everything needed to call this correctly and safely for a 23-parameter mutation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema: it explains the preview-diff mechanism, the config-replacement-with-preserve semantics, and how SSH-key and root-password parameters behave at runtime. It goes beyond restating field-level docs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Change/reconfigure) and resource (an existing Cycle VM), then enumerates the changeable groups (config, root password, SSH keys, metadata) and explicitly excludes what it does not cover (volumes, boot image). It names the sibling reconfigure_container and contrasts with the container reimage, so an agent can distinguish it without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit workflow ('Always call with preview:true first... Get explicit confirmation, then call again without preview'), names alternatives for adjacent needs (reconfigure_container, cycle_control_virtual_machine for restart), and flags the volume gap rather than improvising. When-to-use and when-not-to-use are both covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_instance_commandRun a Command in a Container InstanceA
Destructive
Inspect

Run a shell command inside a running container instance on Cycle, over the instance's SSH gateway, and return the shell output. Use this for one-time setup that must run inside a container — e.g. initializing a MongoDB replica set with 'mongosh --eval "rs.initiate({...})"'.

The instance must be RUNNING (start the container first). If the container has multiple instances, pass 'instance' (id or hostname); otherwise the sole instance is used. Containers reach each other by hostname over the environment's private network, so commands can reference sibling containers (e.g. mongo-1:27017).

This executes an arbitrary command inside your container — it is powerful and mutating. Always call with preview:true first: it echoes back the exact command and target without making any connection, so the user can confirm it is what they intend (nothing is validated — the command may still fail when run). Confirm with the user, then call again without preview. Never run a command without explicit confirmation.

Short-lived SSH credentials are generated for the run and expired immediately after. The command runs in an interactive shell; make destructive commands idempotent where possible.

ParametersJSON Schema
NameRequiredDescriptionDefault
commandYesShell command to run inside the instance, e.g. mongosh --quiet --eval 'rs.initiate({...})'.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
previewNoWhen true, echo the target + command and make NO connection. Use it to confirm intent with the user. Always run this first.
instanceNoInstance id or hostname. Optional when the container has a single instance.
containerYesContainer to run in.
environmentYesEnvironment the container lives in.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.
timeout_secondsNoMax seconds to wait for the command, 1-60 (default 60). 60 is the ceiling because a longer block dies at the MCP transport before this tool can report; a command that outlasts it keeps running in the container.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations flag destructive/openWorld/non-idempotent, and the description adds substantial context beyond them: the preview echo makes no connection and validates nothing, short-lived SSH credentials are generated and expired immediately, the command keeps running after the 60s transport cutoff, and destructive commands should be made idempotent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence, followed by prerequisites, networking context, the safety workflow, and credential lifecycle — each paragraph earns its place with no filler. Preview is reinforced but that reinforcement is warranted for a destructive tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description still names the return value (shell output) and the timeout behavior (command outlasts the wait). Given the 8-param, destructive, open-world nature of the tool, nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description still adds meaning: it explains instance selection logic, why timeout is capped at 60s (MCP transport dies first), and the confirmation semantics of 'preview', going beyond the schema's own text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Run a shell command inside a running container instance') with the transport mechanism (SSH gateway) and return value ('return the shell output'). The container-instance target clearly distinguishes it from the sibling run_vm_command.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('one-time setup that must run inside a container'), prerequisites ('instance must be RUNNING — start the container first'), a multi-instance rule for the optional 'instance' param, and a mandatory workflow (always call preview:true first, confirm with the user, then re-call without preview).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_vm_commandRun a Command on a Virtual MachineA
Destructive
Inspect

Run a shell command on a running virtual machine over its serial-over-SSH console, and return the console output. Use this for one-time in-guest setup — e.g. partitioning and formatting an attached data volume, then mounting it, or configuring networking.

The VM must be RUNNING (start it with cycle_control_virtual_machine first). This connects to the guest's SERIAL CONSOLE through the Cycle gateway — it does not use the VM's network — and logs in with the guest's own credentials: username defaults to 'root', and the password defaults to the VM's generated root password when that is still retrievable (~10 minutes after creation). After that window you MUST pass 'password' (and 'username' if not root); this tool never guesses or stores credentials. If nobody has the guest password any more, reconfigure_virtual_machine can set a new one. Because a serial console is a real tty, the guest echoes input, so the returned output includes some command echo — read it as a console transcript, not clean stdout.

This runs an arbitrary command as root inside the guest — it is powerful and mutating. Always call with preview:true first: it echoes the target and command and makes NO connection, so the user can confirm intent. Confirm with the user, then call again without preview. Never run a command without explicit confirmation. Make destructive commands idempotent where possible.

Serial credentials are minted for the run and expired immediately after. For a human who wants to drive their own interactive session instead, use get_vm_console_access — it returns an ssh command they run themselves; this tool is for programmatic one-shot commands.

ParametersJSON Schema
NameRequiredDescriptionDefault
commandYesShell command to run in the guest, e.g. 'mkfs.ext4 /dev/vdb && mkdir -p /mnt/data && mount /dev/vdb /mnt/data'.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
previewNoWhen true, echo the target + command and make NO connection. Use it to confirm intent with the user. Always run this first.
passwordNoGuest login password. Optional while the VM's generated root password is still retrievable (~10 min after creation); required after that, or when logging in as a non-root user with a different password.
usernameNoGuest login username. Defaults to 'root'.
environmentYesEnvironment the VM lives in.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.
timeout_secondsNoMax seconds to wait for login plus command, 1-60 (default 45). 60 is the ceiling because a longer block dies at the MCP transport before this tool can report. A genuinely slow command (e.g. mkfs on a large volume) may outlast the wait — it keeps running in the guest, so re-run a cheap check command afterwards to confirm it finished rather than raising the timeout.
virtual_machineYesVM to run on.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructive/openWorld/non-idempotent, and the description adds substantial context beyond them: the serial-console transport (not the VM network), credential minting and immediate expiry, the ~10 minute window for the auto-generated root password and the need to pass password/username after that, and the fact that the returned text is an echoed console transcript rather than clean stdout.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then safety, then credentials, then the alternative — a logical progression. It is on the long side and repeats the ~10 minute password window and the preview instruction that already appear in the schema, which is mild redundancy rather than waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fills that gap by explaining that the return value is an echoed console transcript, and it covers the prerequisites, credential lifecycle, timeout ceiling rationale (60s MCP transport limit and re-checking after a slow command), and the preview-confirm flow. An agent has everything needed to invoke it safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all nine parameters, including preview, password, timeout_seconds and conversation_id. The description largely restates the credential-window and timeout behavior already present in the schema; it adds little parameter-level meaning that isn't there, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+mechanism: run a shell command on a running VM over its serial-over-SSH console and return console output. It also names the scope (one-time in-guest setup) and implicitly separates itself from siblings like run_instance_command (containers/instances) and get_vm_console_access (interactive human sessions).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use (one-time in-guest setup such as partitioning/mounting a volume, configuring networking), prerequisites (VM must be RUNNING; start it with cycle_control_virtual_machine first), a required confirmation workflow (preview:true first, then call again), and an explicit alternative for the human-interactive case (get_vm_console_access).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 44 tool updates
    • First observedcapture_stream
    • First observedcheck_dns_propagation
    • First observedcheck_server_provisioning
    • First observedcreate_environment
    • First observedcycle_control_container
    • First observedcycle_control_environment
    • First observedcycle_control_virtual_machine
    • First observeddecommission_server
    • First observeddelete_container
    • First observeddelete_dns_record
    • First observeddelete_environment
    • First observeddelete_image_source
    • First observeddelete_stack
    • First observeddelete_virtual_machine
    • First observeddeploy_application
    • First observeddeploy_servers
    • First observeddeploy_virtual_machine
    • First observeddiagnose
    • First observedget_deployable_server_models
    • First observedget_deployment_status
    • First observedget_hub
    • First observedget_jobs
    • First observedget_logs
    • First observedget_more_tools
    • First observedget_telemetry
    • First observedget_vm_console_access
    • First observedlist_containers
    • First observedlist_dns_zones
    • First observedlist_environments
    • First observedlist_image_sources
    • First observedlist_scoped_variables
    • First observedlist_servers
    • First observedlist_virtual_machines
    • First observedmanage_cluster
    • First observedmanage_dns_record
    • First observedmanage_external_volumes
    • First observedmanage_image_source
    • First observedmanage_scoped_variable
    • First observedmigrate_instances
    • First observedquery_metrics
    • First observedreconfigure_container
    • First observedreconfigure_virtual_machine
    • First observedrun_instance_command
    • First observedrun_vm_command

Publisher details

Operator
Petrichor Holdings, Inc. · Publisher source
Vendor relationship
First-party · Publisher source
Trust center
Unknown
Restrictions
Requires a paid Cycle plan.

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    Enables brand visibility monitoring across major AI platforms like ChatGPT, Claude, Gemini, and Perplexity. It allows users to track visibility scores, analyze competitor data, and receive actionable insights to improve AI-generated brand recommendations.
    16
    24 npm
    1
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Browse IndustryLens's published competitive-intelligence reports and head-to-head competitor comparisons from any AI agent — real, source-backed data.
    MIT
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources