Skip to main content
Glama

Execute a full GUI task

computer_execute_task

Delegate a natural-language task to an autonomous GUI agent for cross-platform computer use; approve sensitive steps by re-calling with a confirmation token.

Instructions

Delegate a natural-language task to the autonomous GUI agent loop (observe → decide → locate → execute → verify → recover). mode=delegate|auto. Returns immediately with a taskId unless wait=true. Sensitive actions park in WAITING_CONFIRMATION with a confirmToken — re-call with the token to approve.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
modeNoauto
taskYesNatural language task, e.g. 打开设置,将Wi-Fi打开
waitNoWait for completion instead of returning a taskId
targetNoUnified device target (spec §14). Defaults to this machine with auto-detected platform.
languageNo
maxStepsNo
confirmTokenNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.2.1

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose two valuable traits: asynchronous return with a taskId unless wait=true, and the WAITING_CONFIRMATION/confirmToken approval flow. It omits failure/error behavior, timeout handling, what maxSteps bounds, and what the taskId enables afterwards, so the safety profile of this powerful autonomous tool is only partially covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly packed sentences, front-loaded with the core purpose and no filler. It is dense rather than bloated, though the parenthetical pipeline enumeration and mode clause make it slightly more compressed than ideally scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex autonomous-loop tool with no output schema and no annotations, the description covers the pipeline, the async return contract, and the confirmation flow well. Gaps remain around the assist/direct modes, the target object, and post-launch usage of the taskId (e.g., computer_get_task), which an agent driving a rich sibling set would benefit from.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 43%, so the description must compensate and it does for wait and confirmToken and partially for mode (only delegate|auto mentioned though the enum has four values). The task, target (nested object), language, and maxSteps parameters get no added semantic detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Delegate a natural-language task to the autonomous GUI agent loop" gives a specific verb (delegate) and resource (GUI agent loop), and enumerates the internal pipeline phases. It clearly distinguishes a high-level delegated task from the low-level siblings, but never names alternatives like computer_action or computer_step explicitly to route the choice.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the delegation use case and mentions mode=delegate|auto plus the wait semantics, giving some context. However, it never states when to prefer this tool over computer_action, computer_run_script, or computer_run_ui_test, and it does not explain the assist/direct modes in the enum either.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.