Skip to main content
Glama

run_subagent

Delegate multi-step coding and refactoring tasks to an autonomous engine that edits local files, verifies results, and rolls back on failure with cost tracking.

Instructions

Runs BoteX — an autonomous code execution engine (agent-agnostic harness). BoteX performs multi-step tasks on local files inside workspace_dir.

TIP: If the user didn't specify a model, DO NOT guess. Use the recommend_models tool first to find the best model for this task!

Key engine features:

  • Outline-First and Context Pruning: ~90% token savings.

  • Pre-write Syntax Check: in-memory code validation before touching disk.

  • Multi-tier Fuzzy Patching: tolerant of CRLF/LF and indentation drift.

  • Hard Zero-Retention on OpenRouter (provider: data_collection=deny) and DLP secret censoring.

  • Automatic Snapshot Rollback on loops or critical failures.

Args: task: Precise description of the coding or refactoring task. files: Optional list of primary files affected by the task. workspace_dir: Project directory path (defaults to current directory). model: Any model from the provider's catalog. Empty = model from config (see 'profile' and botex.config.json). profile: Model profile from botex.config.json (e.g. 'default', 'coding', 'auto-beta', 'fast'). When both 'model' and 'profile' are empty, the capability mode picks the tier via engine.mode_profiles (readonly -> 'fast', destructive/full -> 'coding'). Ignored when 'model' is given explicitly. provider: API provider from the 'providers' section of botex.config.json ('openrouter', 'nvidia'). Empty = provider marked 'default: true' in the configuration. mode: Capability preset: 'readonly' (read only — the contract is an analysis report, DONE requires non-empty findings), 'edit' (read + edit — default), 'destructive' (+delete/move), 'full' (+run_command). Empty = engine.default_mode from config. Explicit allow_* flags may only WIDEN a preset — they never narrow it (readonly can never gain file writes). In headless mode this is pre-authorization — grant it consciously. max_turns: Maximum tool-loop steps (0 = config value). max_tokens: Per-turn completion cap (0 = config value; reasoning models need ~8000+ since thinking shares this budget). max_duration_s: Total wall-clock limit in seconds (0 = config value, 0 disables only when config is also 0). Checked between turns. budget_limit_usd: Daily spend limit in USD (negative = config value, 0 = no limit). allow_destructive: Authorize destructive operations (delete_file/ move_file) for this task. Headless mode cannot confirm mid-run — grant ONLY with the user's consent. allow_exec: Authorize run_command for this task (also requires exec.enabled=true in config). WARNING: this is NOT a sandbox — commands run with operator privileges. Grant only for trusted workspaces with the user's consent. api_key: Optional provider API key passed per-request (highest priority — overrides env/.env/config). Empty = resolved internally by the harness. output_path: Optional required output file (workspace-relative). Sets the file_output contract: DONE is accepted only when the file exists on disk, is non-empty, and passes the syntax gate; a DONE response carrying the payload as text is salvaged to disk. Requires a mutating mode. allow_net: Explicitly authorizes the read-only public web tool for this run. It is never enabled by a capability mode. net_allowed_hosts: Optional per-run host authorization. In net.policy='caller' these replace config defaults; in 'public' they narrow public access; in 'allowlist' they must stay inside net.allowed_hosts. net_allowed_urls: Optional per-run URL authorization rules. A URL ending in '/' authorizes that subtree; otherwise it authorizes the exact URL including its query string. verify_command: Optional allowlisted command that must pass before DONE is accepted — it runs against the workspace as it stands, including when the task made no writes (a passing verifier validates a legitimate no-change result). Requires exec.enabled=true in config and exec authorization. recipe: Optional operational persona / workflow prompt (e.g. 'planner', 'code-explorer', 'reviewer', 'security-reviewer', 'build-resolver', 'tdd').

Returns: A pretty-text report in the text content (unchanged format for legacy clients) plus the full engine result as structuredContent (status, failure_kind, ok, files_touched, exec_ran, rollback_verified, attempts[], cost_usd, ...). ok is true only when status == DONE — i.e. the task contract was verified, not merely claimed by the model.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
modeNo
taskYes
filesNo
modelNo
recipeNo
api_keyNo
profileNo
providerNo
allow_netNo
max_turnsNo
allow_execNo
max_tokensNo
output_pathNo
workspace_dirNo.
max_duration_sNo
verify_commandNo
budget_limit_usdNo
net_allowed_urlsNo
allow_destructiveNo
net_allowed_hostsNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
okNo
stepsNo
statusNo
messageNo
summaryNo
task_idNo
attemptsNo
cost_usdNo
exec_ranNo
duration_sNo
lines_addedNo
failure_kindNo
fallback_fromNo
files_touchedNo
lines_removedNo
rollback_verifiedNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.0.0

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure, and it does so thoroughly. It reveals snapshot rollback, pre-write syntax checks, zero-retention on OpenRouter, the non-sandbox nature of allow_exec, the output_path DONE contract, and that 'ok is true only when status == DONE.' This is far beyond minimal disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long, but the tool is genuinely complex with 20 parameters and multiple safety-sensitive modes. The structure is efficient: a front-loaded TIP, a compact feature list, and a well-organized Args section. Every sentence adds operational value rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, the description is complete: it covers model selection, mode semantics, safety authorizations, output contracts, verification behavior, and return-value semantics. The presence of an output schema reduces the need to explain return fields, yet the description still clarifies the critical ok/status relationship.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate entirely, and it does. Every one of the 20 parameters is explained with defaults, interactions, and warnings—for example, mode's capability presets, budget_limit_usd's negative-value semantics, and net_allowed_urls' trailing-slash subtree rule. This is exemplary parameter documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Runs BoteX — an autonomous code execution engine' that 'performs multi-step tasks on local files inside workspace_dir.' This clearly differentiates it from sibling tools like recommend_models, fetch_url, and start_task by establishing it as the main task-execution harness.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit guidance for a key decision: 'If the user didn't specify a model, DO NOT guess. Use the recommend_models tool first.' It also warns about headless pre-authorization and when to grant destructive/exec permissions. It does not explicitly contrast with start_task, but the context is clear enough for an agent to know when this tool is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.