BoteX
Allows using NVIDIA as a model provider for autonomous code-execution tasks, with configurable profiles and model fallbacks.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@BoteXAdd email validation to auth/validators.py and cover it with tests"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
BoteX — Autonomous Code Execution Harness
BoteX is an agent-agnostic, MCP-native execution harness for autonomous code-editing subagents. Any MCP-compatible client (IDE agents, desktop assistants, orchestrators) can delegate multi-step coding tasks to it.
Key features:
Outline-First I/O — agents read symbol outlines before requesting specific line ranges
Read-Before-Write Gate — prevents hallucinated edits by requiring models to inspect files before patching
Context Pruning — stale file reads are compacted after successful patches
Pre-write syntax validation — code is linted in memory before touching disk
Multi-tier fuzzy patching — tolerant to CRLF/LF and indentation differences
Workspace Memory Vault — portable, persistent markdown/YAML context storage per workspace
Personas & Recipes — specialized workflow modes (planner, reviewer, security-reviewer, build-resolver, tdd)
Snapshot rollback — automatic restore on stagnation or critical failure
Security layer — path traversal guard,
.env/key blocking, secret masking (DLP)Zero Data Retention —
provider.data_collection=denyon every OpenRouter callCost ledger — per-task analytics, budget limits, rich CLI reports with DONE rates
Installation
pip install -r requirements.txtRequires Python 3.10+ and an OPENROUTER_API_KEY. The harness resolves the
key internally — no need to pass it from the MCP client. Lookup order:
OPENROUTER_API_KEYenvironment variable.envin the process working directorypaths.env_filein the config /BOTEX_ENV_FILEenv var (custom dotenv path).envin the project root (next toserver.py)secrets.openrouter_api_keyinbotex.config.local.json(gitignored)
Never put the key in
botex.config.json— that file is committed to the repository. The local config is gitignored and additionally blocked from the subagent's file tools.
So for a self-contained setup just drop a .env next to server.py:
OPENROUTER_API_KEY=sk-or-...Related MCP server: subturn
Running as an MCP server
The botex launcher is context-aware: spawned by an MCP client (piped
stdin) it serves the stdio MCP protocol; run on an interactive terminal with
no arguments it prints help and drops into the REPL. botex serve forces
server mode manually.
Three ways to get it:
botex.cmd/botex.sh— zero-install shims shipped in the repo (they auto-detectpy/python3/python)pip install .— installs a realbotexcommand viapyproject.tomlpython server.py— direct interpreter invocation
Register the launcher as a stdio MCP server in any client:
{
"mcpServers": {
"botex": {
"command": "C:/path/to/harness/botex.cmd"
}
}
}The server exposes task tools run_subagent, start_task,
get_task_status, Memory Vault tools save_memory, search_memory, read_memory,
caller-side fetch_url, and utility tools get_outline, get_stats,
check_health, clean_snapshots, recommend_models.
Delegating a task
{
"task": "Add email validation to auth/validators.py and cover it with tests",
"files": ["auth/validators.py"],
"workspace_dir": "./your-project",
"profile": "coding"
}run_subagent parameters:
param | default | meaning |
| — | task description |
|
| primary target files (hint) |
|
| project root the agent is confined to |
|
| explicit model; empty = resolve from provider config/profile |
|
| model profile from the provider's |
|
| provider from |
|
| capability preset: |
|
| operational persona / workflow prompt ( |
|
| tool-loop step cap; |
|
| per-turn completion cap; |
|
| wall-clock limit in seconds; |
|
| daily spend cap; negative = configured default, |
|
| expose |
|
| expose |
|
| expose |
|
| per-run public host authorization ( |
|
| per-run exact URLs or trailing-slash URL subtrees |
|
| per-request provider key; overrides env/.env/config resolution |
|
| required output file; DONE is rejected until it is written |
|
| allowlisted post-write check; requires exec authorization |
Result contract
run_subagent returns the legacy pretty-text report as its text content
plus the full engine result as structuredContent (a published
outputSchema describes it; text-only clients see no change). Detached
tasks expose the same split via get_task_status: result holds the
pretty text, result_data the structured dict.
Key fields: ok, status, failure_kind, task_id, summary,
files_touched, exec_ran, rollback_verified, steps, duration_s,
cost_usd, lines_added/lines_removed, attempts[], fallback_from
(deprecated alias for the last failed candidate — attempts[] carries
the chain).
ok is true only when status == "DONE" — it means the task
contract was verified, not that the model claimed success or that bytes
changed. A partially-written run under a limit returns ok: false with
the progress still visible in files_touched.
The contract derives from the request: output_path → the file must
exist on disk, non-empty and syntax-valid; mode="readonly" → a
non-empty analysis report; otherwise → an edit task where a no-change
DONE is accepted after one confirmation nudge. Model claims and
files_touched alone never produce DONE.
status | meaning | typical |
| contract verified | — |
| no useful output after strikes/nudges |
|
| explicit |
|
| provider call failed |
|
| a single provider request exceeded its bound |
|
|
|
|
| wall-clock limit hit between turns |
|
| step limit hit |
|
| no-progress loop; changes rolled back |
|
| required output never written/salvaged |
|
|
|
|
| model over the pricing cap (pre-flight) |
|
| daily spend cap reached |
|
| bad provider/mode/policy/verify config |
|
| workspace/security pre-run failure |
|
Model fallback (model_fallbacks)
Each provider profile may list backup slugs under
providers.<name>.model_fallbacks.<profile>. When a candidate fails with
a fallbackable failure_kind (model refusals, no-progress, provider
policy/unavailability, resource limits), the next candidate is tried —
attempts[] records each candidate's model/status/failure_kind/cost.
Failures a different model cannot fix (provider_auth_error,
config_error, BUDGET_EXCEEDED, VERIFICATION_FAILED, workspace
errors) never fall back.
A mutated workspace is never handed to the next candidate dirty:
file mutations fall back only after rollback_task is verified
complete (snapshot mutation registry, normalized paths — move_file
restores both endpoints; run_command executions are unverifiable and
always block fallback). engine.max_task_duration_s caps the whole
candidate chain (0 = per-attempt max_duration_s only).
Capability modes
Modes are named presets over orthogonal capability flags
(read, write, destructive, exec, net). Explicit allow_* flags
may only widen a preset, never narrow it — readonly stays readonly.
Unknown modes fail closed with an error.
mode | capabilities |
| read tools only |
| read + |
| edit + |
| destructive + |
net is intentionally not part of any preset — it is an
exfiltration/prompt-injection axis and must be opted into separately via
allow_net. net.policy controls the destination scope: caller accepts
per-run hosts/URL rules, allowlist is limited to configured entries,
public allows any public host for non-sensitive workspaces, and off
disables network access. All modes keep SSRF/IP-pinning, redirect, port,
size, and text-only guards. Resolution: explicit flags > mode param >
engine.default_mode config > edit. In headless MCP the mode is set at
task start — there is no mid-task prompting; grant destructive/full
only after user consent. CLI: botex run --mode full, or mode <preset>
inside botex repl.
Destructive operations
delete_file and move_file exist but are never exposed to the agent by
default (engine.allow_destructive: false / mode below destructive).
Authorization model:
Headless / MCP — there is no way to ask mid-task, so the caller must pre-authorize:
run_subagent(..., allow_destructive=true)ormode="destructive"/"full"(per task) orengine.allow_destructive: true/BOTEX_ALLOW_DESTRUCTIVE=1(global).Interactive CLI — without the flag, an interactive terminal gets a per-operation
[y/N]prompt;--allow-destructiveskips prompting. In non-interactive contexts (pipes, CI) the tools stay hidden.
Every delete/move snapshots the file first — rollback_task can restore it.
Command execution (run_command)
Warning — this is NOT a sandbox.
run_commandspawns a real process with the operator's privileges. The policy below narrows the surface but cannot make arbitrary execution safe: test runners execute repository code by design. Enable it only for trusted workspaces.
run_command lets the agent verify its own changes (pytest, ruff,
git status/diff, …). It is gated by three layers — all required:
exec.enabled: truein the config (master switch, default off; alsoBOTEX_EXEC_ENABLED=1),the
execcapability —mode="full"orallow_exec=trueper request (headless pre-authorization) — or an interactive[y/N]prompt in CLI,the command policy:
exec.allowlist(binaries only),exec.deny_args(eval flags, git mutations, network tools),exec.timeout_s,exec.max_output_bytes.
Guarantees: shell=False argv execution (no metacharacter expansion),
workspace as cwd, secret-named env vars stripped from the child process,
process-tree kill on timeout, stdout/stderr truncated and secret-masked.
Note: command side effects are not covered by snapshot rollback — files changed by a command are unknown to the harness.
Configuration
Resolution order (lowest → highest priority):
botex.config.json— committed defaultsbotex.config.local.json— gitignored machine-local overridesBOTEX_*environment variables
{
"providers": {
"openrouter": {
"default": true,
"base_url": "https://openrouter.ai/api/v1",
"api_key_env": "OPENROUTER_API_KEY",
"extra_body": {"provider": {"data_collection": "deny"}, "reasoning": {"max_tokens": 3000}},
"models": { "default": "openrouter/auto", "auto-beta": "openrouter/auto-beta", "coding": "openrouter/pareto-code", "fast": "deepseek/deepseek-v4-flash-0731" }
},
"nvidia": {
"default": false,
"base_url": "https://integrate.api.nvidia.com/v1",
"api_key_env": "NVIDIA_API_KEY",
"extra_body": {},
"models": { "default": "meta/llama-3.3-70b-instruct", "coding": "qwen/qwen3-coder-480b-a35b-instruct" }
}
},
"engine": { "max_turns": 15, "max_tokens": 8000, "reasoning_max_tokens": 16000, "request_timeout_s": 300, "max_duration_s": 900, "max_task_duration_s": 0, "temperature": 0.1, "budget_limit_usd": 0.5, "default_mode": "edit", "require_read_before_write": true },
"net": { "enabled": true, "policy": "caller", "allowed_hosts": ["docs.openrouter.ai", "openrouter.ai"], "allowed_urls": [] },
"exec": { "enabled": false, "allowlist": ["pytest", "python", "ruff", "git", "..."], "deny_args": ["python -c", "git push", "..."], "timeout_s": 120 },
"paths": { "snapshot_dir": ".snapshots", "analytics_file": ".agent_analytics.json", "env_file": "" },
"app": { "referer": "" },
"analytics": { "store_full_text": false, "preview_chars": 160 },
"ui": { "language": "auto" }
}This is an abridged view — botex.config.json holds the complete defaults
(fallback chains, full exec allow/deny lists, pricing caps, snapshot quotas).
Two separate reasoning limits exist: providers.<name>.extra_body.reasoning.max_tokens
is sent to the provider and caps hidden thinking server-side, while
engine.reasoning_max_tokens is BoteX's local per-turn max_tokens raise,
applied once a response reports reasoning tokens. ui.language accepts
en, pl, or auto (OS locale; override with BOTEX_LANG).
The product name (BoteX, sent as X-Title) is part of the harness
identity and is not configurable; app.referer is your own
attribution URL for OpenRouter.
Providers. Each provider is an OpenAI-compatible endpoint with its
own base_url, api_key_env, extra_body, and model profiles. The
provider flagged "default": true is used when a request does not name
one; run_subagent(provider="nvidia", profile="coding") selects that
provider's coding profile. API keys resolve per provider:
api_key param > <PROVIDER>_API_KEY env > .env/env_file >
secrets.<provider>_api_key in botex.config.local.json.
Provider fields are strictly scoped: a provider's extra_body goes only
to its own endpoint — OpenRouter's data_collection: deny (ZDR) is
centrally enforced there and cannot be weakened per call — and
attribution headers (X-Title/HTTP-Referer) are sent only to
OpenRouter unless another provider sets "attribution_headers": true.
Provider exception text is secret-masked and truncated to
engine.max_error_chars before it reaches results, MCP responses, or
analytics. Note: data_collection: deny is OpenRouter-specific — other
providers apply their own data policies. reasoning.max_tokens caps
hidden "thinking" on reasoning models so they cannot exhaust the whole
max_tokens budget before emitting output (TOKEN_LIMIT) — raise it in
botex.config.local.json for harder tasks (deep-merge keeps ZDR).
Model tiers. Each provider maps profile names (default, coding,
fast) to its own model slugs — the tier names are portable, the slugs are
provider-specific. When a call names neither model nor profile, the
capability mode picks the tier via engine.mode_profiles
(readonly→fast, destructive/full→coding, edit→default). Before
any API call, the provider's /models catalog also gates the resolved
model: over-cap price → PRICE_EXCEEDED; no tool-call support →
UNSUPPORTED (providers without capability metadata fall under
pricing.on_unknown). Providers that publish no capability fields can
declare them explicitly: providers.<name>.capabilities.<slug> = {"supports_tools": true} — the declaration wins over catalog metadata.
Persistent defaults. botex config shows the effective config;
botex config provider <name> / config model [provider] <slug> /
config profile <name> / config mode <preset> write your defaults to
botex.config.local.json (gitignored) — no need to repeat --profile /
--provider / --mode flags on every call.
Environment overrides: BOTEX_DEFAULT_MODEL, BOTEX_CODING_MODEL,
BOTEX_AUTO_BETA_MODEL (per-profile BOTEX_<PROFILE>_MODEL),
BOTEX_MAX_TURNS, BOTEX_MAX_TOKENS, BOTEX_TEMPERATURE,
BOTEX_BUDGET_USD, BOTEX_DEFAULT_MODE, BOTEX_DEFAULT_PROFILE,
BOTEX_EXEC_ENABLED, BOTEX_SNAPSHOT_DIR, BOTEX_ANALYTICS_FILE,
BOTEX_ENV_FILE, BOTEX_REFERER, BOTEX_ANALYTICS_FULL_TEXT.
Pricing guardrail. Before the first API call, the resolved model is
checked against the provider's published price list
(GET {base_url}/models, cached in gitignored .model_pricing.json).
Over-cap models are rejected locally with PRICE_EXCEEDED — zero cost:
"pricing": {
"enabled": true,
"max_input_per_mtok": 5.0,
"max_output_per_mtok": 5.0,
"on_unknown": "allow",
"cache_ttl_hours": 24
}0 disables a given cap. Router/meta models (openrouter/auto,
openrouter/auto-beta, openrouter/pareto-code) report dynamic pricing —
on_unknown decides whether they pass ("deny" forces explicit priced
models). Providers that do not publish prices (NVIDIA's /models returns
a catalog without pricing fields) fall under on_unknown as well —
their models are currently free-tier anyway. botex models lists configured profiles with live
prices; botex config price input|output <usd> persists caps. Env
overrides: BOTEX_PRICE_MAX_INPUT, BOTEX_PRICE_MAX_OUTPUT.
Analytics privacy. Ledger entries never persist secrets: run
summaries are secret-masked before being written to
.agent_analytics.json, and by default only a truncated preview is
kept — the full text survives solely as a summary_sha256 hash.
Set analytics.store_full_text (or BOTEX_ANALYTICS_FULL_TEXT=1) only
if you accept storing complete masked summaries on disk.
Note on ZDR accounts: if your OpenRouter key enforces Zero Data Retention at the account level, routing only reaches ZDR-compliant endpoints.
openrouter/autoresolves automatically; if a chosen model returns 404zdr-violation, pick another one viamodel/profile.
CLI
Interactive mode
botex repl -w ./my-project --profile coding
# botex> refactor the parser module
# botex> exitThe REPL prompts for destructive-op approval ([y/N]) when the agent
requests delete_file/move_file.
Run a task headlessly (no MCP client)
botex run "Add a --verbose flag to cli.py" -w ./my-project --profile coding
botex run "Fix the failing test" --files tests/test_app.py --max-turns 25
botex run "Remove the legacy module" -w ./proj --allow-destructiveFlags: --files, --workspace/-w, --model, --profile, --provider,
--mode, --max-turns, --max-tokens, --max-duration, --budget,
--api-key, --allow-destructive, --allow-exec, --allow-net,
--net-host (repeatable), --net-url (repeatable), --output-path,
--verify-command.
Exit code is 0 when the result ok field is true — that is, only on
DONE. A run that hits MAX_TURNS_REACHED/TIME_LIMIT after writing
files exits non-zero; the partial progress is still on disk and listed
in files_touched.
Analytics & maintenance
botex --stats # weekly report (tasks, tokens, cost per model)
botex --stats --month # monthly report
botex --history <task_id> # details of a specific run
botex --snapshots # snapshot disk usage
botex --clean-snapshots # prune old snapshots
botex health # environment readiness check(python server.py ... works identically — the launcher only resolves the
interpreter for you.)
Repository layout
server.py MCP server entry point (+ CLI dispatcher, main())
botex.cmd, botex.sh zero-install launchers (auto-detect interpreter)
pyproject.toml pip-installable; exposes the `botex` command
botex/
engine.py autonomous tool loop, context pruning, ZDR
config.py layered configuration loader
file_tools.py outline-first file I/O tools
patch_engine.py fuzzy patching, syntax validation, snapshots
security.py path traversal guard, secret masking, .gitignore awareness
capabilities.py capability modes (readonly/edit/destructive/full)
exec_tools.py policy-controlled run_command (allowlist, timeout, masking)
analytics.py cost/token ledger, budget checks, CLI reports
pricing.py pre-flight model price guardrail (provider catalog)
providers.py provider-neutral adapter: request scoping, response
normalization, key resolution, error sanitization
contracts.py result contracts: status/failure_kind taxonomy,
fallback policy, TaskContract resolution
i18n.py CLI localization (en/pl; MCP responses stay English)
ui.py ANSI styling helpers (CLI only, graceful fallback)
tests/test_botex.py unit & integration suite (python tests/test_botex.py)
tests/test_live_e2e.py opt-in live E2E suite against OpenRouter
(requires OPENROUTER_API_TESTS_KEY; skips otherwise)
docs/TUTORIAL.md technical deep dive
ROADMAP.md deferred features & design rationale
skills/botex/ SKILL.md — consumer-facing usage guide for agents
botex.config.json default configuration
LICENSE, NOTICE Apache-2.0 license + attribution/trademark noticesRuntime artifacts (.snapshots/, .agent_analytics.json,
.model_pricing.json, __pycache__/) are gitignored.
For agents consuming BoteX: skills/botex/SKILL.md is a
progressive-disclosure usage guide — when to delegate, how to scope
mode, pre-authorization rules, and status semantics. Copy it into your
agent's skills directory (e.g. .claude/skills/botex/, .agents/skills/)
so the calling agent knows the contract.
License & Trademark
BoteX is released under the Apache License 2.0 — see LICENSE.
Derivative works must retain the attribution notices in NOTICE.
The BoteX name and logo are trademarks of Jakub Grzesiak (https://jg-webtech.pl). The license covers the code, not the brand:
You may fork, modify, and redistribute the code under Apache-2.0.
You may NOT distribute forks or derivatives under the name "BoteX", nor use the name/logo in a way that suggests affiliation with or endorsement by the BoteX project.
Reasonable, customary use to describe origin ("based on BoteX", "fork of BoteX") is permitted and required by the NOTICE retention clause.
Available Tools
12 toolscheck_healthA
Checks BoteX environment readiness (provider keys, snapshot dir, disk).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It discloses what is actually checked—provider keys, snapshot directory, and disk—and the verb 'Checks' implies a non-mutating inspection. It does not explicitly state side-effect-free behavior, but nothing in the description suggests modification.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short sentence communicates the verb, target, and specific scope without any filler. The parenthetical list of checks is compact and informative, making the description easy to scan and parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless health-check tool with an output schema, this description is complete: it names the exact environmental areas inspected and implies the tool's role. No additional context is needed for an agent to select or invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there are no parameter semantics to clarify. The schema is empty and fully covers the input surface, so the description does not need to compensate for missing parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Checks') and resource ('BoteX environment readiness'), then enumerates the exact components checked: provider keys, snapshot dir, and disk. This clearly distinguishes it from sibling tools that handle tasks, memory, URLs, or model recommendations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: this tool is for verifying whether the BoteX environment is properly set up. It does not explicitly state when not to use it or compare it to alternatives, but its health-check role makes the intended use obvious enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clean_snapshotsC
Purges old or excess backup snapshots from disk.
| Name | Required | Description | Default |
|---|---|---|---|
| max_age_days | No | ||
| max_total_mb | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It indicates this is a destructive operation ('purges') but doesn't disclose whether snapshots are permanently deleted, if there's any recovery mechanism, required permissions, or whether there are safeguards. It lacks details on the impact on the system (e.g., could it delete snapshots in use?).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence with no waste. It front-loads the verb and resource. It is appropriately sized for a simple tool, though it could be slightly more descriptive without losing conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that this is a destructive maintenance tool with two parameters that control deletion criteria, and no annotations to clarify safety, the description lacks critical information about the deletion behavior, parameter interaction, and any safeguards. The output schema exists but is not shown; if it explains the outcome, that helps, but the description alone does not fully equip an agent to call this safely and correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. However, the description doesn't explain that max_age_days and max_total_mb are criteria for deletion. It only says 'old or excess' which loosely maps to age and size, but does not clarify how the parameters interact (e.g., are they both applied, or either/or?) or their units. The parameter names and defaults are self-explanatory, so a baseline of 3 is appropriate, but more explicit linkage would improve it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the action ('purges') and the resource ('backup snapshots'), and clarifies it removes 'old or excess' ones. However, it doesn't specify the exact criteria for what counts as 'old' or 'excess' beyond the parameters, and doesn't differentiate from any sibling tools (none appear directly related, but the description could be more precise about the scope).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives. It doesn't mention any prerequisites, conditions for invocation, or when to avoid using it. The agent is left to infer that it's for maintenance tasks, but there's no explicit context on when this should be called.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fetch_urlA
Fetch one public http(s) text document for caller-side research.
This tool is intentionally separate from BoteX's internal read_url: the orchestrating agent chooses the URL, so arbitrary public hosts do not have to be configured in net.allowed_hosts. SSRF protections, redirect revalidation, content-type checks, byte limits, and secret masking still apply. In net.policy='allowlist', only configured hosts may be fetched.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| max_bytes | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burdenmax and does a solid job: it discloses SSRF protections, redirect revalidation, content-type checks, byte limits, and secret masking. It stops short of describing exact outcomes on failure or blocked content, but the major safety behaviors are transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with no filler. The first sentence states the core operation, the second explains the separation from read_urlholistic, and the third lists safety constraints plus the allowlist exception. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose, safety behaviors, configuration implications, and policy restrictions, which is strong for a fetch tool. It is slightly incomplete because max_bytes is never explained in relation to the parameter, and expected failure modes are only implied. Still, the overall context is sufficient for most correct calls.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description should compensate for the undocumented parameters. It mentions byte limits in general but never maps that to the max_bytes parameter or explains its default semantics. The URL parameter is only implicitly covered by 'public http(s)' in the purpose clause. This is a meaningful gap for a tool with only two parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Fetch one public http(s) text document for caller-side research.' It also explicitly distinguishes the tool from BoteX's internal read_url, explaining that the orchestrating agent chooses the URL. This makes the tool's role clear and differentiates it from related helpers.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description names the alternative (read_url) and explains why fetch_url exists: arbitrary public hosts do not need to be configured in net.allowed_hosts. It also notes the allowlist policy restriction. It does not explicitly list exclusions or say 'use this when X, otherwise use Y,' but the context is clear enough for an agent to decide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_outlineA
Returns a lightweight file skeleton (classes, methods, functions, headers, line numbers) without reading the whole file. Saves ~90% of tokens.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| workspace_dir | No | . |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It clearly states that only a skeleton is returned, the whole file is not read, and token savings are significant. This is strong transparency for a read-style outline operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The first sentence front-loads the core purpose and deliverable, and the second adds a concrete efficiency benefit. Every part earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool itself is simple and the output schema likely covers return values, so the description does not need to explain returns. However, the complete lack of parameter documentation, especially workspace_dir, means the agent may still guess about path resolution. The description is adequate for purpose but not fully complete for invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It implies that 'path' refers to a file, but it does not explain the meaning or relationship of 'workspace_dir', nor how paths are resolved. The parameter semantics remain largely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states exactly what the tool does: it returns a lightweight file skeleton and enumerates the included elements (classes, methods, functions, headers, line numbers). This makes the tool's purpose specific and clearly differentiates it from sibling tools like fetch_url or read_memory.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'without reading the whole file' gives a clear use case: choose this tool when you need an overview or structure and want to save tokens. It does not explicitly name alternatives or when-not-to-use cases, but the intended context is strongly implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_statsA
Returns an analytics summary of executed BoteX tasks (USD costs, token usage, code-line balance). Args: period: 'week' (last 7 days) or 'month' (last 30 days).
| Name | Required | Description | Default |
|---|---|---|---|
| period | No | week |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the behavioral disclosure burden. It reveals the aggregation dimensions and valid period values, but it does not explicitly state that the operation is read-only, has no side effects, or how the data is scoped beyond the period. This is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loads the core purpose, and then provides the parameter guidance needed. No filler or redundant content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With one optional parameter, no nested objects, and an output schema present, the description covers everything needed to call the tool correctly. The period choices are documented, and return-value details are presumably covered by the output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description fully compensates by explaining the only parameter, period, with its allowed values and their exact meanings. This adds substantial value over the bare schema type and default.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Returns'), a resource ('executed BoteX tasks'), and concrete metrics (USD costs, token usage, code-line balance). It clearly differentiates this analytics tool from sibling tools like get_task_status or check_health by stating its summary scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: when an analytics summary over a period is needed. It does not explicitly contrast with alternatives or state when not to use it, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_task_statusA
Return status and result for a task started with start_task.
status is the task lifecycle (QUEUED | RUNNING | DONE | FAILED |
CANCELLED) — DONE here means "finished", not "succeeded"; check
result_data.status / result_data.ok for the engine outcome.
result is the pretty-text report; result_data is the structured
engine result (status, failure_kind, ok, files_touched, exec_ran,
rollback_verified, attempts[], cost_usd, ...).
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It explains the task lifecycle values, clarifies the subtle meaning of DONE as 'finished' rather than 'succeeded', and details the contents of result_data. This is strong transparency beyond the bare return-type information.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured. It leads with the primary purpose, then uses a concise bullet-like list to explain status and result fields. Every sentence adds useful information without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a status-retrieval tool with one parameter and an output schema, the description is complete. It covers the lifecycle semantics, the distinction between textual and structured results, and key result_data fields, so an agent has everything needed to interpret the response correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only task_id with no description, so the description must compensate. It does so by linking task_id to a task started with start_task, giving the parameter practical meaning. It could explicitly describe task_id format or source, but for a single parameter this is sufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Return status and result for a task started with start_task.' It clearly distinguishes the tool from the sibling start_task by focusing on retrieval of task outcomes, and it enumerates the exact fields returned.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear context for use: it is for tasks previously started with start_task. It does not explicitly name alternative tools or exclusion cases, but the association with start_task makes the intended usage unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_memoryA
Reads the full content and metadata of a specific memory entry by ID from the workspace Memory Vault.
Args: memory_id: The ID of the memory entry to retrieve (e.g. 'mem_a1b2c3d4'). workspace_dir: Workspace root directory (defaults to current directory).
| Name | Required | Description | Default |
|---|---|---|---|
| memory_id | Yes | ||
| workspace_dir | No | . |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. The verb 'Reads' makes it clear this is a non-mutating operation, and 'full content and metadata' sets expectations for what is returned. It does not mention error behavior or edge cases, but for a simple read operation the main behavioral trait is adequately disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured: a one-sentence purpose statement followed by a short Args list. Every sentence adds value, and the core action is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given two simple parameters, an output schema, and a read-only operation, the description is nearly complete. It explains what the tool does and what each parameter means. The only gap is the lack of explicit guidance on unknown IDs or alternative tool selection.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The Args section adds real meaning beyond the input schema: memory_id is explained with an example, and workspace_dir is described as the workspace root with its default behavior. Since schema description coverage is 0%, this parameter documentation fully compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Reads'), a specific resource ('memory entry from the workspace Memory Vault'), and the selection mechanism ('by ID'). It also says 'full content and metadata,' which distinguishes this from search_memory, which presumably finds entries without an ID.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'specific memory entry by ID' clearly implies this tool should be used when the caller already knows the memory_id. It does not explicitly mention alternatives or state when not to use it, so it stops short of full guidance, but the context is clear and unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recommend_modelsA
[EN] Fetches the most cost-effective and capable models from OpenRouter based on live benchmarks.
USE THIS TOOL BEFORE calling run_subagent if you are unsure which model ID to use or want to optimize for cost/quality.
[PL] Pobiera rekomendacje modeli z OpenRouter na podstawie benchmarków.
Użyj tego narzędzia ZANIM wywołasz run_subagent, jeśli nie znasz dokładnego ID modelu.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| task_type | No | coding |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the safety burden. It communicates a read-only external fetch from OpenRouter and that results depend on live benchmarks, which covers the main behavioral traits. It does not mention auth/rate-limit implications, but for a simple fetch the key traits are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The EN/PL sections are each two sentences and front-load the core action before the usage note. The duplication is minor and the labels keep it clear.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema and only two optional parameters, so the main missing piece is parameter semantics. Purpose and usage are clear, but an agent cannot know what task_type values are valid or how limit affects results, leaving a real gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage and the description never explains `limit` or `task_type`, nor enumerates valid task types. The only indirect hint is the default `task_type: coding`, which is not enough to compensate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action and object: 'Fetches the most cost-effective and capable models from OpenRouter based on live benchmarks.' It also differentiates its role by pointing to run_subagent, making it clear this tool is a model-recommendation helper rather than a task runner.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly instructs agents to call this tool before run_subagent when uncertain about the model ID or when optimizing cost/quality. This states the condition and the alternative tool, which is sufficient guidance for selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_subagentA
Runs BoteX — an autonomous code execution engine (agent-agnostic harness). BoteX performs multi-step tasks on local files inside workspace_dir.
TIP: If the user didn't specify a model, DO NOT guess. Use the recommend_models tool first to find the best model for this task!
Key engine features:
Outline-First and Context Pruning: ~90% token savings.
Pre-write Syntax Check: in-memory code validation before touching disk.
Multi-tier Fuzzy Patching: tolerant of CRLF/LF and indentation drift.
Hard Zero-Retention on OpenRouter (provider: data_collection=deny) and DLP secret censoring.
Automatic Snapshot Rollback on loops or critical failures.
Args: task: Precise description of the coding or refactoring task. files: Optional list of primary files affected by the task. workspace_dir: Project directory path (defaults to current directory). model: Any model from the provider's catalog. Empty = model from config (see 'profile' and botex.config.json). profile: Model profile from botex.config.json (e.g. 'default', 'coding', 'auto-beta', 'fast'). When both 'model' and 'profile' are empty, the capability mode picks the tier via engine.mode_profiles (readonly -> 'fast', destructive/full -> 'coding'). Ignored when 'model' is given explicitly. provider: API provider from the 'providers' section of botex.config.json ('openrouter', 'nvidia'). Empty = provider marked 'default: true' in the configuration. mode: Capability preset: 'readonly' (read only — the contract is an analysis report, DONE requires non-empty findings), 'edit' (read + edit — default), 'destructive' (+delete/move), 'full' (+run_command). Empty = engine.default_mode from config. Explicit allow_* flags may only WIDEN a preset — they never narrow it (readonly can never gain file writes). In headless mode this is pre-authorization — grant it consciously. max_turns: Maximum tool-loop steps (0 = config value). max_tokens: Per-turn completion cap (0 = config value; reasoning models need ~8000+ since thinking shares this budget). max_duration_s: Total wall-clock limit in seconds (0 = config value, 0 disables only when config is also 0). Checked between turns. budget_limit_usd: Daily spend limit in USD (negative = config value, 0 = no limit). allow_destructive: Authorize destructive operations (delete_file/ move_file) for this task. Headless mode cannot confirm mid-run — grant ONLY with the user's consent. allow_exec: Authorize run_command for this task (also requires exec.enabled=true in config). WARNING: this is NOT a sandbox — commands run with operator privileges. Grant only for trusted workspaces with the user's consent. api_key: Optional provider API key passed per-request (highest priority — overrides env/.env/config). Empty = resolved internally by the harness. output_path: Optional required output file (workspace-relative). Sets the file_output contract: DONE is accepted only when the file exists on disk, is non-empty, and passes the syntax gate; a DONE response carrying the payload as text is salvaged to disk. Requires a mutating mode. allow_net: Explicitly authorizes the read-only public web tool for this run. It is never enabled by a capability mode. net_allowed_hosts: Optional per-run host authorization. In net.policy='caller' these replace config defaults; in 'public' they narrow public access; in 'allowlist' they must stay inside net.allowed_hosts. net_allowed_urls: Optional per-run URL authorization rules. A URL ending in '/' authorizes that subtree; otherwise it authorizes the exact URL including its query string. verify_command: Optional allowlisted command that must pass before DONE is accepted — it runs against the workspace as it stands, including when the task made no writes (a passing verifier validates a legitimate no-change result). Requires exec.enabled=true in config and exec authorization. recipe: Optional operational persona / workflow prompt (e.g. 'planner', 'code-explorer', 'reviewer', 'security-reviewer', 'build-resolver', 'tdd').
Returns:
A pretty-text report in the text content (unchanged format for legacy
clients) plus the full engine result as structuredContent (status,
failure_kind, ok, files_touched, exec_ran, rollback_verified,
attempts[], cost_usd, ...). ok is true only when status == DONE —
i.e. the task contract was verified, not merely claimed by the model.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | ||
| task | Yes | ||
| files | No | ||
| model | No | ||
| recipe | No | ||
| api_key | No | ||
| profile | No | ||
| provider | No | ||
| allow_net | No | ||
| max_turns | No | ||
| allow_exec | No | ||
| max_tokens | No | ||
| output_path | No | ||
| workspace_dir | No | . | |
| max_duration_s | No | ||
| verify_command | No | ||
| budget_limit_usd | No | ||
| net_allowed_urls | No | ||
| allow_destructive | No | ||
| net_allowed_hosts | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | No | |
| steps | No | |
| status | No | |
| message | No | |
| summary | No | |
| task_id | No | |
| attempts | No | |
| cost_usd | No | |
| exec_ran | No | |
| duration_s | No | |
| lines_added | No | |
| failure_kind | No | |
| fallback_from | No | |
| files_touched | No | |
| lines_removed | No | |
| rollback_verified | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, and it does so thoroughly. It reveals snapshot rollback, pre-write syntax checks, zero-retention on OpenRouter, the non-sandbox nature of allow_exec, the output_path DONE contract, and that 'ok is true only when status == DONE.' This is far beyond minimal disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long, but the tool is genuinely complex with 20 parameters and multiple safety-sensitive modes. The structure is efficient: a front-loaded TIP, a compact feature list, and a well-organized Args section. Every sentence adds operational value rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, the description is complete: it covers model selection, mode semantics, safety authorizations, output contracts, verification behavior, and return-value semantics. The presence of an output schema reduces the need to explain return fields, yet the description still clarifies the critical ok/status relationship.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate entirely, and it does. Every one of the 20 parameters is explained with defaults, interactions, and warnings—for example, mode's capability presets, budget_limit_usd's negative-value semantics, and net_allowed_urls' trailing-slash subtree rule. This is exemplary parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Runs BoteX — an autonomous code execution engine' that 'performs multi-step tasks on local files inside workspace_dir.' This clearly differentiates it from sibling tools like recommend_models, fetch_url, and start_task by establishing it as the main task-execution harness.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit guidance for a key decision: 'If the user didn't specify a model, DO NOT guess. Use the recommend_models tool first.' It also warns about headless pre-authorization and when to grant destructive/exec permissions. It does not explicitly contrast with start_task, but the context is clear enough for an agent to know when this tool is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_memoryA
Saves a persistent context note, architectural decision, task handoff, or lesson into the workspace Memory Vault (/.botex/memory/). Masks secrets and provides directory isolation.
Args: title: Short descriptive title of the memory entry. content: Detailed markdown notes, architectural rationale, or handoff context. workspace_dir: Workspace root directory (defaults to current directory). kind: Entry category ('context' | 'decision' | 'handoff' | 'lesson'). tags: Optional list of keyword tags for filtering and discovery. memory_id: Optional custom identifier (alphanumeric/hyphen/underscore). Auto-generated if omitted.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | context | |
| tags | No | ||
| title | Yes | ||
| content | Yes | ||
| memory_id | No | ||
| workspace_dir | No | . |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and adds meaningful behaviors: 'Masks secrets and provides directory isolation', plus persistence to a specific vault path. It does not disclose overwrite behavior or failure modes, but the disclosed traits are genuinely useful beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact: one purpose sentence plus a structured Args block with no filler. It front-loads the primary action and then provides parameter details in a scannable format.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
All six parameters are covered, required fields are clear, and the output schema handles return values. The description gives the agent everything needed to call the tool correctly and confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the Args section compensates fully by clearly documenting every parameter, including kind's allowed values, memory_id format, tags as a list, and workspace_dir default. This is essential for correct invocation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Saves a persistent context note, architectural decision, task handoff, or lesson into the workspace Memory Vault'. This makes the write purpose unmistakable and clearly differentiates it from read/search siblings by emphasizing persistence and a target location.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: whenever a persistent note, decision, handoff, or lesson needs to be stored. It does not explicitly name alternatives or say when not to use it, but the purpose is specific enough that an agent can infer the correct situation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_memoryA
Searches the workspace Memory Vault (/.botex/memory/) by keyword query, tag, or entry kind (context, decision, handoff, lesson).
Args: workspace_dir: Workspace root directory (defaults to current directory). query: Free-text search string matched against title, content, and tags. tag: Filter by specific tag. kind: Filter by entry kind ('context', 'decision', 'handoff', 'lesson').
| Name | Required | Description | Default |
|---|---|---|---|
| tag | No | ||
| kind | No | ||
| query | No | ||
| workspace_dir | No | . |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the storage location and match semantics ('matched against title, content, and tags'), which is useful. However, it does not explicitly state whether the operation is purely read-only, how results are returned, or any limitations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose is front-loaded in a single precise sentence, followed by a compact Args list. Every sentence adds information: location, search dimensions, and parameter semantics. There is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the tool's purpose and all four parameters, and an output schema exists to define return values. It is sufficiently complete for invoking the tool correctly, though it could mention behavior when no filters are supplied or explicitly flag read-only status.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains every parameter: workspace_dir's role and default, query's matching behavior, tag filtering, and kind filtering with allowed enum values. This adds substantial meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Searches the workspace Memory Vault' and identifies the search dimensions ('keyword query, tag, or entry kind'). It clearly distinguishes itself from likely siblings like read_memory by focusing on search rather than direct retrieval.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the use case clear: agent should use this when searching memory by query, tag, or kind. It provides relevant context about the search scope and filters, but it does not explicitly state when not to use it or name alternatives like read_memory.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_taskA
Start a detached BoteX task and return its task_id immediately.
Runs the same engine as run_subagent, but returns instantly with a
task_id for polling via get_task_status. Concurrency is bounded by
engine.max_concurrent_tasks and the queue by engine.max_queued_tasks —
a full queue returns status TOO_MANY_QUEUED.
The capability/limits arguments are identical to run_subagent —
same contracts, same pre-authorization semantics:
mode: capability preset — 'readonly' (read tools only; the contract is an analysis report), 'edit' (read + write; default), 'destructive' (+ delete/move), 'full' (+ run_command). Empty = engine.default_mode.
allow_destructive / allow_exec / allow_net only ever WIDEN the preset, never narrow it — a readonly task cannot gain file writes. allow_exec additionally requires exec.enabled=true in the config. None of these is a sandbox: grant them only with the user's consent.
output_path makes the contract file_output: DONE requires the file to exist, be non-empty, and pass the syntax gate (mutating mode).
verify_command is an allowlisted command that must pass before DONE is accepted — it runs even when the task made no writes. Requires exec authorization.
recipe: Optional operational persona / workflow prompt (e.g. 'planner', 'code-explorer', 'reviewer', 'security-reviewer', 'build-resolver', 'tdd').
max_turns / max_tokens / max_duration_s / budget_limit_usd: 0 (negative for budget) = config value; a non-positive max_turns is clamped to 1.
net_allowed_hosts / net_allowed_urls narrow or replace the net scope per run according to net.policy (see run_subagent).
Returns:
{"task_id": str, "status": "QUEUED"} — poll with get_task_status.
The finished task's result_data carries the full engine result
(status, failure_kind, ok, files_touched, exec_ran,
rollback_verified, attempts[], cost_usd, ...).
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | ||
| task | Yes | ||
| files | No | ||
| model | No | ||
| recipe | No | ||
| api_key | No | ||
| profile | No | ||
| provider | No | ||
| allow_net | No | ||
| max_turns | No | ||
| allow_exec | No | ||
| max_tokens | No | ||
| output_path | No | ||
| workspace_dir | No | . | |
| max_duration_s | No | ||
| verify_command | No | ||
| budget_limit_usd | No | ||
| net_allowed_urls | No | ||
| allow_destructive | No | ||
| net_allowed_hosts | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the full behavioral burden. It discloses detached execution, queue-full behavior, capability preset semantics, widening-only allow flags, the non-sandbox warning, output_path file requirements, verify_command authorization, default/clamping behavior, and the return payload. This is far more transparent than typical tool descriptions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every section earns its place: purpose, engine comparison, queue behavior, argument semantics, and return values. The bullet-list structure makes the dense parameter details scannable, and the most important decision-relevant information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 20 parameters, no schema-level descriptions, and no annotations, the description is remarkably complete. It covers invocation semantics, concurrency and queue limits, return shape, polling route, and safety caveats, while also pointing to run_subagent for shared contracts. An output schema exists, and the description still summarizes the returned task_id/status and result_data fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description compensates with detailed bullet semantics for mode, allow_destructive/allow_exec/allow_net, output_path, verify_command, recipe, limit parameters, and net_allowed_hosts/urls. It also routes shared contracts to run_subagent. The remaining parameters are self-explanatory from their names, types, and default values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific verb and resource: 'Start a detached BoteX task and return its task_id immediately.' It also contrasts with the sibling runner run_subagent ('returns instantly'), so an agent can immediately distinguish this asynchronous submission tool from the synchronous alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the usage context clear: use this tool when you want the same engine as run_subagent but need a task_id to poll via get_task_status. It also gives operational conditions such as bounded concurrency and TOO_MANY_QUEUED. It does not explicitly state a when-not-to-use rule, but the contrast with run_subagent and the polling pointer provide sufficient guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
12 tool updates
v1.0.0- First observed
check_health - First observed
clean_snapshots - First observed
fetch_url - First observed
get_outline - First observed
get_stats - First observed
get_task_status - First observed
read_memory - First observed
recommend_models - First observed
run_subagent - First observed
save_memory - First observed
search_memory - First observed
start_task
TDQS
Scored across 12 tools
Most tools target distinct resources or actions, but run_subagent and start_task are near-duplicates sharing the same engine and argument contracts, differing mainly in sync vs. async execution. The descriptions clearly clarify the intended usage, so confusion is possible but manageable.
All tool names follow a consistent snake_case verb_noun pattern: get_task_status, start_task, save_memory, clean_snapshots, etc. There are no mixed casing styles or vague generic verbs.
Twelve tools is well within the ideal range and each tool serves a distinct operational need: task execution, async polling, model selection, memory vault access, stats, health, and snapshot cleanup. No tool feels superfluous.
The core workflows are well covered: synchronous and detached task execution, status polling, memory persistence and lookup, model recommendation, and environment maintenance. Minor gaps exist, such as no explicit cancel_task and no update/delete for memory entries, but agents can work around these in most scenarios.
Maintenance
Related MCP Connectors
A paid remote MCP for OpenAI Codex agent coordination MCP, built to return verdicts, receipts, usage
Control plane for autonomous software labor. Agents claim objectives over MCP with audit trail.
Build and publish full-stack apps from your coding agent: models, rules, pages, auth, per-app MCP.
System-of-record notebook for AI coding agents: pages, datastores, tasks, skills over MCP.
Related MCP Servers
- AlicenseAqualityBmaintenanceDelegate coding tasks to external AI coding agents in isolated git worktrees with independent verification, enabling any MCP client to orchestrate multi-agent workflows.5MIT
- AlicenseNot gradedqualityBmaintenanceEnables any MCP client to launch and manage subagent sessions in installed coding agents like Codex, Claude Code, Grok, and OpenCode, using your existing logins and chosen models.15 npmMIT
- AlicenseNot gradedqualityAmaintenanceEnables MCP-capable coding assistants to delegate repository investigation, bounded implementation work, and noisy command runs (tests, builds, linters) to sandboxed OpenCode agents. Each role can use an independently selected model, and only concise results are returned to the parent agent, which keeps responsibility for architecture and high-risk operations.MIT
- AlicenseNot gradedqualityBmaintenanceEnables MCP clients like Codex and DSH to invoke DeepSeek subagents for inspect, plan, review, and implement tasks, with write protection, rollback, and status reporting.10 npmMIT