Skip to main content
Glama
thestandard-production

openmausbot-cua-mcp

OpenMausBot CUA MCP

A local Model Context Protocol server that connects Claude, MiniMax, Codex, and other MCP clients to the cua-driver bundled with OpenMausBot.

The server discovers OpenMausBot's current embedded socket on every CUA call. It never listens on a TCP port or installs a background daemon. Admin tools connect outward to OpenMausBot's loopback API, and all tools preserve OpenMausBot's permission model.

Requirements

  • macOS

  • OpenMausBot running with Computer enabled

  • Python 3.11–3.13

  • uv recommended for uvx

Related MCP server: GPT Windows Connector

Install in Claude Desktop

Add this server to claude_desktop_config.json:

{
  "mcpServers": {
    "openmausbot-cua": {
      "command": "uvx",
      "args": [
        "--from",
        "git+https://github.com/thestandard-production/openmausbot-cua-mcp.git",
        "openmausbot-cua-mcp"
      ]
    }
  }
}

Restart Claude Desktop after saving the configuration.

Install in Claude Code

Put the same mcpServers block in your project .mcp.json, or run:

claude mcp add --transport stdio openmausbot-cua -- \
  uvx --from git+https://github.com/thestandard-production/openmausbot-cua-mcp.git \
  openmausbot-cua-mcp

Install in MiniMax

Merge examples/minimax.json into your MiniMax MCP configuration, then start a new session.

Install in Codex

Copy examples/codex.toml into your Codex configuration.

Tools

Tool

Effect

openmausbot_status

Resolve the live connection and report daemon status

openmausbot_list_cua_tools

List tools exposed by the installed cua-driver

openmausbot_describe_cua_tool

Show one tool's description and input schema

openmausbot_call_cua_tool

Invoke a cua-driver tool; may control apps or modify state

openmausbot_get_companion_settings

Read companion settings

openmausbot_set_companion_setting

Atomically update one existing companion setting

openmausbot_cua_skills_status

Check agent skill-pack installation

openmausbot_cua_permissions_status

Check Accessibility and Screen Recording status

openmausbot_cua_check_update

Check for cua-driver updates without installing

Typical sequence:

  1. Call openmausbot_status.

  2. Call openmausbot_list_cua_tools.

  3. Call openmausbot_describe_cua_tool for the chosen tool.

  4. Call openmausbot_call_cua_tool with an arguments object matching that schema.

Admin API (read-only)

The server also exposes bounded administration views from OpenMausBot's local API. Sensitive webhook values and provider account emails are removed before results reach the MCP client.

Tool

Result

openmausbot_api_health

API health, installed app version, and compatibility findings

openmausbot_list_bots

Bots, omitting soul text by default

openmausbot_list_routines

Routines and runs, optionally within a millisecond range

openmausbot_list_webhooks

Webhooks and attempts with sensitive values redacted

openmausbot_usage

Usage and billing summaries for a date range of up to one year

openmausbot_list_models

Provider instances, model options, and effort levels

openmausbot_list_decisions

Recent decisions, with a limit from 1 to 500

openmausbot_export_team

Team export in manifest, package, or backup format

openmausbot_export_team calls OpenMausBot's export endpoint, which is a POST route and therefore requires a paired-device session token. Pairing is always completed through the OpenMausBot app; this package does not create sessions.

The same views are available from omb-ctl. JSON is the default output. Bots and routines also support a compact text table.

omb-ctl status
omb-ctl bots --text
omb-ctl bots --include-soul
omb-ctl routines --since-hours 24 --text
omb-ctl webhooks
omb-ctl usage --from 2026-01-01 --to 2026-01-31 --group-by model
omb-ctl models
omb-ctl decisions --limit 100
omb-ctl export --format manifest

The Admin API client checks 127.0.0.1 ports 8799, 18799, and 28799 when no origin or port override is configured. The following environment variables control this client:

Variable

Purpose

OPENMAUSBOT_URL

API origin override; HTTP is restricted to loopback hosts

OMB_PORT

Loopback API port override used when no URL override is set

OPENMAUSBOT_TOKEN

Paired-device session token

OPENMAUSBOT_TOKEN_KEYCHAIN_SERVICE

macOS Keychain service containing the token

OPENMAUSBOT_TOKEN_KEYCHAIN_ACCOUNT

Optional Keychain account used with the service

OPENMAUSBOT_API_TIMEOUT

API request timeout in seconds; default 10

OPENMAUSBOT_APP_PLIST

App Info.plist override for version detection

OPENMAUSBOT_ENABLE_ADMIN_WRITES

Set to 1 to enable MCP administration write tools

OPENMAUSBOT_MCP_ALLOWLIST

Optional comma-separated MCP server names allowed on bots

Port discovery never sends the token, and a token is only ever sent to an origin you named explicitly with OPENMAUSBOT_URL or OMB_PORT (the same rule as OpenMausBot's own MCP client): a port found by probing could belong to another local process. Reads against a discovered port work without the token; requests that need it fail with a message asking you to set the origin. No request sends a browser Origin header. Server error messages are passed through (bounded).

Admin API (writes)

Administration writes use two independent gates:

  • Every mutation requires a paired-device session token and an explicit API origin through OPENMAUSBOT_URL or OMB_PORT.

  • MCP write tools are disabled unless OPENMAUSBOT_ENABLE_ADMIN_WRITES=1. The omb-ctl CLI does not use that environment gate; it requires --apply for each mutation.

All write tools and CLI commands are dry runs by default. An applied PATCH or POST is followed by a fresh GET and field-by-field verification. If the read-back does not match, the result has "ok": false, "status": "unknown", and lists the mismatched fields. Routine creation is idempotent by exact routine name because the public create endpoint has no idempotency key; duplicate names must be resolved in the app first.

MCP write tools include guarded bot updates and model selection, routine upsert/enable/run/delete, and run cancellation. openmausbot_plan is read-only and plans a desired-state file; applying a desired-state item is intentionally CLI-only.

# Dry run (no mutation)
omb-ctl bot-update research-bot --json '{"title":"Research"}'
omb-ctl routine-enable daily-summary

# Apply and verify
omb-ctl bot-update research-bot --json '{"title":"Research"}' --apply
omb-ctl routine-enable daily-summary --apply

Desired-state files are JSON, with YAML available through the optional yaml extra. Referenced files must be relative to the desired file and stay within its directory. Only fields present in a bot entry are managed.

{
  "version": 1,
  "bots": [
    {
      "name": "research-bot",
      "title": "Research",
      "soul_file": "souls/research.md",
      "soul_append_files": ["souls/common.md"],
      "computer": "off",
      "browser": false,
      "composio": false,
      "mcpServers": [],
      "approvalMode": "ask",
      "allow_loosen": false
    }
  ],
  "routines": [
    {
      "name": "daily-summary",
      "bot": "research-bot",
      "prompt_file": "prompts/daily-summary.md",
      "schedule": {"type": "daily", "time": "09:00", "weekdays": [1, 2, 3, 4, 5]},
      "enabled": true,
      "runOn": "maus",
      "durationMinutes": 30,
      "overlap": "skip"
    }
  ]
}

Bot computer accepts off, browser, cloud, vm, local, or null. OpenMausBot omits unset fields from bot objects: an unset computer means Auto (it may resolve to the local computer) and an unset mcpServers means every configured MCP server. The loosening guard ranks reach as off < browser < cloud = vm < local = Auto, and treats resetting mcpServers to null as widening.

composio (connected apps) and browser (built-in browser) are booleans that OpenMausBot treats as on unless explicitly false. Connected apps are one switch per bot: a bot with composio on can use every connected account, not a chosen subset. Turning either on needs allow_loosen.

cwd sets the bot's working folder (an absolute path, or null to clear it). OpenMausBot pins a thread's folder on its first turn, so a new cwd applies to new threads only. Any change needs allow_loosen, because it changes which files the bot works on.

approvalMode can only be set to ask. Loosen approvals in the OpenMausBot app, where a person sees each change.

Routine schedules are validated against, and normalized to, the shape OpenMausBot stores, so an unchanged desired state plans as a no-op and a read-back compares equal:

type

Fields

Notes

once

at

Epoch milliseconds or RFC3339 with an offset (converted to epoch ms)

daily

time, weekdays

HH:MM; weekdays as 0-6 (Sunday = 0) or names; omitted means every day

interval

everyMinutes, anchorAt, optional weekdays, window, endsAt

5-1440 minutes; anchorAt (first run) is required; all seven weekdays is stored as no restriction

cron

expression, timeZone

Five fields; an IANA timeZone such as Asia/Bangkok or UTC is required

Routine name and prompt are trimmed the same way the app trims them.

Plan the entire file, then select exactly one item to apply:

omb-ctl plan examples/desired.example.json
omb-ctl apply examples/desired.example.json --only bot:research-bot
omb-ctl apply examples/desired.example.json --only bot:research-bot --apply

CLI exit codes are 0 for success or no change, 10 when a plan or dry run contains a change, and 2 for an error, unknown write status, or verification mismatch.

Configuration

The defaults work with a normal OpenMausBot installation. These environment variables are available for custom installations and testing:

Variable

Purpose

OPENMAUSBOT_DATA_DIR

Override the OpenMausBot data directory

OPENMAUSBOT_CUA_CONNECTION_FILE

Override cua-connection.json

OPENMAUSBOT_CUA_DRIVER

Override the cua-driver executable

OPENMAUSBOT_CUA_SOCKET

Override the live Unix socket

OPENMAUSBOT_COMPANION_SETTINGS_FILE

Override companion settings JSON

OPENMAUSBOT_BUNDLE_ID

Override the OpenMausBot bundle identifier

OPENMAUSBOT_CUA_MAX_OUTPUT_CHARS

Bound stdout and stderr returned to the client

Security

The generic call tool can perform actions on the local computer. Its MCP annotations mark it as state-changing and potentially destructive. Keep your MCP client's approval controls enabled and review requested tool calls.

The server only forwards CUA_DRIVER_* keys from OpenMausBot's connection file. It never reads or stores API keys.

See SECURITY.md for the permission boundary and reporting process.

Development

python3 -m venv .venv
.venv/bin/python -m pip install -e '.[dev]'
.venv/bin/ruff check .
.venv/bin/pytest -q
.venv/bin/python -m build

Scope

The initial release targets macOS because OpenMausBot currently exposes its embedded connection through the macOS application data directory. Contributions for verified installations on other platforms are welcome.

OpenMausBot and cua-driver are separate third-party projects. This repository is an independent integration maintained by Production Craft.

Available Tools

25 tools
openmausbot_api_healthA
Read-onlyIdempotent

Return local API health, app version, and compatibility findings.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the description need not repeat these. The description adds a 'local' scope and specific output categories, but does not disclose additional behavioral traits like authentication requirements or potential latency. It neither contradicts nor significantly enriches the annotation-provided safety profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, direct sentence with no filler or redundancy. It front-loads the action and lists the key outputs clearly, making it easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only tool with an output schema, the description sufficiently covers what the tool returns (health, version, compatibility). The output schema provides the detailed structure, so the description need not do more. Nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters and the schema coverage is 100%, so the description is not expected to explain parameter details. Per the rubric, a zero-parameter tool earns a baseline of 4; the description adequately conveys that no input is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Return') and names the resource ('local API health, app version, and compatibility findings'). It clearly distinguishes what the tool produces, though it does not explicitly differentiate from the sibling 'openmausbot_status' which may overlap in purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives such as 'openmausbot_status' or other admin tools. There is no mention of context, prerequisites, or situations where this tool is preferred, leaving the agent to infer.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openmausbot_call_cua_toolA
Destructive

Invoke one cua-driver tool; some tools can control apps or modify local state.

The connected OpenMausBot daemon applies its own permission mode. Call openmausbot_describe_cua_tool first when the input schema is not known.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesExact cua-driver tool name returned by openmausbot_list_cua_tools.
argumentsNoJSON object matching the schema returned by describe_cua_tool.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide destructiveHint=true and idempotentHint=false, and the description adds valuable context by warning that some cua-driver tools 'can control apps or modify local state' and that the daemon enforces its own permission mode. This goes beyond the structured annotations and helps the agent anticipate side effects and permission-related failures.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three short sentences with no filler. The core action is front-loaded, followed by a concise risk warning and a clear prerequisite. Every sentence contributes new information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a dynamic generic invoker, the description covers the essential operational context: what it does, that side effects vary, that daemon permissions apply, and how to discover an unknown schema. With a 100%-covered input schema and an output schema present, no critical guidance is obviously missing, though the description could slightly expand on what happens when the daemon denies permission.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already documents that 'name' comes from openmausbot_list_cua_tools and that 'arguments' must match the schema returned by describe_cua_tool. The description's mention to call describe_cua_tool first is useful operational guidance but does not add meaning beyond the parameter schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific action: 'Invoke one cua-driver tool', which clearly identifies the tool as the execution entry point rather than a listing or describing operation. The warning that 'some tools can control apps or modify local state' adds important scope and helps distinguish this invoker from siblings like openmausbot_list_cua_tools and openmausbot_describe_cua_tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly tells the agent to 'Call openmausbot_describe_cua_tool first when the input schema is not known', which is a clear prerequisite and usage rule. It also notes that the connected daemon applies its own permission mode, signaling that invocation may be gated. It does not explicitly state when not to use the tool, but the guidance is strong and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openmausbot_cancel_runC
Destructive

Plan or cancel one routine run.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYesExact routine run id.
dry_runNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations include destructiveHint=true, but the description does not elaborate on the destructive nature of the action. It does not mention any side effects, such as whether canceling a run is irreversible, whether it affects the routine's future schedule, or whether it requires special permissions. With annotations present, the description adds little beyond the annotations, and for a destructive tool, more transparency would be expected.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short, only 6 words. It is concise, but the structure is overly brief to the point of being vague. It is not front-loaded with the most critical information (that it cancels a run), but it does get to the point quickly. The brevity is not a problem itself, but the ambiguity from the word 'Plan' detracts from its effectiveness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with only 2 parameters, the description is incomplete. It does not explain what 'dry_run' does, which is a key parameter with a default of true, nor does it clarify the difference between canceling a run and other related tools. The output schema exists, so return format is not a concern, but the description lacks guidance on prerequisites (e.g., does the run need to be in a certain state?) and what exactly cancelling means for the routine.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema describes 'run_id' as 'Exact routine run id' and 'dry_run' with a default of true. The description does not add any additional meaning beyond the schema. Coverage is 50%, but since both parameters are documented in the schema, the description does not need to compensate much. However, it would be helpful to explain the effect of 'dry_run' (e.g., whether it simulates the cancellation) and how to obtain a valid run_id.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Plan or cancel one routine run' is somewhat vague because it mentions both 'Plan' and 'cancel' while the tool is specifically named cancel_run. It does state the resource (routine run) and implies the action, but it is ambiguous about whether the tool also plans runs, which could confuse agents. Sibling tools like openmausbot_plan and openmausbot_run_routine_now exist, so clarity on the cancel-only scope would have been better.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives. With many sibling tools like openmausbot_plan, openmausbot_run_routine_now, and openmausbot_delete_routine, it is critical to distinguish between canceling a scheduled run, planning a run, and deleting a routine. No usage context is provided, leaving the agent to infer based on the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openmausbot_cua_check_updateA
Read-onlyIdempotent

Check the configured cua-driver release channel without installing anything.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already signal readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds small context by clarifying that no installation occurs and that the check targets the configured release channel, but it does not address return behavior or any other runtime details. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence communicates the action, scope, and key side-effect disclaimer with no filler. Every phrase adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read-only check with an output schema and safety annotations, the description is sufficient. Nothing about invocation or output format is needed because the output schema covers return structure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there is no parameter documentation burden. The description's reference to the 'configured' release channel confirms the tool operates on existing configuration rather than requiring inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action ('Check'), a specific target ('configured cua-driver release channel'), and explicitly scopes out installation ('without installing anything'). This clearly distinguishes the tool from a hypothetical update/install action and from sibling status/list tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description conveys that this is a safe, non-installing check, so an agent can infer it is appropriate when merely inspecting update availability. It does not explicitly state when to prefer it over siblings or when not to use it, leaving the routing largely implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openmausbot_cua_permissions_statusA
Read-onlyIdempotent

Report Accessibility and Screen Recording permission status.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the operation safe (readOnlyHint=true, idempotentHint=true, destructiveHint=false), so the description does not need to prove the absence of side effects. It adds value by enumerating the specific permission types covered (Accessibility, Screen Recording), which goes beyond the annotation title.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence that communicates purpose and scope with no filler. Every word contributes to the agent's understanding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only status tool with an output schema and safety annotations, the description is functionally complete. It misses only explicit usage context and the macOS framing, though the annotation title supplies the macOS context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

This tool has zero parameters and the schema fully documents an empty property set, so there is no parameter information for the description to add. The baseline for a no-parameter tool applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Report') and a specific resource ('Accessibility and Screen Recording permission status'). It also clearly distinguishes this tool from sibling status tools like cua_skills_status and cua_check_update by naming the exact permission categories being reported.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, when not to use it, or what conditions would make it the right choice. The agent must infer usage entirely from the tool name and general context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openmausbot_cua_skills_statusA
Read-onlyIdempotent

Report cua-driver skill pack installation status for supported agents.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, covering the safety profile. The description adds modest context by framing the operation as a status report, but does not add deeper behavioral detail such as what data is returned or how statuses are computed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. Every word contributes to the tool's purpose and scope.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only, idempotent status tool with an output schema present, the description is sufficient. An agent has everything it needs to select and invoke the tool correctly without additional behavior or parameter details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the description has no parameter documentation burden. The baseline for a zero-parameter tool is 4, and nothing in the description contradicts or fails to support the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Report') and a specific resource ('cua-driver skill pack installation status') with a scope qualifier ('for supported agents'). This clearly distinguishes it from sibling tools such as cua_permissions_status and cua_check_update, which cover different concerns.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context: this tool is for checking installation status of the cua-driver skill pack. It does not explicitly name alternatives or exclusion conditions, but the intent is evident and the zero-parameter signature makes misuse unlikely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openmausbot_delete_routineC
Destructive

Plan or delete one exactly confirmed routine.

ParametersJSON Schema
NameRequiredDescriptionDefault
dry_runNo
routineYesExact routine id or name.
confirm_nameYesExact current routine name.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, so the description adds the confirmation requirement ('exactly confirmed') and hints at a dry-run via 'Plan'. However, it does not explicitly state that dry_run defaults to true and that the tool may not actually delete unless dry_run is set to false. It adds some behavioral context but is vague and could mislead an agent into thinking it always deletes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence, which is concise, but it sacrifices clarity for brevity. It front-loads nothing important and omits critical usage details. It is appropriately sized in length but not in content – being short is not the same as being well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive tool with three parameters and an output schema, the description is severely incomplete. It does not explain the confirmation flow, the consequences of deletion, or the dry_run behavior. An agent would lack enough context to safely invoke this tool without additional research.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67% – routine and confirm_name have descriptions, but dry_run has no description in the schema. The description does not clarify dry_run's role or the importance of setting it to false for actual deletion. It does not add any parameter semantics beyond what the schema provides, and fails to compensate for the dry_run gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Plan or delete one exactly confirmed routine' is ambiguous – it introduces 'Plan' which is not the tool's core purpose (deletion) and is contradicted by the tool name and annotation title. It does mention 'delete' and 'routine' but the inclusion of 'Plan' muddies the intent. It does not differentiate from siblings like upsert_routine or set_routine_enabled.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. It does not mention that this is for deleting routines specifically, nor does it provide any context about prerequisites or situations where it should be avoided. The destructive nature is only hinted by annotations, not explained in the description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openmausbot_describe_cua_toolA
Read-onlyIdempotent

Return the description and JSON input schema for one cua-driver tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesExact cua-driver tool name returned by openmausbot_list_cua_tools.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds that it returns the tool's description and JSON input schema, which is useful context about the output but not additional behavioral detail. No contradictions with annotations; the description aligns with the read-only nature.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, succinct sentence with no filler. It states the action and the resource efficiently. Every word earns its place, and the most important information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple: one required parameter with a clear description, an output schema provided (as indicated by context), and annotations covering safety. The description is sufficient for an agent to correctly invoke it and understand what it will return. There are no missing pieces for a task of this simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description for the 'name' parameter is highly informative: 'Exact cua-driver tool name returned by openmausbot_list_cua_tools.' This tells the agent exactly where to obtain the value, adding meaning beyond the basic type and pattern constraints. Schema coverage is 100%, so the description complements rather than compensates.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Return') and the resource ('description and JSON input schema for one cua-driver tool'). It distinguishes itself from siblings like openmausbot_list_cua_tools (which lists all tools) and openmausbot_call_cua_tool (which invokes a tool). An agent can immediately understand what this tool does and how it differs from related tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: to fetch metadata for a specific tool. The parameter description explicitly states that the name must be 'returned by openmausbot_list_cua_tools', providing a clear prerequisite and indicating that one should list tools first. It doesn't explicitly state when not to use it, but the context is obvious given the sibling names and the parameter guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openmausbot_export_teamB
Read-onlyIdempotent

Export team data using a paired-device session token.

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNoTeam export format.manifest

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only, idempotent, and non-destructive behavior, so the bar for the description is lower. The description adds the requirement of a 'paired-device session token', which is a behavioral precondition not present in the schema or annotations. However, it does not explain what the token is, how to obtain it, or what happens without it. This is a small addition but not comprehensive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that immediately states the action and a key requirement. There is zero filler, and it is front-loaded with the verb. It is as concise as possible while conveying the essential purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple with one optional parameter and an output schema that presumably describes the return structure. Annotations cover safety. The main gap is the unclear 'paired-device session token' — it is mentioned but not explained, and it is not a parameter in the schema. This could confuse an agent about how to supply it. For a simple tool, this is a minor but notable omission; a 3 reflects that it is adequate but not fully self-contained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% for the only parameter (format) with its enum and default fully documented. The description does not add any meaning about the parameter beyond what the schema already states. It mentions 'team data' but not the format choices. Per the rubric, with high coverage, a baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Export team data') and the resource ('team data'), with a specific qualifier ('using a paired-device session token'). It distinguishes itself from sibling tools like list/update by focusing on export, but does not explicitly name alternatives or what makes it different. The annotation title adds a bit more context but the description alone is sufficiently clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus others. It does not mention any alternatives, prerequisites, or conditions for selection. The only hint is the session token requirement, which is not elaborated. An agent is left to infer the appropriate context, which is inadequate for a tool with many siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openmausbot_get_companion_settingsA
Read-onlyIdempotent

Read OpenMausBot's companion settings JSON file.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety and side-effect profile. The description adds no extra behavioral context (e.g., whether the returned JSON is structured a certain way, any caching, or auth needs). With annotations present, the bar is lower, and the description does not contradict them, but it also doesn't enrich beyond the annotations. A 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence that immediately communicates the action and resource. No filler, no redundant phrasing. It is appropriately front-loaded and earns every word.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no parameters and an output schema (indicated by 'Has output schema: true'), the description is sufficient for a simple read operation. It tells the agent what the tool does. It doesn't mention edge cases (e.g., missing file), but the output schema likely covers return structure. For a trivial getter, this is nearly complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and schema description coverage is 100% (trivially, since none exist). The description doesn't need to explain parameters. For 0-param tools, baseline 4 is prescribed, and the description doesn't detract from that. It's a simple read with no input semantics to clarify.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action ('Read') and a well-defined resource ('OpenMausBot's companion settings JSON file'). This distinguishes it from the sibling 'set_companion_setting' by opposing read vs. write operations, though it doesn't explicitly contrast with the status or cua tools. It's specific but not fully differentiated from all siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool should be used when one needs to read the companion settings, but it offers no explicit guidance on when to choose it over alternatives or any prerequisites. There is no mention of when not to use it or when to use the setter instead. For a simple read tool, 'implied usage' is acceptable, which aligns with a score of 3.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openmausbot_list_botsA
Read-onlyIdempotent

List bots without long soul text unless explicitly requested.

ParametersJSON Schema
NameRequiredDescriptionDefault
include_soulNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior. The description adds useful context beyond annotations by disclosing that soul text is omitted unless explicitly requested, which informs the agent about default response size and content. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one tight, front-loaded sentence. It states the action first and then the key behavioral qualifier, with zero wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a simple list tool with one optional parameter, a read-only annotation profile, and an output schema available. The description covers the tool's purpose, the key default behavior, and the parameter's effect, which is sufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description carries the burden for parameter meaning. It directly explains the include_soul parameter by stating soul text is excluded unless explicitly requested, giving the boolean parameter practical semantics beyond the schema's bare title.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource ('List bots') and adds a meaningful behavioral qualifier about excluding long soul text by default. It is clearly distinct from sibling list tools like list_routines or list_webhooks.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for listing bots but gives no explicit guidance about when to choose it over alternatives or when not to use it. The parameter hint ('unless explicitly requested') provides context, but no sibling comparisons or exclusions are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openmausbot_list_cua_toolsA
Read-onlyIdempotent

List every tool exposed by the cua-driver bundled with OpenMausBot.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds useful scoping context ('cua-driver bundled with OpenMausBot') but does not disclose additional behavioral details such as output structure or whether the list is static or dynamically discovered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The entire description is one tight sentence with no filler. The verb and scope are front-loaded, and every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a simple, zero-argument, read-only listing tool with rich annotations and an output schema. Nothing an agent needs to call it correctly is missing, and the description fully covers the operation's scope.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, and the schema is fully empty and unambiguous. With no parameters to explain, the description does not need to compensate for schema gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') and a precise resource ('every tool exposed by the cua-driver bundled with OpenMausBot'). It clearly distinguishes this inventory tool from siblings that describe or call individual CUA tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit guidance about when to use this tool versus alternatives such as openmausbot_describe_cua_tool or openmausbot_call_cua_tool. The intended use is only implied by the word 'List', not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openmausbot_list_decisionsB
Read-onlyIdempotent

List recent decisions.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoNumber of decisions to return.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, indicating a safe read operation. The description adds nothing about behavior beyond the schema's limit parameter (which enforces maximum 500). Since annotations cover the safety and mutation profile, the description doesn't need to repeat that, but it also doesn't add any extra behavioral detail like pagination or ordering. A 4 is given because annotations provide solid coverage and the description is not contradictory.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence, which is efficient. It front-loads the core action and resource. However, it is somewhat terse and could benefit from a brief clarification of what 'decisions' refers to. Still, it earns a 4 for being concise and not padded with filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple: one optional parameter, read-only, idempotent, and no nested objects. An output schema exists, which likely describes the return format, so the description doesn't need to explain that. For a list operation with a well-documented schema and annotations, this is minimally viable. However, the ambiguity of 'decisions' is a gap – an agent might need to know what decisions are (e.g., a decision log, a bot's past choices) to use it correctly. Given the simplicity, a 3 is appropriate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema documents the 'limit' parameter with a clear description ('Number of decisions to return') and default/min/max values. The description mentions 'recent decisions' but doesn't explain how the limit interacts with recency (e.g., are they sorted by date descending? Does a higher limit include older decisions?). With 100% schema coverage, baseline is 3, and the description adds minimal extra semantic value beyond implying a time-based filter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the verb 'list' and the resource 'decisions', but it is quite minimal and doesn't specify what 'decisions' means in this context (e.g., decisions made by a bot, user decisions, etc.). It doesn't differentiate from sibling tools like list_bots or list_routines, though the resource name is unique enough. The title and annotations add a bit more context, but the description itself is vague.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool instead of others. There is no mention of alternatives or typical scenarios. The only hint is the resource name, which distinguishes it from other list tools, but no explicit exclusions or conditions are given. An agent would have to infer usage from the resource name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openmausbot_list_modelsA
Read-onlyIdempotent

List provider instances and models without account details.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, covering the safety profile. The description adds 'without account details,' which is useful context about output scope, but it does not mention pagination, return format, or any other runtime behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. Every word contributes to clarifying the operation and scope.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the zero-parameter interface, existing output schema, and annotations, the description covers the core need: an agent knows what the call does and that it is safe and idempotent. It could briefly define 'provider instances,' but for a no-argument list, the current context is sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline is 4. The description's clause about 'without account details' reinforces that no account credentials are needed, which adds a slight semantic cue beyond the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') and resource ('provider instances and models'), making the operation immediately clear. The qualifier 'without account details' scopes the result furtherable. Among sibling list tools (bots, routines, webhooks), the resource is distinct, so an agent can differentiate it without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies a read-only enumeration use case, but it does not explicitly state when to prefer this tool over alternatives like list_bots or list_routines. It offers no exclusions or comparison to sibling tools, leaving the agent to infer usage from the resource name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openmausbot_list_routinesA
Read-onlyIdempotent

List routines and runs, optionally within a millisecond range.

ParametersJSON Schema
NameRequiredDescriptionDefault
to_msNoOptional range end in Unix ms.
from_msNoOptional range start in Unix ms.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, covering the safety profile. The description adds that runs are included and that a millisecond range is supported, but it does not mention pagination, ordering, or response behavior. This is useful but not rich behavioral context beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler. The verb, resource, and optional constraint all appear immediately, and every word contributes meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool with no required parameters, complete schema documentation, and an output schema, this description is nearly sufficient. The main gap is clarifying whether the millisecond range applies to routines, runs, or both, and how routines relate to runs, but the output schema mitigates return-format concerns.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already fully documented. The description only restates the millisecond range concept without adding semantics beyond what the schema provides, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('List'), a concrete resource ('routines and runs'), and a scope ('optionally within a millisecond range'). It clearly distinguishes this tool from sibling list tools such as list_bots, list_webhooks, and list_decisions by naming the resource.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The usage is implied: use this tool when you need to list routines or runs. However, the description provides no explicit guidance about when to prefer this over sibling tools or any exclusions, so it stops at implied usage rather than clear routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openmausbot_list_webhooksA
Read-onlyIdempotent

List webhooks and attempts with secret material redacted.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds the behavioral note that secret material is redacted, which is important for an agent to know to avoid expecting secrets. Annotations already declare readOnlyHint and idempotentHint, so the description complements these with a security-relevant detail. It also specifies that it lists both webhooks and attempts, which is additional behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence that states the primary action and a key detail (redaction). It is front-loaded and efficient, with no unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (no parameters, read-only, idempotent), the description is largely sufficient. It clarifies the resource (webhooks and attempts) and the redaction policy. The presence of an output schema means the return structure is documented elsewhere. Minor omissions like pagination or scope are not critical for this simple list operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, and the schema is empty with 100% coverage. There is nothing for the description to add regarding parameter semantics. Baseline of 4 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: listing webhooks and their delivery attempts, with secret material redacted. This is specific and distinguishes it from other list tools like list_bots or list_routines. The mention of redaction adds clarity about the returned data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not explicitly provide usage context or conditions for when to use this tool over alternatives. While the tool name and purpose make it obvious that it is for listing webhooks, there is no explicit guidance about scenarios or prerequisites. Given the many sibling tools, a note about when to use this (e.g., to inspect webhook configurations) would improve clarity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openmausbot_planA
Read-onlyIdempotent

Plan desired-state changes without sending any mutation.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesDesired-state JSON or YAML file path.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds 'Plan' which implies a dry-run preview, but does not elaborate on what the plan output contains or how it behaves beyond stating it does not mutate. The value beyond annotations is modest.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that front-loads the core purpose and the non-mutating nature. Every word earns its place; there is no redundancy or unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with one well-documented parameter and an output schema (covering return format), the description provides sufficient information. It states what it does and the key behavioral constraint (no mutation), and the annotations cover safety and idempotency. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already documents the only parameter (path) with a description ('Desired-state JSON or YAML file path.'), and schema description coverage is 100%. The tool description adds no further parameter-specific meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Plan') and a clear resource ('desired-state changes'), and explicitly states the tool does not send any mutation. This clearly distinguishes it from sibling mutation tools like openmausbot_update_bot or openmausbot_upsert_routine, making its purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly states a key exclusion: it does not send mutations, implying it should be used when the agent wants to preview changes before applying them. However, it does not name specific alternative tools for applying changes, so the guidance is clear on 'when not' but not on explicit alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openmausbot_run_routine_nowC

Plan or start one routine immediately.

ParametersJSON Schema
NameRequiredDescriptionDefault
dry_runNo
routineYesExact routine id or name.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the tool as non-read-only and non-idempotent, so the description's job is to add behavioral context. It discloses that a routine can be started (a side effect) but does not explain that dry_run defaults to true (so a bare call plans rather than starts), nor what happens when a run is started. The description is not misleading but is thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. It is concise to the point of underspecification, but structurally it is exactly what a short description should look like.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool that can either dry-run or execute, alongside siblings like openmausbot_plan, openmausbot_cancel_run, and openmausbot_status, this description leaves crucial selection and invocation details implicit. The existence of an output schema covers return values, but the dual-mode behavior and default dry_run need to be stated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema describes 'routine' as an exact id or name, and the tool description's 'Plan or start' hints that dry_run switches between modes. However, the description doesn't explicitly say which value of dry_run triggers planning versus execution, and dry_run has no schema description. With 50% schema coverage, this is partial compensation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Plan or start one routine immediately' names a resource (one routine) and an action, but offers two verbs, leaving unsaid what distinguishes planning from starting and how that maps to the dry_run parameter. It also doesn't differentiate this tool from the sibling openmausbot_plan, which may already cover the planning behavior.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is present: it doesn't say 'use this to execute immediately vs. plan via openmausbot_plan' or mention alternatives such as openmausbot_cancel_run for cleanup. An agent must infer context from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openmausbot_set_bot_modelC
Idempotent

Plan or apply a validated model selection.

ParametersJSON Schema
NameRequiredDescriptionDefault
botYesExact bot id or unique exact name.
modelYesExact offered model id.
effortNoOptional supported effort level.
dry_runNo
instance_idYesProvider instance id.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide idempotentHint=true and readOnlyHint=false, but the description adds little behavioral context. It does not explain the dry_run flow, whether 'plan' performs a side-effect-free check, or what 'apply' changes beyond the implied model assignment. No contradiction with annotations, but also no meaningful disclosure beyond them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The text is short and easy to parse, but brevity comes at the expense of substance. For a tool with five parameters and a plan/apply mode split, one vague sentence is under-specified rather than appropriately concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity—two modes, a dry_run parameter, required instance and model identifiers—the description is far too minimal to let an agent call it correctly. It does not explain the plan/apply distinction, validation steps, or dry_run semantics, even though an output schema exists to cover return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 80%, so the schema carries most of the parameter meaning. The description adds no parameter-specific information, and notably does not explain the undocumented dry_run parameter, which appears important given the 'plan or apply' distinction. This is a baseline score, not a strength.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description says 'Plan or apply a validated model selection,' which names an action and a resource but is vague about what actually happens: does it change the bot's assigned model, or only decide on one? It is not a tautology and it loosely distinguishes from list_models/plan siblings, but the dual 'plan or apply' phrasing leaves the core purpose unclear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives like openmausbot_plan, openmausbot_list_models, or openmausbot_update_bot. There is no mention of prerequisites, the difference between planning and applying, or when a caller should choose this over a sibling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openmausbot_set_companion_settingA
Idempotent

Atomically update one existing key in OpenMausBot's companion settings.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYesAn existing companion setting key.
valueYesNew JSON-compatible value for the setting.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds the atomicity guarantee and the constraint that only one existing key is updated, which are not captured by the annotations alone. The annotations already provide idempotency, non-destructiveness, and non-read-only status, and the description does not contradict them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, tightly written sentence conveys the core behavior with no filler or redundant phrasing. The key qualifiers ('atomically', 'one existing key') are front-loaded, making the description immediately scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter atomic setter, the description, schema, annotations, and output schema provide enough context for correct invocation. A minor gap is the lack of explicit guidance on when to use this versus related companion-setting tools, but the low complexity and strong schema coverage keep the definition complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already fully documents both parameters, including that 'key' must be an existing companion setting key and 'value' is a new JSON-compatible value. With 100% schema coverage, the description's contribution is minimal but acceptable; it adds the atomic update framing without needing to repeat parameter details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('update') and resource ('one existing key in OpenMausBot's companion settings'), making the tool's function immediately clear. It also distinguishes itself from sibling read/list tools like openmausbot_get_companion_settings by focusing on a targeted mutation of a single existing key.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: this is the tool to use when updating an existing companion setting key, and it excludes creating new keys by requiring an existing key. However, it does not explicitly state when to prefer this over alternatives or mention sibling tools such as openmausbot_get_companion_settings for reading settings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openmausbot_set_routine_enabledB
Idempotent

Plan or apply one routine's enabled state.

ParametersJSON Schema
NameRequiredDescriptionDefault
dry_runNo
enabledYes
routineYesExact routine id or name.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish that this is not read-only, is idempotent, and is non-destructive. The description adds a plan/apply mode that annotations do not convey, but it does not explain what planning entails or what side effects applying has beyond the schema's dry_run flag.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with no filler or repetition. The core subject, the routine's enabled state, is front-loaded, and every word contributes to the meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a state-changing tool with a critical dry_run default, the description omits the most important call behavior: by default the tool plans rather than applies. Given the large set of routine-management siblings, this under-specification creates a real risk of misuse.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33%, with enabled and dry_run having no schema descriptions. The description's 'Plan or apply' loosely maps to dry_run, but it never states that dry_run defaults to true or that setting it to false is what actually applies the change.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action ('Plan or apply') and a specific resource ('one routine's enabled state'), which is enough to distinguish it from listing, running, creating, or deleting routines. It is clear without being as explicit as the annotation title 'Enable or disable an OpenMausBot routine'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool versus alternatives like run_routine_now, upsert_routine, or list_routines. The 'Plan or apply' phrasing hints at two modes, but no explicit conditions or exclusions are given to help an agent choose correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openmausbot_statusA
Read-onlyIdempotent

Check discovery and the live cua-driver daemon without changing state.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds context by specifying what is checked ('discovery' and the daemon) and reinforces the no-state-change behavior, but it does not disclose other behavioral traits such as latency, failure modes, or whether it requires a running daemon.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence conveys both the action and the no-state-change guarantee without filler. Every word contributes meaning, and the phrasing is immediately scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple (no parameters) and has an output schema, so return-value details do not need to be in the description. Given the number of sibling status tools, the description could more explicitly explain what makes this status check distinct, but the core information needed to invoke it is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there is no parameter semantics burden. The baseline of 4 applies because the description does not need to explain any input fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Check') and concrete resources ('discovery' and 'the live cua-driver daemon'), which clearly indicates the tool's scope. It does not explicitly contrast with sibling status tools like openmausbot_cua_skills_status or openmausbot_cua_permissions_status, but the mention of 'discovery' and 'daemon' provides enough distinction to avoid obvious confusion.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to choose this tool over the various sibling status tools. The description implies it is for general status checking, but it does not state exclusions, prerequisites, or scenarios where another tool would be more appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openmausbot_update_botC
Idempotent

Plan or apply a guarded bot update.

ParametersJSON Schema
NameRequiredDescriptionDefault
botYesExact bot id or unique exact name.
patchYesBot fields to update: name, title, description, soul, modelSelection, computer, mcpServers, composio, browser, cwd, approvalMode ('ask' only).
dry_runNo
allow_loosenNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate the operation is a non-read-only, non-destructive, idempotent write. The description adds only 'guarded', which suggests safety behavior but does not disclose what that means—whether changes are staged, reversible, or restricted. No permission requirements or side effects are mentioned, so the description contributes little beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence and is front-loaded, but it is more under-specified than elegantly concise. The vague term 'guarded' and the ambiguous 'Plan or apply' reduce the value of the brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a complex update tool with four parameters, a nested patch object, and two booleans that control behavior, yet the description leaves dry_run, allow_loosen, and the meaning of 'guarded' unexplained. An agent would need to infer critical behavior from defaults and parameter names alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%, and the description does not compensate for the gap. 'Plan or apply' weakly maps to dry_run, but allow_loosen is left completely unexplained, and 'guarded' does not clarify the boolean semantics. The patch field list is already in the schema, so the description adds almost no parameter-level value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb and resource: 'update a bot', and hints at two modes, 'Plan or apply', which distinguishes it from a single-action tool. The word 'guarded' adds context but is undefined, so it does not fully clarify what makes this update different from a plain update or from sibling tools like set_bot_model.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not say when to use this tool versus alternatives such as openmausbot_plan or openmausbot_set_bot_model. It also does not explain when dry_run should be true versus false, or what role allow_loosen plays in choosing between planning and applying.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openmausbot_upsert_routineC
Idempotent

Plan or apply a name-idempotent routine upsert.

ParametersJSON Schema
NameRequiredDescriptionDefault
specYesRoutine spec using bot name or id.
dry_runNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare idempotentHint=true and destructiveHint=false, so the safety profile is covered. The description adds the 'name-idempotent' specificity (idempotency keyed on the routine name) and the 'Plan or apply' distinction, which hints at the dry-run mode not present in annotations. No contradiction, but it does not disclose what 'apply' overwrites or how conflicts are resolved.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single terse sentence with no filler, which is structurally efficient. However, the compact phrasing ('Plan or apply a name-idempotent routine upsert') is jargon-heavy and sacrifices clarity for brevity, so it reads as under-specified rather than elegantly concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating upsert with an open-ended nested spec object and a dry-run mode, one cryptic line is inadequate. The output schema covers return values, but an agent cannot confidently construct a valid spec or understand the plan-vs-apply distinction. Substantial behavioral and structural detail is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50% and the description adds no parameter meaning. The spec parameter is an open-ended object (additionalProperties: true) with the thin schema hint 'Routine spec using bot name or id', yet the description never explains what keys spec should contain. 'name-idempotent' implies a name field must exist, but this is never made explicit. The dry_run parameter is left entirely to its name and default.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific operation (upsert) on a specific resource (routine) and adds the meaningful qualifier 'name-idempotent' (re-running with the same name updates rather than duplicates). The title annotation 'Create or update an OpenMausBot routine' reinforces this. However, 'Plan or apply' is ambiguous without being tied to the dry_run parameter, and the phrasing is jargon-dense.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No usage guidance is provided. The description never states when to use this tool versus the many routine-related siblings (set_routine_enabled, run_routine_now, delete_routine, list_routines), nor any exclusions or prerequisites. The agent must infer its role from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openmausbot_usageA
Read-onlyIdempotent

Read usage for a date range of no more than one year.

ParametersJSON Schema
NameRequiredDescriptionDefault
to_dateYesYYYY-MM-DD
group_byNoOptional API grouping dimension.
from_dateYesYYYY-MM-DD

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds the maximum date range, which is a useful behavioral constraint, but it does not elaborate on other behaviors like error handling, rate limits, or authentication. Since annotations carry most of the burden, a 3 is appropriate – the description adds some context beyond the structured data but not rich detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with zero waste. It communicates the verb, resource, and a key constraint efficiently. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values are documented separately. The description covers the purpose and the critical range constraint, which is sufficient for a simple read-only tool. It doesn't mention edge cases or errors, but given the simplicity and annotation coverage, it is complete enough for an agent to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers 100% of parameters with descriptions (e.g., 'YYYY-MM-DD' for dates and 'Optional API grouping dimension' for group_by). The tool description adds the crucial constraint that the date range must be ≤1 year, which is not present in the schema. This goes beyond the baseline and helps the agent validate inputs correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Read') and resource ('usage') with a clear constraint ('date range of no more than one year'). It is unambiguous and distinct from the sibling tools, which cover other operations like health or bot management. It doesn't explicitly name a sibling alternative, but the uniqueness of 'usage' makes differentiation implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool (when needing usage data) but does not explicitly state context or alternatives. The 'no more than one year' constraint is a usage guideline for parameter selection, but there is no explicit 'use this instead of X' guidance. This is adequate but leans on inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 16 tool updatesv0.2.2
    • Addedopenmausbot_api_health
    • Addedopenmausbot_cancel_run
    • Addedopenmausbot_delete_routine
    • Addedopenmausbot_export_team
    • Addedopenmausbot_list_bots
    • Addedopenmausbot_list_decisions
    • Addedopenmausbot_list_models
    • Addedopenmausbot_list_routines
    • Addedopenmausbot_list_webhooks
    • Addedopenmausbot_plan
    • Addedopenmausbot_run_routine_now
    • Addedopenmausbot_set_bot_model
    • Addedopenmausbot_set_routine_enabled
    • Addedopenmausbot_update_bot
    • Addedopenmausbot_upsert_routine
    • Addedopenmausbot_usage
  2. 9 tool updatesv0.1.0
    • First observedopenmausbot_call_cua_tool
    • First observedopenmausbot_cua_check_update
    • First observedopenmausbot_cua_permissions_status
    • First observedopenmausbot_cua_skills_status
    • First observedopenmausbot_describe_cua_tool
    • First observedopenmausbot_get_companion_settings
    • First observedopenmausbot_list_cua_tools
    • First observedopenmausbot_set_companion_setting
    • First observedopenmausbot_status

TDQS

B3/5.0

Scored across 25 tools

Disambiguation4/5

Most tools target distinct resources and actions, with clear separation between listing, mutating, and CUA-driver operations. Minor confusion is possible between openmausbot_api_health and openmausbot_status, and between the generic openmausbot_plan and the per-resource 'plan or apply' mutation tools, but descriptions mostly mitigate this.

Naming Consistency3/5

All tools share the openmausbot_ prefix, which helps, but verb patterns are mixed: list_bots vs get_companion_settings vs usage vs status, and cua_skills_status/cua_permissions_status/cua_check_update deviate from the list_/set_ convention. The naming is readable but not uniformly predictable.

Tool Count3/5

At 25 tools, this sits at the heavy end of the acceptable range. The count is somewhat justified by covering bot management, routines, webhooks, usage, companion settings, and the CUA-driver surface, but it feels like several smaller logical servers have been merged into one.

Completeness3/5

Routine lifecycle is well covered with upsert, enable/disable, run, and delete, and companion settings have get/set. However, webhooks are list-only with no create/update/delete, bots lack create/delete, and the CUA-driver tools only inspect and invoke rather than manage the driver itself, leaving notable gaps.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    A
    maintenance
    Enables MCP-compatible AI clients to invoke CLI-driven agent tools over Streamable HTTP, including shell execution, file operations, patching, image viewing, web search, and nested agent tasks, with permission modes and real-time progress streaming.
    -
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables AI clients to securely control and interact with a local Windows machine through 218 configurable tools for files, Git, processes, Windows UI, browser automation, WSL, Office, recovery, skills, and child MCP servers.
    11 npm
    MIT