Skip to main content
Glama

Update Runtimes

update_runtimes
Destructive

Keep language runtimes current by checking for updates; dry-run by default and apply changes only after confirmation.

Instructions

Update language runtimes. SAFE BY DEFAULT: with apply=False this is a dry run — it returns the update commands that WOULD run without changing anything. Pass apply=True to actually execute them (mise up, rustup update, swiftly update, apt-get upgrade of language packages, npm -g update, uv tool upgrade).

apply=True asks the caller to confirm first (a protocol-level gate, not just the anthropic/requiresUserInteraction _meta hint — see codecalc/confirmation.py); apply=False is never gated, since nothing runs.

PRIVILEGE: the apt manager updates system packages and its command begins with sudo. Those commands do NOT run unless the HOST has set CODECALC_ALLOW_RUNTIME_APPLY=1; without it they are reported as skipped with ok: false and the variable named, and the rest still run. Every entry carries an elevated flag either way. mise/rustup/swiftly/npm/uv touch user-owned toolchains and are never gated.

NETWORK: yes, on both paths. apply=False still asks each manager what the latest version is, which is a remote lookup; apply=True additionally downloads and installs. "Dry run" bounds what changes on disk, not what is sent.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
applyNoFalse (default) is a dry run reporting commands only; True actually runs them and asks for confirmation first
timeoutNoWall-clock seconds allowed for the update commands to complete
languagesNoComma-separated languages to update, e.g. 'python3,node,rust'; empty updates all

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed3 schema fields changedv0.12.0
    • addedInput schema / properties / apply / description
      Added value: +"False (default) is a dry run reporting commands only; True actually runs them and asks for confirmation first"
    • addedInput schema / properties / languages / description
      Added value: +"Comma-separated languages to update, e.g. 'python3,node,rust'; empty updates all"
    • addedInput schema / properties / timeout / description
      Added value: +"Wall-clock seconds allowed for the update commands to complete"
  2. Changed1 schema field changedv0.11.0
    • changedOutput schema / (root)
      Previous value: -nullNew value: +{
      +  "additionalProperties": true,
      +  "title": "update_runtimesDictOutput",
      +  "type": "object"
      +}
  3. Changed6 schema fields changedv0.2.0
    • removedInput schema / additionalProperties
      Removed value: -false
    • addedInput schema / properties / apply / title
      Added value: +"Apply"
    • addedInput schema / properties / languages / title
      Added value: +"Languages"
    • addedInput schema / properties / timeout / title
      Added value: +"Timeout"
    • addedInput schema / title
      Added value: +"update_runtimesArguments"
    • changedOutput schema / (root)
      Previous value: -{
      -  "additionalProperties": true,
      -  "type": "object"
      -}New value: +null
  4. First observedv0.1.0

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (destructiveHint=true, readOnlyHint=false), the description discloses materially more: the CODECALC_ALLOW_RUNTIME_APPLY=1 gate for sudo commands, the protocol-level confirmation on apply=True, and the critical network caveat that dry-run still performs remote lookups. It flags that failed privileged commands return 'ok: false' and an 'elevated' flag. This context substantially exceeds what annotations alone convey, and it is consistent with them — no contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but the length is earned: it is a privileged, mutating, network-active tool whose complexity demands the SAFE BY DEFAULT, PRIVILEGE, and NETWORK sections. It is well structured with clear section headers and front-loads the safety default. It borders on verbose but no sentence is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool of this complexity — mutation, elevation, gating, network side effects, and dual modes — the description covers every operational concern an agent needs: what commands run, when confirmation is required, when commands are skipped (including the variable name), and what 'dry run' does and does not bound. The presence of an output schema relieves it of describing return values, and nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all three parameters (apply, timeout, languages) with defaults and types. The description reinforces apply's dry-run semantics and adds color about which managers map to which languages, but it does not meaningfully extend what the schema provides. Baseline 3 is appropriate since the schema carries the load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb+resource statement, 'Update language runtimes,' then enumerates the exact underlying managers (mise, rustup, swiftly, apt, npm, uv). It distinguishes the two execution modes (dry-run vs apply) and is clearly separable from the sibling runtimes_status, which reports status rather than mutating. The purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives strong internal usage guidance: when to use apply=False (safe dry run), when to use apply=True (actual execution with a confirmation gate), and the privilege conditions under which sudo commands do or do not run. It does not explicitly name sibling tools to avoid (e.g., install_package for single packages or runtimes_status for read-only checks), leaving those exclusions implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.