Skip to main content
Glama

Execute code

execute_code

Run Python, SQL, Scala, or R code on Databricks clusters or serverless compute, then get status or cancel pending executions and return results.

Instructions

Execute Python, SQL, Scala or R code on Databricks compute and return its output.

  • run (code, language[, compute, cluster_id, timeout_seconds]): on a RUNNING classic cluster via the Command Execution API (a fresh execution context per call; no state is kept between calls), or, for Python, on serverless jobs compute (temporary notebook in ~/.dbx_mcp/tmp, one-time run; stdout/stderr captured). Returns status success/failed with output (text or table rows/columns) and error summary/stack trace, and which compute was used. If not finished within timeout_seconds it returns status 'pending' with ids to poll.

  • get_status (cluster_id+context_id+command_id, or run_id): poll a pending execution.

  • cancel (same ids): stop a pending execution. For SQL on a SQL warehouse prefer execute_sql. Classified EXECUTION: code can change data and costs money.

Safety classification: depends on input (EXECUTION, READ_ONLY, WRITE).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
codeNorun: the code to execute.
actionNorun: execute code; get_status: poll a pending execution; cancel: stop it.run
run_idNoget_status/cancel (serverless): run id.
computeNorun: 'cluster' (classic all-purpose cluster; any language) or 'serverless' (Python only, one-time serverless job run; slower to start).cluster
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
languageNorun: code language.python
cluster_idNoCluster id (default DBX_MCP_DEFAULT_CLUSTER_ID); also for get_status/cancel.
command_idNoget_status/cancel (cluster): command id.
context_idNoget_status/cancel (cluster): execution context id.
timeout_secondsNorun: max seconds to wait before returning 'pending' (capped by server).

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNosuccess
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: fresh execution context per call with no state kept between calls, temp notebook path, stdout/stderr capture, success/failed/pending return semantics with ids to poll, and explicit cost/security note ('code can change data and costs money'). The safety classification (EXECUTION/READ_ONLY/WRITE depending on input) adds nuance the static destructiveHint=false does not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the one-line purpose, then uses tight action-scoped bullets. It is dense and slightly long, but nearly every clause earns its place (compute semantics, state behavior, timeout/polling).

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers compute selection, state model, return shapes, timeout/polling lifecycle, and the SQL alternative. With an output schema already present, the return-format detail is bonus rather than a requirement, and nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: it groups parameters by action (which ids apply to get_status/cancel vs run), explains the compute tradeoff, and ties timeout_seconds to the 'pending' return. The only gap is that it does not elaborate on confirm/dry_run beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Execute) and resource (Python/SQL/Scala/R code on Databricks compute) and returns output. It enumerates the three sub-actions (run, get_status, cancel) and explicitly names the sibling execute_sql for the SQL-on-warehouse case, so an agent can distinguish it from siblings without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit routing: 'For SQL on a SQL warehouse prefer execute_sql.' It also tells the agent when to use cluster vs serverless (serverless is Python-only, slower to start) and when to call get_status/cancel (pending executions). Alternatives and conditions are stated, not inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.