Skip to main content
Glama

dbx-mcp: a safety-first Databricks MCP server

dbx-mcp is an open-source Model Context Protocol server. It lets AI agents and assistants work with a Databricks workspace: SQL, clusters and warehouses, notebooks, Jobs, Lakeflow pipelines, Unity Catalog, Volumes, AI/BI dashboards, Genie, model serving, Vector Search, Lakebase and Apps.

It is built on the official Databricks SDK for Python and the official MCP Python SDK. It does not use or copy any other Databricks MCP implementation.

Design goals

  • Safe by default. Every action has a safety class: read, write, destructive, execution or security-sensitive.

    • Destructive and security-sensitive changes use two steps. The first call returns a plan and changes nothing; the change runs only when the call is repeated with confirm=true.

    • Resources whose name or tags mark them as production are protected.

    • The server can run fully read-only.

  • Honest. Every SDK call is checked against the real SDK. Features with no official API raise UNSUPPORTED_OPERATION with an explanation; nothing is faked.

  • Machine-readable. Every tool has a typed input schema and a typed output envelope, plus a human-readable summary.

  • No secret leakage. Responses and logs pass through secret redaction. Credentials are never returned unless a tool exists for that purpose and the user explicitly asks.


Contents

  1. Quick start

  2. Installation

  3. Authentication

  4. Configuration (environment variables)

  5. MCP client setup

  6. Tools

  7. Security model

  8. Example tool calls

  9. Architecture

  10. Development

  11. Testing

  12. Troubleshooting

  13. Known limitations


Related MCP server: Databricks MCP Server

Quick start

git clone <this repo> dbx-mcp && cd dbx-mcp
uv venv && uv pip install -e .            # or: python -m venv .venv && pip install -e .
cp .env.example .env                       # set DATABRICKS_HOST + credentials
dbx-mcp --env-file .env --list-tools       # verify configuration
dbx-mcp --env-file .env --read-only        # start (stdio) in read-only mode

Installation

You need Python 3.10 or newer.

pip install -e .            # core
pip install -e ".[pdf]"     # adds HTML->PDF conversion (xhtml2pdf) for generate_and_upload_pdf
pip install -e ".[dev]"     # tests + linters

For reproducible installs with the exact tested versions, use the pinned files (generated from pyproject.toml):

pip install -r requirements.txt && pip install --no-deps -e .        # runtime (incl. PDF support)
pip install -r requirements-dev.txt && pip install --no-deps -e .    # + tests and linters

To regenerate the pinned files after changing dependencies: uv pip compile pyproject.toml --extra pdf --python-version 3.10 -o requirements.txt (add --extra dev and -o requirements-dev.txt for the dev file).

The installed entry points are dbx-mcp and python -m dbx_mcp.

dbx-mcp [--transport stdio|streamable-http|sse] [--host 127.0.0.1] [--port 8765]
        [--env-file PATH] [--read-only] [--list-tools] [--version]

stdio is the default and the recommended transport. The HTTP transports bind to 127.0.0.1 by default. Exposing them on a network gives anyone who can reach the port your Databricks permissions, so put an authenticating proxy in front.

Authentication

Authentication is handled entirely by the Databricks SDK's unified authentication. The server never reads, stores or returns credentials itself. Supported methods:

Method

Environment

Personal access token

DATABRICKS_HOST, DATABRICKS_TOKEN

OAuth M2M (service principal)

DATABRICKS_HOST, DATABRICKS_CLIENT_ID, DATABRICKS_CLIENT_SECRET

OAuth U2M (browser login)

run databricks auth login --host ... once, then DATABRICKS_CONFIG_PROFILE

Config profile

DATABRICKS_CONFIG_PROFILE (from ~/.databrickscfg)

Azure (CLI, MSI, service principal)

DATABRICKS_HOST plus ARM_* / Azure CLI login

Google Cloud

DATABRICKS_HOST plus GOOGLE_CREDENTIALS / DATABRICKS_GOOGLE_SERVICE_ACCOUNT

Least privilege. The server can do anything the authenticated principal can do. For agents, prefer a dedicated service principal granted only the Unity Catalog privileges and workspace entitlements it needs. Add DBX_MCP_READ_ONLY=true for exploration-only use.

get_current_user and manage_workspace action=info show which identity and workspace are active (never tokens). manage_workspace action=switch_profile reconnects using another profile.

Configuration

Server behaviour is configured with DBX_MCP_* environment variables. All of them are optional.

Variable

Default

Purpose

DBX_MCP_AUTH_MODE

env

env: the server's own credentials (one workspace). request: each HTTP request sends its workspace URL + PAT in headers (multi-workspace; see below). CLI: --auth-mode.

DBX_MCP_ALLOWED_WORKSPACE_HOSTS

Databricks domains

Request mode: allowed host suffixes (e.g. .azuredatabricks.net), or * for any host.

DBX_MCP_REQUEST_CLIENT_CACHE_SIZE

64

Request mode: number of per-credential SDK clients kept in memory.

DBX_MCP_TOOLSETS

all

Comma list of toolsets to enable (see docs/TOOLS.md).

DBX_MCP_DISABLED_TOOLS

Comma list of individual tools to hide.

DBX_MCP_READ_ONLY

false

Allow only read actions (SELECTs are allowed; writes, DDL and code execution are not).

DBX_MCP_BLOCKED_SAFETY_LEVELS

Block classes entirely, e.g. DESTRUCTIVE,SECURITY_SENSITIVE,EXECUTION.

DBX_MCP_REQUIRE_CONFIRMATION

true

Two-step confirm=true for destructive or security-sensitive changes.

DBX_MCP_CONFIRM_EXECUTION

false

Also require confirmation for code/job execution.

DBX_MCP_PROTECTED_NAME_PATTERNS

(?i)(^|[-_ .])prod(uction)?($|[-_ .])

Regexes for resource names and tags that must not be deleted, terminated or changed. none disables.

DBX_MCP_ALLOW_PROTECTED_CHANGES

false

Allow changes to protected resources (still requires confirmation).

DBX_MCP_ALLOWED_VOLUME_PREFIXES

Restrict volume file tools to these /Volumes/... prefixes.

DBX_MCP_ALLOWED_WORKSPACE_PREFIXES

Restrict workspace file tools to these paths.

DBX_MCP_LOCAL_FILE_ROOT

(disabled)

Directory the server may read from or write to for local uploads and downloads.

DBX_MCP_DEFAULT_WAREHOUSE_ID

DATABRICKS_WAREHOUSE_ID

Warehouse for SQL tools.

DBX_MCP_WAREHOUSE_SELECTION

prefer_running

prefer_running (automatic, explained in every response) or configured_only.

DBX_MCP_DEFAULT_CLUSTER_ID

DATABRICKS_CLUSTER_ID

Cluster for execute_code.

DBX_MCP_SQL_MAX_ROWS

1000

Hard cap on rows returned by SQL tools.

DBX_MCP_SQL_WAIT_TIMEOUT_SECONDS

30

How long SQL waits (5-50) before returning a pending statement id.

DBX_MCP_DEFAULT_PAGE_SIZE / DBX_MCP_MAX_PAGE_SIZE

50 / 100

Pagination.

DBX_MCP_MAX_INLINE_DOWNLOAD_BYTES

10485760

Max file bytes returned inline.

DBX_MCP_TOOL_TIMEOUT_SECONDS

300

Per-call timeout.

DBX_MCP_MAX_WAIT_SECONDS

240

Cap for wait=true on long-running operations (must be below the tool timeout).

DBX_MCP_HTTP_TIMEOUT_SECONDS

60

Per HTTP request to Databricks.

DBX_MCP_RETRY_TIMEOUT_SECONDS

300

SDK retry budget for 429/503/transient errors.

DBX_MCP_RATE_LIMIT_PER_SECOND

Client-side request rate limit.

DBX_MCP_MANIFEST_PATH

.databricks_mcp/manifest.json

Project manifest file.

DBX_MCP_LOG_LEVEL

INFO

JSON logs to stderr.

DBX_MCP_DEBUG

false

Include stack traces in errors (development only).

MCP client setup

Claude Code

claude mcp add databricks -- dbx-mcp --env-file /absolute/path/to/.env

Claude Desktop / Cursor / any client using mcpServers JSON

{
  "mcpServers": {
    "databricks": {
      "command": "dbx-mcp",
      "args": ["--env-file", "/absolute/path/to/.env"],
      "env": {
        "DBX_MCP_TOOLSETS": "identity,sql,compute,unity_catalog,volumes",
        "DBX_MCP_READ_ONLY": "true"
      }
    }
  }
}

If dbx-mcp is not on the client's PATH, use the absolute path to the virtualenv's executable, e.g. /path/to/repo/.venv/bin/dbx-mcp (Windows: .venv\\Scripts\\dbx-mcp.exe).

VS Code (.vscode/mcp.json)

{
  "servers": {
    "databricks": { "type": "stdio", "command": "dbx-mcp", "args": ["--env-file", "${workspaceFolder}/.env"] }
  }
}

LangGraph / LangChain agents: see examples/langgraph_agent. It is a working agent that connects through langchain-mcp-adapters and adds a human-approval gate for confirm=true calls.

HTTP transport (for clients that connect by URL): dbx-mcp --transport streamable-http --port 8765, then connect to http://127.0.0.1:8765/mcp.

One server, many workspaces (request-auth mode)

In request-auth mode the server stores no Databricks credentials. Each client sends its own workspace URL and PAT as HTTP headers, so one running server can serve any number of workspaces and users at once:

dbx-mcp --auth-mode request --transport streamable-http --host 0.0.0.0 --port 8765

Client config. Add one entry per workspace; all entries point at the same server:

{
  "mcpServers": {
    "databricks-prod-eu": {
      "type": "http",
      "url": "https://mcp.example.com/mcp",
      "headers": {
        "X-Databricks-Host": "https://adb-1111111111111111.1.azuredatabricks.net",
        "Authorization": "Bearer dapi...",
        "X-Databricks-Warehouse-Id": "optional-default-warehouse"
      }
    },
    "databricks-dev-us": {
      "type": "http",
      "url": "https://mcp.example.com/mcp",
      "headers": {
        "X-Databricks-Host": "https://dbc-2222.cloud.databricks.com",
        "Authorization": "Bearer dapi..."
      }
    }
  }
}

Header

Required

Meaning

X-Databricks-Host

yes

Workspace URL (https://..., root only)

Authorization: Bearer <PAT> (or X-Databricks-Token)

yes

That workspace's personal access token

X-Databricks-Warehouse-Id

no

Default SQL warehouse for this connection

X-Databricks-Cluster-Id

no

Default cluster for execute_code on this connection

How it behaves:

  • Isolation: each request's credentials get their own SDK client, built only from the headers. The server machine's env vars and ~/.databrickscfg are never mixed in.

  • Secrets: clients are cached by a SHA-256 hash of host + token; tokens are never logged or returned.

  • Allowed hosts: only Databricks domains are accepted (*.azuredatabricks.net, *.cloud.databricks.com, *.gcp.databricks.com, ...), which stops the server being used to reach arbitrary or internal URLs. Adjust with DBX_MCP_ALLOWED_WORKSPACE_HOSTS.

  • Not available in this mode: server-side profiles (list_profiles, switch_profile) and server-wide default warehouse/cluster ids. Use the per-connection headers instead.

  • Manifest: entries are scoped per workspace, so tenants only see their own.

  • Transport: request mode needs an HTTP transport; stdio is refused.

  • Use HTTPS beyond localhost. PATs travel in headers, so put the server behind a TLS-terminating reverse proxy (nginx, Caddy, a cloud load balancer). A request without valid headers simply fails; there is no shared server identity to fall back on.

Tools

Tools are grouped into toolsets that can be enabled independently. See docs/TOOLS.md for the complete reference: every parameter, the safety class of every action, and the capability matrix.

Toolset

Tools

identity

get_current_user, manage_workspace

sql

execute_sql, execute_sql_multi, manage_sql_statement, get_table_stats_and_schema

compute

manage_cluster, manage_sql_warehouse, manage_warehouse, list_compute

workspace

execute_code, manage_workspace_files

jobs

manage_jobs, manage_job_runs

pipelines

manage_pipeline, manage_pipeline_run

unity_catalog

manage_uc_objects, manage_uc_grants, manage_uc_storage, manage_uc_connections, manage_uc_tags, manage_uc_security_policies, manage_uc_monitors, manage_uc_sharing, manage_metric_views

volumes

get_volume_folder_details, manage_volume_files

dashboards

manage_dashboard

ai

manage_serving_endpoint, manage_ka, manage_mas, manage_genie, ask_genie

vector_search

manage_vs_endpoint, manage_vs_index, query_vs_index, manage_vs_data

lakebase

manage_lakebase_database, manage_lakebase_branch, manage_lakebase_sync, generate_lakebase_credential

apps

manage_app

manifest

list_tracked_resources, delete_tracked_resource

pdf

generate_and_upload_pdf

Response envelope

Every tool returns:

{
  "status": "success | pending | dry_run | confirmation_required | partial_failure | failed",
  "tool": "manage_cluster", "action": "terminate",
  "summary": "Human-readable one-liner.",
  "safety": ["DESTRUCTIVE"],
  "data": { /* machine-readable result */ },
  "page": { "page_size": 50, "returned": 50, "has_more": true, "next_page_token": "..." },
  "plan": { /* what will change: present for dry_run / confirmation_required */ },
  "warnings": [], "next_steps": [], "request_id": "4f1c..."
}

Errors are MCP tool errors with a category prefix, for example [NOT_FOUND], [PERMISSION_DENIED], [AUTHENTICATION_FAILED], [INVALID_PARAMETER], [CONFLICT], [RATE_LIMITED], [TIMEOUT], [DATABRICKS_SERVICE_ERROR], [UNSUPPORTED_OPERATION] or [BLOCKED_BY_SAFETY_POLICY]. Each category comes with a hint and, when Databricks provides one, a request id.

Security model

Control

Behaviour

Safety classification

Every action is READ_ONLY, WRITE, DESTRUCTIVE, EXECUTION and/or SECURITY_SENSITIVE. The class is shown in the tool description and in MCP tool annotations (readOnlyHint, destructiveHint).

Central enforcement

A single wrapper applies policy, dry-run, confirmation, timeout, redaction and error handling to every tool, so an individual tool cannot skip them. Registration fails if a mutating tool lacks dry_run or confirm.

Two-step confirmation

DESTRUCTIVE and SECURITY_SENSITIVE changes first return confirmation_required with a plan (target, current state, diff, warnings, reversibility). Nothing runs until the call is repeated with confirm=true. Agents are instructed to get user approval first.

Dry run

dry_run=true previews any change.

Read-only mode and blocked classes

DBX_MCP_READ_ONLY, DBX_MCP_BLOCKED_SAFETY_LEVELS.

SQL classification

Statements are classified lexically (comments and literals stripped), so DROP, DELETE, TRUNCATE, UPDATE, MERGE, INSERT OVERWRITE, CREATE OR REPLACE, GRANT/REVOKE, ownership, row-filter and mask changes are detected. Unknown statements count as destructive. This is a guardrail, not a security boundary; real enforcement is Databricks permissions on the principal.

Production protection

Deleting, terminating, stopping or changing resources whose name or tags match DBX_MCP_PROTECTED_NAME_PATTERNS is refused unless explicitly allowed.

Grants

Grant and revoke plans show a before/after diff. ALL_PRIVILEGES needs an explicit flag, and broad principals produce a warning.

Paths

Volume and workspace paths are validated (no .., no relative paths, no control characters, no backslashes), with optional prefix allowlists. Local filesystem access is off unless DBX_MCP_LOCAL_FILE_ROOT is set, and is confined to it.

SQL injection

Values go through bound statement parameters. Identifiers and literals that tools build into DDL are quoted and escaped.

Secrets

Responses (data, summary, warnings) and logs are redacted by key name and by pattern (PATs, JWTs, bearer tokens, password=). Connection options and recipient activation links are stripped. The Lakebase credential token is returned only with reveal_token=true plus confirmation.

Logging

Structured JSON on stderr (stdout is the MCP channel): tool, action, request id, user, duration, outcome and error category. No arguments, tokens or query results are logged.

Errors

Normalized and categorized. Stack traces are shown only with DBX_MCP_DEBUG=true.

Example tool calls

// Who am I?
{"name": "get_current_user", "arguments": {}}

// Explore
{"name": "manage_uc_objects", "arguments": {"action": "list", "object_type": "schema", "catalog_name": "main"}}
{"name": "get_table_stats_and_schema", "arguments": {"name": "main.sales.orders", "stats": "metadata"}}

// Query (parameters are bound server-side)
{"name": "execute_sql", "arguments": {
  "statement": "SELECT region, SUM(amount) AS total FROM main.sales.orders WHERE day >= :d GROUP BY region",
  "parameters": {"d": "2026-01-01"}, "max_rows": 100, "row_format": "objects"}}

// Destructive: the first call returns a plan; repeat with confirm=true after user approval
{"name": "execute_sql", "arguments": {"statement": "DROP TABLE main.sandbox.tmp_orders"}}
// -> {"status": "confirmation_required", "plan": {...}, "next_steps": ["... confirm=true"]}
{"name": "execute_sql", "arguments": {"statement": "DROP TABLE main.sandbox.tmp_orders", "confirm": true}}

// Long-running: returns quickly; poll for status
{"name": "manage_job_runs", "arguments": {"action": "get", "run_id": 123}}

// Pagination
{"name": "manage_cluster", "arguments": {"action": "list", "page_size": 20, "page_token": "eyJvIjogMjB9"}}

// Genie (the answer is model-generated; check the SQL)
{"name": "ask_genie", "arguments": {"space_id": "01ef...", "question": "Top 5 products by revenue last month?"}}

Architecture

src/dbx_mcp/
  __main__.py            CLI entry point (transport, --env-file, --read-only, --list-tools)
  server/
    app.py               builds the MCPServer and registers enabled toolsets
    config.py            Settings from DBX_MCP_* env vars
    context.py           process-wide context (settings, client, safety policy, manifest)
    manifest.py          local JSON project manifest
  databricks/
    client.py            the only place a WorkspaceClient is created (auth, retries, timeouts)
    sql_runner.py        Statement Execution API, warehouse selection, result chunking
  safety/
    levels.py            SafetyLevel enum and semantics
    guard.py             policy, confirmation, production protection, path allowlists
    validation.py        path/identifier validation, quoting, SQL statement classifier
  tools/
    registry.py          @tool decorator + the wrapper applying all cross-cutting behaviour
    common.py            shared parameter types and response helpers
    identity.py sql.py compute.py workspace.py jobs.py pipelines.py volumes.py
    dashboards.py ai.py vector_search.py lakebase.py apps.py manifest.py pdf.py
    unity_catalog/       objects, grants, storage, connections, tags, security_policies,
                         monitors, sharing, metric_views
  models/                typed pydantic response models
  utils/                 errors, logging, redaction, pagination, serialization

Key mechanisms:

  • call_with_spec. Create and update tools accept a spec using Databricks REST field names. It is validated against the real SDK method signature and converted to SDK dataclasses. Unknown fields and invalid enum values are rejected; the SDK alone would silently drop them.

  • Pagination. Every list action returns an opaque next_page_token.

  • Long-running operations return immediately with an id and state (status: pending). An optional wait=true is bounded by DBX_MCP_MAX_WAIT_SECONDS.

  • Retries and rate limits use the SDK's built-in retry/backoff (DBX_MCP_RETRY_TIMEOUT_SECONDS) plus an optional client-side rate limit.

Development

uv venv && uv pip install -e ".[dev,pdf]"
ruff check src tests
python -m pytest                        # unit tests (no Databricks needed)
python scripts/gen_tool_docs.py         # regenerate docs/TOOLS.md after changing tools

Adding a tool:

  1. Write a synchronous function in a module under tools/.

  2. Decorate it with @tool(toolset=..., title=..., safety={action: levels}).

  3. Give it typed Annotated[..., Field(description=...)] parameters, and add dry_run / confirm if it mutates (registration enforces this).

  4. Return ok(...) or paged_response(...).

  5. For destructive or security-sensitive actions, add a preview function that describes exactly what will change, and call ctx().safety.check_protected(...).

  6. Verify every SDK method, field and enum value by introspection before using it.

Testing

python -m pytest tests/unit                    # fast, fully mocked
DBX_MCP_RUN_INTEGRATION=1 INTEGRATION_ENV_FILE=.env python -m pytest tests/integration
  • Unit tests mock the WorkspaceClient with services autospecced from the real SDK classes. A tool that calls a non-existent SDK method or passes a wrong keyword argument fails its tests.

  • Integration tests never run unless DBX_MCP_RUN_INTEGRATION=1 is set. The default suite is read-only: the server is forced into read-only mode, and SQL runs only on an already-running warehouse.

Troubleshooting

Symptom

Fix

[AUTHENTICATION_FAILED] ... cannot configure default credentials

Set DATABRICKS_HOST plus credentials, or pass --env-file. Check with databricks auth describe.

[PERMISSION_DENIED]

The principal lacks a privilege (UC grant, warehouse CAN_USE, cluster CAN_ATTACH_TO, ...). Run get_current_user to see the identity.

[BLOCKED_BY_SAFETY_POLICY] ... read-only mode

Unset DBX_MCP_READ_ONLY, or adjust DBX_MCP_BLOCKED_SAFETY_LEVELS.

... protected/production resource

Intended. Set DBX_MCP_ALLOW_PROTECTED_CHANGES=true or adjust DBX_MCP_PROTECTED_NAME_PATTERNS if appropriate.

SQL says No SQL warehouses are visible

Create or grant a warehouse, or set DBX_MCP_DEFAULT_WAREHOUSE_ID.

SQL returns status: pending

The query is still running. Poll manage_sql_statement action=get.

Responses are truncated

Raise max_rows (up to DBX_MCP_SQL_MAX_ROWS) or add LIMIT/filters.

Client shows no tools or garbled output

Something printed to stdout. The server logs only to stderr; check wrappers or shell profile output.

generate_and_upload_pdf reports missing dependency

pip install "dbx-mcp[pdf]".

Need details of an unexpected error

Set DBX_MCP_DEBUG=true and DBX_MCP_LOG_LEVEL=DEBUG (development only).

Known limitations

See the "Limitations" notes in docs/TOOLS.md and docs/LIMITATIONS.md. Capabilities without a stable official API are reported as UNSUPPORTED_OPERATION instead of being emulated.

License

Apache-2.0. See LICENSE.

Available Tools

45 tools
ask_genieAsk GenieA
Read-only

Ask a natural-language question in a Genie space and return Genie's answer.

Starts a new conversation (or a follow-up when conversation_id is given), waits up to wait_seconds, and returns: the model-generated text answer, the generated SQL with its description, the rows produced by running that SQL (capped), status and ids. If Genie is still working, returns status 'pending' with conversation_id/message_id - call again with those ids (and no question) to poll. The answer and SQL are MODEL-GENERATED, not authoritative data.

Safety classification: EXECUTION+READ_ONLY.

ParametersJSON Schema
NameRequiredDescriptionDefault
max_rowsNoMax result rows to return per query (capped by the server SQL row limit).
questionNoNatural-language question. Omit when polling an existing message_id.
space_idYesGenie space id.
message_idNoPoll mode: with conversation_id, fetch status/result of a previous question.
wait_secondsNoHow long to wait for Genie to finish (default 60s, capped by server max wait).
conversation_idNoContinue this conversation (follow-up question), or poll a message in it.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint, destructiveHint=false, openWorldHint), and the description adds genuinely new behavior: the wait_seconds bound, the 'pending' status and polling loop, and the caveat that answer/SQL are MODEL-GENERATED and not authoritative. The 'EXECUTION+READ_ONLY' label is consistent with the read-only annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and return, then workflow, then a caveat. Multi-line but every sentence carries information (return shape, poll mechanism, model-generated warning); nothing is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present the description needn't restate return values, yet it still summarizes them and explains the pending/poll lifecycle, which is the main risk of misuse. Nothing needed to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents space_id, question, message_id, conversation_id, wait_seconds, and max_rows. The description reinforces the polling semantics (omit question, reuse ids) but adds little parameter-level syntax beyond the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence names a specific verb and resource ('Ask a natural-language question in a Genie space') and states the return. It is distinguishable from execute_sql and the manage_* siblings by scoping to Genie spaces and NL questions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear conditional usage: start vs follow-up with conversation_id, and explicitly says to poll by calling again with the returned ids and no question. It does not name sibling alternatives (e.g. execute_sql) for when a raw SQL path is preferable, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_tracked_resourceStop tracking a resourceA

Remove an entry from the local project manifest (stop tracking it). This does NOT delete the Databricks resource itself - use the matching manage_* tool for that.

Safety classification: WRITE.

ParametersJSON Schema
NameRequiredDescriptionDefault
dry_runNoIf true, validate and return the planned change without executing it.
resource_idYesId of the tracked entry (as shown by list_tracked_resources).
resource_typeYesType of the tracked entry, e.g. job, dashboard, app.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations flag readOnlyHint=false/destructiveHint=false, and the description adds real value by explaining what is and isn't destroyed (manifest entry only, not the resource), which directly explains the non-destructive classification. It omits auth requirements or reversibility of the manifest edit, but adds more than the annotations alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and the important non-deletion caveat in the first sentence. The trailing 'Safety classification: WRITE.' is largely redundant with the annotations, costing a little efficiency but nothing is buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter, output-schema-backed tool, the description covers the action, the scope boundary and the alternative tool. The dry_run parameter is never mentioned in prose, but the schema handles it, leaving the definition essentially complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so resource_id, resource_type and dry_run are already documented with examples and defaults. The description adds no parameter-level detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Remove an entry from the local project manifest') and disambiguates scope with '(stop tracking it)'. It clearly separates itself from the manage_* siblings that mutate the actual Databricks resource.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells the agent when NOT to use it ('This does NOT delete the Databricks resource itself - use the matching manage_* tool for that'), which is the key routing decision. It doesn't mention when to prefer this over list_tracked_resources or any prerequisite sequencing, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

execute_codeExecute codeA

Execute Python, SQL, Scala or R code on Databricks compute and return its output.

  • run (code, language[, compute, cluster_id, timeout_seconds]): on a RUNNING classic cluster via the Command Execution API (a fresh execution context per call; no state is kept between calls), or, for Python, on serverless jobs compute (temporary notebook in ~/.dbx_mcp/tmp, one-time run; stdout/stderr captured). Returns status success/failed with output (text or table rows/columns) and error summary/stack trace, and which compute was used. If not finished within timeout_seconds it returns status 'pending' with ids to poll.

  • get_status (cluster_id+context_id+command_id, or run_id): poll a pending execution.

  • cancel (same ids): stop a pending execution. For SQL on a SQL warehouse prefer execute_sql. Classified EXECUTION: code can change data and costs money.

Safety classification: depends on input (EXECUTION, READ_ONLY, WRITE).

ParametersJSON Schema
NameRequiredDescriptionDefault
codeNorun: the code to execute.
actionNorun: execute code; get_status: poll a pending execution; cancel: stop it.run
run_idNoget_status/cancel (serverless): run id.
computeNorun: 'cluster' (classic all-purpose cluster; any language) or 'serverless' (Python only, one-time serverless job run; slower to start).cluster
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
languageNorun: code language.python
cluster_idNoCluster id (default DBX_MCP_DEFAULT_CLUSTER_ID); also for get_status/cancel.
command_idNoget_status/cancel (cluster): command id.
context_idNoget_status/cancel (cluster): execution context id.
timeout_secondsNorun: max seconds to wait before returning 'pending' (capped by server).

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: fresh execution context per call with no state kept between calls, temp notebook path, stdout/stderr capture, success/failed/pending return semantics with ids to poll, and explicit cost/security note ('code can change data and costs money'). The safety classification (EXECUTION/READ_ONLY/WRITE depending on input) adds nuance the static destructiveHint=false does not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the one-line purpose, then uses tight action-scoped bullets. It is dense and slightly long, but nearly every clause earns its place (compute semantics, state behavior, timeout/polling).

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers compute selection, state model, return shapes, timeout/polling lifecycle, and the SQL alternative. With an output schema already present, the return-format detail is bonus rather than a requirement, and nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: it groups parameters by action (which ids apply to get_status/cancel vs run), explains the compute tradeoff, and ties timeout_seconds to the 'pending' return. The only gap is that it does not elaborate on confirm/dry_run beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Execute) and resource (Python/SQL/Scala/R code on Databricks compute) and returns output. It enumerates the three sub-actions (run, get_status, cancel) and explicitly names the sibling execute_sql for the SQL-on-warehouse case, so an agent can distinguish it from siblings without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit routing: 'For SQL on a SQL warehouse prefer execute_sql.' It also tells the agent when to use cluster vs serverless (serverless is Python-only, slower to start) and when to call get_status/cancel (pending executions). Alternatives and conditions are stated, not inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

execute_sqlExecute SQLA
Destructive

Execute one SQL statement on a Databricks SQL warehouse via the Statement Execution API.

The statement is classified before running: SELECT/SHOW/DESCRIBE are reads; INSERT/CREATE are writes; DROP/DELETE/TRUNCATE/UPDATE/MERGE/OR REPLACE/INSERT OVERWRITE are destructive and GRANT/REVOKE/ownership/row-filter/mask changes are security-sensitive. Destructive and security-sensitive statements require confirm=true. The response separates data.result (columns, rows, truncation) from data.execution (statement id, state, warehouse used and why). Rows are capped by max_rows. If the statement is still running after wait_timeout_seconds the response has status 'pending' - poll with manage_sql_statement.

Safety classification: depends on input (DESTRUCTIVE, EXECUTION, READ_ONLY, SECURITY_SENSITIVE, WRITE).

ParametersJSON Schema
NameRequiredDescriptionDefault
schemaNoDefault schema for unqualified names.
catalogNoDefault catalog for unqualified names.
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
max_rowsNoMaximum rows to return (capped by DBX_MCP_SQL_MAX_ROWS).
statementYesA single SQL statement (SELECT, DDL or DML). Use execute_sql_multi for scripts.
parametersNoNamed parameters referenced as :name in the statement (values are bound server-side, never interpolated).
row_formatNo'arrays' (compact, aligned with columns) or 'objects' (one dict per row).arrays
warehouse_idNoSQL warehouse id. If omitted: DBX_MCP_DEFAULT_WAREHOUSE_ID, else automatic selection (reported in the response).
wait_timeout_secondsNoSeconds to wait (5-50) before returning a pending statement id.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations by disclosing the statement classification scheme (read/write/destructive/security-sensitive and what falls in each bucket), the confirm gating behavior, row capping via max_rows, and the 'pending' status when wait_timeout_seconds elapses. The annotations only broadly declare destructiveHint=true; the description explains exactly which inputs trigger destructive behavior and what the response envelope contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then classification/safety rules, then response shape and timeout behavior. Dense but each sentence carries actionable content. The trailing 'Safety classification' line restates the classification rules already enumerated above, which is mild redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers classification, safety gating, dry-run/confirm flow, timeout and polling, row capping, and warehouse selection for a 10-parameter, high-stakes tool. An output schema exists, yet the description still usefully sketches the data.result/data.execution split, leaving no material gap for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is already 100%, so the baseline is 3. The description adds operational meaning by tying parameters together: max_rows caps returned rows, the wait_timeout/pending path connects to a follow-up tool, and warehouse selection falls back through env var to automatic selection reported in the response. Some of this repeats the schema, keeping it below 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (execute one SQL statement) and the exact backend (Databricks SQL warehouse via the Statement Execution API). It also implicitly distinguishes itself from the sibling execute_sql_multi (scripts) and manage_sql_statement (polling a pending statement by id), so an agent can separate it from neighbors without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete conditional guidance: destructive/security-sensitive statements require confirm=true, and a still-running statement should be polled with manage_sql_statement. It does not explicitly say when NOT to use this tool (e.g. multi-statement scripts), though that alternative is named in the statement parameter description rather than in the prose.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

execute_sql_multiExecute multiple SQL statementsA
Destructive

Execute several SQL statements sequentially, preserving order, and report success/failure per statement with statement-level errors. Stops at the first failure unless continue_on_error=true (remaining statements are reported as 'skipped'). There is no transaction: completed statements are not rolled back. Safety is the union of all statements' classifications (any destructive statement requires confirm=true).

Safety classification: depends on input (DESTRUCTIVE, EXECUTION, READ_ONLY, SECURITY_SENSITIVE, WRITE).

ParametersJSON Schema
NameRequiredDescriptionDefault
schemaNoDefault schema.
scriptNoA SQL script; split on top-level semicolons (comments/literals respected).
catalogNoDefault catalog.
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
row_formatNo'arrays' (compact, aligned with columns) or 'objects' (one dict per row).arrays
statementsNoStatements to run in order. Alternatively pass `script`.
warehouse_idNoSQL warehouse id. If omitted: DBX_MCP_DEFAULT_WAREHOUSE_ID, else automatic selection (reported in the response).
continue_on_errorNoKeep executing after a failed statement (default: stop at first failure).
max_rows_per_statementNoRow cap per statement result.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite destructiveness already flagged in annotations, the description adds substantial context: no transaction/rollback semantics, stop-at-first-failure default, 'skipped' reporting for remaining statements, and the union-of-classifications safety model requiring confirm=true. This goes well beyond what structured fields convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core behavioral rules are front-loaded and dense with no filler. The trailing 'Safety classification: depends on input (...)' line is a somewhat terse tag list, but it is short and does not bloat the definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With annotations covering the safety profile, a full 100% schema description coverage, and an output schema handling return values, the description supplies exactly the behavioral context an agent needs for a multi-statement, non-transactional executor. Nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, establishing a baseline of 3, but the description adds genuine meaning for continue_on_error ('remaining statements reported as skipped') and confirm (required for destructive/security-sensitive actions), clarifying behavior the schema text does not fully capture.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening clause states a specific verb and resource: 'Execute several SQL statements sequentially, preserving order.' This clearly differentiates it from the single-statement execute_sql sibling by emphasizing multiplicity, though it never names the sibling explicitly to route the agent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the continue_on_error and confirm behaviors well, but gives no explicit when-to-use guidance relative to execute_sql or execute_code. The choice of this tool over its single-statement sibling is only implied by the word 'multi'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_and_upload_pdfGenerate and upload PDFA
Destructive

Render HTML (or Markdown / plain text, converted to escaped HTML) to a PDF and upload it to a Unity Catalog Volume path ending in .pdf. Remote URLs, file: links and relative resources in the HTML are blocked (only inline data: URIs are used). Returns the path, size in bytes, page count and SHA-256. Requires the optional xhtml2pdf dependency (pip install "dbx-mcp[pdf]").

Safety classification: depends on input (DESTRUCTIVE, WRITE).

ParametersJSON Schema
NameRequiredDescriptionDefault
htmlNoHTML to render (external resources are never fetched).
textNoPlain text to render instead of html.
titleNoOptional document title (used for markdown/text/HTML fragments).
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
markdownNoMarkdown (headings, lists, code blocks, bold/italic) to render instead of html.
overwriteNoReplace an existing file (DESTRUCTIVE; requires confirm).
destination_pathYesTarget file in a Unity Catalog Volume, e.g. /Volumes/main/reports/files/q3.pdf

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds real context beyond annotations: external resources are blocked (only inline data: URIs), it discloses the return payload (path, size, page count, SHA-256), the optional dependency, and the destructive/write classification. Annotations already cover readOnly/destructive/openWorld hints, so the description's environment and output disclosures are the value-add; slightly more on overwrite semantics would push this to 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and format behaviour, then safety classification. Dense but each sentence carries information. Minor redundancy between the parenthetical normalization note and the mention of html/text/markdown parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be re-explained, but the description still names them helpfully. With annotations and full schema coverage, the description covers inputs, environment restrictions, and safety classification. The confirm/dry_run workflow could be spelled out slightly more for a destructive tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all 8 parameters including confirm, dry_run, and overwrite are already documented. The description reinforces input format conversion (Markdown/plain text to escaped HTML) but adds little syntax beyond the schema, making baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (render to PDF and upload) and resource (Unity Catalog Volume path ending in .pdf), and enumerates accepted input formats. No sibling tool does this, so it is distinguishable without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete context: input format options, destination constraint, dependency requirement. It also implies the confirm/dry_run flow for destructive actions. But it does not name an alternative (e.g., manage_volume_files) or say when NOT to use this vs. generic file tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_lakebase_credentialGenerate Lakebase credentialA

Generate a short-lived OAuth credential for connecting to Lakebase Postgres as the current identity.

kind='provisioned' (instance_names and/or claims) or kind='autoscaling' (endpoint, optional ttl_seconds). By default the token is NOT returned - only its expiration and connection details (host, port 5432, database databricks_postgres, user, sslmode=require). Pass reveal_token=true (with confirm=true) to receive the token in data.token; treat it as a secret and never log or store it.

Safety classification: SECURITY_SENSITIVE+WRITE.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNoprovisioned: Lakebase database instances (w.database). autoscaling: Lakebase autoscaling projects/branches/endpoints (w.postgres).provisioned
claimsNoOptional Unity Catalog claims scoping the token, e.g. [{"permission_set": "READ_ONLY", "resources": [{"table_name": "cat.schema.table"}]}].
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
endpointNoautoscaling: endpoint resource name projects/<p>/branches/<b>/endpoints/<e>.
request_idNoprovisioned: optional idempotency request id.
ttl_secondsNoautoscaling: token lifetime in seconds (300-3600).
reveal_tokenNoReturn the token itself. Default false: only expiry and connection details are returned. Requires confirm=true.
instance_namesNoprovisioned: database instance names the credential is for.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false, destructiveHint=false, and openWorldHint=true. The description adds substantial context beyond that: the SECURITY_SENSITIVE+WRITE safety classification, the fact that the token is withheld by default, exactly what connection details are returned, and that reveal_token requires confirm=true and must be treated as a secret. This is rich disclosure that the structured fields do not provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded in the first sentence, followed by mode routing, default-return behavior, and secret handling. Sentences are dense and mostly earn their place, though the trailing 'Safety classification' tag reads as metadata rather than guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With annotations present, an output schema available, and 100% schema coverage, the remaining burden on the description is the mode selection and the token-reveal safety contract, both of which are covered. An agent has everything needed to call this correctly and safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the per-parameter documentation baseline is 3. The description goes beyond the schema by documenting inter-parameter relationships the schema cannot express: kind selects which of instance_names/claims vs endpoint/ttl_seconds apply, and reveal_token is gated on confirm. That added coupling justifies a 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Generate) and resource (short-lived OAuth credential) with scope (for Lakebase Postgres, as the current identity), which cleanly separates it from siblings like manage_lakebase_database and manage_lakebase_branch. An agent can identify the tool's function without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description routes the caller between the two operating modes (kind='provisioned' vs kind='autoscaling') and tells which parameters apply to each, which is clear contextual guidance. It does not, however, name any sibling alternative or state when-not to use this tool, so it stops short of explicit when/when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_current_userGet current userA
Read-only

Return the Databricks identity this server is authenticated as: username, user id, display name, group memberships, home folder and workspace. Never returns credentials.

Safety classification: READ_ONLY.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=true and destructiveHint=false. The description adds real value beyond them: it discloses the returned identity scopes and explicitly states 'Never returns credentials', which is a security-relevant guarantee the agent could not infer from structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences, front-loaded with the verb and resource, followed by the return scope and the credential caveat. 'Safety classification: READ_ONLY' mildly restates the readOnlyHint annotation but is short and harmless.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no parameters and an existing output schema, the description only needs to orient the agent, which it does. It could have noted failure behavior when the session is unauthenticated, but nothing essential for invoking the tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters and the schema is empty, so there is nothing for the description to disambiguate. Baseline 4 applies; the field enumeration is a return-value concern rather than a parameter one.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Return the Databricks identity this server is authenticated as') and enumerates exactly what that identity includes (username, user id, display name, groups, home folder, workspace). No sibling tool overlaps with identity retrieval, so the agent can distinguish it immediately.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the phrase 'this server is authenticated as' — the tool is for resolving the acting principal — but there is no explicit when-to-use, when-not-to-use, or named alternative. In practice no sibling competes with it, so the missing routing guidance costs little.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_table_stats_and_schemaTable/schema details and statisticsA
Read-only

Inspect a Unity Catalog table (catalog, schema, type, format, columns with types, nullability, comments, partition columns, location, owner, properties, row filter/masks presence) plus optional statistics (file count, size, partitioning, row count). Given a two-part 'catalog.schema' name, lists the tables in that schema (paginated).

Safety classification: depends on input (EXECUTION, READ_ONLY).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesFully qualified table 'catalog.schema.table' for one table, or 'catalog.schema' to list the tables in a schema.
statsNonone: Unity Catalog metadata only. metadata: also DESCRIBE DETAIL (files, size, partitioning) on a warehouse. exact_count: also SELECT COUNT(*) (scans the table). auto: metadata only if a warehouse is already running (never starts one).auto
page_sizeNoMax items to return (server caps this).
page_tokenNonext_page_token from a previous response.
warehouse_idNoSQL warehouse id. If omitted: DBX_MCP_DEFAULT_WAREHOUSE_ID, else automatic selection (reported in the response).

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety baseline is covered. The description goes further by disclosing that exact_count scans the table and that a warehouse may be started (auto never starts one), plus the 'depends on input' safety classification — useful side-effect and cost context beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and the returned-field inventory, then the alternate listing mode and safety note. It is dense but every clause carries information; the long parenthetical field list is borderline verbose but useful for an inspection tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and full schema coverage, the description needn't detail return shapes, and it correctly focuses on the dual mode, stats cost semantics, and safety classification. An agent has enough to call it correctly; only explicit sibling alternatives are missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already carries name, stats enum, pagination, and warehouse resolution details. The description largely restates these (pagination for schema listing, stats categories) without adding format or syntax beyond the structured fields, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Inspect) on a specific resource (Unity Catalog table), enumerates exactly what metadata is returned, and covers the distinct second mode (listing tables in a schema from a two-part name). This is easily separable from the manage_* siblings, which are mutations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The two usage modes (full table name vs catalog.schema) are explained, which is genuinely helpful routing. However, there is no explicit when-to-use/when-not-to-use guidance or named alternative (e.g. execute_sql, manage_uc_objects) for overlapping needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_volume_folder_detailsVolume folder detailsA
Read-only

Inspect a Unity Catalog Volume path. For a directory: entries with type (file/directory), size, modification time and detected format (parquet, csv, json, delta, avro, orc, text, ...), plus summary counts/total size by format; recursive walks sub-directories within max_depth / max_entries caps and reports directories containing _delta_log as Delta tables. For a file: its metadata (size, content type, last modified, format). Entries are paginated.

Safety classification: READ_ONLY.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesVolume path: /Volumes/<catalog>/<schema>/<volume>[/sub/path].
max_depthNoMax directory depth when recursive (1-10).
page_sizeNoMax items to return (server caps this).
recursiveNoAlso list sub-directories (bounded by max_depth/max_entries).
page_tokenNonext_page_token from a previous response.
max_entriesNoMax entries scanned in total (capped at 10000).
detect_deltaNoProbe listed sub-directories for a _delta_log folder (Delta tables).

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety burden is covered, yet the description still adds real behavioral detail: recursion is bounded by max_depth/max_entries caps, the scan is capped at 10000 entries, results are paginated, and directories with _delta_log are reported as Delta tables. This is meaningful context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the verb and resource in the first clause, then branches cleanly into directory and file cases. Dense but every clause about formats, caps and delta detection earns its place; the trailing 'Safety classification: READ_ONLY.' is mildly redundant with the annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, yet the description helpfully enumerates them anyway. Combined with annotations covering safety and a fully documented schema, an agent has nearly everything needed; only sibling routing is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all seven parameters and the baseline is 3. The description still adds interpretive value by explaining what the recursive, max_depth and max_entries parameters jointly govern and what delta detection actually probes for.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (inspect) and resource (a Unity Catalog Volume path) and clearly distinguishes directory vs file behavior. It does not, however, differentiate itself from the obvious sibling manage_volume_files, so the agent must infer the boundary itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains what is returned but never says when to use this tool versus alternatives such as manage_volume_files, nor any precondition or exclusion. Usage is left entirely to inference from a sibling list the description ignores.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_computeList compute resourcesA
Read-only

List compute: all-purpose clusters and SQL warehouses with current state, available node types (cores, memory, GPUs, Photon support) and Databricks Runtime (Spark) versions.

Safety classification: READ_ONLY.

ParametersJSON Schema
NameRequiredDescriptionDefault
filterNoCase-insensitive substring filter on name/id (node_types, spark_versions).
resourceNosummary: clusters + warehouses with states; or one resource type.summary
page_sizeNoMax items to return (server caps this).
page_tokenNonext_page_token from a previous response.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=true, so the 'Safety classification: READ_ONLY' sentence merely restates structured data and earns no credit. The description adds useful context about what the listing exposes (current state, node types, runtime versions), but says nothing about auth needs, result caps, or pagination behavior beyond what the schema already carries.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded and compact: the core scope sentence comes first and reads cleanly. The trailing 'Safety classification: READ_ONLY' line is pure redundancy against the readOnlyHint annotation and should be dropped, which is the only real waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description need not explain return values, and it does adequately convey the two resource families and their attributes. Given 4 parameters at full schema coverage and rich annotations, the remaining gap is the absence of any guidance on narrowing results with the filter or handling pagination.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (filter, resource, page_size, page_token) are already documented with semantics and defaults. The description only loosely echoes the resource categories through its mention of node types and runtime versions, adding little beyond the enum's own description. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') plus the concrete resource set it covers ('all-purpose clusters and SQL warehouses') and enumerates the returned attributes (state, node types, Photon support, Databricks Runtime versions). This cleanly separates it from the manage_cluster and manage_sql_warehouse siblings, which are mutation-oriented.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the verb and the scope statement, but there is no explicit guidance on when to call this versus manage_cluster, manage_sql_warehouse, or get_table_stats_and_schema, and no mention of prerequisites or follow-up tools. An agent can infer the read-only inventory role but gets no routing help.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_tracked_resourcesList tracked resourcesA
Read-only

List resources recorded in the local project manifest (created through this server): type, id, name, creating tool, creation time and workspace. With verify=true, each returned item is checked against Databricks and missing ones are reported. Paginated.

Safety classification: READ_ONLY.

ParametersJSON Schema
NameRequiredDescriptionDefault
verifyNoCheck whether each returned resource still exists in Databricks (one GET per item; supported for job, pipeline, dashboard, app, cluster, warehouse).
page_sizeNoMax items to return (server caps this).
page_tokenNonext_page_token from a previous response.
resource_typeNoFilter by type, e.g. job, pipeline, dashboard, app, cluster, warehouse.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, destructiveHint=false, and openWorldHint, so the safety profile is covered. The description adds real behavioral context beyond them: the manifest-only scope, that verify performs a Databricks existence check per item and reports missing ones, and that results are paginated. Cost side-effects of verify are disclosed, though the trailing 'Safety classification: READ_ONLY' merely restates the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, front-loaded with the purpose and scope before the verify behavior and pagination note; no filler in the body. The final 'Safety classification: READ_ONLY.' sentence is redundant with the readOnlyHint annotation and could be cut, which keeps it short of a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the field list is a bonus rather than a requirement, and the annotations cover safety. The description supplies the scope, the verify side-effect, and pagination, which is everything an agent needs to call this read-only list tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents verify, page_size, page_token, and resource_type. The description reinforces verify semantics ('missing ones are reported') and implies pagination, adding only marginal meaning beyond the schema. Baseline 3 is appropriate when the schema carries the parameter burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (List tracked resources) and immediately defines the scope: resources recorded in the local project manifest created through this server. It enumerates the returned fields (type, id, name, creating tool, creation time, workspace), which makes the tool self-differentiating from siblings like list_compute and delete_tracked_resource without needing to open a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The verify=true clause gives concrete context for when that option is wanted, and mentioning pagination hints at iterative use. However, there is no explicit when-to-use-this-vs-alternatives guidance or prerequisite statement (e.g., use before delete_tracked_resource, only sees resources created via this server). Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_appManage Databricks AppsA
Destructive

Manage Databricks Apps. Actions: create (name, app fields, no_compute), get, list, update (partial: only the given app fields), delete, deploy (source_code_path, mode SNAPSHOT|AUTO_SYNC, extra deployment fields), get_deployment, list_deployments, start, stop. create/deploy/start/stop return immediately with status 'pending' unless wait=true (bounded). logs is not available via the API/SDK. Created apps are tracked in the project manifest.

Safety classification: depends on input (DESTRUCTIVE, EXECUTION, READ_ONLY, SECURITY_SENSITIVE, WRITE).

ParametersJSON Schema
NameRequiredDescriptionDefault
appNocreate/update: App fields (REST names), e.g. description, resources, compute_size, user_api_scopes, budget_policy_id. update changes only the fields given.
modeNodeploy: SNAPSHOT (copy source now) or AUTO_SYNC (keep syncing from source_code_path).
nameNoApp name (lowercase letters, numbers, hyphens).
waitNoWait (bounded) for the operation to reach a steady state.
actionYescreate | get | list | update | delete | deploy | get_deployment | list_deployments | start | stop | logs
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
page_sizeNoMax items to return (server caps this).
deploymentNodeploy: extra AppDeployment fields (e.g. git_source, command, env_vars).
no_computeNocreate: do not start app compute after creation.
page_tokenNonext_page_token from a previous response.
deployment_idNoget_deployment: deployment id.
timeout_secondsNoMax seconds to wait when wait=true (capped by DBX_MCP_MAX_WAIT_SECONDS).
source_code_pathNodeploy: workspace folder with the app source, e.g. /Workspace/Users/me/app.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true/openWorld, but the description adds real value beyond them: safety level varies by action (DESTRUCTIVE/EXECUTION/READ_ONLY/SECURITY_SENSITIVE/WRITE), create/deploy/start/stop return 'pending' unless wait=true (bounded), `logs` is unavailable, and created apps are tracked in the manifest. Return format is left to the output schema, which exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then a scannable action list, then behavioral notes. Dense but every clause carries signal (async semantics, partial-update semantics, safety variance); only mildly verbose overall.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists, return values need not be described, and the description covers actions, async/wait behavior, and per-action safety. The confirmation/dry_run flow is delegated to the schema, which is acceptable but leaves a little room.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaningful action-to-parameter mapping (e.g. deploy = source_code_path + mode SNAPSHOT|AUTO_SYNC, update = partial fields) that orients the agent before reading the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (manage Databricks Apps) and enumerates the full action set, so the agent knows exactly the operation surface without opening the schema. The name and scope cleanly separate it from sibling manage_* tools targeting clusters, warehouses, jobs, etc.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Maps each action to its relevant inputs (create takes name/app/no_compute, deploy takes source_code_path/mode, etc.), which gives clear invocation context. It does not, however, explicitly route between siblings or state when-not-to-use, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_clusterManage clustersA
Destructive

Manage all-purpose clusters.

Actions: list (optionally filtered by state), get, events (recent cluster events), create (spec = Clusters API create body, e.g. {"cluster_name","spark_version","node_type_id", "num_workers" or "autoscale","autotermination_minutes"}), update (partial update: spec holds only the fields to change), resize, start, restart, terminate (stop; restartable) and delete (permanent). restart/terminate/delete require confirm=true and are refused for clusters whose name/tags match the protected (production) patterns. Lifecycle actions return immediately with the current state unless wait=true.

Safety classification: list, get, events = READ_ONLY; create, update, resize, start = WRITE; restart, terminate, delete = DESTRUCTIVE.

ParametersJSON Schema
NameRequiredDescriptionDefault
specNoRequest body fields for create/update, using the Databricks REST API field names (snake_case). Unknown fields are rejected.
waitNoWait (bounded) for the operation to reach a steady state.
actionYesterminate = stop (restartable); delete = permanent removal.
statesNoFor list: filter by states, e.g. ['RUNNING','PENDING'].
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
page_sizeNoMax items to return (server caps this).
cluster_idNoCluster id (all actions except list/create).
page_tokenNonext_page_token from a previous response.
num_workersNoFor resize: fixed worker count.
timeout_secondsNoMax seconds to wait when wait=true (capped by DBX_MCP_MAX_WAIT_SECONDS).
autoscale_max_workersNoFor resize: autoscale maximum.
autoscale_min_workersNoFor resize: autoscale minimum.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses behavior far beyond annotations: confirm gating, refusal for clusters matching protected production name/tag patterns, immediate vs. wait behavior for lifecycle actions, and a full safety classification mapping each action to READ_ONLY/WRITE/DESTRUCTIVE. This is exactly the context a mutation dispatcher needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded and well-sectioned under 'Actions:' and 'Safety classification:', with no filler sentences. It repeats the terminate/delete semantics already in the action enum description, a minor redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 13-parameter, 10-action dispatcher with an output schema present, the description covers every action, the confirm/dry_run/wait contract, and the safety profile. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so the baseline is 3, but the description adds real value by supplying the create spec body example field names ({cluster_name, spark_version, node_type_id, num_workers/autoscale, autotermination_minutes}) and the partial-update semantics for spec, which the schema does not enumerate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Manage all-purpose clusters') and enumerates every action (list, get, events, create, update, resize, start, restart, terminate, delete) with a parenthetical clarifying each. An agent can distinguish it from siblings like list_compute or manage_jobs without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear per-action conditions: confirm=true is required for restart/terminate/delete, wait=true controls whether lifecycle actions return immediately, and list can be state-filtered. It never names alternative sibling tools to route to, so it stops short of explicit when-not/alternatives guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_dashboardManage AI/BI dashboardsA
Destructive

Manage AI/BI (Lakeview) dashboards. Actions: create (display_name, optional parent_path, warehouse_id, serialized_dashboard), get, list (show_trashed), update (draft fields; etag for optimistic concurrency), delete (moves to trash; recoverable), publish (embed_credentials, warehouse_id), unpublish, get_published. Created dashboards are tracked in the project manifest.

Safety classification: depends on input (DESTRUCTIVE, READ_ONLY, SECURITY_SENSITIVE, WRITE).

ParametersJSON Schema
NameRequiredDescriptionDefault
etagNoupdate: etag from get, to fail if the draft changed meanwhile.
actionYescreate | get | list | update (draft) | delete (move to trash) | publish | unpublish | get_published
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
page_sizeNoMax items to return (server caps this).
page_tokenNonext_page_token from a previous response.
parent_pathNocreate: workspace folder for the dashboard, e.g. /Users/me@x.com/dashboards.
dashboard_idNoDashboard id (all actions except create/list).
display_nameNocreate/update: dashboard name.
show_trashedNolist: include dashboards in the trash.
warehouse_idNocreate/update: SQL warehouse for the draft; publish: override warehouse.
dataset_schemaNocreate/update: default schema for all datasets.
dataset_catalogNocreate/update: default catalog for all datasets.
embed_credentialsNopublish: run viewers' queries with the publisher's credentials (SECURITY_SENSITIVE; requires confirm). Default false: viewers use their own credentials.
serialized_dashboardNocreate/update: the dashboard definition (JSON string or object, as exported by Databricks).

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint=false, destructiveHint=true), the description adds real behavioral context: delete moves to trash and is recoverable, update uses etag for optimistic concurrency, publish with embed_credentials is SECURITY_SENSITIVE and requires confirm, and created dashboards are tracked in the project manifest. It also clarifies the safety classification varies by action rather than being fixed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the resource and then a compact action list with parenthetical arguments, followed by two short standalone notes. Dense but every clause carries information; only the safety-classification restatement is somewhat redundant with annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 15-parameter, 8-action tool with an output schema and annotations present, the description covers the action surface, the recoverability of delete, the concurrency mechanism, and credential sensitivity. Pagination and the confirmation/dry_run workflow are left to the schema, which is reasonable given full coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter is already documented in the schema. The description maps parameters to actions (e.g. parent_path and serialized_dashboard for create, etag for update), which adds light orientation but nothing beyond the schema's own descriptions. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Manage) and resource (AI/BI Lakeview dashboards) and enumerates the eight discrete actions with their key arguments. An agent can identify this as the dashboard lifecycle tool and distinguish it from siblings like manage_metric_views or manage_workspace.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The per-action argument mapping implies what each action does, which aids action selection, but there is no explicit guidance on when to use this tool versus siblings, nor any when-not conditions. Usage is inferable but not spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_genieGenie spacesA
Destructive

Manage AI/BI Genie spaces (natural-language-to-SQL over Unity Catalog tables).

Actions:

  • list / get (include_serialized_space for the full definition).

  • create: spec = {warehouse_id, serialized_space (JSON string or object), title, description, parent_path}. Tip: get an existing space with include_serialized_space=true to see the serialized_space format.

  • update: spec with any of title, description, warehouse_id, parent_path, serialized_space (full replacement), etag.

  • delete: move the space to trash (requires confirm).

  • list_conversations / list_messages (conversation_id) / delete_conversation (requires confirm). Use ask_genie to ask questions.

Safety classification: list, get, list_conversations, list_messages = READ_ONLY; create, update = WRITE; delete, delete_conversation = DESTRUCTIVE.

ParametersJSON Schema
NameRequiredDescriptionDefault
specNoRequest body fields for create/update, using the Databricks REST API field names (snake_case). Unknown fields are rejected.
actionYesOperation to perform.
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
space_idNoGenie space id (all actions except list/create).
page_sizeNoMax items to return (server caps this).
page_tokenNonext_page_token from a previous response.
include_allNolist_conversations: include all users' conversations (requires CAN MANAGE).
conversation_idNoConversation id for list_messages/delete_conversation.
include_serialized_spaceNoget: include the serialized space definition (requires CAN EDIT).

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations by classifying each action as READ_ONLY/WRITE/DESTRUCTIVE, disclosing that delete moves to trash, that update's serialized_space is a full replacement, that etag is used for concurrency, that list_conversations requires CAN MANAGE and include_serialized_space requires CAN EDIT, and the two-step confirm flow.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the resource scope, then uses a tight per-action list plus a safety classification line. Every line carries decision-relevant information; nothing is padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-action, 10-parameter tool with an output schema and annotations, the description supplies the action semantics, safety tiers, confirmation protocol, and permissions that structured fields alone would not convey. Nothing an agent needs to select or invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, but the description adds action-to-parameter mapping that the schema does not carry: which spec fields apply to create vs update, that etag belongs to update, and that include_serialized_space is a get-only flag. Useful, though individual parameter meanings largely remain in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (manage AI/BI Genie spaces) and immediately scopes it as natural-language-to-SQL over Unity Catalog tables. It enumerates every action, so an agent can tell it apart from ask_genie and the other manage_* siblings without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes query-asking to ask_genie, tells the agent which actions need confirm, and gives a concrete tip (fetch include_serialized_space=true first to learn the spec format). Usage conditions are stated, not implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_job_runsManage job runsA
Destructive

Start, monitor, inspect, cancel and repair Databricks job runs.

  • submit: spec = one-time run (run_name, tasks [task_key + task type + compute], environments, git_source, timeout_seconds, idempotency_token, ...). Returns run_id with status 'pending'.

  • list (job_id, active_only/completed_only, start_time_from/to), get (run_id: state, per-task states, error messages), wait (run_id, timeout_seconds: bounded poll), get_output (run_id[, task_key]: notebook exit values, logs, errors/stack traces; multi-task runs are expanded per task).

  • cancel (run_id), cancel_all (job_id or all_queued_runs), delete_run (run_id): DESTRUCTIVE, need confirm.

  • repair (run_id, spec: rerun_all_failed_tasks | rerun_tasks, rerun_dependent_tasks, latest_repair_id, job_parameters, ...): EXECUTION.

Safety classification: depends on input (DESTRUCTIVE, EXECUTION, READ_ONLY, SECURITY_SENSITIVE).

ParametersJSON Schema
NameRequiredDescriptionDefault
specNoRequest body fields for create/update, using the Databricks REST API field names (snake_case). Unknown fields are rejected.
waitNosubmit/repair: poll until the run finishes (bounded).
actionYessubmit: one-time run (spec = runs/submit body); list: runs (filters); get: run with task states and errors; get_output: outputs/errors per task; wait: poll until finished (bounded); cancel: one run; cancel_all: all active runs of a job; repair: re-run failed/selected tasks; delete_run: delete a finished run record.
job_idNolist/cancel_all: restrict to this job.
run_idNoRun id (get/get_output/wait/cancel/repair/delete_run).
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
task_keyNoget_output: only this task's output.
page_sizeNoMax items to return (server caps this).
page_tokenNonext_page_token from a previous response.
active_onlyNolist: only active runs.
start_time_toNolist: runs started at/before (epoch ms).
completed_onlyNolist: only completed runs.
all_queued_runsNocancel_all: cancel queued runs (of all jobs when job_id is omitted).
start_time_fromNolist: runs started at/after (epoch ms).
timeout_secondsNoMax seconds to wait (capped by server).

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, and openWorldHint=true, so the safety profile is partly covered. The description does add value by pinning destructiveness to specific actions and describing the confirm-after-confirmation_required workflow plus the bounded-poll behavior of wait. However, it does not describe permissions requirements, rate limits, or what state is left behind after a destructive action, so it remains moderate rather than rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The summary sentence is front-loaded, followed by a compact bulleted map of action-to-parameter semantics and a one-line safety note. For a 9-action, 16-parameter tool the length is justified, though the per-action bullets partly restate the schema's own action enum descriptions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be re-explained, and the description covers action semantics, the confirmation gate, and per-action destructive classification. It is complete enough to call correctly, with only minor gaps such as pagination expectations that the schema already handles.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter is already documented in the schema and the baseline is 3. The description adds some grouping value by mapping actions to their parameters (e.g., spec fields for submit, repair's rerun options, task_key for get_output), but it does not add format or syntax details beyond what the schema already carries.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence names a specific resource (Databricks job runs) with concrete verbs (start, monitor, inspect, cancel, repair), and the bullet list enumerates every action. It is clear what the tool does, but it never names a sibling such as manage_jobs (job definitions) or manage_pipeline_run to disambiguate runs from job-level management, so it stops short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Each action is annotated with its intent and required inputs, and destructive actions (cancel, cancel_all, delete_run) are flagged as needing confirmation while repair is flagged EXECUTION, which tells the agent how to proceed. What is missing is explicit routing guidance against alternative tools when a user wants to manage job definitions rather than runs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_jobsManage jobsA
Destructive

Create, inspect, change, delete and trigger Databricks Lakeflow Jobs.

  • create: spec = JobSettings fields (name, tasks, job_clusters, environments, schedule, trigger, continuous, parameters, email_notifications, webhook_notifications, tags, queue, max_concurrent_runs, timeout_seconds, git_source, run_as, access_control_list, ...). Each task needs task_key and one task type (notebook_task, spark_python_task, python_wheel_task, sql_task, pipeline_task, run_job_task, ...) plus compute (existing_cluster_id, job_cluster_key, new_cluster, or environment_key for serverless).

  • get (job_id), list (name filter, paginated).

  • update (job_id, spec and/or fields_to_remove): partial; top-level fields in spec replace existing ones, tasks/job_clusters are merged by key.

  • reset (job_id, spec): full overwrite of all settings (DESTRUCTIVE, needs confirm).

  • delete (job_id): DESTRUCTIVE, needs confirm.

  • run_now (job_id, spec: job_parameters, notebook_params, python_params, only, queue, performance_target, idempotency_token, ...): returns the run_id immediately (status 'pending'); wait=true polls (bounded). Specs setting run_as/access_control_list are additionally SECURITY_SENSITIVE (confirm required).

Safety classification: depends on input (DESTRUCTIVE, EXECUTION, READ_ONLY, SECURITY_SENSITIVE, WRITE).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNolist only: exact job name filter (server-side).
specNoRequest body fields for create/update, using the Databricks REST API field names (snake_case). Unknown fields are rejected.
waitNorun_now only: poll until the run finishes (bounded).
actionYescreate: new job from spec; get: full job definition; list: jobs (optional name filter); update: partial change (spec = fields to set, fields_to_remove); reset: replace ALL settings with spec; delete: delete the job; run_now: trigger a run (spec = run parameters).
job_idNoJob id (get/update/reset/delete/run_now).
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
page_sizeNoMax items to return (server caps this).
page_tokenNonext_page_token from a previous response.
timeout_secondsNorun_now with wait=true: max seconds to wait (capped by server).
fields_to_removeNoupdate only: top-level settings to remove, or 'tasks/<task_key>' / 'job_clusters/<key>'.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true, readOnlyHint=false, openWorldHint=true, but the description adds rich operational detail: which actions require confirm, the SECURITY_SENSITIVE confirmation for run_as/access_control_list, partial vs full overwrite semantics, and the wait polling behavior for run_now. This goes well beyond structured data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with a clear purpose statement, then uses bullet points to organize action-specific details. It is somewhat long because it enumerates many spec fields, but the structure is efficient and no sentence is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (11 parameters, 7 actions), the description covers all actions, safety confirmations, dry_run, and run_now return behavior. With an output schema present, return values need not be explained, so the definition is complete for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 100% description coverage, so the baseline is 3. However, the description adds significant meaning about the 'spec' object expected fields, task requirements, and update merge semantics, which are not fully captured in the schema's high-level parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb list and resource ('Create, inspect, change, delete and trigger Databricks Lakeflow Jobs'), then enumerates every supported action. It does not explicitly name alternative tools for run management, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The action descriptions give clear context for when to use each mode (e.g., create vs update vs reset), and the safety notes explain prerequisites like confirm and dry_run. No explicit when-not-to-use guidance or sibling alternatives are named, so it stops short of the top rubric.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_kaKnowledge AssistantsA
Destructive

Manage Knowledge Assistants (Agent Bricks document Q&A agents over UC volumes, tables or vector indexes).

Actions:

  • list / get / delete; create (spec: display_name, description, instructions); update (spec fields among display_name, description, instructions).

  • list_sources / get_source / delete_source; add_source (spec: display_name, description, source_type 'files'|'index'|'file_table' plus files={path:'/Volumes/...'} or index={index_name,text_col,doc_uri_col} or file_table={table_name,file_col}); update_source (display_name, description); sync_sources re-ingests non-index sources.

  • list_examples / get_example / add_example (spec: question, guidelines) / update_example / delete_example.

  • get_permissions / update_permissions (spec: access_control_list). Query an assistant through its serving endpoint (manage_serving_endpoint action=query).

Safety classification: list, get, list_sources, get_source, list_examples, get_example = READ_ONLY; create, update, add_source, update_source, add_example, update_example = WRITE; delete, delete_source, delete_example = DESTRUCTIVE; sync_sources = EXECUTION+WRITE; get_permissions = READ_ONLY+SECURITY_SENSITIVE; update_permissions = SECURITY_SENSITIVE+WRITE.

ParametersJSON Schema
NameRequiredDescriptionDefault
specNoRequest body fields for create/update, using the Databricks REST API field names (snake_case). Unknown fields are rejected.
actionYesOperation to perform.
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
page_sizeNoMax items to return (server caps this).
source_idNoKnowledge source id (or full resource name).
example_idNoExample id (or full resource name).
page_tokenNonext_page_token from a previous response.
update_maskNoComma-separated fields to update; defaults to the keys present in spec.
knowledge_assistant_idNoKnowledge Assistant id or resource name 'knowledge-assistants/{id}'.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only give a global destructiveHint=true and openWorldHint=true, which over-warns for the many read actions. The description adds real per-action granularity by classifying each action as READ_ONLY, WRITE, DESTRUCTIVE, EXECUTION+WRITE, or SECURITY_SENSITIVE, and notes that sync_sources re-ingests non-index sources. That is meaningful disclosure beyond the annotations, though it does not describe async timing or failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded, actions are grouped as scannable bullets, and the safety classification is one consolidated line rather than repeated per action. It is dense and long, but the density is driven by 18 actions and 10 parameters, and nearly every clause carries actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a broad multi-action tool with a rich output schema and fully described parameters, the description supplies the missing pieces: action inventory, per-action payload requirements, safety tiers, and the cross-tool route for querying. Nothing an agent needs to select and invoke an action correctly is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3 and the schema alone documents confirm, dry_run, page_size, ids, and update_mask. The description goes beyond that by spelling out the spec payload shape per action (display_name/description/instructions; source_type 'files'|'index'|'file_table' with the nested key names for each), which the generic 'spec' schema object does not convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence gives a specific verb+resource (manage Knowledge Assistants) and immediately characterizes what they are (Agent Bricks document Q&A agents over UC volumes, tables or vector indexes). The action enumeration then maps every sub-operation, and the sibling it is NOT for (querying an assistant) is explicitly routed to manage_serving_endpoint.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The action list plus per-action spec hints tell the agent which operation to pick and what fields each requires. It also names the alternative for the query path (manage_serving_endpoint action=query). However, there is no explicit when-not guidance for overlapping siblings like manage_vs_index or query_vs_index, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_lakebase_branchManage Lakebase branchesA
Destructive

Manage Lakebase autoscaling branches (copy-on-write Postgres branches) and their compute endpoints.

Branch actions: list (project), get, create (project, branch id, optional source_branch, source_branch_time for point-in-time, source_branch_lsn, spec), update (spec, e.g. {"is_protected": true}), delete (soft unless purge=true; the default branch is refused unless allow_default_branch=true), undelete. Endpoint actions: list_endpoints, get_endpoint, create_endpoint (endpoint id + spec with endpoint_type), update_endpoint (e.g. CU limits, {"disabled": true}), delete_endpoint. get_operation polls a long-running operation. Writes return status 'pending' unless wait_seconds.

Safety classification: list, get, list_endpoints, get_endpoint, get_operation = READ_ONLY; create, update, undelete, create_endpoint, update_endpoint = WRITE; delete, delete_endpoint = DESTRUCTIVE.

ParametersJSON Schema
NameRequiredDescriptionDefault
specNoBranchSpec fields (create/update, e.g. {"ttl": "86400s"}, {"no_expiry": true}, {"is_protected": true}) or EndpointSpec fields (create_endpoint/update_endpoint, e.g. {"endpoint_type": "ENDPOINT_TYPE_READ_WRITE", "autoscaling_limit_min_cu": 0.5, "autoscaling_limit_max_cu": 2}). Unknown fields are rejected.
purgeNodelete: hard delete (irreversible). Default is a soft delete restorable with action='undelete'.
actionYesBranch lifecycle, plus compute endpoints (*_endpoint) of a branch.
branchNoBranch id or full name 'projects/<p>/branches/<b>'.
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
projectNoProject id or 'projects/<id>'.
endpointNoEndpoint id or full endpoint name.
page_sizeNoMax items to return (server caps this).
page_tokenNonext_page_token from a previous response.
update_maskNoComma-separated field paths to update. Default: derived from the keys of `spec` (autoscaling resources use 'spec.<field>' paths).
show_deletedNolist: include soft-deleted branches.
wait_secondsNoSeconds to wait for a long-running create/update/delete to finish. 0 (default) returns immediately with status 'pending'. Capped by DBX_MCP_MAX_WAIT_SECONDS and the tool timeout.
source_branchNocreate: parent branch id/name to branch from (default: the project's default branch).
operation_nameNoAutoscaling operation name returned by a previous call (for action='get_operation').
source_branch_lsnNocreate: Postgres LSN of the parent branch to branch from.
source_branch_timeNocreate: point in time of the parent branch (RFC3339, e.g. 2025-01-31T12:00:00Z).
allow_default_branchNodelete: permit deleting the project's default branch (refused otherwise).

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations it discloses a full safety taxonomy (READ_ONLY/WRITE/DESTRUCTIVE per action), reversibility of soft delete via undelete, the default-branch refusal, the 'pending' async status and wait_seconds semantics, and the confirm/dry_run flow. This is far richer than what readOnlyHint/destructiveHint already convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but tightly organized into branch actions, endpoint actions, and a safety classification line, with the goal statement front-loaded. Every clause carries information an agent needs; there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 18-parameter multi-action tool with an output schema present, the description covers action scoping, async semantics, confirmation, and destructive guards. Nothing an agent needs to invoke it correctly is missing, and return values need not be explained given the output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds value by grouping which parameters belong to which action alongside the per-action behaviors they trigger. It does not restate the spec field formats those parameters already document, so it stops at 4 rather than 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence names the specific resource ('Lakebase autoscaling branches (copy-on-write Postgres branches) and their compute endpoints') and the description then enumerates every action, so an agent knows exactly what the tool covers. It clearly separates itself from siblings like manage_lakebase_database and generate_lakebase_credential by scoping to branch/endpoint lifecycle.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It maps every action to the parameters it consumes and states the conditions that gate risky actions ('soft unless purge=true; the default branch is refused unless allow_default_branch=true'). What is missing is explicit routing against sibling tools (e.g. when to prefer manage_lakebase_database), so it stops short of the top score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_lakebase_databaseManage Lakebase databasesA
Destructive

Manage Lakebase (Postgres) databases.

kind='provisioned' manages database instances: list, get, create (spec = DatabaseInstance fields, e.g. {"capacity": "CU_1"}), update (spec = fields to change, e.g. {"stopped": true} or {"capacity": "CU_2"}), delete (force=true also removes point-in-time children). kind='autoscaling' manages projects: list, get, create (spec = Project fields, e.g. {"spec": {"display_name": "My app", "pg_version": 17}}), update (e.g. {"spec": {"display_name": "x"}}), delete (soft unless purge=true), undelete, get_operation. Catalog actions register a Postgres database in Unity Catalog: list_catalogs (provisioned, name = instance), get_catalog, create_catalog (catalog_name, database_name, name/branch), delete_catalog. Compute is billed; long-running work returns status 'pending' unless wait_seconds is set.

Safety classification: list, get, get_operation, list_catalogs, get_catalog = READ_ONLY; create, update, undelete, create_catalog = WRITE; delete, delete_catalog = DESTRUCTIVE.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNoprovisioned: Lakebase database instances (w.database). autoscaling: Lakebase autoscaling projects/branches/endpoints (w.postgres).provisioned
nameNoprovisioned: database instance name. autoscaling: project id or 'projects/<id>'.
specNoRequest body fields for create/update, using the Databricks REST API field names (snake_case). Unknown fields are rejected.
forceNoprovisioned delete: also delete descendant point-in-time instances (otherwise the delete is rejected if any exist).
purgeNoautoscaling delete: hard delete (irreversible). Default is a soft delete restorable with action='undelete'.
actionYesOperation to perform. *_catalog actions register/unregister a Lakebase Postgres database as a Unity Catalog catalog.
branchNoautoscaling create_catalog: branch id or full branch name (default: the project's default branch).
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
page_sizeNoMax items to return (server caps this).
page_tokenNonext_page_token from a previous response.
update_maskNoComma-separated field paths to update. Default: derived from the keys of `spec` (autoscaling resources use 'spec.<field>' paths).
catalog_nameNoUnity Catalog catalog name for *_catalog actions.
show_deletedNoautoscaling list: include soft-deleted projects.
wait_secondsNoSeconds to wait for a long-running create/update/delete to finish. 0 (default) returns immediately with status 'pending'. Capped by DBX_MCP_MAX_WAIT_SECONDS and the tool timeout.
database_nameNocreate_catalog: Postgres database to register.
operation_nameNoAutoscaling operation name returned by a previous call (for action='get_operation').
create_database_if_missingNocreate_catalog: create the Postgres database if it does not exist.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the coarse annotations (destructiveHint=true, readOnlyHint=false), the description adds a per-action safety classification (READ_ONLY/WRITE/DESTRUCTIVE), explains that delete is soft unless purge=true, that force removes point-in-time children, and that long-running work returns 'pending' unless wait_seconds is set. This meaningfully deepens the agent's understanding of destructive and async behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-structured: it front-loads the purpose, then groups details by kind and catalog actions, and ends with a compact safety table. Every sentence contributes to action/kind mapping, async behavior, or safety, and the length is justified by the tool's 11 actions and 18 parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the rich input schema and an output schema, the description need not document every parameter or return value. It covers the essential conceptual model (kinds, actions, safety, async), but it omits any mention of the dry_run/confirm workflow for destructive actions and does not differentiate from sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents all 18 parameters at 100% coverage, so baseline is 3. The description goes further by providing concrete spec examples for both kinds and clarifying how force, purge, and branch interact with specific actions. This adds useful cross-parameter semantics beyond the schema's per-parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource ('Manage Lakebase (Postgres) databases') and then scopes the tool by kind ('provisioned' vs 'autoscaling') and catalog actions. It does not name sibling tools (e.g., manage_lakebase_branch), so an agent must infer scope rather than being told explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It maps each action to a kind and gives concrete examples for create/update specs, plus notes on soft/hard delete and async wait behavior. However, it never states when to choose this tool over sibling tools like manage_lakebase_branch or generate_lakebase_credential, and gives no explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_lakebase_syncManage Lakebase synced tablesA
Destructive

Manage Lakebase synced tables (reverse ETL: Unity Catalog Delta table -> Lakebase Postgres table).

Actions: list (provisioned; instance_name), get, create (table_name + spec with source_table_full_name, primary_key_columns, scheduling_policy SNAPSHOT/TRIGGERED/CONTINUOUS), delete (purge_data=true also drops the Postgres table), trigger (starts the synced table's managed pipeline via pipelines.start_update; not for CONTINUOUS), get_operation (autoscaling). update is not supported by the Databricks API. kind='autoscaling' uses w.postgres synced tables (no list).

Safety classification: list, get, get_operation = READ_ONLY; create, update = WRITE; delete = DESTRUCTIVE; trigger = EXECUTION.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNoprovisioned: Lakebase database instances (w.database). autoscaling: Lakebase autoscaling projects/branches/endpoints (w.postgres).provisioned
specNocreate: synced table spec, e.g. {"source_table_full_name": "main.sales.orders", "primary_key_columns": ["order_id"], "scheduling_policy": "TRIGGERED"} (SNAPSHOT | TRIGGERED | CONTINUOUS; optional new_pipeline_spec / existing_pipeline_id, timeseries_key, create_database_objects_if_missing). Autoscaling specs also take branch and postgres_database. Unknown fields are rejected.
actionYesSynced table operation. trigger starts a sync for TRIGGERED/SNAPSHOT policies.
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
page_sizeNoMax items to return (server caps this).
page_tokenNonext_page_token from a previous response.
purge_dataNoprovisioned delete: also DROP the Postgres table.
table_nameNoFull Unity Catalog name of the synced table: catalog.schema.table.
wait_secondsNoSeconds to wait for a long-running create/update/delete to finish. 0 (default) returns immediately with status 'pending'. Capped by DBX_MCP_MAX_WAIT_SECONDS and the tool timeout.
instance_nameNoprovisioned: database instance (required for list; for create unless the target catalog is a registered database catalog).
operation_nameNoAutoscaling operation name returned by a previous call (for action='get_operation').
logical_database_nameNoprovisioned create: target Postgres database name.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only supply coarse tool-level flags (destructiveHint=true, openWorldHint=true); the description refines them with a per-action safety classification (READ_ONLY/WRITE/DESTRUCTIVE/EXECUTION) and discloses that purge_data drops the Postgres table and that confirm is required after a 'confirmation_required' status. This is real behavioral context beyond the annotations and does not contradict them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the resource and reverse-ETL framing, then breaks actions and safety into scannable segments. Dense but efficient; a few items (scheduling policy names, purge_data) duplicate the schema, costing a point.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description need not describe returns, and it covers the remaining essentials: action prerequisites, the unsupported update, kind branching, destructive scope, and the confirmation/dry-run workflow for a 13-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so the baseline is 3, but the description adds cross-parameter semantics: which params each action needs, that trigger drives pipelines.start_update, and the SNAPSHOT/TRIGGERED/CONTINUOUS scheduling policies. It largely echoes the schema for individual fields, so it stops short of a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific resource (Lakebase synced tables) and clarifies the domain with the parenthetical 'reverse ETL: Unity Catalog Delta table -> Lakebase Postgres table'. The enumerated action list makes it unmistakably distinct from siblings like manage_lakebase_branch and manage_lakebase_database.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives per-action selection criteria: list requires instance_name, create requires table_name + spec, trigger is 'not for CONTINUOUS', update 'is not supported by the Databricks API', and kind='autoscaling' has 'no list'. Explicit when-to-use and when-not guidance for every mode.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_masSupervisor (multi-agent) agentsA
Destructive

Manage Supervisor Agents (Agent Bricks multi-agent orchestrators that route to Genie spaces, Knowledge Assistants, UC functions, UC connections (MCP), apps and volumes).

Actions:

  • list / get / delete; create (spec: display_name, description, instructions); update (spec fields).

  • list_tools / get_tool / delete_tool; add_tool (tool_id + spec: tool_type, description and the matching block, e.g. {'tool_type':'genie_space','genie_space':{'id':'...'},'description':'...'} or {'tool_type':'knowledge_assistant','knowledge_assistant':{'knowledge_assistant_id':'...'}}); update_tool (only description can change).

  • list_examples / get_example / add_example (spec: question, guidelines) / update_example / delete_example.

  • get_permissions / update_permissions (spec: access_control_list). Query a supervisor through its serving endpoint (manage_serving_endpoint action=query).

Safety classification: list, get, list_tools, get_tool, list_examples, get_example = READ_ONLY; create, update, add_tool, update_tool, add_example, update_example = WRITE; delete, delete_tool, delete_example = DESTRUCTIVE; get_permissions = READ_ONLY+SECURITY_SENSITIVE; update_permissions = SECURITY_SENSITIVE+WRITE.

ParametersJSON Schema
NameRequiredDescriptionDefault
specNoRequest body fields for create/update, using the Databricks REST API field names (snake_case). Unknown fields are rejected.
actionYesOperation to perform.
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
tool_idNoTool id (add_tool: the id to assign; others: id or full resource name).
page_sizeNoMax items to return (server caps this).
example_idNoExample id (or full resource name).
page_tokenNonext_page_token from a previous response.
update_maskNoComma-separated fields to update; defaults to the keys present in spec.
supervisor_agent_idNoSupervisor Agent id or resource name 'supervisor-agents/{id}'.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the coarse annotations (readOnlyHint=false, destructiveHint=true, openWorldHint=true) by classifying each individual action as READ_ONLY, WRITE, DESTRUCTIVE, or SECURITY_SENSITIVE, so an agent knows exactly which action triggers confirmation. It also discloses the constraint that update_tool can only change the description, a real behavioral limit not present in annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the resource definition, then grouped into action families with terse bullets and a compact safety line. Slightly long for a description, but the density is justified by 17 actions; little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a composite multi-action tool with an output schema and 100% schema coverage, the description covers action semantics, spec shape, and the confirmation/destructive flow (confirm, dry_run interplay implied). It does not spell out the confirmation workflow in full, but the schema's confirm field carries that.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds concrete spec payload examples (e.g. {'tool_type':'genie_space','genie_space':{'id':'...'}}) that the generic schema field ('Request body fields...') does not provide. It also clarifies that update_mask defaults to spec keys and that unknown fields are rejected.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Starts with a specific verb+resource ('Manage Supervisor Agents') and immediately defines what a Supervisor Agent is (a multi-agent orchestrator routing to Genie spaces, KAs, UC functions/connections, apps, volumes), which sharply distinguishes it from siblings like manage_genie and manage_ka. The per-action breakdown makes the scope unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Enumerates all 17 actions with the spec fields each requires, and routes querying to the sibling 'manage_serving_endpoint action=query', which is genuinely useful cross-tool guidance. It lacks explicit when-not-to-use guidance (e.g. when to pick manage_ka over add_tool), so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_metric_viewsUnity Catalog metric viewsA
Destructive

Manage Unity Catalog metric views (semantic layer) - implemented with documented SQL DDL.

  • create(full_name, yaml_definition): CREATE VIEW ... WITH METRICS LANGUAGE YAML AS $$...$$

  • get(full_name): YAML definition, columns and metadata.

  • list(catalog_name, schema_name): metric views in a schema.

  • update(full_name, yaml_definition): CREATE OR REPLACE - destructive, plan shows old vs new definition.

  • delete(full_name): DROP VIEW - destructive, needs confirm.

  • query(full_name, dimensions, measures, filters?, limit?): SELECT dims, MEASURE(m) ... GROUP BY dims. DDL and queries run on a SQL warehouse (warehouse_id optional).

Safety classification: create = WRITE; get, list = READ_ONLY; update, delete = DESTRUCTIVE+WRITE; query = EXECUTION+READ_ONLY.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoquery: max rows.
actionYescreate / update (CREATE OR REPLACE) / delete a metric view from YAML; get: definition + metadata; list: metric views in a schema; query: SELECT dimensions + MEASURE(measures).
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
filtersNoquery: [{dimension, op, value}] combined with AND; op in =, !=, <>, <, <=, >, >=, LIKE, NOT LIKE, IS NULL, IS NOT NULL. Values are bound as parameters.
measuresNoquery: measure names (wrapped in MEASURE()).
full_nameNoMetric view name catalog.schema.view.
page_sizeNoMax items to return (server caps this).
dimensionsNoquery: dimension names to group by.
page_tokenNonext_page_token from a previous response.
schema_nameNolist: schema.
catalog_nameNolist: catalog.
warehouse_idNoSQL warehouse (default: configured/auto-selected).
yaml_definitionNocreate/update: the metric view YAML (e.g. version, source, dimensions[{name, expr}], measures[{name, expr}], optional filter/joins). Must not contain '$$'.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the aggregate annotations by disambiguating per-action safety: create=WRITE; get/list=READ_ONLY; update/delete=DESTRUCTIVE+WRITE; query=EXECUTION+READ_ONLY. It also discloses that update is CREATE OR REPLACE with a plan showing old vs new, that delete requires confirm, and that DDL/queries execute on a SQL warehouse. This is exactly the added context annotations cannot provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose in one line, then uses a tight per-action bullet list, ending with a compact safety classification line. Every sentence carries signal with no repetition of schema text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 14 params, a full output schema, and annotations present, the description fills all remaining gaps: action semantics, per-action safety, confirmation flow, and execution target. Nothing needed to invoke the tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds value by mapping parameters to actions (e.g., create(full_name, yaml_definition), query(full_name, dimensions, measures, filters?, limit?)) and noting warehouse_id is optional. This clarifies which params apply to which action beyond the per-param schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (managing Unity Catalog metric views / semantic layer) and enumerates every action with its signature, so an agent knows exactly what the tool operates on and which actions exist. It is clearly distinguishable from generic siblings like execute_sql or manage_uc_objects.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Each action is annotated with its intent (create/update/delete from YAML, get returns definition+metadata, list enumerates, query runs SELECT dims + MEASURE). It also states that confirm is required for destructive actions, giving clear per-action context. It stops short of naming when to prefer this over sibling tools like execute_sql, so not a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_pipelineManage pipelinesA
Destructive

Create, inspect, change, clone and delete Lakeflow Spark Declarative Pipelines (DLT).

  • create: spec = pipeline settings (name, catalog, schema, libraries [{notebook:{path}} | {file:{path}} | {glob:{include}}], root_path, serverless, clusters, configuration, continuous, development, channel, edition, photon, notifications, tags, trigger, environment, event_log, run_as, ...).

  • get (pipeline_id), list (name_contains or filter; paginated).

  • update (pipeline_id, spec): the given top-level fields are merged onto the current settings (set a field to null to remove it); uses expected_last_modified to avoid overwriting concurrent edits.

  • delete (pipeline_id[, cascade, force]): DESTRUCTIVE, needs confirm. By default tables are deleted too.

  • clone (pipeline_id, spec: catalog, schema/target, clone_mode='MIGRATE_TO_UC', ...): HMS -> UC copy. Run/monitor updates with manage_pipeline_run. Specs with run_as are SECURITY_SENSITIVE.

Safety classification: depends on input (DESTRUCTIVE, EXECUTION, READ_ONLY, SECURITY_SENSITIVE, WRITE).

ParametersJSON Schema
NameRequiredDescriptionDefault
specNoRequest body fields for create/update, using the Databricks REST API field names (snake_case). Unknown fields are rejected.
forceNodelete: proceed even if resource cleanup fails.
actionYescreate: new pipeline from spec; get: full definition and state; list: pipelines; update: change settings (merged onto the current spec); delete: delete pipeline; clone: copy a Hive-metastore pipeline to Unity Catalog (starts an update on the clone).
filterNolist: raw server filter, e.g. "notebook='/Users/me/nb'" or "name LIKE '%sales%'".
cascadeNodelete: false keeps the pipeline's tables/views (server default true deletes them).
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
page_sizeNoMax items to return (server caps this).
page_tokenNonext_page_token from a previous response.
pipeline_idNoPipeline id (get/update/delete/clone).
name_containsNolist: only pipelines whose name contains this text (server-side LIKE).

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true, readOnlyHint=false, and openWorldHint=true, but the description adds substantial behavioral context beyond them: delete is DESTRUCTIVE and needs confirm; tables are deleted by default unless cascade=false; update merges top-level fields and uses expected_last_modified to avoid concurrent edits; run_as specs are SECURITY_SENSITIVE; and safety classification depends on input.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the overall purpose, then uses concise action bullets for create/get/list/update/delete/clone, followed by routing and safety notes. Every sentence earns its place for a multi-action, 11-parameter tool with complex semantics.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity—six actions, 11 parameters, a nested spec object, and destructive/security-sensitive behaviors—the description covers action semantics, destructive safeguards, update merging, clone migration, and the manage_pipeline_run alternative. With an output schema present, it need not explain return values, so nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema: update merges top-level fields onto current settings and setting a field to null removes it; delete's cascade default is true (deletes tables) while false keeps them; clone is HMS-to-UC with clone_mode='MIGRATE_TO_UC'; and spec uses Databricks REST API snake_case names with unknown fields rejected.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb set and resource: 'Create, inspect, change, clone and delete Lakeflow Spark Declarative Pipelines (DLT).' It then enumerates each action, making clear what the tool does and distinguishing it from manage_pipeline_run, which is explicitly named for run/monitor operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear action-level context (e.g., 'get (pipeline_id)', 'list (name_contains or filter; paginated)', 'update (pipeline_id, spec)') and routes run/monitor work to manage_pipeline_run. It does not explicitly state when not to use this tool versus other pipeline-adjacent siblings, so it falls short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_pipeline_runRun and monitor pipelinesA
Destructive

Run and monitor Spark Declarative Pipeline updates and surface pipeline errors.

  • start (pipeline_id[, full_refresh, refresh_selection, full_refresh_selection, validate_only, parameters, wait, timeout_seconds]): EXECUTION; returns update_id with status 'pending'. Full refreshes are also DESTRUCTIVE (confirm required) because table state is reset.

  • stop (pipeline_id): stops the active update (DESTRUCTIVE, confirm required).

  • get_update / wait (pipeline_id[, update_id] - default latest): state; failed updates include ERROR events.

  • list_updates (pipeline_id): update history, newest first.

  • list_events (pipeline_id[, level, update_id, filter]): event log, newest first.

Safety classification: depends on input (DESTRUCTIVE, EXECUTION, READ_ONLY).

ParametersJSON Schema
NameRequiredDescriptionDefault
waitNostart: poll until the update finishes (bounded).
levelNolist_events: only events of this level.
actionYesstart: start an update; stop: stop the active update; get_update: one update's state (+errors if failed); list_updates: update history; list_events: event log (use level='ERROR' for errors); wait: poll an update until it finishes (bounded).
filterNolist_events: raw filter, e.g. "timestamp > '2025-01-01T00:00:00Z'".
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
page_sizeNoMax items to return (server caps this).
update_idNoget_update/wait (defaults to the latest update); list_events: filter.
page_tokenNonext_page_token from a previous response.
parametersNostart: key/value pipeline parameters.
pipeline_idYesPipeline id.
full_refreshNostart: reset ALL tables before running (DESTRUCTIVE).
validate_onlyNostart: only validate the source code; materialize nothing.
timeout_secondsNoMax seconds to wait (capped by server).
refresh_selectionNostart: tables to refresh (incremental).
full_refresh_selectionNostart: tables to fully refresh (DESTRUCTIVE).

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial context beyond the annotations: full refreshes are flagged DESTRUCTIVE because table state is reset, stop halts an active update, failed updates carry ERROR events, and start returns status 'pending' with an update_id. It even notes the safety class varies by input, which the readOnly/destructive annotations cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded one-line purpose followed by a tight action list and a safety line, with no repeated filler. It is on the longer side for 16 params, but the structure is scannable and each bullet carries distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the description covers the branchy action semantics and destructive cases that the flat schema cannot. The only gap is the absence of explicit routing guidance against sibling pipeline tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so every parameter is already documented, giving the description a baseline of 3. The per-action invocation signatures in the description (e.g., start(...) argument lists) are a mild convenience but do not add syntax or semantics the schema lacks.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific resource (Spark Declarative Pipeline updates) and enumerates exactly the operations it covers via per-action bullets. The distinction from siblings like manage_pipeline (configuration) and manage_job_runs (job runs) is implicit in the resource named, and no sibling covers pipeline run lifecycle.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The bullets give clear context per action and the safety sentence tells the agent when confirmation is required. It stops short of naming an alternative sibling or stating when NOT to use this tool (e.g., vs manage_pipeline for editing), leaving some routing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_serving_endpointModel serving endpointsA
Destructive

Manage and query Databricks Model Serving endpoints.

Actions:

  • list / get: endpoints with state and served entities. Credentials of external-model providers (API keys, secrets, tokens, plaintext env vars) are always stripped.

  • create: spec = create body (config, ai_gateway, tags, route_optimized, budget_policy_id, description, email_notifications, rate_limits, ...); name is a dedicated parameter.

  • update_config: spec = {served_entities, traffic_config, auto_capture_config, served_models}.

  • update_ai_gateway: spec = {guardrails, inference_table_config, rate_limits, usage_tracking_config, fallback_config}.

  • delete: permanently delete (requires confirm).

  • query: invoke the endpoint with request (chat messages / prompt / embeddings input / dataframe).

  • get_build_logs / get_logs: build or server logs for served_model_name (tail, size-capped). create/update_config are long-running: they return status 'pending' unless wait_seconds is set.

Safety classification: list, get, get_build_logs, get_logs = READ_ONLY; create, update_config, update_ai_gateway = WRITE; delete = DESTRUCTIVE; query = EXECUTION.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoServing endpoint name (all actions except list).
specNoRequest body fields for create/update, using the Databricks REST API field names (snake_case). Unknown fields are rejected.
actionYesOperation to perform.
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
requestNoquery: request body. Chat: {'messages': [{'role': 'user', 'content': '...'}], 'max_tokens': 256}; completions: {'prompt': '...'}; embeddings: {'input': ['...']}; custom models: {'dataframe_records': [...]} / {'dataframe_split': {...}} / {'instances': [...]} / {'inputs': ...}. Streaming is not supported.
page_sizeNoMax items to return (server caps this).
page_tokenNonext_page_token from a previous response.
wait_secondsNoOptionally wait up to this many seconds for the operation to finish (capped by the server's max wait). Default: return immediately with status 'pending'.
max_output_charsNoCap on returned query/log text size (default 20000, max 200000).
served_model_nameNoServed model/entity name for get_build_logs/get_logs.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the coarse destructiveHint/readOnlyHint annotations: it assigns a per-action safety classification (READ_ONLY/WRITE/DESTRUCTIVE/EXECUTION), discloses that provider credentials are always stripped on list/get, that delete permanently destroys and needs confirm, that create/update return 'pending' unless wait_seconds is set, and that logs are tailed and size-capped.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded purpose followed by a dense but well-structured bulleted action list where every line carries operational information; the trailing safety classification is compact and useful. Slightly heavy, but nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a nine-action tool with an output schema present, it covers the create/update/delete lifecycle, the dry_run/confirm confirmation flow, asynchronous completion via wait_seconds, pagination-adjacent caps, and per-action risk. An agent has everything needed to invoke the right action correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real value by mapping spec contents to specific actions and clarifying that request is only for query with per-mode shapes. It clarifies role of wait_seconds and max_output_chars beyond raw schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Manage and query Databricks Model Serving endpoints') and then enumerates all nine actions with concrete semantics, so an agent knows exactly what the tool covers. It is clearly distinguishable from the Vector Search sibling manage_vs_endpoint by the 'Model Serving' framing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Each action is given a usage-specific meaning, including which spec keys belong to create vs update_config vs update_ai_gateway, that delete requires confirm, and that wait_seconds controls blocking on long-running calls. There is no explicit when-not or sibling routing (e.g., vs manage_vs_endpoint), but the action-level guidance is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_sql_statementInspect or cancel a SQL statementA
Read-only

Poll a previously submitted SQL statement (status and, once finished, its results) or cancel it. Use after execute_sql returned status 'pending'.

Safety classification: get = READ_ONLY; cancel = EXECUTION.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesget: status and results; cancel: stop it.
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
max_rowsNoMaximum rows to return (capped by DBX_MCP_SQL_MAX_ROWS).
row_formatNo'arrays' (compact, aligned with columns) or 'objects' (one dict per row).arrays
statement_idYesStatement id returned by execute_sql.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description usefully discloses per-action safety ('get = READ_ONLY; cancel = EXECUTION'), which is genuinely more granular than the tool-level annotations. However, this conflicts with readOnlyHint=true and destructiveHint=false, which declare the whole tool non-mutating even though cancel stops a running statement — an agent trusting the annotation could treat a state-changing action as safe. The accurate disclosure softens the penalty, but the mismatch with the declared safety profile is a real defect.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences: the action first, then the workflow trigger, then the safety classification. Every sentence earns its place, nothing is repeated, and the most important routing information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and 100% parameter coverage, return-value explanation is not needed, and the description covers modes, trigger, and safety. The only thin spot is the confirmation/dry_run path implied by those parameters, which the description never touches and the agent must infer from the schema alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so action, confirm, dry_run, max_rows, row_format, and statement_id are all already documented in the schema. The description adds only the outcome of 'get' (status plus results once finished) and confirms 'cancel it', which does not go meaningfully beyond the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource in both modes: 'Poll a previously submitted SQL statement (status and, once finished, its results) or cancel it.' It clearly positions itself relative to execute_sql as the follow-up step, so an agent can place it in the workflow. It does not explicitly distinguish itself from execute_sql_multi or other SQL siblings, which keeps it from a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a precise, actionable trigger: 'Use after execute_sql returned status "pending".' That tells the agent exactly when to reach for this tool. There is no explicit when-not guidance or named alternative, but the trigger is unambiguous enough to route correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_sql_warehouseManage SQL warehousesA
Destructive

Manage SQL warehouses.

Actions: list, get, create (spec e.g. {"name","cluster_size":"2X-Small","max_num_clusters":1, "auto_stop_mins":10,"enable_serverless_compute":true,"warehouse_type":"PRO"}), update (spec holds only fields to change; merged onto the current configuration), start, stop and delete. stop and delete require confirm=true and are refused for production-marked warehouses.

Safety classification: list, get = READ_ONLY; create, update, start = WRITE; stop, delete = DESTRUCTIVE.

ParametersJSON Schema
NameRequiredDescriptionDefault
specNoRequest body fields for create/update, using the Databricks REST API field names (snake_case). Unknown fields are rejected.
waitNoWait (bounded) for the operation to reach a steady state.
actionYes
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
page_sizeNoMax items to return (server caps this).
page_tokenNonext_page_token from a previous response.
warehouse_idNoWarehouse id (all actions except list/create).
timeout_secondsNoMax seconds to wait when wait=true (capped by DBX_MCP_MAX_WAIT_SECONDS).

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only say destructiveHint=true/openWorldHint=true globally; the description adds a per-action safety classification (READ_ONLY/WRITE/DESTRUCTIVE), the confirmation gate and its 'reviewed plan / confirmation_required' precondition, and the production-warehouse refusal rule. That is materially more behavioral context than the annotations carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Information is front-loaded into an action list followed by a safety classification, both quick to scan. The inline JSON spec example adds length but earns its place by showing required field names and casing; still slightly dense for a single paragraph.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return-value detail is unnecessary, and the description covers the action set, destructive prerequisites, and mutation semantics. The one soft spot is that list pagination or error/refusal response shape is left entirely to the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 89%, so the baseline is 3; the description goes beyond it by explaining the spec object's field names and merge semantics for update, and the confirm workflow. It does not add much on wait/timeout/pagination beyond what the schema descriptions state.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource and then enumerates every supported action (list, get, create, update, start, stop, delete), so an agent knows exactly what surface this tool covers. The scope is distinct from generic compute/warehouse siblings, and the create spec example makes the capability concrete.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear per-action conditions: spec is required for create, update's spec holds only changed fields and is merged onto the current config, stop/delete need confirm=true and are refused for production warehouses. It does not explicitly route against sibling tools like manage_warehouse or list_compute, so an alternative-selection note is missing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_uc_connectionsManage Lakehouse Federation connectionsA
Destructive

Manage Unity Catalog Lakehouse Federation connections (Snowflake, PostgreSQL, MySQL, SQL Server, Redshift, BigQuery, Oracle, Teradata, Databricks, ...).

create needs name, connection_type and options; spec may add comment, properties, read_only. update needs the full options map (Databricks replaces it) and spec may set owner/new_name. Credentials in options are sent to Databricks but never returned: responses show only non-secret option keys (host, port, ...). All changes are SECURITY_SENSITIVE; delete is also DESTRUCTIVE.

Safety classification: get, list = READ_ONLY; create, update = SECURITY_SENSITIVE+WRITE; delete = DESTRUCTIVE+SECURITY_SENSITIVE.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoConnection name (required except for list).
specNoRequest body fields for create/update, using the Databricks REST API field names (snake_case). Unknown fields are rejected.
actionYescreate | get | list | update | delete
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
optionsNoConnection options, e.g. {host, port, user, password} (Snowflake also sfWarehouse...). Required for create and update. Secret values are never returned.
page_sizeNoMax items to return (server caps this).
page_tokenNonext_page_token from a previous response.
connection_typeNocreate only: AWS_SECRETS_MANAGER, AZURE_KEY_VAULT, BIGQUERY, CONFLUENCE, DATABRICKS, DYNAMICS365, GA4_RAW_DATA, GITHUB, GLUE, HIVE_METASTORE, HTTP, HUBSPOT, JDBC, META_MARKETING, MYSQL, NETSUITE, ORACLE, OUTLOOK, POSTGRESQL, POWER_BI, REDSHIFT, SALESFORCE, SALESFORCE_DATA_CLOUD, SERVICENOW, SMARTSHEET, SNOWFLAKE, SQLDW, SQLSERVER, TERADATA, TIKTOK_ADS, WORKDAY_RAAS, ZENDESK

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial behavioral detail beyond annotations: create/update requirements, update replacing the full options map, secrets never being returned, and per-action safety classifications (READ_ONLY, SECURITY_SENSITIVE, DESTRUCTIVE). This is exactly the kind of context an agent needs for a multi-action, security-sensitive tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and connection types, then action requirements, then security context. Dense and useful, though the safety classification sentence partially repeats the preceding SECURITY_SENSITIVE/DESTRUCTIVE statement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter, multi-action tool with an output schema, the description covers the critical call-time requirements, secret handling, destructive behavior, and per-action safety. Nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is already 100%, but the description adds cross-parameter semantics: which fields are required per action, how spec relates to options, and that update replaces the options map. This goes beyond the schema's field-level descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: managing Unity Catalog Lakehouse Federation connections, with examples of supported connection types. Distinguishes this tool from the many other manage_uc_* siblings by naming the specific resource domain.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly explains what each action does and the required parameters for create/update, giving strong action-level usage context. It does not explicitly name alternatives or when-not-to-use this tool versus sibling connection-related tools, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_uc_grantsManage Unity Catalog grantsA
Destructive

Show, grant and revoke Unity Catalog privileges on any securable (catalog, schema, table/view, volume, function, external_location, storage_credential, connection, share, metastore, ...).

get returns direct grants, get_effective includes privileges inherited from parents. grant/revoke require principal + privileges and always show the principal's before/after direct privileges in the plan. ALL_PRIVILEGES is rejected unless allow_all_privileges=true; grants to 'account users' are flagged.

Safety classification: get, get_effective = READ_ONLY+SECURITY_SENSITIVE; grant = SECURITY_SENSITIVE+WRITE; revoke = DESTRUCTIVE+SECURITY_SENSITIVE.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesget: direct grants; get_effective: incl. inherited; grant / revoke privileges.
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
full_nameYesFull name of the securable, e.g. 'main.sales.orders' (metastore: the metastore id).
page_sizeNoMax items to return (server caps this).
principalNoUser email, group name or service principal application id. Required for grant/revoke; optional filter for get.
page_tokenNonext_page_token from a previous response.
privilegesNoPrivileges for grant/revoke, e.g. ['SELECT', 'USE_SCHEMA'] (spaces allowed: 'USE CATALOG').
securable_typeYesSecurable type: agent_service, catalog, clean_room, connection, credential, external_location, external_metadata, function, mcp_service, metastore, model, model_provider_service, model_service, pipeline, provider, recipient, schema, share, skill, staging_table, storage_credential, table, volume. Views use 'table'.
allow_all_privilegesNoMust be true to grant ALL_PRIVILEGES.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial context beyond the coarse annotations: per-action safety classification (get/get_effective READ_ONLY, grant WRITE, revoke DESTRUCTIVE), the requirement that grant/revoke echo the principal's before/after direct privileges in a plan, the ALL_PRIVILEGES gating via allow_all_privileges, and the 'account users' flag. This refines the global destructiveHint=true into actionable per-action behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the scope and action semantics, then the behavioral/safety constraints. Dense but every clause carries distinct, non-redundant information (action semantics, gating rules, safety classification) with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is unnecessary, and the description covers the confirmation/dry-run safety flow and privilege gating. It omits any mention of pagination behavior (page_size/page_token) or the inherited-vs-direct distinction for mutating actions, which is a minor gap for a 10-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description reinforces cross-parameter constraints (principal+privileges required for grant/revoke; ALL_PRIVILEGES rejected unless allow_all_privileges) but largely restates what the schema property descriptions already say, adding little new syntax or format detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb set (show/grant/revoke) and resource (Unity Catalog privileges on any securable), and enumerates the securable types so an agent knows the blast radius. It clearly distinguishes its four internal actions (get vs get_effective vs grant vs revoke), which is the key ambiguity for this tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use guidance for the internal actions (get for direct, get_effective for inherited, grant/revoke require principal+privileges) and the confirmation workflow. It does not, however, differentiate this tool from siblings like manage_uc_sharing, manage_uc_security_policies, or manage_uc_tags, leaving some routing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_uc_monitorsData quality monitors (Lakehouse Monitoring)A
Destructive

Manage Unity Catalog data quality monitors (Lakehouse Monitoring) via the Data Quality API.

Actions (full_name identifies the table, or schema with object_type=schema):

  • create(spec): table monitors take DataProfilingConfig fields - output_schema_name (catalog.schema, or output_schema_id), exactly one of snapshot {} | time_series {timestamp_column, granularities: ["AGGREGATION_GRANULARITY_1_DAY", ...]} | inference_log {...}, plus optional schedule {quartz_cron_expression, timezone_id}, slicing_exprs, custom_metrics, baseline_table_name, assets_dir, warehouse_id, notification_settings, skip_builtin_dashboard. Schema monitors take AnomalyDetectionConfig fields (excluded_table_full_names).

  • get, update(spec: only the fields to change), delete (metric tables/dashboard are kept).

  • refresh (starts compute), list_refreshes, get_refresh(refresh_id), cancel_refresh(refresh_id).

  • metrics: profile_metrics_table_name, drift_metrics_table_name, dashboard_id - query them with SQL.

  • query_metrics(metrics_table=profile|drift, sample_rows, warehouse_id?): sample rows of a metric table. Listing all monitors is not available (the SDK marks list_monitor as unimplemented).

Safety classification: create, update, cancel_refresh = WRITE; get, list_refreshes, get_refresh, metrics = READ_ONLY; delete = DESTRUCTIVE+WRITE; refresh = EXECUTION; query_metrics = EXECUTION+READ_ONLY.

ParametersJSON Schema
NameRequiredDescriptionDefault
specNoRequest body fields for create/update, using the Databricks REST API field names (snake_case). Unknown fields are rejected.
actionYescreate/get/update/delete a monitor; refresh: start a metrics refresh; list_refreshes/get_refresh/cancel_refresh; metrics: names of the profile/drift metric tables and dashboard; query_metrics: sample rows from a metric table via a SQL warehouse.
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
full_nameNoMonitored object: table catalog.schema.table (or catalog.schema when object_type=schema).
page_sizeNoMax items to return (server caps this).
page_tokenNonext_page_token from a previous response.
refresh_idNoRefresh id for get_refresh / cancel_refresh.
object_typeNotable: data profiling monitor; schema: anomaly detection monitor.table
sample_rowsNoquery_metrics: rows to return.
warehouse_idNoSQL warehouse for query_metrics.
metrics_tableNoquery_metrics: which metric table to sample.profile

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations only give coarse flags (readOnlyHint=false, destructiveHint=true, openWorldHint=true), while the description supplies a far richer per-action safety classification (create/update/cancel_refresh=WRITE, delete=DESTRUCTIVE+WRITE, refresh=EXECUTION, etc.) and discloses a non-obvious side effect: 'delete (metric tables/dashboard are kept)'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose and action list are front-loaded, with actions and safety classes grouped into scannable bullets. It is dense and long, but nearly every line carries actionable detail, so little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 12-parameter, 10-action tool with an output schema present, the description covers the mutating actions, the refresh lifecycle, where metrics live, and the confirmation/dry-run flow, leaving no obvious gap an agent would need to guess at.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3; the description earns above that by documenting the shape of the free-form spec argument per action (DataProfilingConfig fields, the mutually exclusive snapshot/time_series/inference_log variants, AnomalyDetectionConfig for schema monitors) and clarifying that update takes only the fields to change.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ('Manage Unity Catalog data quality monitors (Lakehouse Monitoring) via the Data Quality API') and then enumerates the exact action surface, so an agent can tell it apart from siblings like manage_metric_views or manage_uc_objects. Scope is further pinned by noting that full_name identifies a table or schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear per-action context (create/get/update/delete, refresh lifecycle, metrics/query_metrics) and even states a when-not: 'Listing all monitors is not available (the SDK marks list_monitor as unimplemented)'. What it lacks is any explicit routing against sibling tools, so it stops short of the 5 bar.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_uc_objectsManage Unity Catalog objectsA
Destructive

Create, inspect, list, update and delete Unity Catalog catalogs, schemas, tables, volumes and functions.

Hierarchy is catalog -> schema -> object: list schemas needs catalog_name; list tables/volumes/functions need catalog_name + schema_name. Identify a target by full_name or by catalog_name/schema_name/name. create/update take spec with Databricks API fields (e.g. catalog: comment, storage_root, properties; volume: volume_type, storage_location, comment; update: comment, owner, new_name, properties). Tables: only EXTERNAL Delta tables can be created via the API (use execute_sql for CREATE TABLE/VIEW); table/function update supports only 'owner'. delete is DESTRUCTIVE; force=true on catalog/schema deletes all contents. Changing owner/isolation_mode is SECURITY_SENSITIVE.

Safety classification: depends on input (DESTRUCTIVE, READ_ONLY, SECURITY_SENSITIVE, WRITE).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoObject name relative to its parent (alternative to full_name).
specNoRequest body fields for create/update, using the Databricks REST API field names (snake_case). Unknown fields are rejected.
forceNodelete only (catalog/schema/function): drop even if not empty - RECURSIVELY deletes all contents.
actionYescreate | get | list | update | delete
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
full_nameNoTarget name: 'catalog', 'catalog.schema' or 'catalog.schema.object' (backticks allowed). For list, may give the parent ('catalog' or 'catalog.schema').
page_sizeNoMax items to return (server caps this).
page_tokenNonext_page_token from a previous response.
object_typeYescatalog | schema | table | volume | function
schema_nameNoParent schema (required to list tables/volumes/functions).
catalog_nameNoParent catalog (required to list schemas/tables/volumes/functions).

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the generic annotations (destructiveHint/openWorldHint/readOnlyHint), the description discloses what gets destroyed (force=true recursively deletes catalog/schema contents), that owner/isolation_mode changes are SECURITY_SENSITIVE, and the confirm-after-confirmation_required workflow. This is meaningful behavioral context the annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but well organized, with the action list front-loaded and the hierarchy/spec/restriction details grouped logically. It is longer than typical, but nearly every sentence carries actionable constraints rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 12-parameter, 5-action, 5-object-type tool, the description covers targeting, hierarchy, spec contents, destructive/sensitive semantics, and the confirmation flow, and an output schema exists to handle return values. Little an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, and the description adds real value by listing concrete spec fields per object type (catalog: comment/storage_root/properties; volume: volume_type/storage_location; update: comment/owner/new_name/properties) that the schema only refers to generically. It also clarifies full_name vs catalog_name/schema_name/name targeting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific set of verbs (create, inspect, list, update, delete) against a specific resource family (Unity Catalog catalogs, schemas, tables, volumes, functions). This clearly separates it from siblings like manage_uc_grants, manage_uc_tags, or manage_uc_storage, which target different UC sub-resources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit usage conditions: hierarchy requirements for listing (catalog_name for schemas; catalog_name + schema_name for tables/volumes/functions) and the when-not case that only EXTERNAL Delta tables can be created via the API, routing table/view creation to execute_sql instead. It also flags table/function update as owner-only.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_uc_security_policiesUnity Catalog row filters, column masks & ABAC policiesA
Destructive

Manage Unity Catalog fine-grained access control.

Actions:

  • get(table_name): current row filter + column masks (from table metadata) and ABAC policies in effect.

  • set_row_filter(table_name, function_name, using_columns) / drop_row_filter(table_name)

  • set_column_mask(table_name, column_name, function_name, using_columns?) / drop_column_mask(table_name, column_name) These run ALTER TABLE DDL on a SQL warehouse (warehouse_id optional).

  • list_policies(securable_type, securable_fullname, include_inherited?) / get_policy(+policy_name)

  • create_policy(securable_type, securable_fullname, policy_name, spec) - spec uses PolicyInfo fields: to_principals, for_securable_type, policy_type (POLICY_TYPE_ROW_FILTER|POLICY_TYPE_COLUMN_MASK), row_filter {function_name, using}, column_mask {function_name, on_column, using}, match_columns, when_condition, except_principals, comment.

  • update_policy(..., policy_name, spec, update_mask?) / delete_policy(..., policy_name) All changes are security-sensitive: call without confirm to get a plan showing current vs new state, then repeat with confirm=true. Change responses include an audit block (who/what/when).

Safety classification: get, list_policies, get_policy = READ_ONLY+SECURITY_SENSITIVE; set_row_filter, set_column_mask, create_policy, update_policy = SECURITY_SENSITIVE+WRITE; drop_row_filter, drop_column_mask, delete_policy = DESTRUCTIVE+SECURITY_SENSITIVE+WRITE.

ParametersJSON Schema
NameRequiredDescriptionDefault
specNoRequest body fields for create/update, using the Databricks REST API field names (snake_case). Unknown fields are rejected.
actionYesget: current row filter, column masks and ABAC policies on table_name; list_policies / get_policy: ABAC policies on a securable; set_row_filter / drop_row_filter / set_column_mask / drop_column_mask: table-bound UDF filters/masks (SQL DDL on a warehouse); create_policy / update_policy / delete_policy: ABAC policies.
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
page_sizeNoMax items to return (server caps this).
page_tokenNonext_page_token from a previous response.
table_nameNoTable full name catalog.schema.table.
column_nameNoColumn for set_column_mask / drop_column_mask.
policy_nameNoABAC policy name (get/update/delete/create).
update_maskNoupdate_policy: comma-separated fields to update (default: the keys present in spec).
warehouse_idNoSQL warehouse for filter/mask DDL (default: configured/auto-selected).
function_nameNoFully qualified SQL UDF catalog.schema.function used as row filter or column mask.
using_columnsNoset_row_filter: table columns passed to the filter UDF, in order ([] for none). set_column_mask: additional columns passed after the masked column (USING COLUMNS).
securable_typeNoABAC policies: type of the securable the policy is defined on.
include_inheritedNolist_policies/get: include policies inherited from parent schema/catalog (get defaults to true).
securable_fullnameNoABAC policies: full name of that catalog / schema / table.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only give a coarse readOnly=false/destructive=true profile for a tool that mixes read-only and destructive actions; the description repairs that gap with an explicit per-action safety classification (READ_ONLY+SECURITY_SENSITIVE vs WRITE vs DESTRUCTIVE). It also discloses the two-step confirm protocol (call without confirm to get a current-vs-new plan, repeat with confirm=true), that filter/mask changes run ALTER TABLE DDL on a SQL warehouse, and that change responses carry an audit block.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Long but organized as a bulleted action list with the confirmation and safety rules front-loaded, so an agent can scan it. A few lines (e.g., the action enum explanation) restate what the schema already encodes, but overall density is high with little waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 16-parameter, 10-action, security-sensitive tool, the description covers what each action needs, the required confirm/dry_run semantics, DDL execution requirements, and the per-action safety tier. Since an output schema exists, return-value detail is correctly omitted, leaving no meaningful gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% (baseline 3), but the description adds real meaning the schema cannot: the `spec` parameter is an untyped object in the schema, and the description enumerates its PolicyInfo fields (to_principals, policy_type enum values, row_filter/column_mask sub-shapes, when_condition, except_principals). It also clarifies using_columns differs between set_row_filter and set_column_mask and that warehouse_id is optional for DDL actions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific scope statement ('Manage Unity Catalog fine-grained access control') and then enumerates all ten actions with their exact arguments, so an agent knows precisely what each verb does. It is cleanly distinguishable from siblings like manage_uc_grants, manage_uc_tags, and manage_uc_objects by its row-filter/mask/ABAC focus.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives per-action context ('get: current row filter, column masks and ABAC policies on table_name', the DDL-backed set/drop actions, the ABAC CRUD actions) and states the confirmation workflow clearly. It does not explicitly route the agent away from neighboring UC tools (grants/tags/objects), so the exclusion guidance is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_uc_sharingDelta Sharing: shares, recipients, providersA
Destructive

Manage Delta Sharing shares, recipients and providers.

  • share: list | get(name) | create(name, spec{comment, storage_root}) | update(name, spec{comment, new_name, owner, storage_root, updates}) | delete | add_objects / remove_objects(name, objects) | get_permissions(name) | update_permissions(name, changes=[{principal, add, remove}]).

  • recipient: list | get | create(name, spec{authentication_type: TOKEN|DATABRICKS|OIDC_FEDERATION|..., data_recipient_global_metastore_id, comment, ip_access_list, expiration_time, owner, properties_kvpairs}) | update | delete | get_permissions (shares it can read) | rotate_token(existing_token_expire_in_seconds).

  • provider: list | get | create(name, spec{authentication_type, recipient_profile_str, comment}) | update | delete | list_shares. All changes are security-sensitive and need confirm=true after reviewing the plan (adding objects or granting recipients is external data exposure). Activation links, tokens and provider credentials are never returned.

Safety classification: depends on input (DESTRUCTIVE, READ_ONLY, SECURITY_SENSITIVE, WRITE).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoShare / recipient / provider name.
specNoRequest body fields for create/update, using the Databricks REST API field names (snake_case). Unknown fields are rejected.
actionYesAll: list, get, create, update, delete. share: add_objects, remove_objects, get_permissions (recipients with access), update_permissions. recipient: get_permissions (shares it can access), rotate_token. provider: list_shares.
changesNoupdate_permissions: [{principal: <recipient>, add: ['SELECT'], remove: [...]}].
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
objectsNoadd_objects/remove_objects: data objects, e.g. {name: 'cat.sch.tbl', data_object_type: 'TABLE', shared_as?, cdf_enabled?, history_data_sharing_status?, partitions?, comment?}; remove_objects also accepts plain names.
resourceYesDelta Sharing object type.
page_sizeNoMax items to return (server caps this).
page_tokenNonext_page_token from a previous response.
include_shared_dataNoshare get: include the shared objects.
existing_token_expire_in_secondsNorotate_token: seconds until the current token expires (0 = immediately).

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and openWorldHint=true, yet the description goes further: it explains the two-step confirm workflow keyed off a 'confirmation_required' status, warns that adding objects/granting recipients is external data exposure, and states that activation links, tokens, and provider credentials are never returned. This is substantial behavioral context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening sentence is front-loaded and the per-resource bullets are scannable, but the action enumeration partially duplicates the action enum description in the schema, adding length without new information. Still well-organized and each remaining sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, and the description covers the mutation/confirmation flow, the redaction policy, and the resource-action matrix. An agent has everything needed to select an action and invoke it safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter including confirm, dry_run, objects, and changes is already documented in the schema. The description's field listings (spec fields, permission changes) largely restate what the enum and property descriptions provide, so it adds only marginal meaning. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (manage) and resource set (Delta Sharing shares, recipients, providers), then enumerates the exact sub-actions per resource type. This clearly separates it from adjacent siblings like manage_uc_connections, manage_uc_grants, or manage_uc_objects.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The per-resource action breakdown effectively routes the agent to the correct action, and it flags that modifications are security-sensitive requiring confirm=true after reviewing the plan. It stops short of naming alternative tools or explicit when-not-to-use conditions, but for a multiplexed tool the action mapping is unusually clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_uc_storageManage UC storage credentials & external locationsA
Destructive

Manage Unity Catalog storage credentials and external locations.

create/update use spec with Databricks API fields - storage_credential: aws_iam_role {role_arn}, azure_managed_identity {access_connector_id}, databricks_gcp_service_account {}, comment, read_only, skip_validation (update also owner, new_name, isolation_mode); external_location: url, credential_name, comment, read_only, skip_validation (update also owner, new_name, isolation_mode). validate tests cloud access (storage_credential: with url or spec.external_location_name; external_location: its own url). All changes are SECURITY_SENSITIVE, delete is also DESTRUCTIVE. Secret fields are never returned.

Safety classification: get, list, validate = READ_ONLY; create, update = SECURITY_SENSITIVE+WRITE; delete = DESTRUCTIVE+SECURITY_SENSITIVE.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNovalidate (storage_credential): cloud URL to test access against.
nameNoName of the storage credential / external location.
specNoRequest body fields for create/update, using the Databricks REST API field names (snake_case). Unknown fields are rejected.
forceNodelete/update: proceed even if dependent objects exist.
actionYescreate | get | list | update | delete | validate
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
resourceYesstorage_credential | external_location
page_sizeNoMax items to return (server caps this).
page_tokenNonext_page_token from a previous response.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=true, and openWorldHint=true, so the safety profile is partially covered. The description goes beyond by stating 'All changes are SECURITY_SENSITIVE, delete is also DESTRUCTIVE' and provides a safety classification table mapping each action to its risk level. It also notes 'Secret fields are never returned', which is critical behavioral context not captured by annotations. This is rich, non-redundant disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose and then organized into action-specific blocks. It's dense but efficient, with no wasted sentences. The safety classification at the end is a helpful summary. Minor deduction for being quite long, but every sentence serves a purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (10 params, 6 actions, security implications) and the presence of an output schema, the description covers all necessary context. It explains the spec fields, validation behavior, safety classifications, and secret field handling. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 10 parameters. The description adds spec field details (aws_iam_role {role_arn}, etc.) which is useful but largely mirrors what the schema's spec description implies. The description doesn't add syntax or format details beyond what the schema provides, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it manages 'Unity Catalog storage credentials and external locations', naming both sub-resources explicitly. The six actions (create, get, list, update, delete, validate) are enumerated, so an agent can distinguish this from siblings like manage_uc_connections or manage_uc_objects. It's a specific verb+resource statement that leaves no ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides field-level guidance for each action's spec (create/update, validate), which is clear contextual guidance. However, it doesn't explicitly state when to use this tool versus alternatives like generate_lakebase_credential or manage_uc_connections. The action semantics are implied but there's no explicit when/when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_uc_tagsManage Unity Catalog tags & commentsA
Destructive

Read, add, update and remove Unity Catalog tags (business metadata, PII classification, ...) and set comments on catalogs, schemas, tables/views, columns and volumes.

Tags use the Entity Tag Assignments API (governed tags may need ASSIGN permission on the tag policy). remove is DESTRUCTIVE and needs confirm. Table/column comments run one safely-quoted DDL statement on a SQL warehouse (warehouse_id optional); catalog/schema/volume comments use the API.

Safety classification: depends on input (DESTRUCTIVE, EXECUTION, READ_ONLY, WRITE).

ParametersJSON Schema
NameRequiredDescriptionDefault
tagsNoadd/update: {tag_key: tag_value}; null value = key-only tag. E.g. {'pii': 'email', 'owner_team': 'sales'}.
actionYesget: tags + comment; add/update: set tag values; remove: delete tag keys; set_comment.
commentNoset_comment: the new comment ('' clears it).
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
tag_keysNoremove: tag keys to remove.
full_nameYes'catalog', 'catalog.schema', 'catalog.schema.table|volume' or 'catalog.schema.table.column' (backticks allowed).
page_sizeNoMax items to return (server caps this).
page_tokenNonext_page_token from a previous response.
entity_typeYescatalog | schema | table (incl. views) | column | volume
warehouse_idNoSQL warehouse for table/column comments (auto-selected if omitted).

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the blanket destructiveHint=true annotation, the description explains that destructiveness is input-dependent, that remove deletes tag keys and needs confirm, that governed tags may require ASSIGN permission on the tag policy, and that comment writes execute a quoted DDL statement on a warehouse. This is exactly the kind of auth/mutation detail annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and scope are front-loaded in the first sentence, followed by the API/permission and DDL notes. The closing 'Safety classification: depends on input' line largely restates what the annotations and prior sentences already imply, a small redundancy, but overall the density is justified for an 11-parameter multi-action tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need no explanation, and 100% schema coverage handles parameter documentation. The description fills the remaining gaps: multi-path execution (API vs DDL), permission requirements, and the destructive-action confirmation loop, so an agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, and the description adds real value on top: it clarifies that warehouse_id is optional and only relevant to the table/column comment path, and that the API vs DDL route depends on entity type. It does not explain the confirm/dry_run interplay beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb set (read, add, update, remove, set comments) applied to a specific resource (Unity Catalog tags and comments) and enumerates the entity types covered (catalogs, schemas, tables/views, columns, volumes). No sibling tool in the list handles UC tags or comments, so the domain alone makes it unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives conditional context for the in-tool paths: remove is destructive and requires confirm, governed tags may need ASSIGN permission, and table/column comments go through a SQL warehouse while catalog/schema/volume comments use the API. It never names an alternative sibling (e.g. execute_sql for comment DDL) or states when to prefer this over them, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_volume_filesManage volume filesA
Destructive

Unity Catalog Volume file operations: list, get_metadata, upload (inline content / content_base64, or local_path; overwrite replaces an existing file and is DESTRUCTIVE), download (returned inline up to DBX_MCP_MAX_INLINE_DOWNLOAD_BYTES - as text when UTF-8, else base64 - or saved to local_path), delete (file), delete_directory (empty dirs; recursive deletes contents after confirmation), create_directory. Paths are validated against traversal and the configured volume allowlist.

Safety classification: depends on input (DESTRUCTIVE, READ_ONLY, WRITE).

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesVolume path: /Volumes/<catalog>/<schema>/<volume>/...
actionYeslist: directory entries; get_metadata: file (or directory) metadata; upload: write a file; download: read a file; delete: delete a file; delete_directory: delete a directory (recursive=true deletes its contents too); create_directory: create a directory (and parents).
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
contentNoupload: UTF-8 text content.
dry_runNoIf true, validate and return the planned change without executing it.
overwriteNoupload: replace an existing file (DESTRUCTIVE; requires confirm).
page_sizeNoMax items to return (server caps this).
recursiveNodelete_directory: also delete all files and sub-directories inside it.
local_pathNoupload: source file / download: destination file, relative to DBX_MCP_LOCAL_FILE_ROOT (local file access is disabled unless that is set).
page_tokenNonext_page_token from a previous response.
content_base64Noupload: binary content, base64-encoded.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite annotations already flagging destructiveHint=true, the description adds substantial context the annotations cannot: the confirm/'confirmation_required' workflow, path validation against traversal and a volume allowlist, the DESTRUCTIVE nature of overwrite, the inline download byte cap with UTF-8-vs-base64 behavior, and that local file access is disabled unless an env var is set.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the resource and scope, then walks the action list and closes with the safety classification. It is dense but every clause carries distinct operational information; nothing is filler for an 11-parameter, 7-action tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values needn't be described, and annotations plus description together cover the safety profile, the confirm flow, per-action semantics, and path validation. Nothing an agent needs to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all 11 parameters and the baseline is 3. The description still adds connective meaning the schema lacks: it groups upload sources (content / content_base64 / local_path), ties overwrite to DESTRUCTIVE+confirm, and explains download's local_path vs inline behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ('Unity Catalog Volume file operations') and enumerates every supported action, so an agent knows exactly what domain this covers. However, it never explicitly distinguishes itself from nearby siblings like get_volume_folder_details or manage_workspace_files, leaving that routing to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Each action carries a parenthetical that clarifies its scope (delete_directory for empty dirs, recursive deletes contents after confirmation, create_directory creates parents), which effectively tells the agent when to pick each. What's missing is any exclusion or alternative-tool guidance versus the file-related siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_vs_dataVector index dataA
Destructive

Read and write the data inside a Vector Search index.

Actions:

  • scan: page through stored rows (last_primary_key to continue).

  • upsert: insert/overwrite rows in a Direct Vector Access index (records or inputs_json).

  • delete: delete rows by primary key from a Direct Vector Access index (requires confirm).

  • sync: trigger a refresh of a Delta Sync index from its source table. Delta Sync indexes cannot be written directly: modify the source table and sync instead.

Safety classification: scan = READ_ONLY; upsert = WRITE; delete = DESTRUCTIVE; sync = EXECUTION+WRITE.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesOperation to perform.
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
recordsNoupsert: rows to write, each including the primary key and the vector column(s).
index_nameYesFull index name catalog.schema.index.
inputs_jsonNoupsert: alternative to records - a JSON array (string or already-parsed array).
num_resultsNoscan: rows to return (default 10, server-capped).
primary_keysNodelete: primary key values of rows to delete.
include_vectorsNoscan: return full embedding vectors (default: summarized as <vector dim=N>).
last_primary_keyNoscan: continue after this primary key (from a previous scan).

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the coarse annotations (readOnlyHint=false, destructiveHint=true) by classifying each action separately: scan=READ_ONLY, upsert=WRITE, delete=DESTRUCTIVE, sync=EXECUTION+WRITE. It also flags that delete requires confirm and that Delta Sync indexes reject direct writes, which the annotations cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded one-line purpose followed by a tight bulleted action list, a single constraint sentence, and a safety table. Every line carries distinct information with no padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter, multi-action tool it covers action semantics, per-action safety, the confirm requirement, and the Delta Sync write restriction; an output schema exists, so return-value detail is unnecessary. An agent has everything needed to choose an action and call it safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents every parameter (records, inputs_json, num_results, primary_keys, last_primary_key, confirm, etc.). The description restates the action-to-parameter mapping ('upsert: ... records or inputs_json', 'last_primary_key to continue') but adds little beyond what the per-parameter descriptions already say, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('read and write the data') and resource ('inside a Vector Search index'), then enumerates four concrete actions with their effects. This clearly separates it from siblings like query_vs_index and manage_vs_index.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Each action is scoped (scan for paging rows, upsert for Direct Vector Access writes, sync to refresh Delta Sync from source), and it gives an explicit when-not: 'Delta Sync indexes cannot be written directly: modify the source table and sync instead.' It stops short of naming a sibling tool as the alternative for pure queries.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_vs_endpointVector Search endpointsA
Destructive

Manage Vector Search endpoints (the compute that hosts vector indexes).

Actions:

  • list / get: endpoint state, type, number of indexes, tags.

  • create: name + endpoint_type; optional spec {budget_policy_id, target_qps, usage_policy_id}. Provisioning is long-running: returns status 'pending' unless wait_seconds is set.

  • update: spec with any of target_qps, budget_policy_id, custom_tags ({key: value} - replaces all tags).

  • delete: permanently delete the endpoint (requires confirm).

Safety classification: list, get = READ_ONLY; create, update = WRITE; delete = DESTRUCTIVE.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoEndpoint name (all actions except list).
specNoRequest body fields for create/update, using the Databricks REST API field names (snake_case). Unknown fields are rejected.
actionYesOperation to perform.
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
page_sizeNoMax items to return (server caps this).
page_tokenNonext_page_token from a previous response.
wait_secondsNoOptionally wait up to this many seconds for the endpoint to come ONLINE (capped by the server's max wait). Default: return immediately with status 'pending'.
endpoint_typeNocreate: endpoint type.STANDARD

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations by classifying per-action safety (list/get = READ_ONLY, create/update = WRITE, delete = DESTRUCTIVE), disclosing that provisioning is long-running and returns 'pending', that delete requires confirm, and that update's custom_tags REPLACES all tags. This is exactly the extra context annotations alone can't convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Bulleted by action with a compact spec summary, then a one-line safety classification. Every line carries information and the scan order (actions first, safety last) is sensible.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be described, and the description covers the remaining gaps an agent needs: long-running provisioning, confirm gating, dry_run semantics, and per-action risk. Nothing material is missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: which spec fields apply to create vs update, that custom_tags overwrites existing tags, the confirm workflow tied to 'confirmation_required', and the wait_seconds default behavior. Only the pagination params (page_size/page_token) are left to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource and parenthetically explains what an endpoint is ('the compute that hosts vector indexes'). The action enumeration (list/get/create/update/delete) makes it immediately distinguishable from siblings like manage_vs_index or query_vs_index.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Each action is annotated with required inputs and behavior, and it explains when to set wait_seconds (only if you want to block for ONLINE) versus taking the immediate 'pending' return. It does not explicitly route the agent when to pick this over manage_vs_index/manage_vs_data, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_vs_indexVector Search indexesA
Destructive

Manage Vector Search indexes.

Actions:

  • list (endpoint_name) / get (index_name): type, primary key, status and readiness.

  • create: index_name + endpoint_name + spec {primary_key, index_type: DELTA_SYNC|DIRECT_ACCESS, index_subtype?, delta_sync_index_spec: {source_table, pipeline_type: TRIGGERED|CONTINUOUS, embedding_source_columns: [{name, embedding_model_endpoint_name}] or embedding_vector_columns, columns_to_sync?} | direct_access_index_spec: {embedding_vector_columns: [{name, embedding_dimension}], schema_json}}.

  • sync: trigger a Delta Sync index refresh. delete: delete the index (requires confirm).

  • update: not supported by the API (recreate the index instead).

Safety classification: list, get = READ_ONLY; create, update = WRITE; delete = DESTRUCTIVE; sync = EXECUTION+WRITE.

ParametersJSON Schema
NameRequiredDescriptionDefault
specNoRequest body fields for create/update, using the Databricks REST API field names (snake_case). Unknown fields are rejected.
actionYesOperation to perform.
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
page_sizeNoMax items to return (server caps this).
index_nameNoFull index name catalog.schema.index (all actions except list).
page_tokenNonext_page_token from a previous response.
endpoint_nameNoVector Search endpoint (required for list and create).

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only give coarse tool-level hints (readOnlyHint=false, destructiveHint=true), but the description adds a per-action safety classification (list/get=READ_ONLY, create/update=WRITE, delete=DESTRUCTIVE, sync=EXECUTION+WRITE) and notes delete requires confirm. This is meaningful context beyond what the structured fields carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the action list and uses bullets, so it scans well. The create bullet is a dense run-on with nested braces, but nearly every element adds actionable detail and little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be documented. The description covers all eight parameters' roles, action requirements, and safety semantics across the tool's complexity. Minor gaps remain around pagination/return behavior, but these are handled by the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds real value by sketching the create spec shape (primary_key, index_type enum values, delta_sync vs direct_access specs, embedding columns) and by clarifying which parameters each action requires. It exceeds the schema without duplicating it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (manage) and resource (Vector Search indexes) and then enumerates every supported action (list, get, create, sync, delete), clearly separating this tool from siblings like manage_vs_endpoint, manage_vs_data, and query_vs_index. An agent can identify scope without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives per-action guidance: which parameters each action needs (list needs endpoint_name, create needs index_name+endpoint_name+spec), and explicitly notes that update is not supported and to recreate instead. It stops short of routing to sibling tools such as query_vs_index for queries, so it's clear but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_warehouseWarehouse status & selectionA
Read-only

Inspect SQL warehouses and the server's warehouse-selection logic. Selection is transparent and configurable: an explicit warehouse_id wins, then DBX_MCP_DEFAULT_WAREHOUSE_ID, then (with DBX_MCP_WAREHOUSE_SELECTION=prefer_running) the best visible warehouse ranked running > starting

stopped, then serverless > pro > classic, then name.

Safety classification: READ_ONLY.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionNolist: warehouses ranked for SQL execution; status: state/health of one warehouse; select: which warehouse SQL tools will use and why.select
warehouse_idNoWarehouse id for status (or to validate in select).
require_runningNoFor select: only accept a RUNNING warehouse.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, matching the description's 'READ_ONLY' claim, so no contradiction. The description adds useful context via the transparent selection precedence rules, which go beyond annotations. But it doesn't explain what 'list' or 'status' returns or any rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose and then explains the selection logic concisely. Every sentence adds value, though the safety classification sentence slightly duplicates annotation information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With annotations covering safety and an output schema present, the description needn't explain return values. It adequately covers the tool's scope and the critical selection precedence, though it could better differentiate from the sibling 'manage_sql_warehouse' for full completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters including the enum for 'action' and the meaning of 'warehouse_id' and 'require_running'. The description reinforces the selection precedence but adds no syntax or format details beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb-and-resource pairing ('Inspect SQL warehouses and the server's warehouse-selection logic'), which is clear. However, it does not distinguish itself from the closely named sibling 'manage_sql_warehouse', which risks agent confusion about which tool handles warehouse operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage through the selection logic ('an explicit warehouse_id wins, then DBX_MCP_DEFAULT_WAREHOUSE_ID...'), which helps an agent understand when selection applies. But it never explicitly states when to use this tool versus 'manage_sql_warehouse' or how 'select' relates to SQL execution tools like execute_sql.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_workspaceWorkspace contextA

Identify or change the Databricks workspace this server talks to: workspace URL, workspace id, active profile and auth type (never tokens), and available config profiles.

Safety classification: info, list_profiles = READ_ONLY; switch_profile = WRITE.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionNoinfo: current workspace/auth context; list_profiles: profiles in ~/.databrickscfg (names and hosts only); switch_profile: reconnect using another profile.info
dry_runNoIf true, validate and return the planned change without executing it.
profileNoProfile name for switch_profile (use 'env' to revert to environment-variable configuration).

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only give a blanket readOnlyHint=false, which would mislead for the two read actions; the description corrects this with a per-action safety classification (info/list_profiles = READ_ONLY, switch_profile = WRITE). It also adds that tokens are never returned. Good added context beyond annotations, though error/rate-limit behavior is unstated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two front-loaded sentences with no filler: the scope comes first, the safety classification second. It is tight and information-dense, though the safety line reads a bit like a field tag rather than prose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values needn't be described, and annotations plus the per-action safety line cover the behavioral essentials for a small 3-param tool. Only minor gaps (side effects of switch_profile, persistence of the new profile) remain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the enum values and dry_run/profile semantics are already fully documented in the schema. The description names the context fields but adds no syntax or format detail beyond what the schema provides, so the baseline of 3 holds.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb pair (identify or change) plus the exact resource (the Databricks workspace this server talks to) and enumerates the context fields. It is clearly differentiated from compute/cluster/warehouse siblings, none of which touch connection context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Enumerates the three actions and their intent (info = current context, list_profiles = config profiles, switch_profile = reconnect), and the safety classification steers an agent toward the safe reads. It does not, however, explicitly state when to prefer this over siblings or any preconditions/exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_workspace_filesManage workspace files and notebooksA
Destructive

Manage Databricks workspace files, notebooks and folders (Workspace API).

  • list (path[, recursive]), get_status (path): metadata (type, language, size, object_id).

  • export (path[, format, local_path]): text content inline (UTF-8) or base64 for binary; capped by DBX_MCP_MAX_INLINE_DOWNLOAD_BYTES; with local_path the file is written under DBX_MCP_LOCAL_FILE_ROOT.

  • import (path, content | content_base64 | local_path[, language, format, overwrite]): create or update a file or notebook (10 MB limit). Notebooks: language=PYTHON|SQL|SCALA|R with format SOURCE (default when language is set) or JUPYTER (.ipynb content). overwrite=true is DESTRUCTIVE (confirm required).

  • mkdirs (path): create directory and parents.

  • delete (path[, recursive]): DESTRUCTIVE, confirm required.

Safety classification: depends on input (DESTRUCTIVE, READ_ONLY, WRITE).

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesAbsolute workspace path, e.g. /Users/me@x.com/project/nb.
actionYeslist: directory contents; get_status: object metadata; export: download content (inline or to local_path); import: upload/create/update a file or notebook; mkdirs: create directories; delete: delete object (recursive for non-empty directories).
formatNoexport/import format. Export default: SOURCE for notebooks, AUTO otherwise. Import default: SOURCE when language is set (notebook), else AUTO (file, or notebook if the content has a notebook header). RAW is import-only.
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
contentNoimport: text content (UTF-8).
dry_runNoIf true, validate and return the planned change without executing it.
languageNoimport: notebook language (required for a single SOURCE notebook).
overwriteNoimport: replace an existing object (DESTRUCTIVE); export: replace a local file.
page_sizeNoMax items to return (server caps this).
recursiveNolist: walk sub-directories (files only); delete: delete non-empty directory.
local_pathNoimport: read from / export: write to this path relative to DBX_MCP_LOCAL_FILE_ROOT.
page_tokenNonext_page_token from a previous response.
content_base64Noimport: binary content, base64-encoded.
create_parentsNoimport: create missing parent directories.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the blanket destructiveHint=true annotation, the description pinpoints exactly which operations are destructive (import with overwrite, delete), states that overwrite/delete require confirmation, and discloses operational limits (DBX_MCP_MAX_INLINE_DOWNLOAD_BYTES, 10 MB import cap, DBX_MCP_LOCAL_FILE_ROOT scoping, inline UTF-8 vs base64 export). This is substantive behavioral context the annotations cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One framing sentence followed by tight per-action bullets, front-loading the overall purpose and then the safety classification. For a six-action, 14-parameter tool the length is proportionate and every line carries action-specific information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given high complexity, an output schema (which removes the need to describe return values), and a confirmation workflow, the description covers all actions, destructive-warning semantics, defaults, and resource scoping. An agent has everything needed to select an action and invoke it safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema coverage the baseline is 3, but the description adds genuine meaning by grouping parameters per action (which flags apply to export vs import vs list) and explaining the language/format interaction and defaults (SOURCE for single notebook, JUPYTER for .ipynb). It layers action-scoped semantics on top of the per-parameter schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource (Databricks workspace files, notebooks, folders via the Workspace API) and enumerates each of the six actions with its distinct effect. An agent can distinguish workspace-file operations from the sibling manage_volume_files/manage_workspace without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Each action line pairs the verb with its parameters and scope (e.g., 'list (path[, recursive])', 'delete (path[, recursive]): DESTRUCTIVE'), giving clear per-action context and explicitly flagging when overwrite/delete require confirmation. It stops short of naming alternatives or when-not-to-use conditions versus sibling tools, so it does not reach a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

query_vs_indexQuery a vector indexA
Read-only

Run a similarity / hybrid / full-text search against a Vector Search index.

Returns the matching records as a list of {column: value} objects, their scores (the 'score' column), the column list, facets (if requested) and query information. Pass the returned next_page_token as page_token to continue.

Safety classification: EXECUTION+READ_ONLY.

ParametersJSON Schema
NameRequiredDescriptionDefault
columnsNoColumns to return (required unless page_token is given).
filtersNoFilter object (sent as filters_json), e.g. {'category': 'news', 'id >': 5, 'tag': ['a', 'b']}.
optionsNoExtra query_index fields: score_threshold, query_columns, sort_columns, facets, columns_to_rerank, reranker. Unknown fields are rejected.
index_nameYesFull index name catalog.schema.index.
page_tokenNonext_page_token from a previous query_vs_index response to fetch the next page.
query_textNoText query (indexes with a model-computed embedding, or HYBRID/FULL_TEXT).
query_typeNoSearch type (default ANN).
num_resultsNoResults to return (default 10, server-capped).
query_vectorNoQuery embedding (Direct Access or self-managed-embedding indexes).
endpoint_nameNoEndpoint name (optional, used with page_token).

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNo
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint, destructiveHint=false, openWorldHint), and the description still adds value by disclosing the return shape (list of {column: value} objects plus scores, column list, facets) and the pagination continuation contract via next_page_token. The explicit 'EXECUTION+READ_ONLY' classification reinforces rather than merely restates the annotations. It stops short of covering permissions, quotas, or result-size limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded in one sentence, followed by return format and pagination guidance. Dense and mostly waste-free, though the separate 'Safety classification' line is somewhat redundant against the provided annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema carrying return values and 100% schema description coverage, the description only needs to add wrapping context — and it does, covering search modes, output shape, and pagination. It remains thin on when this tool is preferable to sibling read/query tools, but nothing required to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter (including enums, defaults, and the filters/options shapes) is already documented in the schema. The description adds essentially no parameter syntax or semantics beyond it, which is the expected baseline when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource — running similarity, hybrid, or full-text search against a Vector Search index — and the three named modes match the query_type enum. Its read/query function is cleanly separable from the sibling management tools (manage_vs_index, manage_vs_data, manage_vs_endpoint).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: the description notes query_text applies to model-computed-embedding or HYBRID/FULL_TEXT indexes and the schema notes query_vector applies to Direct Access/self-managed indexes, which helps mode selection. However, there is no explicit when-to-use/when-not statement or routing to an alternative tool (e.g., execute_sql) for ad-hoc retrieval.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 45 tool updatesv0.1.0
    • First observedask_genie
    • First observeddelete_tracked_resource
    • First observedexecute_code
    • First observedexecute_sql
    • First observedexecute_sql_multi
    • First observedgenerate_and_upload_pdf
    • First observedgenerate_lakebase_credential
    • First observedget_current_user
    • First observedget_table_stats_and_schema
    • First observedget_volume_folder_details
    • First observedlist_compute
    • First observedlist_tracked_resources
    • First observedmanage_app
    • First observedmanage_cluster
    • First observedmanage_dashboard
    • First observedmanage_genie
    • First observedmanage_job_runs
    • First observedmanage_jobs
    • First observedmanage_ka
    • First observedmanage_lakebase_branch
    • First observedmanage_lakebase_database
    • First observedmanage_lakebase_sync
    • First observedmanage_mas
    • First observedmanage_metric_views
    • First observedmanage_pipeline
    • First observedmanage_pipeline_run
    • First observedmanage_serving_endpoint
    • First observedmanage_sql_statement
    • First observedmanage_sql_warehouse
    • First observedmanage_uc_connections
    • First observedmanage_uc_grants
    • First observedmanage_uc_monitors
    • First observedmanage_uc_objects
    • First observedmanage_uc_security_policies
    • First observedmanage_uc_sharing
    • First observedmanage_uc_storage
    • First observedmanage_uc_tags
    • First observedmanage_volume_files
    • First observedmanage_vs_data
    • First observedmanage_vs_endpoint
    • First observedmanage_vs_index
    • First observedmanage_warehouse
    • First observedmanage_workspace
    • First observedmanage_workspace_files
    • First observedquery_vs_index

TDQS

A3.9/5.0

Scored across 45 tools

Disambiguation4/5

Most tools have clearly distinct resource targets (clusters, warehouses, pipelines, UC, Vector Search, Lakebase), and detailed descriptions clarify action scopes. A few pairs overlap, such as manage_sql_warehouse vs manage_warehouse, and list_compute duplicates listing functionality found in resource-specific tools, so occasional misselection is possible.

Naming Consistency4/5

All names use snake_case with a consistent verb_noun/manage_noun pattern across the large tool set. Abbreviations (ka, mas, uc, vs) are used consistently within families, though some names rely on them and one tool is noun-only (list_compute).

Tool Count2/5

45 tools exceeds the 25+ threshold for being too many, which creates a large surface for agents to search and select from. Even though each tool multiplexes many actions and Databricks is a broad platform, the count is well beyond a comfortably scoped server.

Completeness4/5

The surface covers broad Databricks lifecycle operations across jobs, pipelines, UC objects/grants/tags, Vector Search, Lakebase, volumes, and workspace files. Minor gaps remain, such as secrets, repos, and user/group management, but most core workflows can be completed or worked around.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables LLM-powered tools to interact with Databricks clusters, jobs, notebooks, SQL warehouses, and Unity Catalog through the Model Completion Protocol. Provides comprehensive access to Databricks REST API functionality including cluster management, job execution, workspace operations, and data catalog operations.
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI assistants to interact with Databricks workspaces programmatically, providing comprehensive tools for cluster management, notebook operations, job orchestration, Unity Catalog data governance, user management, permissions control, and FinOps cost analytics.
    492 npm
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables AI assistants to interact with Databricks workspaces, running SQL queries, managing jobs, and exploring schemas via the Model Context Protocol.
    1
    GPL 3.0