Skip to main content
Glama

Model serving endpoints

manage_serving_endpoint
Destructive

Create, update, query, delete, and monitor Databricks Model Serving endpoints, with confirmation steps and secret redaction.

Instructions

Manage and query Databricks Model Serving endpoints.

Actions:

  • list / get: endpoints with state and served entities. Credentials of external-model providers (API keys, secrets, tokens, plaintext env vars) are always stripped.

  • create: spec = create body (config, ai_gateway, tags, route_optimized, budget_policy_id, description, email_notifications, rate_limits, ...); name is a dedicated parameter.

  • update_config: spec = {served_entities, traffic_config, auto_capture_config, served_models}.

  • update_ai_gateway: spec = {guardrails, inference_table_config, rate_limits, usage_tracking_config, fallback_config}.

  • delete: permanently delete (requires confirm).

  • query: invoke the endpoint with request (chat messages / prompt / embeddings input / dataframe).

  • get_build_logs / get_logs: build or server logs for served_model_name (tail, size-capped). create/update_config are long-running: they return status 'pending' unless wait_seconds is set.

Safety classification: list, get, get_build_logs, get_logs = READ_ONLY; create, update_config, update_ai_gateway = WRITE; delete = DESTRUCTIVE; query = EXECUTION.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
nameNoServing endpoint name (all actions except list).
specNoRequest body fields for create/update, using the Databricks REST API field names (snake_case). Unknown fields are rejected.
actionYesOperation to perform.
confirmNoSet to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions.
dry_runNoIf true, validate and return the planned change without executing it.
requestNoquery: request body. Chat: {'messages': [{'role': 'user', 'content': '...'}], 'max_tokens': 256}; completions: {'prompt': '...'}; embeddings: {'input': ['...']}; custom models: {'dataframe_records': [...]} / {'dataframe_split': {...}} / {'instances': [...]} / {'inputs': ...}. Streaming is not supported.
page_sizeNoMax items to return (server caps this).
page_tokenNonext_page_token from a previous response.
wait_secondsNoOptionally wait up to this many seconds for the operation to finish (capped by the server's max wait). Default: return immediately with status 'pending'.
max_output_charsNoCap on returned query/log text size (default 20000, max 200000).
served_model_nameNoServed model/entity name for get_build_logs/get_logs.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
dataNo
pageNo
planNo
toolYes
actionNo
safetyNo
statusNosuccess
summaryYes
warningsNo
next_stepsNoSuggested follow-up calls.
request_idNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the coarse destructiveHint/readOnlyHint annotations: it assigns a per-action safety classification (READ_ONLY/WRITE/DESTRUCTIVE/EXECUTION), discloses that provider credentials are always stripped on list/get, that delete permanently destroys and needs confirm, that create/update return 'pending' unless wait_seconds is set, and that logs are tailed and size-capped.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded purpose followed by a dense but well-structured bulleted action list where every line carries operational information; the trailing safety classification is compact and useful. Slightly heavy, but nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a nine-action tool with an output schema present, it covers the create/update/delete lifecycle, the dry_run/confirm confirmation flow, asynchronous completion via wait_seconds, pagination-adjacent caps, and per-action risk. An agent has everything needed to invoke the right action correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real value by mapping spec contents to specific actions and clarifying that request is only for query with per-mode shapes. It clarifies role of wait_seconds and max_output_chars beyond raw schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Manage and query Databricks Model Serving endpoints') and then enumerates all nine actions with concrete semantics, so an agent knows exactly what the tool covers. It is clearly distinguishable from the Vector Search sibling manage_vs_endpoint by the 'Model Serving' framing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Each action is given a usage-specific meaning, including which spec keys belong to create vs update_config vs update_ai_gateway, that delete requires confirm, and that wait_seconds controls blocking on long-running calls. There is no explicit when-not or sibling routing (e.g., vs manage_vs_endpoint), but the action-level guidance is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.