Model serving endpoints
manage_serving_endpointCreate, update, query, delete, and monitor Databricks Model Serving endpoints, with confirmation steps and secret redaction.
Instructions
Manage and query Databricks Model Serving endpoints.
Actions:
list / get: endpoints with state and served entities. Credentials of external-model providers (API keys, secrets, tokens, plaintext env vars) are always stripped.
create: spec = create body (config, ai_gateway, tags, route_optimized, budget_policy_id, description, email_notifications, rate_limits, ...); name is a dedicated parameter.
update_config: spec = {served_entities, traffic_config, auto_capture_config, served_models}.
update_ai_gateway: spec = {guardrails, inference_table_config, rate_limits, usage_tracking_config, fallback_config}.
delete: permanently delete (requires confirm).
query: invoke the endpoint with
request(chat messages / prompt / embeddings input / dataframe).get_build_logs / get_logs: build or server logs for
served_model_name(tail, size-capped). create/update_config are long-running: they return status 'pending' unlesswait_secondsis set.
Safety classification: list, get, get_build_logs, get_logs = READ_ONLY; create, update_config, update_ai_gateway = WRITE; delete = DESTRUCTIVE; query = EXECUTION.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Serving endpoint name (all actions except list). | |
| spec | No | Request body fields for create/update, using the Databricks REST API field names (snake_case). Unknown fields are rejected. | |
| action | Yes | Operation to perform. | |
| confirm | No | Set to true ONLY after the user has reviewed the plan returned by a previous call with status 'confirmation_required'. Required for destructive/security-sensitive actions. | |
| dry_run | No | If true, validate and return the planned change without executing it. | |
| request | No | query: request body. Chat: {'messages': [{'role': 'user', 'content': '...'}], 'max_tokens': 256}; completions: {'prompt': '...'}; embeddings: {'input': ['...']}; custom models: {'dataframe_records': [...]} / {'dataframe_split': {...}} / {'instances': [...]} / {'inputs': ...}. Streaming is not supported. | |
| page_size | No | Max items to return (server caps this). | |
| page_token | No | next_page_token from a previous response. | |
| wait_seconds | No | Optionally wait up to this many seconds for the operation to finish (capped by the server's max wait). Default: return immediately with status 'pending'. | |
| max_output_chars | No | Cap on returned query/log text size (default 20000, max 200000). | |
| served_model_name | No | Served model/entity name for get_build_logs/get_logs. |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| data | No | ||
| page | No | ||
| plan | No | ||
| tool | Yes | ||
| action | No | ||
| safety | No | ||
| status | No | success | |
| summary | Yes | ||
| warnings | No | ||
| next_steps | No | Suggested follow-up calls. | |
| request_id | No |