DeepSeek Chat Completion
deepseek_chatGenerate AI responses by sending chat messages to DeepSeek V4 models, with multi-turn sessions, function calling, thinking mode, JSON output, and automatic cost tracking.
Instructions
Chat with DeepSeek V4 models. deepseek-v4-flash (fast, economical) and deepseek-v4-pro (most capable), both 1M context with optional chain-of-thought thinking mode. deepseek-chat and deepseek-reasoner are deprecated aliases, still accepted for backward compatibility (resolve to v4-flash) but slated for removal; prefer the v4 names. Features: multi-turn sessions (session_id), function calling (tools parameter), thinking mode, JSON output mode, multimodal input (when enabled), automatic cost tracking, and model fallback with circuit breaker resilience.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Model to use. deepseek-v4-flash (default, fast/economical) or deepseek-v4-pro (most capable), both 1M context, up to 384K output. Non-thinking by default for speed; pass thinking:{type:"enabled"} to reason. Deprecated aliases (still accepted, prefer v4 names): deepseek-chat -> v4-flash non-thinking, deepseek-reasoner -> v4-flash thinking. | deepseek-v4-flash |
| tools | No | Array of tool definitions for function calling. Each tool has type "function" and a function object with name, description, and parameters (JSON Schema). | |
| stream | No | Enable streaming mode. Returns full response after streaming completes. | |
| messages | Yes | Array of conversation messages. Each message has role (system/user/assistant/tool) and content (string or array of content parts for multimodal). Tool messages require tool_call_id. | |
| thinking | No | Toggle chain-of-thought thinking mode. Use {type: "enabled"} to reason, {type: "disabled"} for a fast direct answer (the default here). When enabled, temperature/top_p are ignored. | |
| json_mode | No | Enable JSON output mode. The model will output valid JSON. Include the word "json" in your prompt for best results. Supported by both models. | |
| max_tokens | No | Maximum tokens to generate. V4 models support up to 384000 output tokens. | |
| session_id | No | Session ID for multi-turn conversations. When provided, previous messages from this session are prepended to the current messages. If the session does not exist, it is created automatically. Omit for stateless single-turn requests. | |
| temperature | No | Sampling temperature (0-2). Higher = more random. Default: 1.0. Ignored when thinking mode is enabled. | |
| tool_choice | No | Controls which tool the model calls. "auto" (default), "none", "required", or {type:"function",function:{name:"..."}} | |
| response_schema | No | JSON Schema to validate the model output against. Implies JSON output mode. The server validates the parsed result and, on failure, issues up to RESPONSE_SCHEMA_MAX_RETRIES repair retries (feeding the validation error back). The returned content is the first schema-valid object, or the last attempt with schema.valid=false and schema.error set. | |
| reasoning_effort | No | Reasoning effort while thinking mode is active: "high" (default) or "max". Only applies when thinking is enabled. |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| model | Yes | ||
| usage | Yes | ||
| schema | No | ||
| content | Yes | ||
| request | Yes | ||
| cost_usd | No | ||
| fallback | No | ||
| effective | No | ||
| session_id | No | ||
| tool_calls | No | ||
| routed_from | No | ||
| finish_reason | Yes | ||
| json_parse_error | No | ||
| reasoning_content | No |