Skip to main content
Glama
arikusi

Deepseek MCP Server

by arikusi

DeepSeek Chat Completion

deepseek_chat

Generate AI responses by sending chat messages to DeepSeek V4 models, with multi-turn sessions, function calling, thinking mode, JSON output, and automatic cost tracking.

Instructions

Chat with DeepSeek V4 models. deepseek-v4-flash (fast, economical) and deepseek-v4-pro (most capable), both 1M context with optional chain-of-thought thinking mode. deepseek-chat and deepseek-reasoner are deprecated aliases, still accepted for backward compatibility (resolve to v4-flash) but slated for removal; prefer the v4 names. Features: multi-turn sessions (session_id), function calling (tools parameter), thinking mode, JSON output mode, multimodal input (when enabled), automatic cost tracking, and model fallback with circuit breaker resilience.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
modelNoModel to use. deepseek-v4-flash (default, fast/economical) or deepseek-v4-pro (most capable), both 1M context, up to 384K output. Non-thinking by default for speed; pass thinking:{type:"enabled"} to reason. Deprecated aliases (still accepted, prefer v4 names): deepseek-chat -> v4-flash non-thinking, deepseek-reasoner -> v4-flash thinking.deepseek-v4-flash
toolsNoArray of tool definitions for function calling. Each tool has type "function" and a function object with name, description, and parameters (JSON Schema).
streamNoEnable streaming mode. Returns full response after streaming completes.
messagesYesArray of conversation messages. Each message has role (system/user/assistant/tool) and content (string or array of content parts for multimodal). Tool messages require tool_call_id.
thinkingNoToggle chain-of-thought thinking mode. Use {type: "enabled"} to reason, {type: "disabled"} for a fast direct answer (the default here). When enabled, temperature/top_p are ignored.
json_modeNoEnable JSON output mode. The model will output valid JSON. Include the word "json" in your prompt for best results. Supported by both models.
max_tokensNoMaximum tokens to generate. V4 models support up to 384000 output tokens.
session_idNoSession ID for multi-turn conversations. When provided, previous messages from this session are prepended to the current messages. If the session does not exist, it is created automatically. Omit for stateless single-turn requests.
temperatureNoSampling temperature (0-2). Higher = more random. Default: 1.0. Ignored when thinking mode is enabled.
tool_choiceNoControls which tool the model calls. "auto" (default), "none", "required", or {type:"function",function:{name:"..."}}
response_schemaNoJSON Schema to validate the model output against. Implies JSON output mode. The server validates the parsed result and, on failure, issues up to RESPONSE_SCHEMA_MAX_RETRIES repair retries (feeding the validation error back). The returned content is the first schema-valid object, or the last attempt with schema.valid=false and schema.error set.
reasoning_effortNoReasoning effort while thinking mode is active: "high" (default) or "max". Only applies when thinking is enabled.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
modelYes
usageYes
schemaNo
contentYes
requestYes
cost_usdNo
fallbackNo
effectiveNo
session_idNo
tool_callsNo
routed_fromNo
finish_reasonYes
json_parse_errorNo
reasoning_contentNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed8 schema fields changedv2.3.0
    • changedInput schema / properties / model / description
      Previous value: -"Model to use. deepseek-v4-flash (default, fast/economical) or deepseek-v4-pro (most capable), both 1M context, up to 384K output. Non-thinking by default for speed; pass thinking:{type:\"enabled\"} to reason. Aliases: deepseek-chat -> v4-flash non-thinking, deepseek-reasoner -> v4-flash thinking."New value: +"Model to use. deepseek-v4-flash (default, fast/economical) or deepseek-v4-pro (most capable), both 1M context, up to 384K output. Non-thinking by default for speed; pass thinking:{type:\"enabled\"} to reason. Deprecated aliases (still accepted, prefer v4 names): deepseek-chat -> v4-flash non-thinking, deepseek-reasoner -> v4-flash thinking."
    • addedInput schema / properties / response_schema
      Added value: +{
      +  "additionalProperties": {},
      +  "description": "JSON Schema to validate the model output against. Implies JSON output mode. The server validates the parsed result and, on failure, issues up to RESPONSE_SCHEMA_MAX_RETRIES repair retries (feeding the validation error back). The returned content is the first schema-valid object, or the last attempt with schema.valid=false and schema.error set.",
      +  "propertyNames": {
      +    "type": "string"
      +  },
      +  "type": "object"
      +}
    • addedOutput schema / properties / effective
      Added value: +{
      +  "additionalProperties": false,
      +  "properties": {
      +    "model": {
      +      "type": "string"
      +    },
      +    "temperature": {
      +      "type": "number"
      +    },
      +    "thinking": {
      +      "type": "boolean"
      +    }
      +  },
      +  "required": [
      +    "model",
      +    "thinking"
      +  ],
      +  "type": "object"
      +}
    • addedOutput schema / properties / fallback
      Added value: +{
      +  "additionalProperties": false,
      +  "properties": {
      +    "fallbackModel": {
      +      "type": "string"
      +    },
      +    "originalModel": {
      +      "type": "string"
      +    },
      +    "reason": {
      +      "type": "string"
      +    }
      +  },
      +  "required": [
      +    "originalModel",
      +    "fallbackModel",
      +    "reason"
      +  ],
      +  "type": "object"
      +}
    • addedOutput schema / properties / json_parse_error
      Added value: +{
      +  "type": "string"
      +}
    • addedOutput schema / properties / request
      Added value: +{
      +  "additionalProperties": false,
      +  "properties": {
      +    "cache_hit_tokens": {
      +      "type": "number"
      +    },
      +    "cache_miss_tokens": {
      +      "type": "number"
      +    },
      +    "completion_tokens": {
      +      "type": "number"
      +    },
      +    "cost_usd": {
      +      "type": "number"
      +    },
      +    "fallback_used": {
      +      "type": "boolean"
      +    },
      +    "finish_reason": {
      +      "type": "string"
      +    },
      +    "model": {
      +      "type": "string"
      +    },
      +    "prompt_tokens": {
      +      "type": "number"
      +    },
      +    "temperature": {
      +      "type": "number"
      +    },
      +    "thinking": {
      +      "type": "boolean"
      +    },
      +    "total_tokens": {
      +      "type": "number"
      +    },
      +    "wire_model": {
      +      "type": "string"
      +    }
      +  },
      +  "required": [
      +    "model",
      +    "wire_model",
      +    "thinking",
      +    "fallback_used",
      +    "finish_reason",
      +    "prompt_tokens",
      +    "completion_tokens",
      +    "total_tokens",
      +    "cache_hit_tokens",
      +    "cache_miss_tokens",
      +    "cost_usd"
      +  ],
      +  "type": "object"
      +}
    • addedOutput schema / properties / schema
      Added value: +{
      +  "additionalProperties": false,
      +  "properties": {
      +    "attempts": {
      +      "type": "number"
      +    },
      +    "error": {
      +      "type": "string"
      +    },
      +    "valid": {
      +      "type": "boolean"
      +    }
      +  },
      +  "required": [
      +    "valid",
      +    "attempts"
      +  ],
      +  "type": "object"
      +}
    • changedOutput schema / required
      Previous value: -[
      -  "content",
      -  "model",
      -  "usage",
      -  "finish_reason"
      -]New value: +[
      +  "content",
      +  "model",
      +  "usage",
      +  "request",
      +  "finish_reason"
      +]
  2. Changed9 schema fields changedv2.0.0
    • changedInput schema / properties / max_tokens / description
      Previous value: -"Maximum tokens to generate. deepseek-chat: max 8192, deepseek-reasoner: max 65536"New value: +"Maximum tokens to generate. V4 models support up to 384000 output tokens."
    • changedInput schema / properties / max_tokens / maximum
      Previous value: -65536New value: +384000
    • changedInput schema / properties / model / default
      Previous value: -"deepseek-chat"New value: +"deepseek-v4-flash"
    • changedInput schema / properties / model / description
      Previous value: -"Model to use. Both run DeepSeek V3.2 (128K context). deepseek-chat: non-thinking mode (max 8K output), deepseek-reasoner: thinking mode (max 64K output)"New value: +"Model to use. deepseek-v4-flash (default, fast/economical) or deepseek-v4-pro (most capable), both 1M context, up to 384K output. Non-thinking by default for speed; pass thinking:{type:\"enabled\"} to reason. Aliases: deepseek-chat -> v4-flash non-thinking, deepseek-reasoner -> v4-flash thinking."
    • changedInput schema / properties / model / enum
      Previous value: -[
      -  "deepseek-chat",
      -  "deepseek-reasoner"
      -]New value: +[
      +  "deepseek-v4-flash",
      +  "deepseek-v4-pro",
      +  "deepseek-chat",
      +  "deepseek-reasoner"
      +]
    • addedInput schema / properties / reasoning_effort
      Added value: +{
      +  "description": "Reasoning effort while thinking mode is active: \"high\" (default) or \"max\". Only applies when thinking is enabled.",
      +  "enum": [
      +    "high",
      +    "max"
      +  ],
      +  "type": "string"
      +}
    • changedInput schema / properties / thinking / description
      Previous value: -"Enable thinking mode. When enabled, temperature/top_p/frequency_penalty/presence_penalty are automatically ignored. Use {type: \"enabled\"} to activate."New value: +"Toggle chain-of-thought thinking mode. Use {type: \"enabled\"} to reason, {type: \"disabled\"} for a fast direct answer (the default here). When enabled, temperature/top_p are ignored."
    • addedOutput schema / properties / cost_usd
      Added value: +{
      +  "type": "number"
      +}
    • addedOutput schema / properties / routed_from
      Added value: +{
      +  "type": "string"
      +}
  3. First observedv1.5.0

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does add useful behavioral context: deprecated aliases slated for removal, model fallback with circuit breaker resilience, automatic cost tracking, and optional thinking mode. It does not cover every runtime behavior, but the traits disclosed go beyond a bare feature list.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded with the core purpose, then efficiently lists model variants and capabilities. It is dense but every sentence contributes useful differentiating information, with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a high-complexity tool with 12 parameters, nested objects, and an output schema, the description gives a strong high-level map: model choice, context length, aliases, and the major feature areas. The schema fills in the rest, so the agent is not left without critical calling context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already fully documents all 12 parameters. The description restates some parameter concepts (tools, thinking, session_id, json_mode) but adds no syntax or usage detail beyond what the input schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: "Chat with DeepSeek V4 models," which clearly identifies this as a chat-completion tool. The feature list (multi-turn sessions, function calling, thinking mode, JSON mode) further clarifies scope, though it does not explicitly contrast with deepseek_fim or deepseek_sessions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage from the verb "Chat" and names relevant capabilities like session_id and tools, giving an agent a reasonable sense of when to call it. However, it provides no explicit guidance about when to prefer deepseek_fim or deepseek_sessions, nor any exclusions or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools