douyin-mcp-server
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| DOUYIN_ASR_MODE | No | Set to 'local' to use local Whisper ASR instead of cloud SiliconFlow. Optional; default is cloud mode. | |
| DEEPSEEK_API_KEY | No | Optional DeepSeek API Key for post-processing transcripts (removing filler words, punctuation, segmentation). | |
| DOUYIN_MODEL_HUB | No | Optional cache directory for local Whisper models. If unset, standard Hugging Face cache is used. | |
| SILICONFLOW_API_KEY | No | SiliconFlow API Key required for cloud ASR mode. Can be set in WebUI or as an environment variable. Default cloud mode uses SiliconFlow SenseVoice. |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| get_douyin_download_linkC | 解析抖音链接;HTTP 失败时自动使用浏览器获取元数据。 |
| extract_douyin_textC | 提取文案。cloud 需要 SILICONFLOW_API_KEY;local 是可选离线模式。 |
| recognize_audio_fileC | 识别本地音频;默认 SiliconFlow,也可选择本地 Whisper。 |
| parse_douyin_video_infoC | 解析视频 ID、标题、下载地址与解析方式。 |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
| douyin_text_extraction_guide |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 4 tools
get_douyin_download_link and parse_douyin_video_info both resolve Douyin links and return download-related metadata, so agents may struggle to pick between them. recognize_audio_file and extract_douyin_text also have adjacent ASR/text-extraction purposes, though descriptions help somewhat.
All names use snake_case with a verb_object pattern: recognize_audio_file, get_douyin_download_link, extract_douyin_text, parse_douyin_video_info. The douyin prefix is used consistently where relevant.
4 tools is well within the 3-15 range and each maps to a core stage: audio recognition, link resolution, text extraction, and video info parsing. No obvious filler tools.
Core workflows for parsing a Douyin link, obtaining a download link, extracting text, and recognizing local audio are covered. Minor gaps include no direct video download/upload or comment/profile retrieval, but these are outside the apparent focus.