Skip to main content
Glama

抖音文案提取器(Windows 维护版)

把抖音视频链接粘贴进 WebUI,自动获取音频并用 SenseVoice 生成逐字稿。HTTP 快速解析失效时,程序会自动使用电脑已有的 Chrome 或 Edge,不需要用户理解页面内部结构。

本项目基于 yzfly/douyin-mcp-server 继续维护。感谢原作者和贡献者;本 fork 保留原项目的 Apache-2.0 许可与归属说明。

为什么有这个 fork

原项目停止维护后,抖音页面结构、精选链接和浏览器签名流程已经变化。这个版本集中修复新版链接与浏览器降级,并把默认 ASR 改为轻量、快速的 SiliconFlow SenseVoice,清除了维护者工作站的绝对路径依赖。

能力

v1.0

普通 /video/ 链接

✅

jingxuan?modal_id=

✅

aweme_id / item_ids / 短链跳转

✅

HTTP 快速解析

✅

Chrome → Edge → Playwright Chromium 降级

✅

SiliconFlow 云 ASR

✅ 默认

本地 Whisper

✅ 可选

Windows 一键启动

✅

固定 D:\AI_Tools

❌ 已移除

Related MCP server: Wanyi Watermark Remover

Windows 快速开始

需要先安装 uv、Node.js 18+、FFmpeg 以及 Chrome 或 Edge。

git clone https://github.com/maplezhang123/douyin-mcp-server.git
cd douyin-mcp-server
.\start.bat

start.bat 会检查依赖并启动 WebUI,不会永久修改系统环境。启动 WebUI 后访问 http://localhost:8080。

手动启动:

uv sync --extra web
npm install
uv run --extra web python web/app.py

API Key 配置

SiliconFlow(云端模式必需)

在 SiliconFlow 创建 Key,然后在 WebUI 的 “SiliconFlow API Key” 输入框填写。Key 只保存在浏览器 localStorage,并发送给本机服务。

也可在当前 PowerShell 设置:

$env:SILICONFLOW_API_KEY="sk-..."

默认使用 FunAudioLLM/SenseVoiceSmall。上传前转换为 16 kHz、单声道、64 kbps MP3;超出单次限制时自动切片。

DeepSeek(可选)

DeepSeek 只负责去口水词、补标点和分段:

$env:DEEPSEEK_API_KEY="sk-..."

没有 DeepSeek Key 时仍会正常返回并保存 SenseVoice 原始逐字稿。DeepSeek 请求失败也不会丢失已完成的 ASR 结果。

环境变量示例见 .env.example。服务不会把完整 Key 写入日志或输出文件。

切换到本地 Whisper

本地模式是可选的离线高级模式,不是默认依赖:

uv sync --extra web --extra local
$env:DOUYIN_ASR_MODE="local"
uv run --extra web --extra local python web/app.py

也可直接在 WebUI 选择“本地 Whisper”。首次使用会按 faster-whisper 的标准方式下载模型;默认使用 small / CPU int8。可用 DOUYIN_MODEL_HUB 指定缓存目录,未设置时使用标准 Hugging Face 缓存,不依赖固定盘符。

工作流程

  1. 从普通、精选、API 参数或短链接中识别视频 ID。

  2. 先尝试 HTTP 解析和下载;失败后自动启动 Chrome、Edge,最后才尝试 Playwright 自带 Chromium。

  3. FFmpeg 准备音频。

  4. 默认上传 SiliconFlow SenseVoice;本地模式使用 faster-whisper。

  5. 有 DeepSeek Key 时整理文案;没有或整理失败时返回原始逐字稿。

输出位于 output//:

audio.m4a
transcript-raw.md
copy.md
metadata.json
run-report.json

WebUI 展示标题、视频 ID、原始逐字稿、可选整理文案、耗时、解析方式和 ASR 模型。

命令行与 MCP

uv run python scripts/douyin_downloader.py -l "抖音链接" -a info
$env:SILICONFLOW_API_KEY="sk-..."
uv run python scripts/douyin_downloader.py -l "抖音链接" -a extract
uv run --extra local python scripts/douyin_downloader.py -l "抖音链接" -a extract --asr-mode local

MCP 入口为 douyin-mcp-server,提供 extract_douyin_text、parse_douyin_video_info、get_douyin_download_link 和 recognize_audio_file。

常见问题与排错

HTTP 页面中没有 videoInfoRes / _ROUTER_DATA 这是页面结构变化导致的快速解析失败。程序会自动进入浏览器降级;只要浏览器成功,无需处理该内部错误。

未检测到 Node.js 安装 Node.js 18+,重新打开终端,用 node -v 确认。

Playwright package 不存在 运行 npm install。项目优先使用 Chrome/Edge,通常不必额外下载约 200 MB 的 Chromium。

未检测到 Chrome/Edge 安装 Chrome 或 Edge。确需自带浏览器时运行 npx playwright install chromium。

未检测到 ffmpeg 运行 winget install Gyan.FFmpeg,重新打开终端,再用 ffmpeg -version 确认。

SiliconFlow Key 未填写、无效或网络失败 确认 WebUI 输入或 SILICONFLOW_API_KEY,检查余额与网络。自动测试不会调用真实 API。

faster-whisper-medium 权重不存在 默认流程不使用 medium,也不要求下载 GB 级模型。本地模式建议先用 small;模型可自动下载或通过 DOUYIN_MODEL_HUB 指定缓存。

DeepSeek 失败 结果仍包含原始逐字稿。检查可选的 DEEPSEEK_API_KEY 后可重试。

测试

uv run python -m unittest discover -s tests -v
uv run python -c "import douyin_mcp_server.server, scripts.douyin_downloader, web.app"

测试使用 mock,不访问抖音,也不消耗付费 API。

来源、许可与免责声明

Based on / forked from: yzfly/douyin-mcp-server.

项目按 Apache License 2.0 发布。仅供学习与研究;请遵守平台条款、版权与当地法律,不要处理无权使用的内容。

Available Tools

4 tools
extract_douyin_textC

提取文案。cloud 需要 SILICONFLOW_API_KEY;local 是可选离线模式。

ParametersJSON Schema
NameRequiredDescriptionDefault
ctxNo
contextNorewrite
asr_modeNocloud
share_linkYes
local_modelNosmall

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses that cloud mode requires SILICONFLOW_API_KEY and that local mode is an optional offline mode. However, it omits other important behavior such as extraction output format, failure conditions, or limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short clauses and is front-loaded with the core action before mode details. It is concise, though its extreme brevity reflects under-specification rather than optimal restraint.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter tool with no annotations and 0% schema description coverage, the description is incomplete. Output schema exists, so return values need not be explained, but parameter meaning and usage context are still missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for five parameters, so the description must explain them. It loosely maps to asr_mode's cloud/local values but never names or explains share_link, context, asr_mode, or local_model. The compensation is far from sufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description says '提取文案' (extract text/copy), which gives a verb+object but does not specify that it extracts text from a Douyin share link. It does not distinguish this tool from siblings like recognize_audio_file or parse_douyin_video_info. The mode note helps but remains vague about the exact resource.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It mentions cloud vs. local mode and the API-key requirement for cloud, but gives no guidance on when to choose this tool over alternatives. No when-to-use or when-not-to-use guidance is provided relative to sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

parse_douyin_video_infoC

解析视频 ID、标题、下载地址与解析方式。

ParametersJSON Schema
NameRequiredDescriptionDefault
share_linkYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, yet it says nothing about network fetch behavior, failure modes for invalid or expired share links, rate limits, or whether parsing is local vs remote. It only lists outputs, which the output schema already covers.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with no filler, front-loading the action verb and the field list. It is efficient, though the brevity is partly under-specification rather than discipline.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be re-explained, but the description leaves the only input parameter undocumented and offers no guidance on the tool's place among three overlapping siblings. For a network-fetch parsing tool this is thin.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single required parameter share_link has 0% schema description coverage and is never mentioned in the description. The description enumerates outputs, not the input, so nothing compensates for the undocumented parameter — the agent must guess whether it accepts a full URL, a share code, or plain text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The verb '解析' (parse) plus the explicit enumeration of extracted fields (视频 ID、标题、下载地址、解析方式) makes the purpose concrete. However, it never distinguishes itself from the similar siblings extract_douyin_text or get_douyin_download_link, which clearly overlap in scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus its siblings, nor any precondition about the share link format. An agent seeing four Douyin-related tools gets no routing guidance from this description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

recognize_audio_fileC

识别本地音频;默认 SiliconFlow,也可选择本地 Whisper。

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNosmall
asr_modeNocloud
file_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It hints at two backends (cloud SiliconFlow vs local Whisper) but never states that cloud mode requires network/API credentials, what audio formats or size limits apply, or how failures are surfaced – all material for an ASR tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words. It is efficiently structured, though the brevity is partly under-specification rather than disciplined concision.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. However, with zero annotations and 0% parameter coverage, the description should compensate by explaining the mode/model choices and any input constraints, and it does not.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and none of the three parameters are documented in the schema. The description loosely maps to 'asr_mode' (cloud vs local) but never names the parameter or enumerates valid values, and the 'model' parameter (default 'small') is entirely unaddressed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('识别本地音频' – recognize local audio), which is unambiguous and clearly distinct from the Douyin video siblings. It stops short of explicitly differentiating itself from those siblings, but the resource makes the boundary obvious.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The mention of 'default SiliconFlow, also can choose local Whisper' describes backend selection, not when to use this tool versus alternatives or prerequisites. There is no guidance on when to prefer local vs cloud mode, or what inputs are acceptable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv1.0.0
    • First observedextract_douyin_text
    • First observedget_douyin_download_link
    • First observedparse_douyin_video_info
    • First observedrecognize_audio_file

TDQS

B3.2/5.0

Scored across 4 tools

Disambiguation3/5

get_douyin_download_link and parse_douyin_video_info both resolve Douyin links and return download-related metadata, so agents may struggle to pick between them. recognize_audio_file and extract_douyin_text also have adjacent ASR/text-extraction purposes, though descriptions help somewhat.

Naming Consistency5/5

All names use snake_case with a verb_object pattern: recognize_audio_file, get_douyin_download_link, extract_douyin_text, parse_douyin_video_info. The douyin prefix is used consistently where relevant.

Tool Count5/5

4 tools is well within the 3-15 range and each maps to a core stage: audio recognition, link resolution, text extraction, and video info parsing. No obvious filler tools.

Completeness4/5

Core workflows for parsing a Douyin link, obtaining a download link, extracting text, and recognizing local audio are covered. Minor gaps include no direct video download/upload or comment/profile retrieval, but these are outside the apparent focus.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers