Skip to main content
Glama

DeepSRT MCP 服务器

模型上下文协议 (MCP) 服务器通过与 DeepSRT 的 API 集成提供 YouTube 视频摘要功能。

特征

  • 生成 YouTube 视频摘要

  • 支持叙述和要点摘要模式

  • 多语言支持(默认:zh-tw)

  • 与支持 MCP 的环境无缝集成

Related MCP server: youtube-mcp

工作原理

  1. 内容缓存

    • 必须先通过 DeepSRT 打开视频,以确保内容已缓存在服务中

    • 首次观看会触发 DeepSRT 服务中的缓存过程

  2. MCP摘要检索

    • 通过 MCP 请求摘要时,内容由 DeepSRT 的 CDN 边缘位置提供

    • 这确保了快速有效地传递摘要

  3. 预缓存内容

    • 一些视频可能已根据之前的用户请求缓存在系统中

    • 虽然您或许可以获取这些预缓存视频的摘要,但无法保证其可用性

    • 为获得最佳效果,请确保首先通过 DeepSRT 打开视频

%%{init: {'theme': 'dark', 'themeVariables': { 'primaryColor': '#2496ED', 'secondaryColor': '#38B2AC', 'tertiaryColor': '#1F2937', 'mainBkg': '#111827', 'textColor': '#E5E7EB', 'lineColor': '#4B5563', 'noteTextColor': '#E5E7EB'}}}%%
sequenceDiagram
    participant User
    participant DeepSRT
    participant Cache as DeepSRT Cache/CDN
    participant MCP as MCP Client

    Note over User,MCP: Step 1: Initial Caching
    User->>DeepSRT: Open video through DeepSRT
    DeepSRT->>Cache: Process and cache content
    Cache-->>DeepSRT: Confirm cache storage
    DeepSRT-->>User: Display video/content

    Note over User,MCP: Step 2: MCP Summary Retrieval
    MCP->>Cache: Request summary via MCP
    Cache-->>MCP: Return cached summary from edge location

    Note over User,MCP: Alternative: Pre-cached Content
    rect rgba(31, 41, 55, 0.6)
        MCP->>Cache: Request summary for pre-cached video
        alt Content exists in cache
            Cache-->>MCP: Return cached summary
        else Content not cached
            Cache-->>MCP: Cache miss
        end
    end

安装

安装 Claude Desktop

  1. 首先,搭建服务器:

npm install
npm run build
  1. 将服务器配置添加到您的 Claude Desktop 配置文件:

  • 在 macOS 上: ~/Library/Application Support/Claude/claude_desktop_config.json

  • 在 Windows 上: %APPDATA%/Claude/claude_desktop_config.json

{
  "mcpServers": {
    "deepsrt-mcp": {
      "command": "node",
      "args": [
        "/path/to/deepsrt-mcp/build/index.js"
      ]
    }
  }
}

为 Cline 安装

只需在聊天中要求 Cline 安装:

“嘿,请帮我从https://github.com/DeepSRT/deepsrt-mcp安装这个 MCP 服务器”

Cline 将自动为您安装deepsrt-mcp并更新您的cline_mcp_settings.json 。

用法

服务器提供以下工具:

获取摘要

获取 YouTube 视频的摘要。

参数:

  • videoId (必需):YouTube 视频 ID

  • lang (可选):语言代码(例如 zh-tw)- 默认为 zh-tw

  • mode (可选):摘要模式(“叙述”或“项目符号”) - 默认为叙述

示例用法

使用 Claude Desktop:

// The MCP tool will fetch the video summary
const result = await mcp.use_tool("deepsrt-mcp", "get_summary", {
  videoId: "dQw4w9WgXcQ",
  lang: "zh-tw",
  mode: "narrative"
});

使用Cline:

const result = await mcp.use_tool("deepsrt", "get_summary", {
  videoId: "dQw4w9WgXcQ",
  lang: "zh-tw",
  mode: "bullet"
});

发展

安装依赖项:

npm install

启动开发服务器:

npm run dev

为生产而构建:

npm run build

演示

常问问题

问:我收到404错误,为什么?

答:这是因为视频摘要未缓存在 CDN 边缘位置,您需要使用 DeepSRT chrome 扩展程序打开此视频以将其缓存在 CDN 网络中,然后才能使用 MCP 获取该摘要。

您可以使用 cURL 来验证缓存状态

curl -s 'https://worker.deepsrt.com/transcript' \
-i --data '{"arg":"v=VafNvIcOs5w","action":"summarize","lang":"zh-tw","mode":"narrative"}' | grep -i "^cache-status"
cache-status: HIT

如果您看到cache-status: HIT则内容已缓存在 CDN 边缘位置,并且您的 MCP 服务器不应该收到404 。

Available Tools

2 tools
get_summaryC

Get summary for a YouTube video

ParametersJSON Schema
NameRequiredDescriptionDefault
videoIdYesYouTube video ID
langNoTarget language (default: zh-tw)zh-tw
modeNoSummary mode (default: narrative)narrative

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states what the tool does but doesn't mention any behavioral traits such as rate limits, authentication needs, error handling, or what the summary output looks like (e.g., format, length). This leaves significant gaps for a tool that likely interacts with external APIs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that directly states the tool's purpose without any wasted words. It is appropriately sized and front-loaded, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations and output schema, the description is incomplete. It doesn't address key contextual aspects like the summary format, potential errors, or how it differs from the sibling tool. For a tool with external dependencies (YouTube API), more information on behavior and constraints is needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters (videoId, lang, mode) with descriptions and defaults. The description adds no additional meaning beyond what the schema provides, such as explaining what 'narrative' vs 'bullet' modes entail or how the lang parameter affects the summary.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get summary') and resource ('for a YouTube video'), making the purpose immediately understandable. It distinguishes from the sibling tool 'get_transcript' by focusing on summaries rather than transcripts, though it doesn't explicitly mention this distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like 'get_transcript'. The description lacks context about prerequisites, limitations, or scenarios where this tool is preferred, leaving the agent with no usage direction beyond the basic purpose.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_transcriptB

Get transcript for a YouTube video with timestamps

ParametersJSON Schema
NameRequiredDescriptionDefault
videoIdYesYouTube video ID or full YouTube URL
langNoPreferred language code for captions (default: en)en

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions 'with timestamps,' which adds some context about the output format, but fails to address critical aspects like rate limits, authentication needs, error handling, or whether the operation is read-only or has side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that front-loads the core purpose without unnecessary words. Every part of the sentence contributes directly to understanding the tool's function, making it highly concise and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (2 parameters, no annotations, no output schema), the description is minimally adequate. It covers the basic purpose and output feature (timestamps) but lacks details on behavioral traits, usage guidelines, and output structure, leaving gaps for the agent to navigate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, clearly documenting both parameters. The description adds no additional parameter semantics beyond what the schema provides, such as examples or constraints. With high schema coverage, the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Get transcript for a YouTube video with timestamps.' It specifies the verb ('Get'), resource ('transcript'), and key feature ('with timestamps'), but doesn't explicitly differentiate from the sibling tool 'get_summary' beyond the resource type.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like 'get_summary.' It lacks context about use cases, prerequisites, or exclusions, leaving the agent to infer usage based solely on the tool name and purpose.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv1.0.0
    • Addedget_summary
    • Addedget_transcript

TDQS

B3.1/5.0

Scored across 2 tools

Disambiguation5/5

The two tools have clearly distinct purposes: one retrieves a summary, the other retrieves a transcript with timestamps. There is no overlap in functionality, making it easy for an agent to select the correct tool based on the need for either a concise overview or detailed textual content.

Naming Consistency5/5

Both tools follow a consistent verb_noun pattern (get_summary, get_transcript), using the same verb 'get' and descriptive nouns. This uniformity makes the tool set predictable and easy to understand.

Tool Count2/5

With only two tools, the server feels thin for a domain like YouTube video processing. While the tools cover basic retrieval, the scope is limited, lacking operations such as search, analysis, or management of video content, which could be expected from a more comprehensive server.

Completeness2/5

The tool set is severely incomplete for a YouTube-focused server. It only provides retrieval functions (summary and transcript), missing essential operations like video search, metadata fetching, comment handling, or any CRUD capabilities, leaving significant gaps in coverage for typical agent workflows.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    An MCP server that enables interaction with the YouTube Data API, allowing users to search videos, get video and channel details, analyze trends, and fetch video transcripts.
    -
  • A
    license
    A
    quality
    C
    maintenance
    An MCP server that enables the extraction of transcripts and detailed metadata from YouTube videos. It allows users to retrieve video information like titles and descriptions, as well as transcripts with optional timestamps and language selection.
    2
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    MCP server that fetches YouTube video transcripts and optionally summarizes them. Supports multiple transcript formats (text, JSON, SRT, WebVTT), multi-language retrieval, and flexible YouTube URL parsing.
    6
    6
    MIT