DeepSRT MCP Server
OfficialDeepSRT MCP サーバー
DeepSRT の API との統合を通じて YouTube ビデオ要約機能を提供するモデル コンテキスト プロトコル (MCP) サーバー。
特徴
YouTube動画の要約を生成する
物語形式と箇条書き形式の要約モードの両方をサポート
多言語サポート(デフォルト:zh-tw)
MCP対応環境とのシームレスな統合
Related MCP server: youtube-mcp
仕組み
コンテンツキャッシュ
コンテンツがサービスにキャッシュされていることを確認するために、ビデオはまずDeepSRT経由で開く必要があります。
この最初の閲覧により、DeepSRTサービスのキャッシュプロセスが開始されます。
MCPサマリー検索
MCPを通じて要約をリクエストすると、コンテンツはDeepSRTのCDNエッジロケーションから提供されます。
これにより、要約を迅速かつ効率的に配信できます。
事前キャッシュされたコンテンツ
一部のビデオは、以前のユーザーリクエストからすでにシステムにキャッシュされている可能性があります。
事前にキャッシュされた動画の要約を取得できる場合もありますが、利用可能かどうかは保証されません。
最良の結果を得るには、まずDeepSRTで動画を開いてください。
%%{init: {'theme': 'dark', 'themeVariables': { 'primaryColor': '#2496ED', 'secondaryColor': '#38B2AC', 'tertiaryColor': '#1F2937', 'mainBkg': '#111827', 'textColor': '#E5E7EB', 'lineColor': '#4B5563', 'noteTextColor': '#E5E7EB'}}}%%
sequenceDiagram
participant User
participant DeepSRT
participant Cache as DeepSRT Cache/CDN
participant MCP as MCP Client
Note over User,MCP: Step 1: Initial Caching
User->>DeepSRT: Open video through DeepSRT
DeepSRT->>Cache: Process and cache content
Cache-->>DeepSRT: Confirm cache storage
DeepSRT-->>User: Display video/content
Note over User,MCP: Step 2: MCP Summary Retrieval
MCP->>Cache: Request summary via MCP
Cache-->>MCP: Return cached summary from edge location
Note over User,MCP: Alternative: Pre-cached Content
rect rgba(31, 41, 55, 0.6)
MCP->>Cache: Request summary for pre-cached video
alt Content exists in cache
Cache-->>MCP: Return cached summary
else Content not cached
Cache-->>MCP: Cache miss
end
endインストール
Claude Desktop へのインストール
まず、サーバーを構築します。
npm install
npm run buildClaude Desktop 構成ファイルにサーバー構成を追加します。
macOSの場合:
~/Library/Application Support/Claude/claude_desktop_config.jsonWindows の場合:
%APPDATA%/Claude/claude_desktop_config.json
{
"mcpServers": {
"deepsrt-mcp": {
"command": "node",
"args": [
"/path/to/deepsrt-mcp/build/index.js"
]
}
}
}Cline のインストール
チャットでClineにインストールを依頼するだけです:
「ねえ、この MCP サーバーをhttps://github.com/DeepSRT/deepsrt-mcpからインストールしてください」
Cline はdeepsrt-mcpを自動的にインストールし、 cline_mcp_settings.jsonを更新します。
使用法
サーバーは次のツールを提供します。
要約を取得する
YouTube ビデオの概要を取得します。
パラメータ:
videoId(必須): YouTube 動画 IDlang(オプション): 言語コード (例: zh-tw) - デフォルトは zh-twmode(オプション):要約モード(「narrative」または「bullet」) - デフォルトはnarrative
使用例
Claude Desktop の使用:
// The MCP tool will fetch the video summary
const result = await mcp.use_tool("deepsrt-mcp", "get_summary", {
videoId: "dQw4w9WgXcQ",
lang: "zh-tw",
mode: "narrative"
});Cline の使用:
const result = await mcp.use_tool("deepsrt", "get_summary", {
videoId: "dQw4w9WgXcQ",
lang: "zh-tw",
mode: "bullet"
});発達
依存関係をインストールします:
npm install開発サーバーを起動します:
npm run dev生産用にビルド:
npm run buildデモ
よくある質問
Q: 404エラーが発生しますが、なぜでしょうか?
A: これは、ビデオの概要が CDN エッジ ロケーションにキャッシュされていないためです。MCP を使用して概要を取得する前に、DeepSRT Chrome 拡張機能を使用してこのビデオを開き、CDN ネットワークにキャッシュする必要があります。
次のようにcURLを使用してキャッシュの状態を確認できます。
curl -s 'https://worker.deepsrt.com/transcript' \
-i --data '{"arg":"v=VafNvIcOs5w","action":"summarize","lang":"zh-tw","mode":"narrative"}' | grep -i "^cache-status"
cache-status: HITcache-status: HITと表示される場合、コンテンツは CDN エッジ ロケーションにキャッシュされており、MCP サーバーは404取得しません。
Available Tools
2 toolsget_summaryC
Get summary for a YouTube video
| Name | Required | Description | Default |
|---|---|---|---|
| videoId | Yes | YouTube video ID | |
| lang | No | Target language (default: zh-tw) | zh-tw |
| mode | No | Summary mode (default: narrative) | narrative |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states what the tool does but doesn't mention any behavioral traits such as rate limits, authentication needs, error handling, or what the summary output looks like (e.g., format, length). This leaves significant gaps for a tool that likely interacts with external APIs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without any wasted words. It is appropriately sized and front-loaded, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and output schema, the description is incomplete. It doesn't address key contextual aspects like the summary format, potential errors, or how it differs from the sibling tool. For a tool with external dependencies (YouTube API), more information on behavior and constraints is needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters (videoId, lang, mode) with descriptions and defaults. The description adds no additional meaning beyond what the schema provides, such as explaining what 'narrative' vs 'bullet' modes entail or how the lang parameter affects the summary.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Get summary') and resource ('for a YouTube video'), making the purpose immediately understandable. It distinguishes from the sibling tool 'get_transcript' by focusing on summaries rather than transcripts, though it doesn't explicitly mention this distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like 'get_transcript'. The description lacks context about prerequisites, limitations, or scenarios where this tool is preferred, leaving the agent with no usage direction beyond the basic purpose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_transcriptB
Get transcript for a YouTube video with timestamps
| Name | Required | Description | Default |
|---|---|---|---|
| videoId | Yes | YouTube video ID or full YouTube URL | |
| lang | No | Preferred language code for captions (default: en) | en |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions 'with timestamps,' which adds some context about the output format, but fails to address critical aspects like rate limits, authentication needs, error handling, or whether the operation is read-only or has side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core purpose without unnecessary words. Every part of the sentence contributes directly to understanding the tool's function, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (2 parameters, no annotations, no output schema), the description is minimally adequate. It covers the basic purpose and output feature (timestamps) but lacks details on behavioral traits, usage guidelines, and output structure, leaving gaps for the agent to navigate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, clearly documenting both parameters. The description adds no additional parameter semantics beyond what the schema provides, such as examples or constraints. With high schema coverage, the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Get transcript for a YouTube video with timestamps.' It specifies the verb ('Get'), resource ('transcript'), and key feature ('with timestamps'), but doesn't explicitly differentiate from the sibling tool 'get_summary' beyond the resource type.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'get_summary.' It lacks context about use cases, prerequisites, or exclusions, leaving the agent to infer usage based solely on the tool name and purpose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v1.0.0- Added
get_summary - Added
get_transcript
TDQS
Scored across 2 tools
The two tools have clearly distinct purposes: one retrieves a summary, the other retrieves a transcript with timestamps. There is no overlap in functionality, making it easy for an agent to select the correct tool based on the need for either a concise overview or detailed textual content.
Both tools follow a consistent verb_noun pattern (get_summary, get_transcript), using the same verb 'get' and descriptive nouns. This uniformity makes the tool set predictable and easy to understand.
With only two tools, the server feels thin for a domain like YouTube video processing. While the tools cover basic retrieval, the scope is limited, lacking operations such as search, analysis, or management of video content, which could be expected from a more comprehensive server.
The tool set is severely incomplete for a YouTube-focused server. It only provides retrieval functions (summary and transcript), missing essential operations like video search, metadata fetching, comment handling, or any CRUD capabilities, leaving significant gaps in coverage for typical agent workflows.
Maintenance
Related MCP Connectors
An MCP server that gives any LLM or agent clean YouTube transcripts on demand: a single video, a whole channel, or a playlist, plus AI cleanup of auto-generated captions. API-key auth, credit-based, same backend as the public v1 API. Get a free API key with 25 free credits at youtubetranscriptdownload.com/account.
MCP server for RiverScript, an AI transcription platform - fetches transcripts shared via a link.
An MCP server that integrates with Discord to provide AI-powered features.
MCP server for OpenAI Sora AI video generation
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceAn MCP server that enables interaction with the YouTube Data API, allowing users to search videos, get video and channel details, analyze trends, and fetch video transcripts.-
- AlicenseAqualityCmaintenanceAn MCP server that enables the extraction of transcripts and detailed metadata from YouTube videos. It allows users to retrieve video information like titles and descriptions, as well as transcripts with optional timestamps and language selection.2MIT
- AlicenseAqualityAmaintenanceMCP server that fetches YouTube video transcripts and optionally summarizes them. Supports multiple transcript formats (text, JSON, SRT, WebVTT), multi-language retrieval, and flexible YouTube URL parsing.66MIT
- AlicenseAqualityDmaintenanceA powerful MCP server for summarizing YouTube videos with transcript fetching, intelligent summarization, key point extraction, and metadata retrieval.4MIT