DeepSRT MCP Server
OfficialDeepSRT MCP 서버
DeepSRT의 API와 통합하여 YouTube 비디오 요약 기능을 제공하는 MCP(Model Context Protocol) 서버입니다.
특징
YouTube 동영상에 대한 요약 생성
내러티브 및 요점 요약 모드 모두 지원
다국어 지원(기본값: zh-tw)
MCP 지원 환경과의 원활한 통합
Related MCP server: youtube-mcp
작동 원리
콘텐츠 캐싱
콘텐츠가 서비스에 캐시되도록 하려면 먼저 DeepSRT를 통해 비디오를 열어야 합니다.
이 초기 보기는 DeepSRT 서비스의 캐싱 프로세스를 트리거합니다.
MCP 요약 검색
MCP를 통해 요약을 요청하는 경우 콘텐츠는 DeepSRT의 CDN 에지 위치에서 제공됩니다.
이를 통해 요약의 빠르고 효율적인 전달이 보장됩니다.
사전 캐시된 콘텐츠
일부 비디오는 이전 사용자 요청으로 시스템에 이미 캐시되었을 수 있습니다.
미리 캐시된 비디오에 대한 요약을 가져올 수는 있지만 가용성은 보장되지 않습니다.
최상의 결과를 얻으려면 먼저 DeepSRT를 통해 비디오를 열어야 합니다.
지엑스피1
설치
Claude Desktop 설치
먼저 서버를 빌드합니다.
npm install
npm run buildClaude Desktop 구성 파일에 서버 구성을 추가합니다.
macOS의 경우:
~/Library/Application Support/Claude/claude_desktop_config.jsonWindows의 경우:
%APPDATA%/Claude/claude_desktop_config.json
{
"mcpServers": {
"deepsrt-mcp": {
"command": "node",
"args": [
"/path/to/deepsrt-mcp/build/index.js"
]
}
}
}Cline 설치
채팅에서 Cline에게 설치를 요청하세요.
"안녕하세요, https://github.com/DeepSRT/deepsrt-mcp 에서 이 MCP 서버를 설치해 주세요."
Cline은 자동으로 deepsrt-mcp 설치하고 cline_mcp_settings.json 업데이트합니다.
용법
서버는 다음과 같은 도구를 제공합니다.
요약 받기
YouTube 동영상에 대한 요약을 가져옵니다.
매개변수:
videoId(필수): YouTube 비디오 IDlang(선택 사항): 언어 코드(예: zh-tw) - 기본값은 zh-tw입니다.mode(선택 사항): 요약 모드("내러티브" 또는 "글머리 기호") - 기본적으로 내러티브로 설정됨
사용 예
Claude Desktop 사용:
// The MCP tool will fetch the video summary
const result = await mcp.use_tool("deepsrt-mcp", "get_summary", {
videoId: "dQw4w9WgXcQ",
lang: "zh-tw",
mode: "narrative"
});클라인을 사용하여:
const result = await mcp.use_tool("deepsrt", "get_summary", {
videoId: "dQw4w9WgXcQ",
lang: "zh-tw",
mode: "bullet"
});개발
종속성 설치:
npm install개발 서버 시작:
npm run dev프로덕션을 위해 빌드:
npm run build데모
자주 묻는 질문
질문: 404 오류가 발생하는 이유는 무엇인가요?
답변: 이는 비디오 요약이 CDN 에지 위치에 캐시되지 않기 때문입니다. MCP를 사용하여 요약을 얻으려면 먼저 DeepSRT 크롬 확장 프로그램을 사용하여 비디오를 열어 CDN 네트워크에 캐시해야 합니다.
cURL을 사용하면 캐시 상태를 다음과 같이 확인할 수 있습니다.
curl -s 'https://worker.deepsrt.com/transcript' \
-i --data '{"arg":"v=VafNvIcOs5w","action":"summarize","lang":"zh-tw","mode":"narrative"}' | grep -i "^cache-status"
cache-status: HITcache-status: HIT 표시되면 콘텐츠가 CDN 에지 위치에 캐시되므로 MCP 서버는 404 받아서는 안 됩니다.
Available Tools
2 toolsget_summaryC
Get summary for a YouTube video
| Name | Required | Description | Default |
|---|---|---|---|
| videoId | Yes | YouTube video ID | |
| lang | No | Target language (default: zh-tw) | zh-tw |
| mode | No | Summary mode (default: narrative) | narrative |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states what the tool does but doesn't mention any behavioral traits such as rate limits, authentication needs, error handling, or what the summary output looks like (e.g., format, length). This leaves significant gaps for a tool that likely interacts with external APIs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without any wasted words. It is appropriately sized and front-loaded, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and output schema, the description is incomplete. It doesn't address key contextual aspects like the summary format, potential errors, or how it differs from the sibling tool. For a tool with external dependencies (YouTube API), more information on behavior and constraints is needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters (videoId, lang, mode) with descriptions and defaults. The description adds no additional meaning beyond what the schema provides, such as explaining what 'narrative' vs 'bullet' modes entail or how the lang parameter affects the summary.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Get summary') and resource ('for a YouTube video'), making the purpose immediately understandable. It distinguishes from the sibling tool 'get_transcript' by focusing on summaries rather than transcripts, though it doesn't explicitly mention this distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like 'get_transcript'. The description lacks context about prerequisites, limitations, or scenarios where this tool is preferred, leaving the agent with no usage direction beyond the basic purpose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_transcriptB
Get transcript for a YouTube video with timestamps
| Name | Required | Description | Default |
|---|---|---|---|
| videoId | Yes | YouTube video ID or full YouTube URL | |
| lang | No | Preferred language code for captions (default: en) | en |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions 'with timestamps,' which adds some context about the output format, but fails to address critical aspects like rate limits, authentication needs, error handling, or whether the operation is read-only or has side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core purpose without unnecessary words. Every part of the sentence contributes directly to understanding the tool's function, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (2 parameters, no annotations, no output schema), the description is minimally adequate. It covers the basic purpose and output feature (timestamps) but lacks details on behavioral traits, usage guidelines, and output structure, leaving gaps for the agent to navigate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, clearly documenting both parameters. The description adds no additional parameter semantics beyond what the schema provides, such as examples or constraints. With high schema coverage, the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Get transcript for a YouTube video with timestamps.' It specifies the verb ('Get'), resource ('transcript'), and key feature ('with timestamps'), but doesn't explicitly differentiate from the sibling tool 'get_summary' beyond the resource type.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'get_summary.' It lacks context about use cases, prerequisites, or exclusions, leaving the agent to infer usage based solely on the tool name and purpose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v1.0.0- Added
get_summary - Added
get_transcript
TDQS
Scored across 2 tools
The two tools have clearly distinct purposes: one retrieves a summary, the other retrieves a transcript with timestamps. There is no overlap in functionality, making it easy for an agent to select the correct tool based on the need for either a concise overview or detailed textual content.
Both tools follow a consistent verb_noun pattern (get_summary, get_transcript), using the same verb 'get' and descriptive nouns. This uniformity makes the tool set predictable and easy to understand.
With only two tools, the server feels thin for a domain like YouTube video processing. While the tools cover basic retrieval, the scope is limited, lacking operations such as search, analysis, or management of video content, which could be expected from a more comprehensive server.
The tool set is severely incomplete for a YouTube-focused server. It only provides retrieval functions (summary and transcript), missing essential operations like video search, metadata fetching, comment handling, or any CRUD capabilities, leaving significant gaps in coverage for typical agent workflows.
Maintenance
Related MCP Connectors
An MCP server that gives any LLM or agent clean YouTube transcripts on demand: a single video, a whole channel, or a playlist, plus AI cleanup of auto-generated captions. API-key auth, credit-based, same backend as the public v1 API. Get a free API key with 25 free credits at youtubetranscriptdownload.com/account.
MCP server for RiverScript, an AI transcription platform - fetches transcripts shared via a link.
An MCP server that integrates with Discord to provide AI-powered features.
MCP server for OpenAI Sora AI video generation
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceAn MCP server that enables interaction with the YouTube Data API, allowing users to search videos, get video and channel details, analyze trends, and fetch video transcripts.-
- AlicenseAqualityCmaintenanceAn MCP server that enables the extraction of transcripts and detailed metadata from YouTube videos. It allows users to retrieve video information like titles and descriptions, as well as transcripts with optional timestamps and language selection.2MIT
- AlicenseAqualityAmaintenanceMCP server that fetches YouTube video transcripts and optionally summarizes them. Supports multiple transcript formats (text, JSON, SRT, WebVTT), multi-language retrieval, and flexible YouTube URL parsing.66MIT
- AlicenseAqualityDmaintenanceA powerful MCP server for summarizing YouTube videos with transcript fetching, intelligent summarization, key point extraction, and metadata retrieval.4MIT