Skip to main content
Glama

klax_get_video_transcript

Fetches subtitles or STT transcripts for lecture videos from indexed material IDs or local paths, with optional timecode filtering. Falls back to local speech-to-text when subtitle files are missing.

Instructions

색인된 강의자료(material_id) 또는 로컬 동영상 경로의 자막/STT 전사를 타임코드 구간으로 조회합니다.

material_id_or_video_id가 로컬 색인 저장소(SQLite)에 존재하면 저장된 time:HH:MM:SS 청크를 사용하고, 존재하지 않으면 로컬 동영상 경로로 간주해 같은 파일명의 자막(.vtt/.srt) 또는 로컬 STT(faster-whisper/whisper)로 즉시 전사를 시도합니다. start_time/end_time(HH:MM:SS 또는 MM:SS)을 지정하면 해당 구간과 겹치는 청크만 반환합니다. STT 라이브러리 미설치 시 unsupported 상태와 사유를 반환하며, 원격 URL 다운로드나 외부 클라우드 API는 호출하지 않습니다.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
end_timeNo
start_timeNo
material_id_or_video_idYes

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden. It discloses the fallback behavior (local video path if not indexed), the STT attempt with faster-whisper/whisper, the unsupported status when STT library is missing, and explicitly states it does not download remote URLs or call external cloud APIs. This is substantial behavioral disclosure beyond what annotations would provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single paragraph but front-loaded with the core purpose, then systematically explains fallback, time filtering, and constraints. Every sentence adds value; there is no filler. Slightly long but appropriately detailed for a tool with nuanced behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (indexed vs local, STT fallback, time-range filtering) and the presence of an output schema, the description covers the key scenarios and error conditions (unsupported STT). It does not exhaustively cover edge cases like invalid local paths, but the main use cases are well documented.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It explains the dual meaning of material_id_or_video_id (indexed ID or local path), and specifies the accepted time formats (HH:MM:SS or MM:SS) for start_time/end_time. It also clarifies how time ranges filter chunks. This adds meaning well beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves subtitles/STT transcripts for indexed lecture materials or local video paths, with optional time-range filtering. It uses a specific verb (조회/query) and resource (transcript), and while it doesn't explicitly name sibling tools, the scope is unambiguous and distinct from any other sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use the tool (when transcript of indexed material or local video is needed) and how it chooses between the two sources based on existence in the index. It does not mention alternatives or exclusions, but the context is clear enough for an agent to decide.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.