yueying
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| HF_HOME | No | Hugging Face cache (Whisper weights live here) | |
| PYTHONUTF8 | No | set to 1 on Windows to avoid mojibake | |
| HF_ENDPOINT | No | mirror, e.g. https://hf-mirror.com | huggingface.co |
| YUEYING_LANG | No | language of report.md written by the server (en/zh) | en |
| YUEYING_MODEL | No | default for the model parameter | auto |
| YUEYING_DEVICE | No | auto / cuda / cpu | auto |
| YUEYING_OUT_DIR | No | root folder for results (absolute, ~ ok) | ~/yueying_out |
| YUEYING_MAX_JOBS | No | pipelines running at once per server | 1 |
| YUEYING_JOB_TIMEOUT | No | hard limit per video, seconds | 7200 |
| YUEYING_KEEP_SOURCE | No | 1 keeps the downloaded ≤720p source in _download/ (enables exact-moment frames for URLs) | |
| YUEYING_COOKIES_FROM_BROWSER | No | default browser for cookies (chrome, edge, firefox, …) |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| watch_videoA | Turn a video into a timestamped transcript plus a keyframe overview so you can summarize it, answer questions about it, extract steps, commands or code, or write notes. |
| get_transcriptA | Read part of an already-watched video's transcript with [mm:ss] timestamps. Use when the watch_video overview was truncated, when the user asks about a specific time range, or to export subtitles (format='srt'). Returns at most max_chars; when truncated the last line gives next_start so you can continue from there. Paragraph format is the cheapest. |
| search_transcriptA | Find where something is said in an already-watched video. Each hit shows the time, the surrounding sentences, and the nearest keyframe number and contact-sheet number so you can follow up with get_frame_at or get_transcript. Use this instead of paging the whole transcript when the user asks 'when does he mention X' or 'find the part about Y'. |
| get_framesA | See what is on screen in an already-watched video. Returns contact-sheet images (3x3 keyframes in time order, every tile labelled '#number mm:ss' bottom-left) or individual keyframes. Read the contact sheets first to get the visual storyline, then request single frames only when you need to read code, slides or UI text. At most 3 images per call (default 2), downscaled to max_width; page with start/count. Every image is preceded by its absolute file path so hosts that can read files may open the full-size original instead. |
| get_frame_atA | Look closely at one moment of an already-watched video, e.g. to read code, a slide, a chart or a UI. For local files that still exist the exact frame at that time is extracted from the video; otherwise the nearest cached keyframe is returned and the caption says so. Also returns the transcript paragraphs spoken around that time. Returns one image (~100 KB at 960 px). |
| list_videosA | List videos already processed by yueying on this machine (newest first) and jobs currently running, with video_id, title, duration, text source, date, folder and size. Use when the user refers to a video watched earlier, to get a video_id for the other tools, or to see how much disk space results use. Instant and read-only. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 6 tools
Each tool targets a distinct step in the video processing workflow: create (watch_video), read transcript (get_transcript), search (search_transcript), visual overview (get_frames), precise frame lookup (get_frame_at), and listing (list_videos). There is no meaningful overlap between any two tools; the boundaries are clear.
All tool names follow a consistent imperative verb_noun snake_case convention (watch_video, get_transcript, list_videos, search_transcript, get_frames, get_frame_at). Even 'get_frame_at' is a predictable extension of the get_ pattern, so naming is uniform and easy to guess.
Six tools is well-scoped for a video analysis server. Each tool covers a necessary operation without redundancy, and the count is appropriate for the domain.
The core workflow is well covered: ingest a video, retrieve transcript segments, search within them, and view visual content. The only notable gap is the absence of a delete/cleanup tool to remove processed videos or cached results, though list_videos does expose disk usage as a workaround.