Skip to main content
Glama

web_video

Read-only

Extract captions and chapters from video pages into a single [m:ss] timeline; use instead of web_fetch on video pages. Returns content, captions, chapters, notes; flags DRM streams.

Instructions

Read a video: captions and chapters as one [m:ss] timeline, from a watch page, an embedded player, tracks or the stream manifest. Use instead of web_fetch on a video page, which returns only the description. Returns content, captions (manual, auto or asr), chapters, notes; an encrypted stream is reported as drm, never read.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesA video page, a page embedding a player, or a player address.
langNoCaption language to prefer, e.g. "es".
cursorNoThe `cursor` of a truncated result.
framesNoKeyframes to capture where the picture changes, up to 24; paths land in the timeline.
profileNoProfile saved by web_login whose cookies to use.
out_fileNoWrite the timeline here and return the path.
max_tokensNoCap on the timeline returned. Default 25000.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv1.1.0

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, covering safety. The description adds behavioral detail beyond that: it states the exact return fields (content, captions, chapters, notes) and the critical behavior that an encrypted stream is reported as 'drm' and never read. This is meaningful additional context, though not exhaustive (e.g., no mention of pagination or auth requirements).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no filler. The core function is front-loaded, the alternative is stated succinctly, and the return fields are listed. Every sentence earns its place, and the structure is clean.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (7 parameters, no output schema), the description covers the essential context: what it does, when to use it, what it returns, and a key limitation (DRM). It does not explain every parameter, but those are covered by the schema. The return field list compensates for the missing output schema. Minor gaps like pagination behavior are implied by the cursor parameter, so the description is largely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% – every parameter has a clear description in the schema. The tool description does not add extra parameter-level meaning beyond what the schema already provides, so the baseline of 3 applies. It does not introduce any new parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Read'), a clear resource ('video'), and the output format (captions and chapters as a [m:ss] timeline). It also distinguishes itself from the sibling web_fetch by explicitly noting that web_fetch only returns the description, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Use instead of web_fetch on a video page, which returns only the description,' providing a direct alternative and the condition for choosing it. It also lists the accepted input sources (watch page, embedded player, <video> tracks, stream manifest), which clarifies the scope of use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.