Skip to main content
Glama

process_url

Download one public video or audio URL from YouTube, TikTok, or direct links, then extract transcript, keyframes, and OCR locally in a single job.

Instructions

Download ONE public video/audio URL once (this is the only tool that uses the network), then run the same LOCAL pipeline as process_media: transcript, keyframes, OCR, wall-clock, optional diarization. reused describes the pipeline cache, NOT network activity: a new URL is downloaded to compare bytes even when reused=true; a local rebuild can have reused=false with no download. Read origin.network and origin.reused_url_mapping for the download outcome. An indexed job with an unreadable manifest can rebuild from its verified local source; manifest_recovery_note explains recovery. list_jobs omits unreadable manifests. Lost provider metadata stays unknown unless refresh=true. Supported: direct https:// links to a media file (mp4/mov/webm/mkv/ogv/m4a/mp3/wav/ogg/flac), one public YouTube video (watch, youtu.be, shorts, a completed live), and any public video PAGE yt-dlp can read — Instagram (public reels/posts), TikTok, Wikimedia Commons, pages with an HTML5/HLS player, other sites as far as their yt-dlp extractor works anonymously (Vimeo does not). Not supported: playlists, channels, active live streams, private/members-only/age-restricted/DRM videos, cookies/logins; sites that hide a video behind a login or a bot wall fail with a clear reason. The downloaded source is kept inside the job, so extract_frame works later without network; the raw URL is never stored (only a hash, the provider id/host and a bounded title). A repeat call on the same URL serves the stored job without touching the network unless refresh=true. Job ids stay content hashes: the same video from two URLs is one job. YouTube and other pages need the optional [url] extra. The provider's upload date is NOT the recording start: wall_clock uses usable container creation metadata or recorded_at; download mtime is never used. When NOT to use: for local files (process_media), or to re-fetch data you already processed (use the retrieval tools). Examples:

  • process_url(url="https://youtu.be/nHfGfEiVdE8") — one public YouTube video, defaults are right

  • process_url(url="https://www.youtube.com/watch?v=ID&list=PL...") — the playlist part is ignored: ONE video

  • process_url(url="https://cdn.example.com/recordings/standup.mp4") — direct https link to a media file

  • process_url(url="https://www.tiktok.com/@nasa/video/7…") — a public video page; origin.provider names the site

  • meeting from a link: process_url(url=..., diarize=true, num_speakers=3, vocabulary="Vera, Tom, OKR")

  • non-English narration: process_url(url=..., model="large-v3-turbo", language="ru")

  • known recording start: process_url(url=..., recorded_at="2026-09-05T14:00:00+02:00") — enables t_wall

  • the video changed on the provider → process_url(url=..., refresh=true): new download, maybe a new job_id

  • re-anchor or change the model on a stored URL job → process_url(url=..., recorded_at=..., force=true), no download

  • error mentions [url] → run uvx --python ">=3.11,<3.14" "talkthrough-mcp[diarization,url]" and restart

  • playlist / channel / live / private URL → clear error; pass a single public video URL instead

  • "bot check"/"sign-in" refusal on Instagram/TikTok → the site blocked anonymous access; report it, no workaround

  • origin.published_at is the provider's upload time, not when the recording was made — never use it as t_wall

  • after success continue with get_transcript / search / get_moment on the job_id — never re-download

  • anti-example: a file on disk → process_media(path=...); process_url is only for https URLs

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYes
forceNo
modelNo
diarizeNo
refreshNo
languageNo
vocabularyNo
recorded_atNo
num_speakersNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv0.4.0

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only give the coarse profile (openWorld, non-destructive, non-readonly); the description adds far more: only this tool touches the network, cache semantics of reused vs refresh, that the source is kept in-job so extract_frame works offline, that the raw URL is never stored, manifest recovery behavior, and the wall_clock vs published_at distinction. This is behavior the agent cannot infer from the annotations or schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded purpose, then clearly labeled Supported / Not supported / When NOT to use / Examples sections. It is long but the length is justified by the complexity; some of the reuse/refresh prose repeats itself and a couple of example lines are redundant, so not a perfect 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained; the description instead covers what an agent still needs — network side effects, caching/reuse rules, failure modes (playlist/live/private/bot wall), missing-extra remediation, and where to continue afterwards. Nothing material is left unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 9 params, so the description must carry the burden and largely does — examples demonstrate model, language, recorded_at, diarize, num_speakers, vocabulary, refresh, and force with their intent (e.g. refresh re-downloads, force re-anchors without download). A few params (notably url format constraints) are covered indirectly rather than defined, keeping this short of a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific verb+resource: 'Download ONE public video/audio URL once ... then run the same LOCAL pipeline as process_media.' It explicitly distinguishes itself from process_media (local files) and the retrieval tools (re-fetch), so an agent can place it precisely among its siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is an explicit 'When NOT to use' section routing local files to process_media and re-fetching to retrieval tools, plus a Supported / Not supported enumeration and an anti-example. Conditions for refresh=true and force=true are spelled out with the exact outcomes they select.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.