Skip to main content
Glama
nohosa001-pixel

CleanWeb x402 — Smart Web Scraping & YouTube AI Agent

clean_youtube_transcript

Extracts high-precision subtitles, timestamped transcripts, and comprehensive AI summaries for public YouTube videos using Google Gemini Flash intelligence.

Instructions

Extracts high-precision subtitles, timestamped transcripts, and comprehensive AI summaries for public YouTube videos using Google Gemini Flash intelligence.

Usage Guidelines:

  • Use this tool to ingest YouTube lecture, tutorial, tech talk, or podcast transcripts into agent workflows.

  • Returns: Video metadata (title, channel, URL), AI Knowledge Summary, and cleaned transcript.

  • Do NOT use for general web pages or articles (use clean_web_content).

  • Do NOT use for PDF documents or papers (use clean_pdf_research).

  • Do NOT use for private, unlisted, age-restricted, or live streams without existing closed captions.

  • If captions are missing or auto-captions fail, the tool reports a detailed fallback error.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesPublic YouTube video URL (standard watch, short youtu.be, or Shorts format).
langNoComma-separated ISO 639-1 language priority codes for transcript extraction (e.g., 'ko,en', 'en', 'ja').ko,en
auth_token_or_txNoOptional x402 micropayment authorization token or EVM transaction hash.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed7 schema fields changedv1.2.6
    • addedInput schema / properties / auth_token_or_tx / description
      Added value: +"Optional x402 micropayment authorization token or EVM transaction hash."
    • addedInput schema / properties / lang / description
      Added value: +"Comma-separated ISO 639-1 language priority codes for transcript extraction (e.g., 'ko,en', 'en', 'ja')."
    • addedInput schema / properties / lang / examples
      Added value: +[
      +  "ko,en",
      +  "en",
      +  "ja,en"
      +]
    • addedInput schema / properties / lang / pattern
      Added value: +"^[a-z]{2}(,[a-z]{2})*$"
    • addedInput schema / properties / url / description
      Added value: +"Public YouTube video URL (standard watch, short youtu.be, or Shorts format)."
    • addedInput schema / properties / url / examples
      Added value: +[
      +  "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
      +  "https://youtu.be/dQw4w9WgXcQ",
      +  "https://www.youtube.com/shorts/abcdef12345"
      +]
    • addedInput schema / properties / url / pattern
      Added value: +"^https?:\\/\\/(www\\.)?(youtube\\.com\\/(watch\\?v=|shorts\\/)|youtu\\.be\\/)[\\w-]+.*$"
  2. Addedv1.2.5

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose return contents, unsupported inputs (private/unlisted/age-restricted/live), and fallback error reporting. However, it is silent on a major behavioral trait implied by the schema: the x402 micropayment/auth token requirement, which an agent must satisfy before calling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core capability, then uses labeled sections for usage and returns, which is easily scannable. Slightly padded by marketing phrasing like 'Google Gemini Flash intelligence' and 'high-precision', which don't affect selection.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return details are optional, yet the description still summarizes outputs and covers input restrictions and failure modes thoroughly. The one real omission is the payment/authorization context tied to auth_token_or_tx.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with the url pattern, lang default and examples, and auth_token_or_tx all documented in the schema itself. The description adds no syntax or format detail beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (extracts) and resource (subtitles, timestamped transcripts, AI summaries) for public YouTube videos. Explicitly differentiates itself from the named sibling tools clean_web_content and clean_pdf_research.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides an explicit when-to-use list (lectures, tutorials, tech talks, podcasts) plus three concrete 'Do NOT use' exclusions that route the agent to the correct alternative and warn about unsupported video types. The fallback-error behavior is also stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.