Skip to main content
Glama
nohosa001-pixel

CleanWeb x402 — Smart Web Scraping & YouTube AI Agent

clean_pdf_research

Extract structured text, sections, and metadata from online research PDFs and whitepapers via direct HTTP/HTTPS URLs. Returns title, page counts, word count, and parsed content for analysis.

Instructions

Parses and extracts structured plain text, sections, and academic metadata from online PDF whitepapers and research papers.

Usage Guidelines:

  • Use this tool to ingest scientific papers (e.g., arXiv), technical documentation, or financial reports.

  • Constraint: Target document must be a direct HTTP/HTTPS URL pointing to a PDF file under 15MB.

  • Returns: Title, total/parsed page count, word count, and extracted text.

  • Do NOT use for general HTML web pages (use clean_web_content).

  • Do NOT use for YouTube videos (use clean_youtube_transcript).

  • Do NOT use for password-protected, DRM-encrypted, or scanned image-only PDFs without OCR.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesDirect HTTP/HTTPS URL pointing to an online PDF document.
max_pagesNoMaximum number of pages to parse (1 to 100, default: 30) to control token budget.
auth_token_or_txNoOptional x402 micropayment authorization token or EVM transaction hash.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed7 schema fields changedv1.2.6
    • addedInput schema / properties / auth_token_or_tx / description
      Added value: +"Optional x402 micropayment authorization token or EVM transaction hash."
    • addedInput schema / properties / max_pages / description
      Added value: +"Maximum number of pages to parse (1 to 100, default: 30) to control token budget."
    • addedInput schema / properties / max_pages / maximum
      Added value: +100
    • addedInput schema / properties / max_pages / minimum
      Added value: +1
    • addedInput schema / properties / url / description
      Added value: +"Direct HTTP/HTTPS URL pointing to an online PDF document."
    • addedInput schema / properties / url / examples
      Added value: +[
      +  "https://arxiv.org/pdf/1706.03762.pdf",
      +  "https://bitcoin.org/bitcoin.pdf"
      +]
    • addedInput schema / properties / url / pattern
      Added value: +"^https?:\\/\\/[^\\s/$.?#].[^\\s]*\\.pdf(\\?.*)?$"
  2. Addedv1.2.5

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It discloses the 15MB size limit, that password-protected/DRM/scan-only PDFs won't work without OCR, and summarizes return values. However, it doesn't disclose failure behavior on oversized inputs, rate limits, or how partial page parsing is signaled (only 'total/parsed page count' hints at it).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded opening sentence and clearly labeled sections (Usage Guidelines, Constraint, Returns, Do NOT use). Slightly verbose with the multiple Do-NOT bullets, but every bullet adds routing value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema disclosed in the prompt beyond a mention, the description compensates by naming what is returned (title, page counts, word count, text). Combined with the constraint and negative routing, an agent has everything needed to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so the schema documents url, max_pages, and auth_token_or_tx. The description adds meaningful constraints beyond the schema: the required direct-PDF URL form and the 15MB limit. It doesn't explain when to supply auth_token_or_tx, but the schema description covers that.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'parses and extracts structured plain text, sections, and academic metadata' from PDFs. Distinguishes itself from siblings clean_web_content and clean_youtube_transcript by naming them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use it (scientific papers, technical docs, financial reports) and provides three Do-NOT-use clauses naming the correct alternative sibling for each case (HTML → clean_web_content, YouTube → clean_youtube_transcript, DRM/scanned → OCR caveat). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.