Skip to main content
Glama
nohosa001-pixel

CleanWeb x402 — Smart Web Scraping & YouTube AI Agent

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault

No arguments

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}
prompts
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}
experimental
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
get_payment_infoA

Returns complete Web3 x402 micropayment configuration, supported multi-chain USDC contract addresses, EVM network chain IDs (Polygon: 137, Base: 8453, Arbitrum: 42161), recipient wallet address, and pricing tiers for all CleanWeb Studio agent tools.

Usage Guidelines:

  • Use this tool to discover network parameters and deposit requirements before making x402 paid queries.

  • Returns: Structured pricing markdown table, contract addresses, and pre-funded vault endpoints.

  • Do NOT use for checking individual wallet balances (use get_vault_balance).

clean_web_contentA

Scrapes and converts any target web page into clean, LLM-ready structured Markdown, stripping ads, cookie banners, navigation clutter, modals, and script noise.

Usage Guidelines:

  • Use this tool to ingest real-time web articles, blogs, and documentation into LLM context windows.

  • Returns: Clean markdown body, page title, word count, and extraction metadata.

  • Do NOT use for YouTube video parsing (use clean_youtube_transcript).

  • Do NOT use for PDF whitepapers or academic papers (use clean_pdf_research).

  • Do NOT use for paywalled, login-required, or bot-blocked sites.

clean_youtube_transcriptA

Extracts high-precision subtitles, timestamped transcripts, and comprehensive AI summaries for public YouTube videos using Google Gemini Flash intelligence.

Usage Guidelines:

  • Use this tool to ingest YouTube lecture, tutorial, tech talk, or podcast transcripts into agent workflows.

  • Returns: Video metadata (title, channel, URL), AI Knowledge Summary, and cleaned transcript.

  • Do NOT use for general web pages or articles (use clean_web_content).

  • Do NOT use for PDF documents or papers (use clean_pdf_research).

  • Do NOT use for private, unlisted, age-restricted, or live streams without existing closed captions.

  • If captions are missing or auto-captions fail, the tool reports a detailed fallback error.

clean_pdf_researchA

Parses and extracts structured plain text, sections, and academic metadata from online PDF whitepapers and research papers.

Usage Guidelines:

  • Use this tool to ingest scientific papers (e.g., arXiv), technical documentation, or financial reports.

  • Constraint: Target document must be a direct HTTP/HTTPS URL pointing to a PDF file under 15MB.

  • Returns: Title, total/parsed page count, word count, and extracted text.

  • Do NOT use for general HTML web pages (use clean_web_content).

  • Do NOT use for YouTube videos (use clean_youtube_transcript).

  • Do NOT use for password-protected, DRM-encrypted, or scanned image-only PDFs without OCR.

get_vault_balanceA

Checks the remaining pre-funded USDC balance, total usage, and session status for an agent wallet address or session key.

Usage Guidelines:

  • Use this tool before executing heavy tasks to verify sufficient balance for zero-latency execution.

  • Returns: Agent address, available balance in USDC, total deposited, total consumed, queries handled, and session key.

  • Do NOT use for querying pricing or chain parameters (use get_payment_info).

oracle_groundingA

Executes real-time web search, noise-free markdown extraction, Gemini AI JSON structuring, and cryptographically signs the result with an on-chain verifiable EIP-712 attestation (0.035 USDC).

Usage Guidelines:

  • Use this tool when an autonomous agent or smart contract requires verified, tamper-proof ground truth from the live web.

  • Returns: Structured facts JSON, human/LLM readable summary markdown, source URLs, and EIP-712 cryptographic signature (v, r, s).

  • Smart contracts can verify this off-chain or on-chain using CleanWebOracleVerifier.sol.

verify_oracle_attestationA

Verifies an EIP-712 cryptographic attestation produced by CleanWeb Oracle off-chain without consuming gas.

Usage Guidelines:

  • Use this tool to mathematically verify that data received from CleanWeb Oracle has not been tampered with.

  • Returns: Verification status (valid/invalid), recovered signer address, expected oracle address, and message.

clean_text_rawB

Extracts pure, tag-free plain text optimized for RAG embedding and vector indexing (0.001 USDC).

Usage Guidelines:

  • Use this tool when ingesting raw webpage text directly into vector databases (Pinecone, Chroma, Qdrant).

  • Returns: Page title, word count, clean text, and token metrics.

map_siteA

Discovers domain sitemap or traverses internal anchor links to return a canonical URL tree (0.002 USDC).

Usage Guidelines:

  • Use this tool to map an entire website or documentation site before scraping.

  • Firecrawl /map equivalent for autonomous web navigation agents.

  • Returns: Target URL, domain, total URL count, list of discovered URLs, and sitemap detection status.

search_web_quickB

Performs fast real-time keyword web search returning titles, links, and text snippets (0.002 USDC).

Usage Guidelines:

  • Tavily / s.jina.ai competitor designed specifically for LLM autonomous agent retrieval.

  • Returns: Concise search result list with title, verified URL, and text snippet.

extract_json_schemaA

Extracts schema-constrained structured JSON data from any webpage using Gemini AI (0.030 USDC).

Usage Guidelines:

  • Use when an agent needs structured attributes (e.g. pricing, specs, event dates) directly from a URL.

  • Returns: Clean JSON dictionary matching the requested schema description.

deep_research_topicA

Performs multi-source web crawling and AI synthesis to generate an executive research briefing (0.150 USDC).

Usage Guidelines:

  • Use when an agent needs a comprehensive deep dive into a topic with verified source citations.

  • Returns: Full executive briefing markdown with source citations and key takeaways.

clean_batch_scrapeA

Concurrently scrapes and extracts clean markdown from up to 10 URLs in parallel with high-speed async processing (0.005 USDC).

Usage Guidelines:

  • Use when an agent needs to perform multi-source research across multiple search results simultaneously.

  • Returns: Formatted summary and content preview of parsed web documents.

get_pass_statusA

Checks the active subscription status, remaining query quota, and validity period for an agent EVM wallet address (0x...) or Agent VIP Pass Token.

Usage Guidelines:

  • Use to verify micropayment allowance or query entitlements before dispatching heavy scrape batches.

  • Returns: Pass tier, remaining balance, and expiration timestamp.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

A4/5.0

Scored across 14 tools

Disambiguation5/5

Every tool targets a distinct resource and operation: separate scrapers for web, YouTube, PDF, batch, and raw text; distinct payment/status/balance tools; clear separation between quick search, deep research, and oracle-grounded search. The descriptions explicitly cross-reference other tools with 'Do NOT use' guidance, eliminating ambiguity.

Naming Consistency4/5

Most names follow a consistent verb_noun pattern (get_payment_info, clean_web_content, map_site, verify_oracle_attestation). Minor deviations include 'oracle_grounding' (noun_noun) and 'clean_batch_scrape' (verb-adjective-noun), which are still understandable but slightly break the pattern.

Tool Count5/5

14 tools is well within the ideal 3–15 range and each tool earns its place by covering a distinct aspect of web research, extraction, oracle verification, and payment management. The count feels substantial but not bloated.

Completeness5/5

The tool surface is remarkably comprehensive for its stated domain: single and batch scraping, YouTube and PDF handling, raw text extraction, quick search, deep research, site mapping, JSON schema extraction, oracle signing/verification, and payment/pass/balance queries. No obvious dead ends or missing lifecycle operations are evident.

Maintenance

ActivityActive
ResponsivenessNo issues