Skip to main content
Glama
nohosa001-pixel

CleanWeb x402 — Smart Web Scraping & YouTube AI Agent

⚡ Multi-Chain & Solana x402 Autonomous AI Agent Suite & Spend Firewall (v2.7.0)

"High-Throughput Clean Web Markdown, YouTube Intelligence & Cryptographic Web3 Oracles for Autonomous AI Agents & Swarms."
Zero credit cards required. Pure B2A (Business-to-Agent) USDC micropayments & pre-funded vaults on Solana, Polygon, Base, and Arbitrum with Solscan Verified IDLs & EIP-712 spend policy attestations.

Version PyPI Tests MCP Multi-Chain License: MIT


🌐 Live Production Gateway & Endpoints


Related MCP server: fetch

🤖 1-Minute Framework Integration (ElizaOS, LangChain & CrewAI)

Equip your autonomous agents with x402 clean web extraction in under 1 minute:

# ElizaOS (v2) Autonomous Web3 Agent Action Provider
from examples.elizaos_plugin_example import ElizaOSCleanWebAction

action = ElizaOSCleanWebAction()
res = action.handler("https://paulgraham.com/greatwork.html")
print(res["markdown"])  # Clean markdown with 87% token reduction
# LangChain Agent Tool Example
from examples.langchain_agent_example import clean_web_tool

result = clean_web_tool.invoke({"url": "https://paulgraham.com/greatwork.html"})
print(result)  # Clean markdown with 87% token noise reduction
# CrewAI Researcher Tool Example
from examples.crewai_agent_example import clean_web_for_crewai

content = clean_web_for_crewai("https://paulgraham.com/greatwork.html")

⚡ 10-Second Instant Test (Zero Auth / 3 Free Calls Included)

No wallet, credit card, or API key needed. Attach an X-Agent-Nonce header to test immediately:

# Instant JSON Request via curl
curl -X POST "https://x402-cleanweb-agent-7qxtp3324q-du.a.run.app/api/v1/clean-web" \
     -H "Content-Type: application/json" \
     -H "X-Agent-Nonce: instant-trial-uuid-001" \
     -d '{"url": "https://news.ycombinator.com"}'
# Zero-dependency Python snippet
import urllib.request, json

req = urllib.request.Request(
    "https://x402-cleanweb-agent-7qxtp3324q-du.a.run.app/api/v1/clean-web",
    data=json.dumps({"url": "https://news.ycombinator.com"}).encode(),
    headers={"Content-Type": "application/json", "X-Agent-Nonce": "instant-trial-01"}
)
res = json.loads(urllib.request.urlopen(req).read())
print(f"Tokens Saved: {res.get('estimated_token_savings', '87%')}")
print("Clean Markdown Preview:\n", res["clean_markdown"][:300])

💡 Mathematical Token Arbitrage Proof (Economic Rationality)

Autonomous agents route web traffic through x402 because it mathematically saves the operator money on every query:

  • Direct Scraping Cost: 80KB HTML ≈ 20,000 tokens ($0.05 USD input cost on GPT-4o / Claude 3.5 Sonnet).

  • x402 CleanWeb Cost: $0.001 API fee + 2,600 clean tokens ($0.0065 USD) = $0.0075 USD Total.

  • Net Savings: $0.0425 USD saved per web query (85% net dollar reduction).

  • Service Fee ROI: 4,250% ROI on the $0.001 fee.

  • Live Arbitrage Endpoint: GET /api/v1/agent/arbitrage-roi?input_tokens=50000

{
  "status": "economically_optimal",
  "routing_recommendation": "ROUTE_VIA_X402",
  "financial_analysis": {
    "cost_raw_direct_usd": 0.125,
    "cost_with_x402_usd": 0.01725,
    "net_dollar_savings": 0.10775,
    "net_savings_percentage": "86.2%",
    "service_fee_roi_percentage": "10775.0%"
  }
}

🔮 14 Production Agent Tools & Micro-Pricing (USDC)

Tool Name

Endpoint / Function

Cost (USDC)

What It Does

🌐 Clean Web

GET/POST /api/v1/clean-web

0.001 USDC

Strips HTML boilerplate, ads, scripts; returns ad-free Markdown (87% token reduction)

📝 Clean Text

GET /api/v1/clean-text

0.001 USDC

Ultra-lightweight raw plain text for vector/RAG embeddings

🗺️ Site Mapper

GET /api/v1/map-site

0.002 USDC

Recursive sitemap and internal URL tree discovery

🔍 Agent Search

GET /api/v1/search

0.002 USDC

Fast real-time keyword search & verified snippets (Tavily engine)

📑 PDF Extractor

GET /api/v1/clean-pdf

0.005 USDC

Formula- and table-preserved academic paper & research PDF extractor

📦 Batch Clean

POST /api/v1/batch-clean

0.005 USDC

Concurrent parallel scraping for up to 10 URLs in a single request

🎬 YouTube AI

GET /api/v1/clean-youtube

0.010 USDC

Gemini 3.6 Flash multimodal video intelligence & timestamped audio transcript

📊 Extract JSON

POST /api/v1/extract-json

0.030 USDC

Converts arbitrary web content into strict schema-constrained JSON

🔮 Web3 Oracle

POST /api/v1/oracle/grounding

0.035 USDC

Live Search + Clean-to-JSON + EIP-712 Cryptographic On-Chain Attestation

🛡️ Oracle Verify

POST /api/v1/oracle/verify

0.000 USDC

Verifies ECDSA signature of CleanWeb Oracle attestations on-chain

🧠 Deep Research

GET /api/v1/deep-research

0.150 USDC

Multi-source synthesized AI executive research briefing

💼 Vault Deposit

POST /api/v1/vault/deposit

0.000 USDC

Multi-chain USDC vault prefunding for sub-5ms gasless executions

💳 Vault Balance

GET /api/v1/vault/balance

0.000 USDC

Real-time query of credit pass balance and validity

🎫 Pass Status

GET /api/v1/pass-status

0.000 USDC

Inspects status and remaining credits of a prepaid pass


💻 1-Click MCP Setup (Claude Desktop, Cursor, Windsurf)

Connect CleanWeb directly to Claude Desktop, Cursor, or any Model Context Protocol client using uvx:

{
  "mcpServers": {
    "polygon-x402-cleanweb": {
      "command": "uvx",
      "args": ["x402-cleanweb-agent"]
    }
  }
}

Or configure via local clone:

{
  "mcpServers": {
    "polygon-x402-cleanweb": {
      "command": "python",
      "args": ["-u", "mcp_server.py"],
      "env": {
        "POLYGON_RPC_URL": "https://polygon-bor-rpc.publicnode.com",
        "SERVER_WALLET_ADDRESS": "0xA185B43fDD19619f99952AAed6eabf1029bF36a1",
        "USDC_CONTRACT_ADDRESS": "0x3c499c542cEF5E3811e1192ce70d8cC03d5c3359"
      }
    }
  }
}

🤖 Multi-Agent Framework Integrations

Drop-in ready integrations for every major autonomous agent framework via GET /api/v1/agent/integrations/{framework}:

🦜 LangChain / LangGraph

from langchain.tools import tool
import requests

@tool
def clean_web(url: str) -> str:
    """Fetches a URL and returns ad-free, token-optimized Markdown (87% token savings)."""
    res = requests.get(
        "https://x402-cleanweb-agent-7qxtp3324q-du.a.run.app/api/v1/clean-web",
        params={"url": url},
        headers={"X-Agent-Nonce": "langchain-agent-session"}
    )
    return res.json().get("clean_markdown", "")

👥 CrewAI

from crewai.tools import tool
import requests

@tool("CleanWeb Tool")
def clean_web(url: str) -> str:
    """Strips HTML boilerplate and returns pure Markdown to save LLM context window."""
    res = requests.post(
        "https://x402-cleanweb-agent-7qxtp3324q-du.a.run.app/api/v1/clean-web",
        json={"url": url},
        headers={"X-Agent-Nonce": "crewai-agent-session"}
    )
    return res.json().get("clean_markdown", "")

🤖 Microsoft AutoGen

import requests

def clean_web_tool(url: str) -> str:
    res = requests.get(
        "https://x402-cleanweb-agent-7qxtp3324q-du.a.run.app/api/v1/clean-web",
        params={"url": url},
        headers={"X-Agent-Nonce": "autogen-session"}
    )
    return res.json().get("clean_markdown", "")

# assistant.register_for_llm(name="clean_web", description="Clean web markdown")(clean_web_tool)

🏛️ Verified On-Chain Smart Contracts & Accounts (Polygon, Base, Arbitrum, Solana)

CleanWeb operates verified, deterministic smart contracts and native settlement accounts for on-chain EIP-712 oracle verification, autonomous agent pre-funded vault management, and Solana SPL USDC micropayments.

Contract / Account

Polygon Mainnet (137) ✔

Base Mainnet (8453) ✔

Arbitrum One (42161) ✔

Solana Mainnet-Beta (101) ✔

AgentPaymentVault

0x45ecBf...1861

0x28292D...76DD

0x28292D...76DD

7oZ16Y...Wi3y (Solscan)

CleanWebOracleConsumer

0xAECbfB...6D66

0x2394d8...dabe

0x2394d8...dabe

21ZR1Q...bkTL (Solscan)

CleanWebOracleVerifier

0x18fA45...Dc46

0x3eD259...0740

0x3eD259...0740

9nVrym...iopC (Solscan)

Treasury Recipient Wallet

0xA185B4...36a1

0xA185B4...36a1

0xA185B4...36a1

411ksM...9qp (Solscan)

Native USDC Token Tracker

0x3c499c...3359

0x833589...2913

0xaf88d0...5831

EPjFWd...YTDt1v (Solscan)


🧪 Production Readiness Review: 43/43 Tests Passed (100%)

Verified across all 38+ FastAPI routes, 14 agent tools, multi-chain settlement vaults, and EIP-712 attestations:

python -m pytest tests/ -v
tests/test_agent_discovery.py::test_well_known_manifests PASSED              [  2%]
tests/test_agent_discovery.py::test_agent_capabilities_reflection PASSED     [  4%]
tests/test_agent_discovery.py::test_agent_pricing_catalog PASSED             [  6%]
tests/test_agent_discovery.py::test_agent_arbitrage_roi PASSED               [  9%]
tests/test_agent_discovery.py::test_agent_framework_integrations PASSED      [ 11%]
tests/test_diagnostics.py::test_run_full_diagnostic PASSED                  [ 13%]
tests/test_diagnostics.py::test_diagnostics_endpoint PASSED                 [ 16%]
tests/test_enhanced_agent_features.py (16 tests) PASSED                      [ 53%]
tests/test_oracle_grounding.py (3 tests) PASSED                              [ 60%]
tests/test_payment_comprehensive.py (4 tests) PASSED                         [ 69%]
tests/test_phase1_cleaners.py (4 tests) PASSED                               [ 79%]
tests/test_phase2_payments.py (5 tests) PASSED                               [ 90%]
tests/test_routes_coverage.py::test_all_api_routes_operational PASSED         [ 93%]
tests/test_treasury.py (3 tests) PASSED                                      [100%]

================== 43 passed, 2 warnings in 85.93s (100%) ==================

CleanWeb Studio operates strictly under the Transformative Non-Expressive Text/Data Mining (TDM) Fair Use doctrine for autonomous AI reasoning.

  1. No Financial or Investment Warranty: All oracle feeds, search results, and scraped contents are provided AS-IS. CleanWeb Studio assumes zero liability for DeFi smart contract liquidations, prediction market settlements (e.g. Polymarket), or trading losses resulting from external web misreporting or LLM hallucinations.

  2. Caller Responsibility: The autonomous agent deployer/caller is solely responsible for respecting target website intellectual property and applicable laws.

  3. Zero-Data Retention (GDPR Compliant): All fetched HTML and raw assets are processed ephemerally in RAM and purged immediately upon response transmission.

  4. 100% OFAC Sanctions Filtering: All micropayment addresses are verified against OFAC Specially Designated Nationals (SDN) lists.

  5. Full Legal Endpoints:

    • GET /api/v1/legal/terms: Official Machine-to-Machine Terms of Service

    • GET /api/v1/legal/disclaimer: Full Legal & Financial Disclaimer


📜 License

Distributed under the MIT License. See LICENSE for more information.

Available Tools

14 tools
clean_batch_scrapeA

Concurrently scrapes and extracts clean markdown from up to 10 URLs in parallel with high-speed async processing (0.005 USDC).

Usage Guidelines:

  • Use when an agent needs to perform multi-source research across multiple search results simultaneously.

  • Returns: Formatted summary and content preview of parsed web documents.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlsYesList of target URLs to scrape in parallel (up to 10).
auth_token_or_txNoOptional x402 auth token, vault key, or tx hash.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It does add useful context: concurrent and async processing, a cost of 0.005 USDC, and a return format of summary and content preview. However, it omits potential failure modes, rate limits, or side effects beyond the network operation. For a scraping tool, the lack of explicit read-only or error-handling information leaves some gaps, but the provided details are helpful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-organized, with the core function stated first, followed by a clear usage guideline and return summary. It avoids unnecessary filler and is front-loaded with the most critical information. The structure aids quick comprehension.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema and the tool's relatively simple purpose, the description covers the essential context: when to use it, what it does, and the output format. The only minor gap is the lack of explicit differentiation from sibling tools, but the batch aspect is clear from the 'up to 10 URLs' and the usage guideline. Overall, an agent can correctly invoke this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema covers 100% of the parameters with descriptions: 'urls' lists up to 10 URLs, and 'auth_token_or_tx' is optional. The tool description adds no new meaning beyond restating the limit ('up to 10 URLs') that is already in the schema. Baseline 3 is appropriate since the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: it concurrently scrapes and extracts clean markdown from up to 10 URLs. The verb 'scrapes' and resource 'URLs' are specific, and the 'up to 10' limit distinguishes it from single-URL tools like clean_web_content. The purpose is unambiguous and immediately actionable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description includes an explicit usage guideline: 'Use when an agent needs to perform multi-source research across multiple search results simultaneously.' This provides clear context for when to invoke the tool, though it does not explicitly list alternatives or when not to use it. The guidance is sufficient for typical scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clean_pdf_researchA

Parses and extracts structured plain text, sections, and academic metadata from online PDF whitepapers and research papers.

Usage Guidelines:

  • Use this tool to ingest scientific papers (e.g., arXiv), technical documentation, or financial reports.

  • Constraint: Target document must be a direct HTTP/HTTPS URL pointing to a PDF file under 15MB.

  • Returns: Title, total/parsed page count, word count, and extracted text.

  • Do NOT use for general HTML web pages (use clean_web_content).

  • Do NOT use for YouTube videos (use clean_youtube_transcript).

  • Do NOT use for password-protected, DRM-encrypted, or scanned image-only PDFs without OCR.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesDirect HTTP/HTTPS URL pointing to an online PDF document.
max_pagesNoMaximum number of pages to parse (1 to 100, default: 30) to control token budget.
auth_token_or_txNoOptional x402 micropayment authorization token or EVM transaction hash.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It discloses the 15MB size limit, that password-protected/DRM/scan-only PDFs won't work without OCR, and summarizes return values. However, it doesn't disclose failure behavior on oversized inputs, rate limits, or how partial page parsing is signaled (only 'total/parsed page count' hints at it).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded opening sentence and clearly labeled sections (Usage Guidelines, Constraint, Returns, Do NOT use). Slightly verbose with the multiple Do-NOT bullets, but every bullet adds routing value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema disclosed in the prompt beyond a mention, the description compensates by naming what is returned (title, page counts, word count, text). Combined with the constraint and negative routing, an agent has everything needed to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so the schema documents url, max_pages, and auth_token_or_tx. The description adds meaningful constraints beyond the schema: the required direct-PDF URL form and the 15MB limit. It doesn't explain when to supply auth_token_or_tx, but the schema description covers that.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'parses and extracts structured plain text, sections, and academic metadata' from PDFs. Distinguishes itself from siblings clean_web_content and clean_youtube_transcript by naming them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use it (scientific papers, technical docs, financial reports) and provides three Do-NOT-use clauses naming the correct alternative sibling for each case (HTML → clean_web_content, YouTube → clean_youtube_transcript, DRM/scanned → OCR caveat). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clean_text_rawB

Extracts pure, tag-free plain text optimized for RAG embedding and vector indexing (0.001 USDC).

Usage Guidelines:

  • Use this tool when ingesting raw webpage text directly into vector databases (Pinecone, Chroma, Qdrant).

  • Returns: Page title, word count, clean text, and token metrics.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe target website URL to extract raw text from.
auth_token_or_txNoOptional x402 auth token or tx hash.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It usefully discloses the monetary cost (0.001 USDC) and that the operation returns page title, word count, text and token metrics, which suggests a read-only extraction. It still omits whether the optional auth_token_or_tx is ever required, any rate limits, or failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well front-loaded: purpose in the first clause with price attached, then a short labeled usage block. Every line is short and scannable, with only a minor redundancy in restating return fields that an output schema already covers.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is not strictly needed, and cost context is a genuine addition. The main gap is sibling disambiguation within a crowded family of clean_* tools and any auth requirement for the optional token parameter, leaving the definition adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the url and auth_token_or_tx parameters are already documented in the schema. The description adds no syntax, format, or auth semantics beyond it, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Extracts ... plain text') and adds the intended downstream use ('optimized for RAG embedding and vector indexing'). However, it never distinguishes itself from the sibling clean_web_content, which appears to overlap heavily, so an agent cannot tell the two apart from the text alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives implied context ('when ingesting raw webpage text directly into vector databases') which is a real usage signal. But it offers no when-not guidance and, critically, does not route the agent to or away from clean_web_content, clean_pdf_research, or clean_youtube_transcript, so the alternative-selection decision is left unresolved.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clean_web_contentA

Scrapes and converts any target web page into clean, LLM-ready structured Markdown, stripping ads, cookie banners, navigation clutter, modals, and script noise.

Usage Guidelines:

  • Use this tool to ingest real-time web articles, blogs, and documentation into LLM context windows.

  • Returns: Clean markdown body, page title, word count, and extraction metadata.

  • Do NOT use for YouTube video parsing (use clean_youtube_transcript).

  • Do NOT use for PDF whitepapers or academic papers (use clean_pdf_research).

  • Do NOT use for paywalled, login-required, or bot-blocked sites.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe target HTTP or HTTPS website URL to scrape and convert to markdown.
auth_token_or_txNoOptional x402 micropayment authorization token or EVM transaction hash.
respect_robots_txtNoWhether to enforce target domain robots.txt Disallow rules (compliance-mode).

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does well: it discloses the cleaning behavior (ads, cookie banners, nav clutter stripped), the return shape, and the failure boundary (paywalled/login-required/bot-blocked sites). It stops short of mentioning that some sites may require the micropayment auth token, which the schema hint implies.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then a tightly scoped bulleted usage block. Every line earns its place and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists and the description still summarizes the return payload (markdown body, title, word count, metadata). For a single-URL scraping tool, nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters are already documented in the schema, including the x402 token and robots.txt flag. The description adds no further parameter semantics, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (scrapes/converts) and resource (web page), and explicitly enumerates what it strips. It distinguishes itself from siblings clean_youtube_transcript and clean_pdf_research by name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use (real-time articles, blogs, docs into LLM context) and three explicit when-NOT-to-use clauses naming the correct alternative tool for each. Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clean_youtube_transcriptA

Extracts high-precision subtitles, timestamped transcripts, and comprehensive AI summaries for public YouTube videos using Google Gemini Flash intelligence.

Usage Guidelines:

  • Use this tool to ingest YouTube lecture, tutorial, tech talk, or podcast transcripts into agent workflows.

  • Returns: Video metadata (title, channel, URL), AI Knowledge Summary, and cleaned transcript.

  • Do NOT use for general web pages or articles (use clean_web_content).

  • Do NOT use for PDF documents or papers (use clean_pdf_research).

  • Do NOT use for private, unlisted, age-restricted, or live streams without existing closed captions.

  • If captions are missing or auto-captions fail, the tool reports a detailed fallback error.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesPublic YouTube video URL (standard watch, short youtu.be, or Shorts format).
langNoComma-separated ISO 639-1 language priority codes for transcript extraction (e.g., 'ko,en', 'en', 'ja').ko,en
auth_token_or_txNoOptional x402 micropayment authorization token or EVM transaction hash.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose return contents, unsupported inputs (private/unlisted/age-restricted/live), and fallback error reporting. However, it is silent on a major behavioral trait implied by the schema: the x402 micropayment/auth token requirement, which an agent must satisfy before calling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core capability, then uses labeled sections for usage and returns, which is easily scannable. Slightly padded by marketing phrasing like 'Google Gemini Flash intelligence' and 'high-precision', which don't affect selection.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return details are optional, yet the description still summarizes outputs and covers input restrictions and failure modes thoroughly. The one real omission is the payment/authorization context tied to auth_token_or_tx.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with the url pattern, lang default and examples, and auth_token_or_tx all documented in the schema itself. The description adds no syntax or format detail beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (extracts) and resource (subtitles, timestamped transcripts, AI summaries) for public YouTube videos. Explicitly differentiates itself from the named sibling tools clean_web_content and clean_pdf_research.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides an explicit when-to-use list (lectures, tutorials, tech talks, podcasts) plus three concrete 'Do NOT use' exclusions that route the agent to the correct alternative and warn about unsupported video types. The fallback-error behavior is also stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

deep_research_topicA

Performs multi-source web crawling and AI synthesis to generate an executive research briefing (0.150 USDC).

Usage Guidelines:

  • Use when an agent needs a comprehensive deep dive into a topic with verified source citations.

  • Returns: Full executive briefing markdown with source citations and key takeaways.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesResearch topic or complex question to investigate.
max_sourcesNoNumber of top web sources to synthesize (1 to 5).
auth_token_or_txNoOptional x402 auth token or tx hash.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses the cost (0.150 USDC) and the return shape (executive briefing markdown with citations and takeaways), but says nothing about auth requirements (the x402 param), latency, failure modes, or source-verification behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well front-loaded: the core purpose and price lead, followed by labeled 'Usage Guidelines' and 'Returns' sections. Slightly redundant to restate the return format when an output schema exists, but the description is compact and free of filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be re-explained, and the cost plus primary use case are covered. The notable gap is the absence of annotations and any description of auth/permission behavior around auth_token_or_tx, which an agent paying per call would benefit from knowing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents query, max_sources (1-5), and auth_token_or_tx. The description adds no syntax, format, or defaulting guidance beyond what the schema provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific compound action (multi-source web crawling + AI synthesis) and a concrete deliverable (executive research briefing), which is clearly differentiated from the 'quick' sibling search_web_quick by the word 'deep'/comprehensive. It stops short of explicitly naming that sibling, so an agent must infer the contrast from the adjective alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Usage Guidelines' section gives a clear contextual trigger ('use when an agent needs a comprehensive deep dive ... with verified source citations'). However it never states when NOT to use it, nor names search_web_quick as the cheap/fast alternative, so the routing decision is left implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

extract_json_schemaA

Extracts schema-constrained structured JSON data from any webpage using Gemini AI (0.030 USDC).

Usage Guidelines:

  • Use when an agent needs structured attributes (e.g. pricing, specs, event dates) directly from a URL.

  • Returns: Clean JSON dictionary matching the requested schema description.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesTarget webpage URL to extract JSON from.
auth_token_or_txNoOptional x402 auth token or tx hash.
schema_descriptionYesDescription or format of the fields to extract (e.g., 'price, product_name, in_stock').

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden. It usefully discloses the underlying model (Gemini AI) and the cost (0.030 USDC), which is real behavioral value. However, it says nothing about auth/payment requirements despite an auth_token_or_tx parameter, nor about rate limits or failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded one-line purpose followed by clearly labeled Usage Guidelines and Returns sections. Every line is relevant; it is tight and scannable with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value detail is not required. The description covers purpose, usage context, cost, and return shape, which is sufficient for invocation. The main missing piece is any note on the payment/auth flow implied by auth_token_or_tx.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters are already documented in the schema, making baseline 3 appropriate. The description restates the 'schema_description' concept and the JSON return but adds no syntax or format detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb and resource: 'Extracts schema-constrained structured JSON data from any webpage,' and adds distinguishing detail (Gemini AI, 0.030 USDC). It does not explicitly differentiate itself from close siblings like clean_web_content, search_web_quick, or deep_research_topic, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

An explicit 'Use when...' condition with concrete examples (pricing, specs, event dates) directly from a URL. It lacks any when-not-to-use guidance or named alternatives among the many sibling web/content tools, so it stops short of the top score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_pass_statusA

Checks the active subscription status, remaining query quota, and validity period for an agent EVM wallet address (0x...) or Agent VIP Pass Token.

Usage Guidelines:

  • Use to verify micropayment allowance or query entitlements before dispatching heavy scrape batches.

  • Returns: Pass tier, remaining balance, and expiration timestamp.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_wallet_or_tokenYesAgent EVM wallet address (0x...) or Pass Token (e.g. 'WELCOME100').

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It discloses the return values (Pass tier, remaining balance, expiration timestamp) and the input types, but does not mention potential side effects, rate limits, or error conditions. For a read-only status check, this is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded with the core purpose, followed by a usage guideline and return summary. It is efficient, though the 'Returns:' line is somewhat redundant with the output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has a single parameter, an output schema, and no nested objects, the description is nearly complete. It explains the purpose, usage context, and return values. It could mention error cases or what happens if the wallet/token is invalid, but that is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents the parameter. The description adds a bit of context by mentioning the example 'WELCOME100' and the 0x... format, but it largely repeats what the schema provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool checks subscription status, remaining query quota, and validity period for an agent EVM wallet address or VIP Pass Token. It uses a specific verb ('Checks') and resource, and the usage guideline clarifies its role in verifying micropayment allowance before heavy scrape batches, distinguishing it from siblings like get_payment_info.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use the tool: 'Use to verify micropayment allowance or query entitlements before dispatching heavy scrape batches.' It does not explicitly name alternatives or exclusions, but the context is clear enough for an agent to select it appropriately among siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_payment_infoA

Returns complete Web3 x402 micropayment configuration, supported multi-chain USDC contract addresses, EVM network chain IDs (Polygon: 137, Base: 8453, Arbitrum: 42161), recipient wallet address, and pricing tiers for all CleanWeb Studio agent tools.

Usage Guidelines:

  • Use this tool to discover network parameters and deposit requirements before making x402 paid queries.

  • Returns: Structured pricing markdown table, contract addresses, and pre-funded vault endpoints.

  • Do NOT use for checking individual wallet balances (use get_vault_balance).

ParametersJSON Schema
NameRequiredDescriptionDefault
tierNoOptional specific pricing tier to inspect.
chainNoOptional specific chain name (polygon, base, arbitrum).

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden well. It states the tool is informational ('Returns', 'discover ... before making x402 paid queries') and discloses return format (structured markdown table, contract addresses, vault endpoints), implying a read-only preflight operation. It does not explicitly say 'does not initiate payment' or mention auth/rate limits, but those are less material for a config-retrieval tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The key return information is front-loaded in the first sentence, and the usage/do-not-use bullets are concise. Minor redundancy exists between the opening 'Returns...' and the 'Returns:' bullet, but the bullet adds useful return-format details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only configuration tool with no required parameters and an output schema, this is complete: it covers what data is returned, which chains are supported, why an agent would call it, and what not to use it for. No additional context is needed to select or invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both optional parameters already have clear descriptions in the schema, so the baseline applies. The tool description adds chain IDs (137, 8453, 42161) and context about pricing tiers, but it does not significantly expand on the tier or chain parameter inputs beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence names a specific action and result: returns complete Web3 x402 micropayment configuration, USDC contract addresses, chain IDs, recipient wallet, and pricing tiers for CleanWeb Studio tools. This clearly distinguishes it from sibling tools, especially get_vault_balance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit guidance: use this tool to discover network parameters and deposit requirements before x402 paid queries, and explicitly warns against using it for wallet balances, directing to get_vault_balance. This is exactly the when/when-not/alternative guidance the rubric asks for.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_vault_balanceA

Checks the remaining pre-funded USDC balance, total usage, and session status for an agent wallet address or session key.

Usage Guidelines:

  • Use this tool before executing heavy tasks to verify sufficient balance for zero-latency execution.

  • Returns: Agent address, available balance in USDC, total deposited, total consumed, queries handled, and session key.

  • Do NOT use for querying pricing or chain parameters (use get_payment_info).

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_address_or_keyYesEthereum/Polygon address (0x...) or session key (sk_...) to query balance for.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. The verb 'Checks' implies a read-only operation, and the return fields are listed, but it does not explicitly state safety traits (e.g., read-only, non-destructive) or authentication/rate-limit requirements. Partial transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with purpose, then cleanly organized into usage guidelines and returns. Every sentence adds value, and there is no redundant or filler text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only balance check with one parameter and an output schema, the description covers purpose, usage guidelines, and return fields. The output schema handles return values, so the description is complete without needing to explain them further.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents the single parameter thoroughly. The description restates 'for an agent wallet address or session key' but adds no additional syntax, format, or edge-case detail beyond the schema. Baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (checks) and resource (remaining pre-funded USDC balance, total usage, session status), scoped to an agent wallet address or session key. It distinguishes itself from the sibling get_payment_info by explicitly excluding pricing and chain parameters.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use guidance ('Use this tool before executing heavy tasks to verify sufficient balance') and when-not-to-use guidance ('Do NOT use for querying pricing or chain parameters (use get_payment_info)'), naming the alternative tool directly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

map_siteA

Discovers domain sitemap or traverses internal anchor links to return a canonical URL tree (0.002 USDC).

Usage Guidelines:

  • Use this tool to map an entire website or documentation site before scraping.

  • Firecrawl /map equivalent for autonomous web navigation agents.

  • Returns: Target URL, domain, total URL count, list of discovered URLs, and sitemap detection status.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesTarget website domain or URL to map (e.g., 'https://docs.github.com' or 'example.com').
max_linksNoMaximum internal URLs to discover (default: 50, max: 100).
auth_token_or_txNoOptional x402 auth token or tx hash.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full behavioral burden. It discloses a cost (0.002 USDC) and the return shape, which is useful, but says nothing about rate limits, auth requirements (the auth_token_or_tx param is silently optional), failure behavior for unreachable sites, or the distinction between sitemap discovery and anchor traversal modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the verb and resource in the first sentence, then uses a labeled 'Usage Guidelines' block and a 'Returns' line. Slightly redundant ('Firecrawl /map equivalent' is a marketing reference rather than operational guidance), but overall tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists and the description still summarizes the return fields (target URL, domain, URL count, URL list, sitemap status), which is helpful but not strictly required. The main gap is that auth/payment semantics for the optional auth_token_or_tx parameter are left unexplained despite an advertised per-call cost.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (url, max_links, auth_token_or_tx) are already documented in the schema with examples, defaults, and bounds. The description adds no parameter-level detail beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ('Discovers domain sitemap or traverses internal anchor links to return a canonical URL tree'), which is clear and unambiguous. It does not, however, differentiate itself from named siblings like search_web_quick or clean_web_content, which an agent might reasonably confuse with site traversal.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to use it to map an entire website or documentation site before scraping, which gives a clear pre-scrape context. It stops short of naming an alternative tool or stating when NOT to use it, so it's clear context without exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

oracle_groundingA

Executes real-time web search, noise-free markdown extraction, Gemini AI JSON structuring, and cryptographically signs the result with an on-chain verifiable EIP-712 attestation (0.035 USDC).

Usage Guidelines:

  • Use this tool when an autonomous agent or smart contract requires verified, tamper-proof ground truth from the live web.

  • Returns: Structured facts JSON, human/LLM readable summary markdown, source URLs, and EIP-712 cryptographic signature (v, r, s).

  • Smart contracts can verify this off-chain or on-chain using CleanWebOracleVerifier.sol.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesSearch topic or question to ground with live web sources (e.g., 'Latest Fed interest rate decision').
max_sourcesNoNumber of web sources to synthesize (1 to 5, default: 3).
auth_token_or_txNoOptional x402 micropayment authorization token or EVM transaction hash.
target_schema_jsonNoOptional JSON string defining the expected schema or fields (e.g., '{"rate": "float", "date": "string"}').

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the 0.035 USDC cost, the cryptographic signing, the EIP-712 scheme, the verifier contract, and how the signature can be checked on-chain or off-chain. It omits rate limits, failure/refund behavior, and auth flow specifics, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core capability sentence is front-loaded and the remaining content is organized as scannable bullets. The pipeline enumeration is dense but each clause adds real information; nothing is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a paid, signing oracle tool with an output schema and full param coverage, the description supplies cost, return shape, and verification path. It could still say more about the payment/auth prerequisite and error modes, but an agent has enough to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents query, max_sources, auth_token_or_tx, and target_schema_json. The description adds only the cost figure (0.035 USDC) that loosely relates to the payment param, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a concrete multi-step pipeline (web search, markdown extraction, Gemini JSON structuring, EIP-712 signing) that clearly states what the tool does. It implicitly distinguishes itself from sibling tools like search_web_quick or clean_web_content by adding the signed-attestation step, but it never names a sibling to route the agent explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear when-to-use condition: 'when an autonomous agent or smart contract requires verified, tamper-proof ground truth from the live web.' That is good context, but there are no exclusions or named alternatives (e.g., search_web_quick vs deep_research_topic), so the agent must infer which sibling to pick.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_web_quickB

Performs fast real-time keyword web search returning titles, links, and text snippets (0.002 USDC).

Usage Guidelines:

  • Tavily / s.jina.ai competitor designed specifically for LLM autonomous agent retrieval.

  • Returns: Concise search result list with title, verified URL, and text snippet.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesSearch query or keyword for real-time web discovery.
max_resultsNoMaximum results to return (default: 5, max: 10).
auth_token_or_txNoOptional x402 auth token or tx hash.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full burden. It usefully discloses real-time latency, per-call cost (0.002 USDC), and that returned URLs are 'verified', and the operation is a non-destructive read. It omits rate limits, whether the auth_token_or_tx parameter is required for a paid call, and any failure behavior, which a no-annotation mutation-capable service should ideally state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Content is front-loaded: purpose, cost, then a compact 'Usage Guidelines' block. Every sentence carries information, though framing two short lines as 'Usage Guidelines' is slightly padded and could be a single clause.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return fields need not be spelled out, and the description stays appropriately brief on that front. Yet with zero annotations, the description leaves gaps around authentication (the optional x402 token), pricing verification, and how this tool relates to the deep-research siblings, so an agent must still guess at the invocation context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents query, max_results, and auth_token_or_tx. The description adds no syntax, format, or constraint detail beyond the schema, only the cost and the shape of results. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'fast real-time keyword web search' returning titles, links, and snippets, plus a cost figure. This is far more concrete than a tautology. However, it does not distinguish itself by name from siblings like deep_research_topic or oracle_grounding, leaving the agent to infer the boundary from the word 'quick' alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The line 'designed specifically for LLM autonomous agent retrieval' describes a general audience but gives no when/when-not conditions and names no in-toolset alternative (Tavily and s.jina.ai are external competitors, not routing targets). Usage is implied by 'fast' and 'keyword' contrasting with deeper research tools, but it is never made explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_oracle_attestationA

Verifies an EIP-712 cryptographic attestation produced by CleanWeb Oracle off-chain without consuming gas.

Usage Guidelines:

  • Use this tool to mathematically verify that data received from CleanWeb Oracle has not been tampered with.

  • Returns: Verification status (valid/invalid), recovered signer address, expected oracle address, and message.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesThe original query string that was attested.
data_hashYesThe SHA-256 data hash of the canonical JSON payload (0x...).
signatureYesThe 65-byte hex signature (0x...) from the attestation.
timestampYesThe Unix epoch timestamp from the attestation.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses that verification is off-chain and gas-free and that it returns a validity verdict plus signer/oracle/message, but it omits error behavior, what an invalid result implies, and any permission or rate assumptions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core operation, then a short usage block and return summary. It is tight overall, though the 'Returns' line partly duplicates information the output schema already carries.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a four-required-parameter verification tool with both a full-coverage input schema and an output schema, the description is nearly sufficient. The only real gap is distinguishing it from the sibling oracle_grounding.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already explains all four parameters (query, data_hash, signature, timestamp). The description adds no format or edge-case details beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ('Verifies') and a precise resource ('EIP-712 cryptographic attestation produced by CleanWeb Oracle'), plus the notable constraint that it runs off-chain. It does not, however, differentiate itself from the sibling 'oracle_grounding', so an agent cannot route between the two from this text alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear when-to-use condition ('to mathematically verify that data received from CleanWeb Oracle has not been tampered with'), which is concrete and actionable. It stops short of naming alternatives or stating when not to use it (e.g., vs. oracle_grounding), so it falls below the top tier.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv1.2.8
    • Addedclean_batch_scrape
    • Addedget_pass_status
    • Changedget_payment_info2 fields changed
      • addedInput schema / properties / chain
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Optional specific chain name (polygon, base, arbitrum).",
        +  "title": "Chain"
        +}
      • addedInput schema / properties / tier
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Optional specific pricing tier to inspect.",
        +  "title": "Tier"
        +}
  2. 11 tool updatesv1.2.6
    • Changedclean_pdf_research7 fields changed
      • addedInput schema / properties / auth_token_or_tx / description
        Added value: +"Optional x402 micropayment authorization token or EVM transaction hash."
      • addedInput schema / properties / max_pages / description
        Added value: +"Maximum number of pages to parse (1 to 100, default: 30) to control token budget."
      • addedInput schema / properties / max_pages / maximum
        Added value: +100
      • addedInput schema / properties / max_pages / minimum
        Added value: +1
      • addedInput schema / properties / url / description
        Added value: +"Direct HTTP/HTTPS URL pointing to an online PDF document."
      • addedInput schema / properties / url / examples
        Added value: +[
        +  "https://arxiv.org/pdf/1706.03762.pdf",
        +  "https://bitcoin.org/bitcoin.pdf"
        +]
      • addedInput schema / properties / url / pattern
        Added value: +"^https?:\\/\\/[^\\s/$.?#].[^\\s]*\\.pdf(\\?.*)?$"
    • Addedclean_text_raw
    • Changedclean_web_content5 fields changed
      • addedInput schema / properties / auth_token_or_tx / description
        Added value: +"Optional x402 micropayment authorization token or EVM transaction hash."
      • addedInput schema / properties / respect_robots_txt
        Added value: +{
        +  "default": false,
        +  "description": "Whether to enforce target domain robots.txt Disallow rules (compliance-mode).",
        +  "title": "Respect Robots Txt",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / url / description
        Added value: +"The target HTTP or HTTPS website URL to scrape and convert to markdown."
      • addedInput schema / properties / url / examples
        Added value: +[
        +  "https://en.wikipedia.org/wiki/Web_scraping",
        +  "https://news.ycombinator.com/"
        +]
      • addedInput schema / properties / url / pattern
        Added value: +"^https?:\\/\\/[^\\s/$.?#].[^\\s]*$"
    • Changedclean_youtube_transcript7 fields changed
      • addedInput schema / properties / auth_token_or_tx / description
        Added value: +"Optional x402 micropayment authorization token or EVM transaction hash."
      • addedInput schema / properties / lang / description
        Added value: +"Comma-separated ISO 639-1 language priority codes for transcript extraction (e.g., 'ko,en', 'en', 'ja')."
      • addedInput schema / properties / lang / examples
        Added value: +[
        +  "ko,en",
        +  "en",
        +  "ja,en"
        +]
      • addedInput schema / properties / lang / pattern
        Added value: +"^[a-z]{2}(,[a-z]{2})*$"
      • addedInput schema / properties / url / description
        Added value: +"Public YouTube video URL (standard watch, short youtu.be, or Shorts format)."
      • addedInput schema / properties / url / examples
        Added value: +[
        +  "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
        +  "https://youtu.be/dQw4w9WgXcQ",
        +  "https://www.youtube.com/shorts/abcdef12345"
        +]
      • addedInput schema / properties / url / pattern
        Added value: +"^https?:\\/\\/(www\\.)?(youtube\\.com\\/(watch\\?v=|shorts\\/)|youtu\\.be\\/)[\\w-]+.*$"
    • Addeddeep_research_topic
    • Addedextract_json_schema
    • Changedget_vault_balance3 fields changed
      • addedInput schema / properties / agent_address_or_key / description
        Added value: +"Ethereum/Polygon address (0x...) or session key (sk_...) to query balance for."
      • addedInput schema / properties / agent_address_or_key / examples
        Added value: +[
        +  "0x71C...397",
        +  "sk_cleanweb_agent_01"
        +]
      • addedInput schema / properties / agent_address_or_key / pattern
        Added value: +"^(0x[a-fA-F0-9]{40}|sk_[a-zA-Z0-9_-]+)$"
    • Addedmap_site
    • Addedoracle_grounding
    • Addedsearch_web_quick
    • Addedverify_oracle_attestation
  3. 11 tool updatesv1.2.5
    • Addedclean_pdf_research
    • Addedclean_web_content
    • Addedclean_youtube_transcript
    • Removeddeep_research_briefing
    • Removedextract_json_schema
    • Removedfetch_batch_clean_markdown
    • Removedfetch_clean_web_content
    • Removedfetch_pdf_markdown
    • Removedfetch_plain_text
    • Removedfetch_youtube_transcript
    • Addedget_vault_balance
  4. 3 tool updatesv1.2.1
    • Addeddeep_research_briefing
    • Addedextract_json_schema
    • Addedfetch_batch_clean_markdown
  5. 5 tool updatesv1.2.0
    • First observedfetch_clean_web_content
    • First observedfetch_pdf_markdown
    • First observedfetch_plain_text
    • First observedfetch_youtube_transcript
    • First observedget_payment_info

TDQS

A4/5.0

Scored across 14 tools

Disambiguation5/5

Every tool targets a distinct resource and operation: separate scrapers for web, YouTube, PDF, batch, and raw text; distinct payment/status/balance tools; clear separation between quick search, deep research, and oracle-grounded search. The descriptions explicitly cross-reference other tools with 'Do NOT use' guidance, eliminating ambiguity.

Naming Consistency4/5

Most names follow a consistent verb_noun pattern (get_payment_info, clean_web_content, map_site, verify_oracle_attestation). Minor deviations include 'oracle_grounding' (noun_noun) and 'clean_batch_scrape' (verb-adjective-noun), which are still understandable but slightly break the pattern.

Tool Count5/5

14 tools is well within the ideal 3–15 range and each tool earns its place by covering a distinct aspect of web research, extraction, oracle verification, and payment management. The count feels substantial but not bloated.

Completeness5/5

The tool surface is remarkably comprehensive for its stated domain: single and batch scraping, YouTube and PDF handling, raw text extraction, quick search, deep research, site mapping, JSON schema extraction, oracle signing/verification, and payment/pass/balance queries. No obvious dead ends or missing lifecycle operations are evident.

Maintenance

ActivityActive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    Pay-per-use clean web reader for AI agents. URL in, markdown plus metadata out, in milliseconds. Settled per-call in USDC over x402 — no signup, no API keys.
    1
    89 npm
    3
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Fetches web pages and converts them to markdown for LLM consumption, supporting chunked reading and raw content extraction.
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    MCP server for headless browser automation using Puppeteer, enabling AI to navigate, click, fill forms, take screenshots, and execute JavaScript on web pages.
    -
  • A
    license
    Not graded
    quality
    D
    maintenance
    MCP server providing 17 keyless, pay-per-use web-data tools with signed-provenance receipts, enabling AI agents to autonomously fetch, extract, and verify web content on Base mainnet.
    30 npm
    MIT