opticparse
This server provides AI-powered web scraping and real-time phishing detection for autonomous agents.
opticparse_scrape: Extract token-optimized markdown or structured JSON from any live webpage using multimodal vision, bypassing anti-bot protections and dynamic JavaScript without CSS selectors.
Supports natural-language extraction queries and optional JSON Schema for structured outputs.
Ideal for feeding clean webpage content into LLM context windows and RAG pipelines.
phishvision_detect: Audit any URL for zero-day phishing kits, smart-contract wallet drainers, credential harvesters, and brand impersonation in under 1.6 seconds.
Returns a security verdict and threat score to help agents decide whether to halt navigation.
Provides a pre-built extraction template for monitoring Airbnb listing and pricing data.
Provides a pre-built extraction template for Amazon price tracking and undercut alerts.
Supplies real-time threat detection and brand impersonation forensics for Binance-related URLs.
Supplies real-time threat detection and brand impersonation forensics for Coinbase-related URLs.
Provides a template for Docker Hub vulnerability alerts.
Provides a pre-built extraction template for flash sale auto-cart and product monitoring on Flipkart.
Provides real-time visual safety evaluations and brand impersonation detection for Google domains.
Provides a template for npm package vulnerability alerting.
Provides real-time threat detection and brand impersonation forensics for PayPal-related URLs.
Provides a pre-built extraction template for new product alerts on Shopify stores.
Provides a pre-built extraction template for aggregating Zillow rent estimates.
⭐ Support Open Source: If you find OpticParse or PhishVision useful, please give us a Star on GitHub! It helps us maintain free edge scrapers and datasets for everyone.
⚡ Feature & Architecture Comparison
Looking for a Firecrawl alternative, Crawl4AI alternative, or a Zero-Day Phishing & Visual Shield for your autonomous AI agents? Here is how OpticParse + PhishVision compares to traditional scraping engines and legacy security blacklists:
Feature / Capability | OpticParse & PhishVision | Firecrawl | Crawl4AI | Jina Reader | Google Safe Browsing / VirusTotal |
Primary Extraction Engine | Multimodal AI Vision (Zero-CSS) | HTML / Markdown | DOM Tree / XPath | Regex / Markdown | N/A (Security Only) |
Visual Prompt Injection Shield ( | 🛡️ Built-in (Zero-Opacity & CSS Sanitization) | ❌ None (Vulnerable) | ❌ None (Vulnerable) | ❌ None | N/A |
0-Day Phishing & Wallet Drainer Shield | 🛡️ Built-in (Instant Day 0) | ❌ None | ❌ None | ❌ None | ⚠️ Historical Blacklist (12-48h delay) |
Machine Discovery Standard (A2A) | ✅ Native | ❌ None | ❌ None | ❌ None | ❌ None |
In-IDE Developer Free Trial | ✅ 200 Requests Out-of-the-Box | ❌ Mandatory Credit Card | ❌ Self-hosted Setup | ⚠️ Limited Free Tier | ⚠️ Strict Rate Limits |
In-Terminal Instant Checkout | ✅ ASCII QR in Console + .env Save | ❌ Web Portal Only | ❌ None | ❌ Web Portal Only | ❌ Enterprise Contract |
Agent Ecosystem Toolkits | LangChain, ElizaOS, AgentKit, MCP | LangChain, LlamaIndex | Generic Python | Generic HTTP | None (Raw API) |
Payment & Settlement Flexibility | USDC (Base/Polygon) / $0.01 M2M | Stripe ($16+/mo SaaS) | None (Open Source) | Stripe ($$$) | Enterprise ($$$) |
Token Optimization Noise Reduction | 96% Token Reduction | ~80% | ~70% | ~85% | N/A |
Related MCP server: Hyperbrowser MCP Server
📌 Why This Stack Exists: The Dual-Engine Agent Architecture
Giving an autonomous AI agent unrestricted internet access introduces three critical vulnerabilities:
Context Window & DOM Fragility (OpticParse solves this): Raw HTML consumes 90%+ of LLM token context, while rigid CSS/XPath selectors break the moment websites change frontend frameworks.
Visual Prompt Injection & Screen Poisoning (ToxicCanvas solves this): Adversarial sites embed invisible CSS, zero-opacity layers, and prompt injection strings to hijack multimodal vision models (Claude Computer Use, GPT-4o, Operator). ToxicCanvas sanitizes DOM tokens before ingestion.
The Autonomous Link Trap (PhishVision solves this): When agents follow links autonomously, malicious actors exploit them with zero-day credential harvesting kits, fake Web3 dApps, and wallet drainers.
OpticParse + PhishVision provides the complete solution:
The Shield (PhishVision + ToxicCanvas): Inspects target URLs in real-time for SSL age, brand spoofing, and visual injection attacks before the agent executes navigation.
The Engine (OpticParse): Visually renders authenticated pages at the edge, converting dynamic JavaScript into token-optimized clean Markdown (96% noise reduction) without brittle CSS selectors.
Embodied AI & Edge Hardware Compatible: Lightweight token streams structured for low-power edge processors (NVIDIA Jetson, embedded robotics telemetry) that cannot waste local battery and GPU cycles running headless browsers (Learn more).
Operating 150 autonomous extraction & threat pipelines across 13 industries, the network auto-harvests 1,250+ verified intelligence records every 24 hours at the global edge on Cloudflare Workers, R2, and D1.
🏛️ Autonomous Agent Dual-Shield Workflow
┌─────────────────────────────────────────────────────────────┐
│ Autonomous AI Agent (LangChain / LlamaIndex) │
└──────────────────────────────┬──────────────────────────────┘
│ Target URL
▼
┌──────────────────────────────┐
│ PhishVision Security Shield │
│ - Heuristic Brand Distance │
│ - Zero-Day Phishing Kits │
│ - Web3 Drainer Signatures │
└──────────────┬───────────────┘
│
┌────────────────┴────────────────┐
│ Verdict │
[ MALICIOUS ] [ SAFE ]
│ │
▼ ▼
┌───────────────────────────┐ ┌───────────────────────────┐
│ Halt & Shield Agent Loop │ │ OpticParse Edge Extractor │
│ (Protect API & Wallet) │ │ - Zero-CSS Visual Parse │
└───────────────────────────┘ │ - 96% Token Reduction │
└─────────────┬─────────────┘
│ Clean Markdown
▼
┌───────────────────────────┐
│ LLM Context Window / RAG │
└───────────────────────────┘🏛️ System Architecture
┌─────────────────────────────────────────────────────────────────┐
│ OpticParse System Architecture │
├─────────────────────────────────────────────────────────────────┤
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ 150 Scraping │──▶│ Cloudflare │──▶│ Cloudflare │ │
│ │ Pipelines │ │ Workers Edge │ │ D1 + R2 Lake │ │
│ └──────────────┘ └──────────────┘ └──────┬───────┘ │
│ │ │
│ ┌───────────────────────────┼───────┐ │
│ │ Distribution │ │ │
│ ▼ ▼ ▼ ▼ │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌─────┐ │
│ │ Kaggle │ │ Hugging │ │ RapidAPI │ │Ocean│ │
│ │ Hub │ │ Face Hub │ │ Gateway │ │ NFTs│ │
│ └──────────┘ └──────────┘ └──────────┘ └─────┘ │
│ │
│ ┌──────────────────────────────────────────────────────────┐ │
│ │ Anthropic Model Context Protocol (MCP) Server for Agents │ │
│ └──────────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────┘🎯 150-Pipeline Industry Coverage Matrix
Industry Vertical | Active Pipelines | Core Extractions & Capabilities |
🛡️ Cybersecurity & Threat Intel | 70 | Zero-day phishing kits, typosquats, crypto-drainers, brand impersonation |
🛒 E-Commerce & Retail Arbitrage | 15 | Amazon undercut alerts, Shopify inventory snipers, price elasticity |
💼 B2B Growth & Lead Signals | 45 | Executive hiring telemetry, YC startup surge, wage trends, SEC filings |
⚡ Finance, Crypto & DEX Arbitrage | 20 | Liquidity pool spreads, token mint monitors, funding rate surveillance |
📊 Verified Adversarial Stress Benchmarks (Live September 2026 Audit)
Both engines were independently benchmarked against live adversarial threats and enterprise-grade anti-bot defenses:
1. OpticParse: Zero-CSS Visual Extraction Stress Test
Tested live on production backend against complex dynamic JavaScript, nested schemas, and anti-bot protected targets:
Target Page | Challenge Architecture | Result | Latency | Data Extracted |
Cloudflare Official Plans | Enterprise Bot-Shield & obfuscated DOM | 100% PASS | 42.1s | Extracted complete tiered pricing structure ( |
Stripe Global Pricing | Heavy React hydration & dynamic cards | 100% PASS | 38.2s | Extracted standard card processing fees ( |
Hacker News Frontpage | Real-time community feed & score metadata | 100% PASS | 42.7s | Extracted top 5 stories with real ranks and point totals |
GitHub Trending Hub | Nested repository schemas & live star counts | 100% PASS | 32.3s | Extracted trending repositories ( |
Zero-CSS Extraction Accuracy: 4 / 4 (100.0% Success Rate)
DOM Fragility: Zero broken CSS/XPath selectors; resilient to class renaming and layout refactors.
💡 Understanding the ~35s Latency: Why Multimodal Vision Beats Brittle 2-Second Scrapers
Traditional scrapers finish in 2 seconds because they only regex raw HTML text. The moment a site deploys dynamic React hydration, anti-bot protection, or changes CSS class names, traditional scrapers break and return empty data or 403 Forbidden errors.
OpticParse prioritizes zero-failure autonomous agent execution: It spins up a full headless Chromium browser, bypasses anti-bot layers, captures a high-resolution visual buffer, and has a Multimodal Vision LLM reason over the page like a human.
Result: 100% schema accuracy with zero broken CSS selectors, guaranteed for autonomous background agents where data integrity matters more than raw millisecond speed.
2. PhishVision: 0-Day Adversarial Threat & Crypto Drainer Audit
Audited against 50 live zero-day malicious URLs from the OpenPhish global threat feed (created within hours of test) + 20 difficult authentic authentication portals:
Security Benchmark Metric | PhishVision Score | Traditional Baseline (Google Safe Browsing / VirusTotal on Day 0) |
0-Day Threat Catch Rate | 68.0% (34 / 50) | ~10% – 15% (Fails because new domains lack historical reports) |
Authentic Auth Portals (Clean) | 95.0% (19 / 20) | >90% (Industry standard for benign enterprise login portals) |
Median Execution Latency | 1.58 seconds | 5 – 15 seconds (Heavy commercial sandbox scanners) |
Verified Live Threats Neutralized On First Contact:
🛡️ Uniswap Wallet Drainer:
uniswap-interface.vercel.app(BLOCKED)🛡️ Trezor Hardware Seed Stealer:
sso-trezor-com-start-x-auth.typedream.app(BLOCKED)🛡️ TrustWallet Crypto Harvester:
trust-wallet-liart.vercel.app(BLOCKED)🛡️ Compromised WordPress Injections:
bruceleephilosophy.com/texts/(BLOCKED)🛡️ Apple Brand Impersonator:
apple-fruit.xyz(BLOCKED)
Verified Authentic Portals Cleared as Safe:
accounts.google.com•github.com/login•dashboard.stripe.com/login•appleid.apple.com•reddit.com/login•auth.openai.com(All CLEARED with zero false alarm interruptions).
🐍 Python & LangChain Quickstart (PyPI)
Install the official Python SDK or LangChain multi-agent toolkit:
# Core Python SDK (OpticParse AI Scraper + PhishVision)
pip install opticparse-py
# LangChain & CrewAI Tool Integration
pip install langchain-opticparse
# LlamaIndex Tool Integration
pip install llama-index-tools-opticparseLangChain Dual-Shield Agent Usage:
from langchain_opticparse import OpticParseTool, PhishVisionTool
# Initialize tools
phish_shield = PhishVisionTool(api_key="your_api_key")
optic_scraper = OpticParseTool(api_key="your_api_key")
target_url = "https://example.com/pricing"
# 1. Pre-flight security audit: Block phishing kits & wallet drainers
threat_audit = phish_shield.run({"url": target_url})
if threat_audit.get("verdict") == "MALICIOUS":
print(f"🚨 Security Alert: Agent halted. Threat detected (Score: {threat_audit['threat_score']}/100)")
else:
# 2. Extract clean visual Markdown without brittle CSS selectors
clean_data = optic_scraper.run({
"url": target_url,
"query": "Extract tier prices, plan limits, and feature comparisons"
})
print(clean_data)⚡ 200 Free Trial Extractions & In-IDE Activation:
Every IP receives 200 free trial requests out of the box. When your free trial completes, the terminal automatically prints an ASCII QR code + 1-click checkout link. Upon paying on-chain ($10 Starter / $50 Growth / $200 Scale / $0.05 M2M via MetaMask or mobile wallet), your live API key is automatically minted, saved directly to your local.env, and your code resumes execution seamlessly.
LlamaIndex Usage:
from llama_index.tools.opticparse import OpticParseToolSpec
from llama_index.core.agent import FunctionCallingAgentWorker
# Initialize tool spec and extract documents
tool_spec = OpticParseToolSpec(api_key="your_api_key")
docs = tool_spec.extract(url="https://example.com", query="Extract specs")
print(docs[0].text)
# Convert directly to agent tool list
tools = tool_spec.to_tool_list()Python Agent Usage:
from opticparse import OpticParse
client = OpticParse(api_key="YOUR_API_KEY")
# 1. Clean Markdown Extraction for LLM RAG pipelines
res = client.extract_markdown("https://news.ycombinator.com")
print(res["markdown"])
# 2. Autonomous Zero-Day Threat Inspection
safety = client.detect_phishing("https://suspicious-dapp-claim.xyz")
print(f"Verdict: {safety['verdict']} | Threat Score: {safety['threat_score']}/100")🤖 ElizaOS Autonomous Agent Plugin (npm)
Install the official ElizaOS autonomous agent plugin:
npm install opticparse-eliza-pluginElizaOS Usage:
import { AgentRuntime } from "@elizaos/core";
import { opticParsePlugin } from "opticparse-eliza-plugin";
export const webScoutAgent = {
name: "WebScoutAgent",
plugins: [opticParsePlugin],
settings: {
secrets: {
OPTICPARSE_API_KEY: process.env.OPTICPARSE_API_KEY
}
}
};🤖 Model Context Protocol (MCP) Integration
Connect Claude Desktop, Cursor IDE, or AutoGen directly to OpticParse in 1 click:
1. Install via Smithery
npx -y @smithery/cli install @parastejpal987-cmyk/opticparse --client claude2. Manual Configuration (claude_desktop_config.json)
{
"mcpServers": {
"opticparse": {
"command": "python",
"args": ["-m", "mcp_server"],
"env": {
"OPTICPARSE_API_KEY": "YOUR_API_KEY"
}
}
}
}🛠️ Exposed AI Agent Tools:
opticparse_scrape: Vision-based structured data extraction from any web URL.phishvision_detect: Real-time phishing and brand impersonation heuristic scanner.search_lessons: Query indexed threat telemetry records.
📮 Postman Public API Collections
Test the live endpoints instantly in Postman with 200 free trial requests built-in:
OpticParse Visual Web Scraper Collection: Run in Postman — Includes visual extraction, token-optimized RAG markdown parsing, and live system health checks.
PhishVision AI Threat Shield Collection: Includes 0-day phishing scanning, brand impersonation detection, and permit2 drainer audits in <1.6s.
Offline Import: Both collections are available directly in this repository under
postman/:postman/opticparse_api_collection.jsonpostman/phishvision_api_collection.json
🛡️ Embed Dynamic Scrapeability & Security Badges
Showcase that your open-source repository, documentation, or SaaS application is scrapeable by autonomous AI agents or verified safe against zero-day phishing:
1. OpticParse Scrapeability Badge
[](https://opticparse.com)Output:
2. PhishVision 0-Day Safe Badge
[](https://opticparse.com)Output:
📦 Master Public Datasets (Hugging Face & Kaggle)
All master datasets are public, verified, and streamable in Apache Parquet & CSV format:
Dataset Name | Records | Format | Direct Access |
Master 150-Template Catalog |
|
| |
PhishVision Threat Intelligence |
|
| |
E-Commerce & Retail Arbitrage |
|
| |
B2B Growth & Financial Signals |
|
|
💻 Load in Python (1-Line Quickstart):
import pandas as pd
# Load master threat intelligence dataset
df_threat = pd.read_csv("https://opticparse.com/threat_intel_dataset.csv")
print(f"Loaded {len(df_threat)} live threat vectors")
print(df_threat.head())⚡ Commercial RapidAPI Gateway
Pay-As-You-Go developer access with sub-350ms global edge latency at $0.008/request:
01-ai-web-scraper: Autonomous AI vision web scraper02-dom-poisoner: Streaming adversarial anti-scraping tag injector03-edge-proxy: Global unblockable fetcher proxy04-rate-limit-bypasser: Residential edge IP rotator05-realtime-sentiment: Live threat & telemetry database feed
👉 Target Gateway URL: https://opticparse-rapidapi-gateway.parastejpal987.workers.dev
🦊 Autonomous AI Agent Micropayments (HTTP 402 Machine Paywall)
For autonomous bots, multi-agent frameworks (ElizaOS, AutoGPT, CrewAI), and automated scrapers with no human in the loop, OpticParse supports instant on-chain settlement:
Price:
$0.01 USDCper requestSupported Chains: Base, Polygon, Arbitrum
Treasury Address:
0xd458E709e7d54fd3659EF66624A621Cde74EDD27Machine Manifest:
https://opticparse.com/.well-known/agent.json(Google A2A Standard)
🤖 1-Line Autonomous Agent Request:
curl -X POST https://opticparse-edge.parastejpal987.workers.dev/api/edge/scrape \
-H "X-Payment-TxHash: <YOUR_CONFIRMED_USDC_TX_HASH>" \
-H "Content-Type: application/json" \
-d '{"url": "https://news.ycombinator.com"}'No account creation, no credit card, and zero KYC required. Every verified on-chain transfer is verified cryptographically via global RPC nodes and grants immediate edge execution.
🛡️ Interactive Threat DB Directory (21 Brands)
Explore live brand safety evaluations and visual impersonation checks:
🔍 Common Search Index & Developer Integrations
OpticParse & PhishVision are natively integrated across the agent and web extraction ecosystem:
Alternative To: Firecrawl, Crawl4AI, Jina AI Reader, ScrapeGraphAI, Playwright Python, Puppeteer, Google Safe Browsing API, VirusTotal API, URLScan.io.
Autonomous Agent Compatibility: LangChain (
langchain-opticparse), LlamaIndex (llama-index-tools-opticparse), ElizaOS (opticparse-eliza-plugin), Coinbase AgentKit (opticparse-agentkit-action), Model Context Protocol (mcp_server.py).Use Cases: AI Web Scraping, Multimodal LLM Vision RAG, Zero-Day Phishing Detection, Web3 Crypto Drainer Defense, Real-Time SKU & Price Intelligence, Continuous Snappy Parquet Datasets.
🤝 Open Source & Licensing
OpticParse developer tools, client SDKs, and MCP servers are proudly released under the MIT License.
Author: Paras Tejpal (@parastejpal)
Official Website: https://opticparse.com
Available Tools
2 toolsopticparse_scrapeB
Extract structured, token-optimized data from any live web page using AI Multimodal Vision. Bypasses Cloudflare Turnstile, anti-bot mechanisms, and dynamic JavaScript rendering without brittle CSS selectors. Perfect for LLM context windows and RAG pipelines.
| Name | Required | Description | Default |
|---|---|---|---|
| target_url | Yes | The fully-qualified HTTP/HTTPS URL of the webpage to scrape and extract content from. | |
| response_schema | No | Optional JSON Schema definition to enforce a strict structured output format on the extracted result. | |
| extraction_query | Yes | Natural language instructions specifying what data fields, tables, or text to extract from the webpage. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
It discloses meaningful behaviors such as bypassing Cloudflare Turnstile, anti-bot mechanisms, and dynamic JavaScript rendering, while avoiding CSS selectors. However, with no annotations at all, the description carries the full burden for behavior and does not clarify output format, failure modes, rate limits, or legal/auth constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three focused sentences with the core capability front-loaded. The final 'perfect for LLM context windows and RAG pipelines' is slightly promotional but still conveys appropriate use cases without significant bloat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the tool description should clarify what the agent can expect back. 'Structured, token-optimized data' is a helpful hint but does not define whether the result is JSON, text, or an object respecting response_schema. Overall adequate for a straightforward invocation but incomplete on return and error behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
All three parameters already have descriptive schema entries, so the description adds no extra paramater-level meaning. It reinforces the general extraction behavior but does not clarify specifics like how response_schema interacts with the output or what form extraction_query should take beyond 'natural language'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states a specific action ('Extract structured, token-optimized data') and a specific resource ('any live web page'), making the tool's purpose obvious. It doesn't explicitly contrast with the sibling phishvision_detect, but the extraction-focused language is enough to distinguish it from a detection tool. The naming and description align well.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives useful context by highlighting anti-bot bypass, dynamic JS rendering, and suitability for LLM/RAG pipelines, which implies when it should be used. It does not explicitly state when to prefer an alternative or when not to use it, leaving the routing decision partially inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
phishvision_detectA
Audit and inspect any URL for real-time zero-day phishing campaigns, smart contract wallet drainers, credential harvesting kits, and brand impersonation attacks using visual layout heuristics in under 1.6 seconds.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The target domain or fully qualified URL to audit for security threats, drainers, and malicious vectors. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does disclose useful behavioral details: it uses visual layout heuristics and completes in under 1.6 seconds. However, it does not clarify side effects such as whether the tool actively fetches or renders the target URL, whether it is strictly read-only, or what happens when the URL is malicious or unreachable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The entire description is one front-loaded sentence that opens with the primary action and resource, then adds precise threat categories, method, and a performance bound. There is no filler or redundant restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with no output schema and no annotations, the description covers the main selection and invocation needs: target input, detection scope, method, and speed. It could be more complete by indicating the return shape or verdict format, but that is a minor gap given the low complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents url as 'The target domain or fully qualified URL to audit for security threats, drainers, and malicious vectors.' The description reinforces the purpose but does not add new parameter-level meaning such as format restrictions or accepted schemes.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb-resource pair ('Audit and inspect any URL') and enumerates concrete detection targets: phishing campaigns, wallet drainers, credential harvesting kits, and brand impersonation. This clearly differentiates it from the sibling opticparse_scrape, which by name is oriented toward extraction rather than security inspection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when the tool applies: real-time URL security auditing with visual layout analysis. It does not explicitly state exclusions or name opticparse_scrape as the alternative, so it stops short of a 5, but the use case is unmistakable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.2- Changed
opticparse_scrape3 fields changed- changed
Input schema / properties / extraction_query / descriptionPrevious value: -"Instructions for what information to extract from the webpage."New value: +"Natural language instructions specifying what data fields, tables, or text to extract from the webpage." - changed
Input schema / properties / response_schema / descriptionPrevious value: -"Optional JSON schema to enforce on the extracted data."New value: +"Optional JSON Schema definition to enforce a strict structured output format on the extracted result." - changed
Input schema / properties / target_url / descriptionPrevious value: -"The URL of the webpage to scrape."New value: +"The fully-qualified HTTP/HTTPS URL of the webpage to scrape and extract content from."
- Changed
phishvision_detect1 field changed- changed
Input schema / properties / url / descriptionPrevious value: -"The URL to inspect for threats."New value: +"The target domain or fully qualified URL to audit for security threats, drainers, and malicious vectors."
- Removed
search_lessons
3 tool updates
v0.1.0- First observed
opticparse_scrape - First observed
phishvision_detect - First observed
search_lessons
TDQS
Scored across 2 tools
opticparse_scrape extracts page content for LLM/RAG use, while phishvision_detect audits URLs for security threats. Their purposes, outputs, and trigger conditions are clearly distinct, so an agent should not confuse them.
Both names follow a product_verb pattern and use snake_case, but the product prefix changes from opticparse to phishvision. This breaks the expected consistency for a server named opticparse and makes the second tool feel unrelated.
Two tools is on the thin side for a server that claims to cover web scraping and phishing detection. Each tool serves a distinct high-level purpose, but the overall surface feels minimal rather than well-rounded.
The two tools cover the core stated capabilities: extracting structured page data and detecting phishing URLs. However, there are notable gaps such as no batch processing, no non-URL input support, and no way to refine or iterate on extraction results.
Maintenance
Related MCP Connectors
Screenshot, PDF and HTML-to-image rendering API so Claude and Cursor can see any web page.
Screenshot, PDF and HTML-to-image rendering API so Claude and Cursor can see any web page.
Web scraping, code review, content gen, sentiment. Zero Core Tools.
Web scraping for AI agents. Converts URLs to clean, LLM-ready Markdown with anti-bot bypass.
Related MCP Servers
- AlicenseAqualityBmaintenanceA Model Context Protocol (MCP) integration that provides Claude Desktop with autonomous browser automation capabilities. This agent enables Claude to interact with web content, manipulate DOM elements, execute JavaScript, and perform API requests.133 npm41TypeScriptMozilla Public 2.0
- AlicenseNot gradedqualityDmaintenanceEnables web scraping, crawling, structured data extraction, and browser automation through multiple AI agents including OpenAI's CUA, Anthropic's Claude Computer Use, and Browser Use.4 npmMIT
- FlicenseNot gradedqualityDmaintenanceProvides browser automation, AI-powered analysis, visual processing, web scraping, automated test generation, and DevTools analysis capabilities. Supports multiple AI providers (OpenAI, Anthropic, Google, Ollama) for intelligent web interaction and data extraction.-
- AlicenseNot gradedqualityCmaintenanceThe Zero-Setup Local Browser MCP. Enables AI agents to control web browsers via CDP with zero vision tokens and high-speed DOM mapping.17 npmMIT