stagenth · 网页数据
Server Details
Web scraping to clean Markdown with JS rendering, multi-page crawl, structured extract, sitemaps.
Glama couldn't complete the latest health check. If this server requires authentication, missing or expired test credentials may be the cause. A test profile lets Glama authenticate for health checks and discover tools; it is separate from your personal connections.
If you are the author, claim ownership, then add or update a test profile under Admin → Test Profile.
- Status
- Unhealthy
- Uptime
- 58.3% over 40 days
- Last Tested
- Transport
- Streamable HTTP
- URL
TDQS
Scored across 4 tools
Each tool has a clearly distinct role: site_map discovers URLs, scrape_url fetches a single page, crawl fetches multiple pages, and extract returns structured JSON. The overlap between scrape_url and crawl is clearly addressed in descriptions by single vs multi-page, so an agent can easily select the right tool.
The names are mostly intuitive, with two verbs (crawl, extract) and two noun phrases (scrape_url, site_map). The mix of verb and noun forms is a slight inconsistency, but all names are lowercase with underscores, so the pattern is still predictable.
With only 4 tools, each serves a distinct step in the web data workflow: discovery, single-page fetch, multi-page crawl, and structured extraction. This is an appropriate scope for the server's purpose, neither too sparse nor overwhelming.
The tool set covers the full pipeline from discovering URLs (site_map) to fetching content (scrape_url, crawl) to extracting structured data (extract). There are no obvious gaps for the server's stated purpose of web data collection and transformation.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
- First observed
crawl - First observed
extract - First observed
scrape_url - First observed
site_map
Related MCP Connectors
Turn any URL into clean Markdown and structured data. Scrape, crawl, search and extract.
PDFs, JavaScript pages and ordinary HTML as clean Markdown. Follows robots.txt per RFC 9309.
Cloud scraping & crawling API for AI agents. Turn any URL into clean, LLM-ready markdown.
Extract public webpages as Markdown, map sites, and run bounded crawl jobs. API key required.
Related MCP Servers
- FlicenseNot gradedqualityCmaintenanceCrawl entire websites, sitemaps, and documentation portals into clean, LLM-ready Markdown without ads, cookie popups, or noise for RAG pipelines and AI agents.-
- AlicenseNot gradedqualityCmaintenanceScrapes webpages and converts them to markdown using AI-powered interaction to automatically handle cookie banners, CAPTCHAs, paywalls, and other blocking elements before extracting clean content.3 npm48Apache 2.0
- AlicenseNot gradedqualityDmaintenanceConverts any webpage into clean, LLM-ready Markdown, removing noise and supporting JavaScript rendering.MIT
- AlicenseNot gradedqualityAmaintenanceEnables fetching web pages as clean Markdown with automatic escalation from direct requests to residential proxies and headless browsers to overcome blocks, plus built-in politeness handling like robots.txt and rate limiting.-
Glama MCP Gateway
Add one secure layer between your agents and this server.