Web Scraper MCP Server
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| LOG_LEVEL | No | 日志级别 (debug, info, warn, error) | |
| REQUEST_TIMEOUT | No | HTTP 请求超时时间(毫秒) | |
| PUPPETEER_TIMEOUT | No | Puppeteer 超时时间(毫秒) |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Server capabilities have not been inspected yet.
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| scrape_imagesC | 从指定URL爬取网站上的所有图片并保存到本地 |
| scrape_textC | 从指定URL爬取网站的文本内容并保存为Markdown文件 |
| list_imagesC | 列出所有已下载的图片信息 |
| list_textsC | 列出所有已提取的文本文件 |
| cleanup_imagesC | 清理所有下载的图片文件 |
| cleanup_textsC | 清理所有提取的文本文件 |
| get_scraping_statusC | 获取爬虫状态信息 |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 7 tools
Every tool has a clearly distinct purpose with no ambiguity. The tools are cleanly divided between scraping operations (scrape_images, scrape_text), listing operations (list_images, list_texts), cleanup operations (cleanup_images, cleanup_texts), and status monitoring (get_scraping_status). There is no overlap in functionality.
The naming follows a perfectly consistent verb_noun pattern throughout. All tools use snake_case with clear action prefixes (scrape_, list_, cleanup_, get_) followed by specific resource nouns (images, texts, scraping_status). There are no deviations in naming conventions.
The 7 tools are well-scoped for a web scraping server. Each tool earns its place by covering distinct aspects of the scraping workflow: initiating scrapes, listing results, cleaning up files, and checking status. This is an appropriate number for the domain without being too thin or bloated.
The tool surface provides excellent coverage for core web scraping operations with clear CRUD-like patterns (create via scrape, read via list, delete via cleanup). The only minor gap is the lack of update operations for existing scraped content, but agents can work around this by re-scraping. Status monitoring is included, which is valuable for workflow management.