Scrapezy MCP Server
The Scrapezy MCP Server enables AI models to extract structured data from websites.
Extract Structured Data: Extract specific data from websites by providing a URL and a natural language prompt describing what to extract
Structured Output: Return extracted data in a structured format
Support for Specific Data Types: Handle requests for data like product information, prices, and descriptions
Integration with Claude Desktop: Seamlessly integrate with Claude Desktop for AI-powered extraction tasks
Debugging Support: Utilize the MCP Inspector for debugging server communication
API Key Flexibility: Configure using either environment variables or command-line arguments
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Scrapezy MCP Serverextract product details from https://amazon.com/dp/B08N5WRWNW"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
@scrapezy/mcp MCP Server
A Model Context Protocol server for Scrapezy that enables AI models to extract structured data from websites.
Features
Tools
extract_structured_data- Extract structured data from a websiteTakes URL and prompt as required parameters
Returns structured data extracted from the website based on the prompt
The prompt should clearly describe what data to extract from the website
Related MCP server: ScrapeGraph MCP Server
Installation
Installing via Smithery
To install Scrapezy MCP Server for Claude Desktop automatically via Smithery:
npx -y @smithery/cli install @Scrapezy/mcp --client claudeManual Installation
npm install -g @scrapezy/mcpUsage
API Key Setup
There are two ways to provide your Scrapezy API key:
Environment Variable:
export SCRAPEZY_API_KEY=your_api_key npx @scrapezy/mcpCommand-line Argument:
npx @scrapezy/mcp --api-key=your_api_key
To use with Claude Desktop, add the server config:
On MacOS: ~/Library/Application Support/Claude/claude_desktop_config.json
On Windows: %APPDATA%/Claude/claude_desktop_config.json
{
"mcpServers": {
"scrapezy": {
"command": "npx @scrapezy/mcp --api-key=your_api_key"
}
}
}Example Usage in Claude
You can use this tool in Claude with prompts like:
Please extract product information from this page: https://example.com/product
Extract the product name, price, description, and available colors.Claude will use the MCP server to extract the requested structured data from the website.
Debugging
Since MCP servers communicate over stdio, debugging can be challenging. We recommend using the MCP Inspector, which is available as a package script:
npm run inspectorThe Inspector will provide a URL to access debugging tools in your browser.
License
MIT
Available Tools
1 toolextract-structured-dataC
Extract structured data from a website.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Prompt to extract data from the website | |
| url | Yes | URL of the website to extract data from |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool extracts data but doesn't describe how it behaves: e.g., whether it makes HTTP requests, handles authentication, respects robots.txt, has rate limits, or returns errors. This leaves the agent guessing about operational details, which is a significant gap for a tool with no annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without unnecessary words. It's front-loaded and wastes no space, making it easy to parse quickly. Every word earns its place by conveying the core action and target, though it could benefit from more detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (data extraction from websites) and lack of annotations and output schema, the description is incomplete. It doesn't explain what 'structured data' means, how extraction works, what the return format might be, or any behavioral traits. This leaves the agent with insufficient information to use the tool effectively in real-world scenarios.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with clear descriptions for both parameters: 'url' as the website URL and 'prompt' as the extraction prompt. The description adds no additional semantic context beyond the schema, such as examples of valid prompts or URL formats. Since the schema does the heavy lifting, the baseline score of 3 is appropriate, as the description doesn't compensate but doesn't detract either.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Extract structured data from a website' clearly states the action (extract) and resource (structured data from a website), but it's somewhat vague about what 'structured data' entails. It doesn't specify the type of data (e.g., tables, product info, contact details) or extraction method, leaving room for interpretation. No sibling tools exist for comparison, so differentiation isn't applicable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives, prerequisites, or limitations. It lacks context such as whether it's for scraping, parsing HTML, or using APIs, and doesn't mention any constraints like website accessibility or rate limits. Without siblings, it doesn't need to distinguish from them, but still offers no usage instructions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v1.0.0- First observed
extract-structured-data
TDQS
Scored across 1 tool
With only one tool, there is no possibility of ambiguity or overlap between tools. The tool's purpose is clearly defined as extracting structured data from websites, making it distinct by default.
Since there is only one tool, naming consistency is inherently perfect. The tool name 'extract-structured-data' follows a clear verb_noun pattern, and there are no other tools to compare it against for inconsistency.
A single tool is generally too few for a server's purpose, especially for a scraping domain which often requires multiple operations like fetching, parsing, or handling different data types. This minimal set feels thin and limits functionality.
For a scraping server, the tool surface is severely incomplete. It lacks essential operations such as fetching raw HTML, handling pagination, managing sessions, or error handling, which are critical for effective web scraping workflows.
Related MCP Connectors
A Model Context Protocol server for Wix AI tools
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…
MCP server for Firecrawl — web search, scraping, and biomedical/arXiv paper search.
Model Context Protocol server for Studex tools, notifications, and profile integrations
Related MCP Servers
- AlicenseAqualityDmaintenanceA Model Context Protocol server that provides web content fetching and conversion capabilities.4293 npm4MIT
- AlicenseAqualityDmaintenanceA production-ready Model Context Protocol server that enables language models to leverage AI-powered web scraping capabilities, offering tools for transforming webpages to markdown, extracting structured data, and executing AI-powered web searches.8113MIT
- AlicenseNot gradedqualityDmaintenanceA server that implements the Model Context Protocol, providing a standardized way to connect AI models to different data sources and tools.8 npm11MIT
- FlicenseNot gradedqualityDmaintenanceA Model Context Protocol server that provides AI models with structured access to external data and services, acting as a bridge between AI assistants and applications, databases, and APIs in a standardized, secure way.2-