MCP Server Steel Scraper
Provides a stateful tool 'search' that opens Google search results for a query, enabling web search and interaction with search result pages.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MCP Server Steel Scrapervisit https://example.com and return the page title in plain text"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MCP Server Steel Scraper
A simple Model Context Protocol (MCP) server that wraps the steel-dev API for visiting websites with browser automation.
Quick Start
Install the package:
npm install -g @jharding_npm/mcp-server-steel-scraperAdd to your MCP client configuration:
{ "mcpServers": { "steel-scraper": { "command": "npx", "args": ["@jharding_npm/mcp-server-steel-scraper", "--mode=both"], "env": { "STEEL_API_URL": "http://localhost:3000" } } } }Start using the stateless
visit_with_browsertool, or the stateful interactive tools.
Related MCP server: ma-browser
Features
Dual Modes: Run stateless scraping, stateful interaction, or both via
--mode=stateless|stateful|both(default:both)Stateless Tool:
visit_with_browser- Visit websites using steel-dev APIStateful Tools: Create sessions and interact with pages (navigate, click, type, scroll, snapshot)
Flexible Return Types: HTML, markdown, readability, or cleaned HTML
Local/Remote Support: Works with local or remote steel-dev instances
Browser Automation: Screenshot capture, PDF generation, proxy support
Smart Length Management: Single
maxLengthparameter with intelligent defaults and automatic content/metadata splitClean Output by Default: Minimal metadata output perfect for 7B models and summarization
Verbose Mode: Optional full metadata when detailed information is needed
TypeScript: Fully typed implementation
Installation
Option 1: NPM Package (Recommended)
Install the package globally to use it with npx:
npm install -g @jharding_npm/mcp-server-steel-scraperOr use it directly with npx without installing:
npx @jharding_npm/mcp-server-steel-scraperOption 2: Local Development
Clone this repository:
git clone <repository-url>
cd mcp-server-steel-scraperInstall dependencies:
npm installBuild the project:
npm run buildConfiguration
The server uses environment variables for configuration:
STEEL_API_URL: The steel-dev API endpoint (default:http://localhost:3000)STEEL_TIMEOUT: Request timeout in milliseconds (default:30000)STEEL_RETRIES: Number of retry attempts (default:3)STEEL_LOCAL: Set totruewhen using a local Steel instance for stateful sessionsSTEEL_BASE_URL: Base URL for the Steel Sessions API (default:https://api.steel.dev, orhttp://localhost:3000whenSTEEL_LOCAL=true)STEEL_API_KEY: Required for cloud mode stateful sessionsSTEEL_SESSION_TIMEOUT_MS: Session timeout in milliseconds (default:900000)STEEL_GLOBAL_WAIT_SECONDS: Optional delay after each stateful action (default:0)STEEL_IDLE_TIMEOUT_MS: Auto-release idle sessions after this many milliseconds (default:600000, set to0to disable)
Copy env.example to .env and modify as needed:
cp env.example .envUsage
Running the Server
# Development mode
npm run dev
# Auto-rebuild on changes (recommended for npm link workflows)
npm run build:watch
# Production mode
npm start
# Only stateless scraping tools
npm start -- --mode=stateless
# Only stateful interactive tools
npm start -- --mode=stateful
# Both tool sets (default)
npm start -- --mode=bothMCP Client Configuration
Add this server to your MCP client configuration. Here are examples for popular LLM clients:
For Claude Desktop / Cline / Other MCP Clients (NPM Package)
{
"mcpServers": {
"steel-scraper": {
"command": "npx",
"args": ["@jharding_npm/mcp-server-steel-scraper", "--mode=stateless"],
"env": {
"STEEL_API_URL": "http://localhost:3000"
}
}
}
}To expose the stateful interactive tools, add --mode=stateful or --mode=both to the args array.
For Continue.dev (NPM Package)
{
"mcpServers": {
"steel-scraper": {
"command": "npx",
"args": ["@jharding_npm/mcp-server-steel-scraper", "--mode=stateless"],
"env": {
"STEEL_API_URL": "http://localhost:3000"
}
}
}
}For Cursor IDE (NPM Package)
{
"mcpServers": {
"steel-scraper": {
"command": "npx",
"args": ["@jharding_npm/mcp-server-steel-scraper", "--mode=stateless"],
"env": {
"STEEL_API_URL": "http://localhost:3000"
}
}
}
}For Remote Steel-dev Instance (NPM Package)
{
"mcpServers": {
"steel-scraper": {
"command": "npx",
"args": ["@jharding_npm/mcp-server-steel-scraper", "--mode=stateless"],
"env": {
"STEEL_API_URL": "https://your-steel-dev-instance.com"
}
}
}
}Alternative: Using Global Installation
If you've installed the package globally with npm install -g @jharding_npm/mcp-server-steel-scraper, you can use:
{
"mcpServers": {
"steel-scraper": {
"command": "mcp-server-steel-scraper",
"env": {
"STEEL_API_URL": "http://localhost:3000"
}
}
}
}For Local Development (using absolute path)
{
"mcpServers": {
"steel-scraper": {
"command": "node",
"args": ["/path/to/mcp-server-steel-scraper/dist/index.js"],
"env": {
"STEEL_API_URL": "http://localhost:3000"
}
}
}
}Tool Usage
The server provides one tool: visit_with_browser
Parameters
url(required): The URL to visitformat(optional): Content formats to extract -["html"]for raw HTML source (may be very large),["markdown"]for clean formatted text converted from HTML (recommended for reading),["readability"]for Mozilla Readability format,["cleaned_html"]for cleaned HTML. You can request multiple formats (default:["markdown"])screenshot(optional): Take a screenshot of the page (returns base64 encoded image) (default:false)pdf(optional): Generate a PDF of the page (returns base64 encoded PDF) (default:false)proxyUrl(optional): Proxy URL to use for the request (e.g.,"http://proxy:port")delay(optional): Delay in seconds to wait after page load before scraping (default:0)logUrl(optional): URL to send logs to for debugging purposesmaxLength(optional): Maximum characters to return. Smart defaults: markdown=8000, readability=10000, html=15000, cleaned_html=12000. For markdown, automatically reserves space for metadataverboseMode(optional): Return full metadata instead of clean content-focused output (default: false). Use when you need detailed visit information
Example Usage
// Basic website visit
{
"tool": "visit_with_browser",
"arguments": {
"url": "https://example.com"
}
}
// Advanced visit with multiple formats
{
"tool": "visit_with_browser",
"arguments": {
"url": "https://example.com",
"format": ["markdown", "html"],
"screenshot": true,
"delay": 2
}
}
// Simple visit with smart defaults (perfect for 7B models)
{
"tool": "visit_with_browser",
"arguments": {
"url": "https://example.com",
"format": ["markdown"]
}
}
// Custom length limit (automatically handles content vs metadata split)
{
"tool": "visit_with_browser",
"arguments": {
"url": "https://en.wikipedia.org/wiki/Long_Article",
"format": ["markdown"],
"maxLength": 5000
}
}
// Verbose mode when you need detailed visit information
{
"tool": "visit_with_browser",
"arguments": {
"url": "https://example.com",
"format": ["markdown"],
"maxLength": 8000,
"verboseMode": true
}
}
// With proxy and PDF generation
{
"tool": "visit_with_browser",
"arguments": {
"url": "https://example.com",
"format": ["readability"],
"pdf": true,
"proxyUrl": "http://proxy:8080"
}
}Stateful Interactive Tools
When running with --mode=stateful or --mode=both, the server exposes stateful tools that let the LLM interact with a live page.
Stateful sessions are created via the Steel Sessions API and connected over CDP (Chrome DevTools Protocol).
Available Tools
session_create- Create a new Steel session and connectsession_release- Release the current sessionnavigate- Navigate to a URLsearch- Open Google search results for a queryclick- Click an element by labeltype- Type into an element by labelscroll_down/scroll_up- Scroll the pagego_back- Navigate backwait- Wait a few seconds for dynamic contentsnapshot- Annotated screenshot + labels listsnapshot_unmarked- Screenshot without labelspage_content- Return page HTML or text
Example Session
// Create a session
{
"tool": "session_create",
"arguments": { "timeoutMs": 900000 }
}
// Navigate
{
"tool": "navigate",
"arguments": { "url": "https://example.com" }
}
// Get an annotated snapshot (labels + image)
{
"tool": "snapshot",
"arguments": {}
}
// Click a labeled element
{
"tool": "click",
"arguments": { "label": 3 }
}
// Type into a labeled input
{
"tool": "type",
"arguments": { "label": 5, "text": "hello", "replaceText": true }
}Smart Length Management
The server automatically handles content length optimization:
Unified Length Control: Single
maxLengthparameter handles both content and metadataAutomatic Content/Metadata Split: For markdown, reserves 10% for metadata, uses 90% for content
Smart Defaults: Reasonable defaults when no length is specified (markdown=8000, text=10000, html=15000, json=5000)
Better Truncation: Avoids double-truncation issues that could result in incomplete content
Conversion Detection: Automatically detects when HTML-to-markdown conversion may have failed
Warning System: Provides warnings when content appears truncated or incomplete
How It Works
// Simple usage - uses smart defaults
{
"url": "https://example.com",
"format": ["markdown"]
// Automatically uses 8000 characters, reserves 800 for metadata, 7200 for content
}
// Custom length - automatically splits appropriately
{
"url": "https://example.com",
"format": ["markdown"],
"maxLength": 5000
// Uses 5000 total, reserves 500 for metadata, 4500 for content
}This approach ensures you get complete, properly formatted content while maintaining simple, intuitive parameter management.
Handling Large Pages (Like Amazon)
For large, complex pages like Amazon.com, follow these best practices:
Recommended Approach for Complex Pages
{
"tool": "visit_with_browser",
"arguments": {
"url": "https://www.amazon.com",
"format": ["readability"], // Most reliable for complex pages
"maxLength": 5000, // Reasonable limit for large pages
"delay": 3 // Wait for main content to load
}
}Format Comparison for Large Pages
HTML: Returns raw HTML source (can be 900,000+ characters for Amazon)
Readability: Mozilla Readability format (most reliable, good for complex pages)
Markdown: Converts HTML to clean, readable text (may fail on complex pages like Amazon)
Cleaned HTML: Cleaned HTML with better structure
Note: Markdown conversion may fail on complex, JavaScript-heavy pages like Amazon. Use ["readability"] for the most reliable results.
Troubleshooting
If you get HTML instead of Markdown:
The steel-dev API may not support markdown conversion for that page type
Try using
format: ["readability"]instead for better text extractionComplex pages with heavy JavaScript may not convert properly
If you get truncated content:
The page may be too large for the specified
maxLengthTry increasing
maxLengthor using a longerdelayConsider using
format: ["readability"]for more reliable truncation
For Dynamic Content
Use delay parameter to wait for content to load:
{
"tool": "visit_with_browser",
"arguments": {
"url": "https://www.amazon.com",
"format": ["markdown"],
"delay": 5, // Wait 5 seconds for content to load
"maxLength": 10000 // Longer content for complex pages
}
}Clean Output by Default
The server is designed with 7B models in mind, providing clean, content-focused output by default:
Content Summarization: Perfect for weaker models that need to summarize web content
Content Analysis: Ideal for processing large amounts of text
Context Optimization: Maximizes the content-to-metadata ratio automatically
How It Works
Default Mode (clean output):
# Article Title
This is the actual content...Verbose Mode (verboseMode: true):
SUCCESS: Successfully scraped https://example.com
Method: full-browser-automation (stealth browser, anti-detection)
Format: markdown
Status Code: 200
Processing Time: 1250ms
Content Length: 5000 characters
Content Type: text/html
Timestamp: 2024-01-15T10:30:00.000Z
Title: Article Title
Description: Article description
Language: en
Screenshot: Available (base64)
Links Found: 15
SCRAPED CONTENT:
# Article Title
This is the actual content...Benefits of Clean Output
Maximum Content Space: Removes ~200-300 characters of metadata overhead
Cleaner Output: Direct content without verbose headers
Better for 7B Models: Focuses the model's attention on the actual content
Preserves Warnings: Still shows important warnings if conversion issues occur
Recommended Usage
For summarization tasks, use the default clean output:
{
"tool": "visit_with_browser",
"arguments": {
"url": "https://article-to-summarize.com",
"format": ["markdown"],
"maxLength": 10000 // Automatically optimizes content vs metadata split
}
}Steel-dev API Requirements
This MCP server expects a steel-dev API instance running with the following endpoints:
POST /scrape- Main scraping endpointGET /health- Health check endpoint (optional)GET /info- API information endpoint (optional)POST /v1/sessions- Create a stateful browser sessionPOST /v1/sessions/{id}/release- Release a stateful session
Expected Request Format
{
"url": "https://example.com",
"format": ["html", "markdown"],
"screenshot": true,
"pdf": false,
"proxyUrl": "http://proxy:8080",
"delay": 2,
"logUrl": "https://logs.example.com"
}Expected Response Format
{
"content": {
"html": "<html>...</html>",
"markdown": "# Title\nContent..."
},
"metadata": {
"title": "Page Title",
"description": "Page description",
"statusCode": 200,
"timestamp": "2024-01-15T10:30:00.000Z"
},
"links": [
{"url": "https://example.com/link1", "text": "Link Text"}
],
"screenshot": "base64...",
"pdf": "base64..."
}Development
Project Structure
src/
├── index.ts # Main MCP server implementation
├── steel-api.ts # Steel-dev API wrapper
└── config.ts # Configuration managementScripts
npm run build- Build TypeScript to JavaScriptnpm run start- Run the built servernpm run dev- Run in development mode with tsx
Adding New Features
Modify the tool schema in
src/index.tsUpdate the
SteelAPIclass insrc/steel-api.tsif neededRebuild and test
Error Handling
The server includes comprehensive error handling:
Network errors are caught and returned as error responses
Invalid parameters are validated
Steel-dev API errors are properly forwarded
Timeout handling for long-running requests
License
MIT
Available Tools
1 toolvisit_with_browserA
Visit any website using full browser automation (stealth mode, anti-detection). Returns page content in your chosen format: 'html' for raw HTML source, 'markdown' for clean formatted text (recommended for reading), 'readability' for Mozilla Readability format, or 'cleaned_html' for cleaned HTML. Supports screenshot and PDF generation. Automatically handles JavaScript rendering and provides clean output by default.
| Name | Required | Description | Default |
|---|---|---|---|
| No | Generate a PDF of the page (returns base64 encoded PDF) | ||
| url | Yes | The complete URL to scrape (must include http:// or https://) | |
| delay | No | Delay in seconds to wait after page load before scraping | |
| format | No | Content formats to extract: 'html'=raw HTML source (may be very large), 'markdown'=clean formatted text converted from HTML (recommended for reading), 'readability'=Mozilla Readability format, 'cleaned_html'=cleaned HTML. You can request multiple formats. | |
| logUrl | No | URL to send logs to for debugging purposes | |
| proxyUrl | No | Proxy URL to use for the request (e.g., 'http://proxy:port') | |
| maxLength | No | Maximum characters to return (optional). Smart defaults: markdown=8000, readability=10000, html=15000, cleaned_html=12000. For markdown, automatically reserves space for metadata. | |
| screenshot | No | Take a screenshot of the page (returns base64 encoded image) | |
| verboseMode | No | Return full metadata instead of clean content-focused output (optional, default: false). Use when you need detailed scraping information. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses stealth mode/anti-detection, automatic JavaScript rendering, and clean output by default, and notes screenshot/PDF return base64. It omits failure behavior, timeouts, and any rate-limiting or auth constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then capabilities, then return formats in three compact sentences. Minor redundancy in re-listing formats that the schema enum already enumerates, but no wasted filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Nine parameters and no output schema, yet the description covers the primary capability set, default output behavior, and the shape of non-text returns (base64). Error/edge-case behavior is the only notable gap; overall complete enough to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every parameter including format meanings and maxLength smart defaults. The description largely mirrors that (format enum meanings) rather than adding semantics beyond it, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Visit any website using full browser automation' and enumerates the output formats. Very clear what the tool does, though with no sibling tools there is nothing to differentiate against, capping it below the top mark.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage guidance is embedded in parameter advice ('recommended for reading' for markdown, 'Use when you need detailed scraping information' for verboseMode) but there is no explicit when-to-use/when-not framing for the tool overall, and no alternatives exist to route to.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v1.0.3- First observed
visit_with_browser
TDQS
Scored across 1 tool
Only one tool exists, so there is no possibility of confusion with other tools.
With a single tool, naming consistency is inherently perfect as there are no other names to compare against.
The server has only one tool, which is far too few for a robust web scraping service; typically multiple tools would be expected for different formats or actions.
The tool covers visiting a website and returning content in various formats, including screenshots and PDFs, but lacks separate tools for actions like extracting specific data or handling sessions, which might be needed for a full scraping workflow.
Maintenance
Related MCP Connectors
Turn any public website into an MCP server for agents to search, read and navigate.
Scrape, crawl and search the web for AI agents via MCP.
Stealth web browser for agents: search, fetch, click, download and type in persistent MCP sessions.
Hyperbrowser MCP — wraps the Hyperbrowser AI-agent browsing API
Related MCP Servers
AlicenseNot gradedqualityCmaintenanceAn open-source MCP server that provides browser automation capabilities to external AI systems, enabling navigation, DOM interaction, and web content extraction.20Apache 2.0- AlicenseNot gradedqualityAmaintenanceMCP server that lets AI agents use your real browser as an API, accessing any website with your login state, no keys or scrapers needed.1MIT
- AlicenseNot gradedqualityDmaintenanceA stealth-enhanced browser automation MCP server for AI agents to interact with websites while bypassing anti-bot detection mechanisms like Cloudflare and reCAPTCHA.9MIT
- AlicenseAqualityBmaintenanceAn MCP server that enables AI assistants to visually inspect and interact with rendered web pages via a persistent headless Chromium browser, supporting navigation, screenshots, clicks, viewport resizing, and console log retrieval.81MIT