mcp-webextrator
# MCP WebExtrator Server
<!-- mcp-name: io.github.AceDataCloud/mcp-webextrator -->
A Model Context Protocol (MCP) server for web rendering and structured content extraction
via the AceDataCloud WebExtrator platform.
## Features
- **Structured extraction**: Pull structured data out of any URL via WebExtrator
- **Web rendering**: Render dynamic JavaScript pages and capture the rendered output
- **Asynchronous tasks**: Submit extract / render jobs and poll for results
- **Batch task lookup**: Query multiple task results in one call
## Installation
```bash
pip install mcp-webextrator
```
## Configuration
Set your AceDataCloud API token:
```bash
export ACEDATACLOUD_API_TOKEN=your_token_here
```
Get your token from [https://platform.acedata.cloud](https://platform.acedata.cloud).
## Usage
### stdio mode (default)
```bash
mcp-webextrator
```
### HTTP mode
```bash
mcp-webextrator --transport http --port 8000
```
## Available Tools
| Tool | Description |
|------|-------------|
| `webextrator_extract` | Extract structured content from a URL |
| `webextrator_render` | Render a dynamic web page and return the rendered output |
| `webextrator_get_task` | Get the status / result of an extract or render task |
| `webextrator_get_tasks_batch` | Batch-fetch the status / result of multiple tasks |
| `webextrator_get_usage_guide` | Get the API usage guide |
## Documentation
<!-- canonical-documentation -->
[Documentation](https://platform.acedata.cloud/documents/webextrator)
## License
MIT — see [LICENSE](LICENSE).
TDQS
Scored across 5 tools
Extract and render both involve fetching a URL, but their output types are clearly differentiated: structured content vs. raw rendered HTML. Get_task and get_tasks_batch are likewise distinct by single vs. batch retrieval, with descriptions that reinforce when each should be used.
All tools follow the same webextrator_verb_noun pattern with consistently descriptive verbs: extract, render, get_usage_guide, get_task, get_tasks_batch. The naming is uniform, predictable, and easy to navigate.
Five tools is a well-scoped set for a web extraction server: two core operations, two async retrieval methods, and a usage guide. Each tool serves a clear purpose without unnecessary bloat.
The tool surface covers the full lifecycle: submit extract/render tasks, poll a single task, and retrieve tasks in batch. The usage guide fills discoverability gaps, and the async=true path is well integrated with task retrieval, leaving no obvious dead ends.