Skip to main content
Glama
AceDataCloud

mcp-webextrator

by AceDataCloud
README.md
# MCP WebExtrator Server

<!-- mcp-name: io.github.AceDataCloud/mcp-webextrator -->

A Model Context Protocol (MCP) server for web rendering and structured content extraction
via the AceDataCloud WebExtrator platform.

## Features

- **Structured extraction**: Pull structured data out of any URL via WebExtrator
- **Web rendering**: Render dynamic JavaScript pages and capture the rendered output
- **Asynchronous tasks**: Submit extract / render jobs and poll for results
- **Batch task lookup**: Query multiple task results in one call

## Installation

```bash
pip install mcp-webextrator
```

## Configuration

Set your AceDataCloud API token:

```bash
export ACEDATACLOUD_API_TOKEN=your_token_here
```

Get your token from [https://platform.acedata.cloud](https://platform.acedata.cloud).

## Usage

### stdio mode (default)

```bash
mcp-webextrator
```

### HTTP mode

```bash
mcp-webextrator --transport http --port 8000
```

## Available Tools

| Tool | Description |
|------|-------------|
| `webextrator_extract` | Extract structured content from a URL |
| `webextrator_render` | Render a dynamic web page and return the rendered output |
| `webextrator_get_task` | Get the status / result of an extract or render task |
| `webextrator_get_tasks_batch` | Batch-fetch the status / result of multiple tasks |
| `webextrator_get_usage_guide` | Get the API usage guide |

## Documentation

<!-- canonical-documentation -->
[Documentation](https://platform.acedata.cloud/documents/webextrator)

## License

MIT — see [LICENSE](LICENSE).

TDQS

A4.3/5.0

Scored across 5 tools

Disambiguation4/5

Extract and render both involve fetching a URL, but their output types are clearly differentiated: structured content vs. raw rendered HTML. Get_task and get_tasks_batch are likewise distinct by single vs. batch retrieval, with descriptions that reinforce when each should be used.

Naming Consistency5/5

All tools follow the same webextrator_verb_noun pattern with consistently descriptive verbs: extract, render, get_usage_guide, get_task, get_tasks_batch. The naming is uniform, predictable, and easy to navigate.

Tool Count5/5

Five tools is a well-scoped set for a web extraction server: two core operations, two async retrieval methods, and a usage guide. Each tool serves a clear purpose without unnecessary bloat.

Completeness5/5

The tool surface covers the full lifecycle: submit extract/render tasks, poll a single task, and retrieve tasks in batch. The usage guide fills discoverability gaps, and the async=true path is well integrated with task retrieval, leaving no obvious dead ends.

Maintenance

ActivityActive
ResponsivenessNo issues