Skip to main content
Glama
dev55acc-ai

website-content-crawler

by dev55acc-ai

Website Content Crawler

爬取一个 URL 列表,并将页面内容作为结构化 JSON 返回。以 Apify ActorMCP server 方式运行,以便 Agent 可以直接调用它。

在线演示:https://website-content-crawler.vercel.app

运行成本

在 Apify 上按事件付费:

事件

价格

Actor 启动

$0.00005

返回页面记录

$0.001

平台计算费用会在收益结算前扣除;单位经济效益参见 PRICING.md

Related MCP server: cleanfetch

输出

每次运行都会返回以下完全相同的结构。输出 schema 会强制保证该结构。

{
  "url": "https://example.com",
  "status": 200,
  "title": "Example Domain",
  "description": null,
  "h1": "Example Domain",
  "markdown": "# Example Domain\n\nThis domain is for use in documentation examples without needing permission. Avoid use in operations.\n\nLearn more\n\nLearn more (https://iana.org/domains/example)",
  "links": [
    "https://iana.org/domains/example"
  ],
  "wordCount": 23,
  "crawledAt": "2026-08-23T09:55:20.140Z",
  "error": null
}

目录结构

  • src/ — Actor 源代码(Crawlee Cheerio 爬虫)

  • .actor/ — Apify Actor 规范:actor.jsonINPUT_SCHEMA.jsonoutput_schema.jsondataset_schema.jsonDockerfile

  • mcp-server/ — MCP 包装器,将 Actor 暴露为 crawl_website 工具

  • site/ + tools/ — 根据真实运行记录构建在线演示页面

本地运行

npm install
npm test

MCP 服务器:

cd mcp-server && npm install && node smoke.mjs

发布到 Apify Store

当账户激活后:apify push,按事件付费,listing 使用本 README。

A
license - permissive license
Not graded
quality - not tested
C
maintenance

Maintenance

Maintainers
Response time
Release cycle
Releases (12mo)
Commit activity

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Servers

  • A
    license
    B
    quality
    D
    maintenance
    Enables LLMs to retrieve and process web content by fetching URLs and converting HTML to markdown format. Supports chunked reading of large pages and can access both public websites and local networks.
    1
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Enables AI agents to read web pages reliably, returning clean markdown content, hyperlinks, and metadata without navigation or ad noise.
    3
    14
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables AI agents to fetch and extract clean, readable content from web pages, and search within pages for specific queries, without needing a full browser.
    1
    MIT

View all related MCP servers

Related MCP Connectors

  • Read a URL as clean markdown, screenshot a website, url to PDF. Web access for agents, no signup.

  • Web scraping for AI agents. Converts URLs to clean, LLM-ready Markdown with anti-bot bypass.

  • Read any web page as clean Markdown for AI agents: fetch, search, metadata, links. SSRF-safe.

View all MCP Connectors

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/dev55acc-ai/website-content-crawler'

If you have feedback or need assistance with the MCP directory API, please join our Discord server