Skip to main content
Glama

🕷️ LookaCrawler

免费、开源、token 高效的 Firecrawl 本地替代方案,为 LLM 提供原生 Model Context Protocol (MCP) 服务器。

License: MIT Bun TypeScript MCP Ready GitHub Stars Local-First


📌 为什么选择 LookaCrawler?

面向大型语言模型(LLM)的网页爬取默认是有缺陷的:现代网页包含大量 HTML 冗余(脚本、跟踪像素、嵌套的 div、导航栏、样式表),每页浪费数千 token。

LookaCrawler 是一款开源的 token 优化型本地爬虫,可去除 >73% 至 90% 的网页冗余,提取干净的 Markdown,通过 stealth Playwright 驱动绕过反机器人屏障,并提供原生 Model Context Protocol (MCP) 服务器,可直接用于 Claude Desktop、Cursor 和 Antigravity


Related MCP server: Scraper MCP

🥊 对比:LookaCrawler 与替代方案

功能

🕷️ LookaCrawler

🔥 Firecrawl (Cloud)

⚡ Jina Reader

价格 / 成本

$0.00(100% 免费开源)

$16 至 $99+/月

限速 API

Token 削减

>73% 至 90% 的激进裁剪

标准 Markdown

基础 Markdown

数据隐私

100% 本地(零遥测)

云服务商

云 API

MCP 集成

原生一键服务器(stdio 与 SSE)

社区封装

Stealth 与反机器人

真实 Chrome + Stealth 指纹

云代理

基础请求头

本地 SQLite 缓存

内置(24 小时 TTL 缓存)

Redis / 付费插件

JS SPA 支持

Playwright + Chrome 池

云端无头浏览器

无头浏览器


🚀 主要特性

  • **Token 经济优先:**在将 Markdown 传递给您的 LLM 之前,自动裁剪脚本、样式、内联 SVG、跟踪标签、导航、页脚和冗余表单。

  • 双混合爬取引擎:

    • fast:超快的原生 HTTP GET,带退避重试。遇到反机器人挑战时,自动升级为 deep

    • deep:无头 Playwright 引擎,启动真实 Google Chrome 并应用 stealth 补丁(清除 navigator.webdriver、伪造 WebGL、剥离 CDP 泄漏),从而透明地爬取受 Cloudflare/Turnstile 保护的页面。

  • **原生 MCP 服务器协议:**支持 JSON-RPC 2.0 stdio 与 SSE 传输,可直接用于 Claude Desktop、Cursor、Windsurf 和 Antigravity。

  • **本地 SQLite 缓存:**将提取的 Markdown 存储在 crawler_cache.sqlite 中,以消除重复的网络请求。

  • **结构化 JSON 与元数据提取:**提取 Open Graph 标签(og:titleog:description)、发布日期、canonical URL 和自定义 CSS 选择器。


🔌 一键 MCP 设置(Claude Desktop 与 Cursor)

将 LookaCrawler 添加到您的 claude_desktop_config.json 或 Cursor MCP 设置中:

{
  "mcpServers": {
    "lookacrawler": {
      "command": "bun",
      "args": ["run", "/absolute/path/to/lookacrawler/index.ts"]
    }
  }
}

现在,您可以提示 Claude 或 Cursor:

"使用 LookaCrawler 爬取 https://example.com/docs 并提取 API 文档。"


📦 快速开始与 CLI 用法

1. 安装

需要 Bun 1.1+(或 Node.js 20+):

# Clone the repository
git clone https://github.com/lucasmartins-ai/lookacrawler.git
cd lookacrawler

# Install dependencies
bun install

2. CLI 命令

# Single URL fast Markdown extraction
bun run cli.ts extract https://news.ycombinator.com --mode fast

# Headless Playwright deep extraction with CSS selector target
bun run cli.ts extract https://example.com --mode deep --selector "main" --json

# Batch concurrent multi-URL crawling
bun run cli.ts batch https://site1.com https://site2.com --concurrency 4

# Structured JSON schema extraction
bun run cli.ts structured https://example.com --schema '{"title":"h1","links":"a"}'

# Start MCP Server via SSE on port 3000
bun run cli.ts serve --transport sse --port 3000

3. Docker 部署

# Build and run Docker container
docker build -t lookacrawler .
docker run -p 3000:3000 lookacrawler

🧪 架构与测试

Incoming URL ──► [Local SQLite Cache Check] ──(Hit)──► Return Cached Markdown
                       │ (Miss)
                       ▼
            [Fast HTTP GET Request] ──(Blocked?)──► [Auto-Escalate to Deep Stealth]
                       │                                      │
                       ▼                                      ▼
            [HTML DOM Tree Parser] ◄──────────────────────────┘
                       │
                       ▼
       [Aggressive Token Noise Pruner]
       (Strips SVG, Nav, Ads, Tracking, CSS, JS)
                       │
                       ▼
         [Mozilla Readability Engine]
                       │
                       ▼
         [Turndown Markdown Converter] ──► Return Clean LLM Markdown

运行测试套件:

bun test

⭐ Star 与支持

如果 LookaCrawler 为您节省了 API 费用和 token 成本:

  • 为这个仓库标星,帮助其他开发者发现它!

  • 💡 提交 Issue / PR,提出新的 stealth 绕过手段或爬虫功能。


由 LookADev 构建

lookacrawlerLookADev 构建和维护,这是一家专注于 AI 智能体、Web 架构和 token 优化的工程工作室。

启动项目 → lookadev.com · 邮箱:lucas@lookadev.com

📄 许可证

开源软件,根据 MIT 许可证 授权。

Available Tools

3 tools
batch_extract_web_contentA

Batch extract token-optimized Markdown content from multiple website URLs concurrently with aggregate token statistics.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoCrawl mode: 'fast' (native HTTP fetch) or 'deep' (headless Playwright browser).fast
urlsYesArray of target website URLs to extract.
proxyNoOptional HTTP/SOCKS5 proxy URL.
cookiesNoOptional custom HTTP cookies key-value dictionary.
headersNoOptional custom HTTP request headers key-value dictionary.
concurrencyNoMaximum parallel HTTP/browser crawl worker concurrency (default: 3).
max_retriesNoMaximum retry attempts per URL (default: 3).
css_selectorNoOptional CSS selector to filter DOM node across all target URLs.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral disclosure burden. It does disclose concurrency, output format, and token statistics, which are genuinely useful. However, it does not mention failure behavior, retries, partial failures, rate limits, or what happens when a URL cannot be fetched.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, compact sentence that front-loads the core purpose, then adds concurrency and statistical output details. There is no filler, and every phrase adds distinguishing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description must supply more contextual completeness. It gives the essential purpose and concurrency trait, but it omits output shape, retry/failure semantics, and any usage comparison to sibling tools. For an 8-parameter tool with 100% schema coverage, this is still not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% parameter description coverage, so the schema already documents all parameters well. The description adds no parameter-specific meaning beyond stating that the tool works on multiple URLs; therefore the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the action ('batch extract'), resource ('web content from multiple website URLs'), output format ('token-optimized Markdown'), and a distinctive trait ('concurrently with aggregate token statistics'). This distinguishes it from the likely single-URL sibling `extract_web_content` and from structured extraction (`extract_structured_data`).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The word 'batch' and phrase 'multiple website URLs' imply this tool is for multi-URL scenarios, which provides useful context. However, there is no explicit guidance about when to prefer this over `extract_web_content` or `extract_structured_data`, and no stated exclusion conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

extract_structured_dataC

Extract page metadata (OG tags, canonical URL, author, dates) and custom CSS selector JSON schema mapping from a website.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesTarget website URL to extract content and metadata from.
modeNoCrawl mode: 'fast' (native fetch) or 'deep' (Playwright Chromium).fast
proxyNoOptional HTTP/SOCKS5 proxy URL.
schemaNoOptional key-value map of property names to CSS selectors (e.g. { title: 'h1', price: '.price' }).
cookiesNoOptional custom HTTP cookies key-value dictionary.
headersNoOptional custom HTTP request headers key-value dictionary.
max_retriesNoMaximum retry attempts (default: 3).
css_selectorNoOptional CSS selector to scope content before processing.
include_metadataNoWhether to extract Open Graph tags, canonical URL, author, and date metadata (default: true).

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure, but it only states what is extracted and from where. It does not disclose that this likely performs a live network fetch, any side effects, permission or rate-limit considerations, or distinctions between the 'fast' and 'deep' modes in terms of behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single focused sentence with no superfluous words. It leads with the main action and resources, making it efficient and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 9 parameters, nested objects, no output schema, and sibling tools to disambiguate, this one-sentence description is insufficiently complete. It omits usage guuidelines, behavioral expectations, and return-value structure, leaving meaningful gaps for an agent trying to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema covers all 9 parameters with detailed descriptions, including examples, so the baseline is 3. The description adds little beyond the schema: it lists example metadata fields (OG tags, canonical URL, author, dates) which clarifies output scope, but does not elaborate on parameter syntax or interplay.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb ('Extract') and resources ('page metadata', 'custom CSS selector JSON schema mapping'), making the tool's purpose understandable. It does not explicitly differentiate from sibling tools like extract_web_content or batch_extract_web_content, so it misses the top tier by a narrow margin.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided about when to use this tool versus the sibling tools extract_web_content or batch_extract_web_content. The description does not mention any conditions, alternatives, or exclusions, leaving the agent to infer usage entirely from names.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

extract_web_contentC

Extract token-optimized clean Markdown content from a target website for LLM consumption.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesTarget website URL to extract content from.
modeNoCrawl mode: 'fast' (native HTTP fetch) or 'deep' (headless Playwright browser with JS execution).fast
proxyNoOptional HTTP/SOCKS5 proxy URL (e.g. 'http://proxy.example.com:8080').
cookiesNoOptional custom HTTP cookies key-value dictionary.
headersNoOptional custom HTTP request headers key-value dictionary.
max_retriesNoMaximum retry attempts for transient errors or rate limits (default: 3).
css_selectorNoOptional CSS selector to scope content extraction to a specific HTML node.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full burden of behavioral disclosure. It only states the output is token-optimized clean Markdown; it does not reveal that the tool makes live network requests, that 'deep' mode executes JavaScript via a headless browser, how failures or rate limits are handled, or any side effects. The schema documents the mode options, but the description itself leaves key behavioral traits undisclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that communicates the primary purpose and output format without wasted words. It is appropriately concise, though a little more detail about usage trade-offs would have made it richer without harming structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 7 parameters, nested objects, and no output schema or annotations, yet the description only covers the surface purpose. It omits important operational context such as dynamic content handling, proxy/cookie use cases, retry behavior, and how to decide between 'fast' and 'deep' modes. An agent receives the raw schema but not sufficient high-level orientation for a tool of this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter already has a clear description in the schema. The tool description adds little beyond 'clean Markdown', which indirectly hints at the purpose of css_selector scoping but does not meaningfully extend parameter understanding. Baseline 3 is appropriate because the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description has a clear verb ('Extract'), a specific resource ('content from a target website'), and a defined output format ('clean Markdown content'). It conveys the tool's core purpose well and inherently contrasts with extract_structured_data by promising Markdown. However, it does not explicitly distinguish itself from batch_extract_web_content, leaving the single-URL versus batch distinction to be inferred from the sibling name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided about when to use this tool versus its siblings. There is no mention of batch_extract_web_content for bulk jobs or extract_structured_data for non-Markdown outputs. The 'for LLM consumption' phrase gives some context but does not help an agent choose among alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 3 tool updatesv0.1.0
    • First observedbatch_extract_web_content
    • First observedextract_structured_data
    • First observedextract_web_content

TDQS

B3.3/5.0
Disambiguation4/5

The three tools have clear boundaries: single content extraction, batch content extraction, and structured metadata/CSS mapping. The only soft spot is that batch and single share the same core operation, but their singular-vs-batch distinction prevents real ambiguity.

Naming Consistency5/5

All tools use lowercase snake_case and follow a predictable extract_<object> naming pattern; batch_ is a standard parallelism modifier on the same verb. There are no mixed conventions or vague verbs.

Tool Count4/5

Three tools is appropriate for a narrowly scoped extraction server: one direct, one batched, and one structured metadata. It is not bloated, and slightly minimal but credible for this purpose.

Completeness3/5

Core extraction workflow is covered: single page, batch pages, and structured metadata. However, the crawler name implies link discovery or site traversal, and no tool enumerates links or sitemaps, so agents need URLs supplied beforehand. This is a notable but not deabilitating gap.

Maintenance

ActivityMaintained
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    A Model Context Protocol server that enables web scraping, crawling, and content extraction capabilities through integration with Firecrawl.
    8
    28,853
    2
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    A context-optimized web scraping server that converts HTML to markdown/text and applies CSS selectors server-side, reducing token usage by 70-90% while providing AI tools with clean, filtered web content.
    7
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Web content extraction for AI agents. 10 tools: scrape, crawl, map, batch, extract, summarize, diff, brand, search, research. Uses TLS fingerprinting to bypass anti-bot without a headless browser. Outputs LLM-optimized markdown with 67% fewer tokens than raw HTML.
    10
    2,316
    AGPL 3.0

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/lucasmartins-ai/lookacrawler'

If you have feedback or need assistance with the MCP directory API, please join our Discord server