Skip to main content
Glama
Msparihar

MCP Server Firecrawl

by Msparihar

Firecrawl MCP 서버

Firecrawl API를 사용하여 웹 스크래핑, 콘텐츠 검색, 사이트 크롤링 및 데이터 추출을 위한 MCP(Model Context Protocol) 서버입니다.

특징

  • 웹 스크래핑 : 사용자 정의 옵션을 사용하여 모든 웹 페이지에서 콘텐츠 추출

    • 모바일 장치 에뮬레이션

    • 광고 및 팝업 차단

    • 콘텐츠 필터링

    • 구조화된 데이터 추출

    • 다양한 출력 형식

  • 콘텐츠 검색 : 지능형 검색 기능

    • 다국어 지원

    • 위치 기반 결과

    • 사용자 정의 가능한 결과 제한

    • 구조화된 출력 형식

  • 사이트 크롤링 : 고급 웹 크롤링 기능

    • 깊이 제어

    • 경로 필터링

    • 속도 제한

    • 진행 상황 추적

    • 사이트맵 통합

  • 사이트 매핑 : 사이트 구조 맵 생성

    • 하위 도메인 지원

    • 검색 필터링

    • 링크 분석

    • 시각적 계층 구조

  • 데이터 추출 : 여러 URL에서 구조화된 데이터 추출

    • 스키마 검증

    • 일괄 처리

    • 웹 검색 강화

    • 사용자 정의 추출 프롬프트

Related MCP server: MCP Firecrawl Server

설치

지엑스피1

빠른 시작

  1. 개발자 포털 에서 Firecrawl API 키를 받으세요

  2. API 키를 설정하세요:

    Unix/Linux/macOS(bash/zsh):

    export FIRECRAWL_API_KEY=your-api-key

    Windows(명령 프롬프트):

    set FIRECRAWL_API_KEY=your-api-key

    윈도우(PowerShell):

    $env:FIRECRAWL_API_KEY = "your-api-key"

    대안: .env 파일 사용(개발에 권장):

    # Install dotenv
    npm install dotenv
    
    # Create .env file
    echo "FIRECRAWL_API_KEY=your-api-key" > .env

    그런 다음 코드에서 다음을 수행합니다.

    import dotenv from 'dotenv';
    dotenv.config();
  3. 서버를 실행합니다:

    mcp-server-firecrawl

완성

클로드 데스크톱 앱

MCP 설정에 추가:

{
  "firecrawl": {
    "command": "mcp-server-firecrawl",
    "env": {
      "FIRECRAWL_API_KEY": "your-api-key"
    }
  }
}

클로드 VSCode 확장

MCP 구성에 추가:

{
  "mcpServers": {
    "firecrawl": {
      "command": "mcp-server-firecrawl",
      "env": {
        "FIRECRAWL_API_KEY": "your-api-key"
      }
    }
  }
}

사용 예

웹 스크래핑

// Basic scraping
{
  name: "scrape_url",
  arguments: {
    url: "https://example.com",
    formats: ["markdown"],
    onlyMainContent: true
  }
}

// Advanced extraction
{
  name: "scrape_url",
  arguments: {
    url: "https://example.com/blog",
    jsonOptions: {
      prompt: "Extract article content",
      schema: {
        title: "string",
        content: "string"
      }
    },
    mobile: true,
    blockAds: true
  }
}

사이트 크롤링

// Basic crawling
{
  name: "crawl",
  arguments: {
    url: "https://example.com",
    maxDepth: 2,
    limit: 100
  }
}

// Advanced crawling
{
  name: "crawl",
  arguments: {
    url: "https://example.com",
    maxDepth: 3,
    includePaths: ["/blog", "/products"],
    excludePaths: ["/admin"],
    ignoreQueryParameters: true
  }
}

사이트 매핑

// Generate site map
{
  name: "map",
  arguments: {
    url: "https://example.com",
    includeSubdomains: true,
    limit: 1000
  }
}

데이터 추출

// Extract structured data
{
  name: "extract",
  arguments: {
    urls: ["https://example.com/product1", "https://example.com/product2"],
    prompt: "Extract product details",
    schema: {
      name: "string",
      price: "number",
      description: "string"
    }
  }
}

구성

자세한 설정 옵션은 구성 가이드를 참조하세요.

API 문서

자세한 엔드포인트 사양은 API 문서를 참조하세요.

개발

# Install dependencies
npm install

# Build
npm run build

# Run tests
npm test

# Start in development mode
npm run dev

예시

더 많은 사용 예를 보려면 예제 디렉토리를 확인하세요.

오류 처리

서버는 강력한 오류 처리를 구현합니다.

  • 지수 백오프를 통한 속도 제한

  • 자동 재시도

  • 자세한 오류 메시지

  • 디버그 로깅

보안

  • API 키 보호

  • 요청 검증

  • 도메인 허용 목록

  • 속도 제한

  • 안전 오류 메시지

기여하다

기여 지침은 CONTRIBUTING.md를 참조하세요.

특허

MIT 라이센스 - 자세한 내용은 라이센스를 참조하세요.

Available Tools

2 tools
extractC

Extracts structured data from URLs

ParametersJSON Schema
NameRequiredDescriptionDefault
urlsYesURLs to extract from
promptNoExtraction guidance prompt
schemaNoData structure schema
ignoreSitemapNoIgnore sitemap.xml during processing
enableWebSearchNoUse web search for additional data
includeSubdomainsNoInclude subdomains in processing

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description must disclose behavior. It does not mention network requests, rate limits, authentication, error handling, or whether it is idempotent. Only states 'extracts', which is vague.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Very short at 7 words, but under-specified. Lacks structure such as sections or examples. Conciseness would be appropriate if complete, but here it is insufficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (6 parameters, nested objects, sibling tools, no output schema), the description is extremely incomplete. No information on return format, usage patterns, or edge cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description does not add any extra meaning beyond the schema's parameter descriptions. No elaboration on how parameters interact.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool extracts structured data from URLs, but does not differentiate from sibling 'map'. The verb and resource are clear, but the scope is vague as 'structured data' is not defined.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives like 'map', no prerequisites or exclusions provided. The description is insufficient for an agent to decide usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mapC

Maps a website's structure

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesBase URL to map
limitNoMaximum links to return
searchNoSearch query for mapping
timeoutNoRequest timeout
sitemapOnlyNoOnly use sitemap.xml for mapping
ignoreSitemapNoIgnore sitemap.xml during mapping
includeSubdomainsNoInclude subdomains in mapping

TDQS

C2.4/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description must disclose behavioral traits. It only says 'Maps a website's structure' with no details on whether it crawls, destructiveness, rate limits, or return format. This is insufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, short sentence, which is efficient but under-specified for a tool with 7 parameters. It could include key details without sacrificing conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (7 parameters, no output schema, no annotations), the description lacks essential context such as return values, process details, or limitations. It is highly incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage for all 7 parameters, so the baseline is 3. The description adds no additional meaning beyond the schema, making it adequate but not improved.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Maps') and resource ('a website's structure'), making the core purpose clear. However, it does not differentiate from the sibling tool 'extract', which could lead to ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives (e.g., 'extract'). The description lacks context for selection or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv1.0.1
    • First observedextract
    • First observedmap

TDQS

C2.7/5.0

Scored across 2 tools

Disambiguation5/5

The two tools have clearly distinct purposes: one extracts structured data from URLs, the other maps a website's structure. There is no overlap in functionality.

Naming Consistency5/5

Both tool names are single imperative verbs ('extract' and 'map'), following a consistent and concise naming pattern.

Tool Count3/5

With only 2 tools, the server is on the lower end of acceptable scope. It covers basic functionality but feels minimal for a web crawling service.

Completeness3/5

The server provides core data extraction and site mapping, but lacks advanced features like recursive crawling or search, which are notable gaps for the domain.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    D
    quality
    D
    maintenance
    A server that provides tools to scrape websites and extract structured data from them using Firecrawl's APIs, supporting both basic website scraping in multiple formats and custom schema-based data extraction.
    2
    3
    -
  • A
    license
    A
    quality
    D
    maintenance
    A Model Context Protocol server that enables web scraping, crawling, and content extraction capabilities through integration with Firecrawl.
    8
    117,089 npm
    2
    MIT