Skip to main content
Glama

Book Crawler MCP

一个自动化的 PDF 书籍爬取和 Gumroad 发布工具,使用 AI 生成商品描述。

功能

工具

描述

fetch_arxiv_papers

从 arXiv 获取学术论文

fetch_github_repos

从 GitHub 获取热门仓库

download_pdf

下载 PDF 文件

extract_pdf_content

提取 PDF 内容和元数据

generate_gumroad_listing

使用 AI 生成 Gumroad 商品描述

publish_to_gumroad

发布商品到 Gumroad

auto_crawl_and_publish

一键全自动爬取并发布

list_gumroad_products

查看已发布的商品

Related MCP server: Fetcher MCP

安装

1. 克隆并安装依赖

git clone https://github.com/LiZhuBin/book-crawler-mcp.git
cd book-crawler-mcp
npm install
npm run build

2. 配置环境变量

复制 .env.example.env 并填入你的密钥:

cp .env.example .env
# Gumroad API Token
# 获取方式:https://gumroad.com/settings -> Advanced
GUMROAD_ACCESS_TOKEN=your_token_here

# Anthropic API Key (用于 AI 生成商品描述)
# 获取方式:https://console.anthropic.com/settings/keys
ANTHROPIC_API_KEY=your_key_here

使用

MCP 配置

在 Claude 的 MCP 配置文件中添加:

{
  "mcpServers": {
    "book-crawler": {
      "command": "npx",
      "args": ["book-crawler-mcp"],
      "env": {
        "GUMROAD_ACCESS_TOKEN": "your_token",
        "ANTHROPIC_API_KEY": "your_key"
      }
    }
  }
}

使用示例

1. 从 arXiv 获取论文

{
  "categories": ["cs.AI", "cs.LG", "cs.CL"],
  "limit": 10
}

2. 下载并提取 PDF 内容

{
  "url": "https://arxiv.org/pdf/2301.00001.pdf"
}

3. 生成 Gumroad 商品描述

{
  "pdfContent": "提取的 PDF 内容...",
  "title": "论文标题",
  "author": "作者名"
}

4. 发布到 Gumroad

{
  "name": "商品名称",
  "description": "商品描述 (Markdown)",
  "price": 1999,
  "currency": "USD"
}

5. 全自动模式

{
  "source": "arxiv",
  "categories": ["cs.AI", "cs.LG"],
  "autoPublish": true,
  "priceRange": {
    "min": 999,
    "max": 4999
  }
}

工作流程

┌─────────────────────────────────────────────────────────────┐
│  1. fetch_arxiv_papers / fetch_github_repos                │
│     → 获取热门资源列表                                       │
└─────────────────────────────────────────────────────────────┘
                            ↓
┌─────────────────────────────────────────────────────────────┐
│  2. download_pdf                                            │
│     → 下载 PDF 文件到本地                                     │
└─────────────────────────────────────────────────────────────┘
                            ↓
┌─────────────────────────────────────────────────────────────┐
│  3. extract_pdf_content                                     │
│     → 提取文本内容和元数据                                    │
└─────────────────────────────────────────────────────────────┘
                            ↓
┌─────────────────────────────────────────────────────────────┐
│  4. generate_gumroad_listing                               │
│     → AI 生成商品标题、描述、定价                              │
└─────────────────────────────────────────────────────────────┘
                            ↓
┌─────────────────────────────────────────────────────────────┐
│  5. publish_to_gumroad                                      │
│     → 上传文件并发布商品                                      │
└─────────────────────────────────────────────────────────────┘

开发

# 安装依赖
npm install

# 开发模式
npm run dev

# 构建
npm run build

# 运行
npm start

注意事项

  1. 版权: 确保你有权分发爬取的内容

  2. API 限制: 遵守各平台的 API 调用频率限制

  3. 定价责任: AI 生成的价格仅供参考,请自行判断

License

MIT

Available Tools

1 tool
fetch_github_reposB

Fetch trending or popular repositories from GitHub

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of results (default: 10)
queryNoSearch query (e.g., "machine learning book")
languageNoFilter by language (e.g., Python, JavaScript)
minStarsNoMinimum stars required

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full responsibility for behavioral disclosure. It does not mention read-only nature, rate limits, authentication requirements, or response format, leaving significant gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no waste. While concise, it could include slightly more context (e.g., return value hints) without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and only a minimal description, the tool lacks information about return values, pagination, sorting, or error handling. It is incomplete for a tool with 4 optional parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with all parameters described in the input schema. The description adds no additional meaning beyond the schema, so a baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool fetches trending or popular repositories from GitHub, providing a verb+resource. However, it lacks specificity on what 'trending or popular' means (e.g., by stars or recent activity), preventing a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidelines for when to use this tool versus alternatives are provided, but since no sibling tools are listed, context is minimally adequate. The description implies usage for discovering repositories but lacks exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

B3.2/5.0
Disambiguation5/5

With only one tool, there is no possibility of overlap or confusion between tools. The agent will always select the correct tool.

Naming Consistency5/5

A single tool follows a clear verb_noun pattern (fetch_github_repos), and there are no other tools to introduce inconsistency.

Tool Count2/5

The server has only one tool, which feels insufficient and misaligned with the server name 'book-crawler-mcp'. The tool appears unrelated to book crawling, suggesting poor scoping.

Completeness1/5

The single tool covers only fetching GitHub repos, leaving any book-crawling domain completely uncovered. The surface is severely incomplete.

Maintenance

ActivityInactive
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • -
    license
    C
    quality
    Not graded
    maintenance
    A server that allows fetching web page content using Playwright headless browser with AI-powered capabilities for efficient information extraction.
    2
    4,976
    7
  • F
    license
    A
    quality
    D
    maintenance
    A server designed for processing PDF documents, enabling text extraction, table data retrieval, and metadata collection from local files. It allows users to scan directories for PDFs and read specific pages, specifically optimized for thesis literature analysis.
    3
  • F
    license
    Not graded
    quality
    D
    maintenance
    An intelligent web crawling server that uses Cloudflare's headless browser to render dynamic pages and Workers AI to extract relevant links based on natural language queries. It enables AI assistants to search and filter website content while providing secure access through GitHub OAuth authentication.
    3

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/LiZhuBin/book-crawler-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server