book-crawler-mcp
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@book-crawler-mcpauto crawl arXiv cs.AI papers"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Book Crawler MCP
一个自动化的 PDF 书籍爬取和 Gumroad 发布工具,使用 AI 生成商品描述。
功能
工具 | 描述 |
| 从 arXiv 获取学术论文 |
| 从 GitHub 获取热门仓库 |
| 下载 PDF 文件 |
| 提取 PDF 内容和元数据 |
| 使用 AI 生成 Gumroad 商品描述 |
| 发布商品到 Gumroad |
| 一键全自动爬取并发布 |
| 查看已发布的商品 |
Related MCP server: Fetcher MCP
安装
1. 克隆并安装依赖
git clone https://github.com/LiZhuBin/book-crawler-mcp.git
cd book-crawler-mcp
npm install
npm run build2. 配置环境变量
复制 .env.example 为 .env 并填入你的密钥:
cp .env.example .env# Gumroad API Token
# 获取方式:https://gumroad.com/settings -> Advanced
GUMROAD_ACCESS_TOKEN=your_token_here
# Anthropic API Key (用于 AI 生成商品描述)
# 获取方式:https://console.anthropic.com/settings/keys
ANTHROPIC_API_KEY=your_key_here使用
MCP 配置
在 Claude 的 MCP 配置文件中添加:
{
"mcpServers": {
"book-crawler": {
"command": "npx",
"args": ["book-crawler-mcp"],
"env": {
"GUMROAD_ACCESS_TOKEN": "your_token",
"ANTHROPIC_API_KEY": "your_key"
}
}
}
}使用示例
1. 从 arXiv 获取论文
{
"categories": ["cs.AI", "cs.LG", "cs.CL"],
"limit": 10
}2. 下载并提取 PDF 内容
{
"url": "https://arxiv.org/pdf/2301.00001.pdf"
}3. 生成 Gumroad 商品描述
{
"pdfContent": "提取的 PDF 内容...",
"title": "论文标题",
"author": "作者名"
}4. 发布到 Gumroad
{
"name": "商品名称",
"description": "商品描述 (Markdown)",
"price": 1999,
"currency": "USD"
}5. 全自动模式
{
"source": "arxiv",
"categories": ["cs.AI", "cs.LG"],
"autoPublish": true,
"priceRange": {
"min": 999,
"max": 4999
}
}工作流程
┌─────────────────────────────────────────────────────────────┐
│ 1. fetch_arxiv_papers / fetch_github_repos │
│ → 获取热门资源列表 │
└─────────────────────────────────────────────────────────────┘
↓
┌─────────────────────────────────────────────────────────────┐
│ 2. download_pdf │
│ → 下载 PDF 文件到本地 │
└─────────────────────────────────────────────────────────────┘
↓
┌─────────────────────────────────────────────────────────────┐
│ 3. extract_pdf_content │
│ → 提取文本内容和元数据 │
└─────────────────────────────────────────────────────────────┘
↓
┌─────────────────────────────────────────────────────────────┐
│ 4. generate_gumroad_listing │
│ → AI 生成商品标题、描述、定价 │
└─────────────────────────────────────────────────────────────┘
↓
┌─────────────────────────────────────────────────────────────┐
│ 5. publish_to_gumroad │
│ → 上传文件并发布商品 │
└─────────────────────────────────────────────────────────────┘开发
# 安装依赖
npm install
# 开发模式
npm run dev
# 构建
npm run build
# 运行
npm start注意事项
版权: 确保你有权分发爬取的内容
API 限制: 遵守各平台的 API 调用频率限制
定价责任: AI 生成的价格仅供参考,请自行判断
License
MIT
Links
Available Tools
1 toolfetch_github_reposB
Fetch trending or popular repositories from GitHub
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of results (default: 10) | |
| query | No | Search query (e.g., "machine learning book") | |
| language | No | Filter by language (e.g., Python, JavaScript) | |
| minStars | No | Minimum stars required |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full responsibility for behavioral disclosure. It does not mention read-only nature, rate limits, authentication requirements, or response format, leaving significant gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no waste. While concise, it could include slightly more context (e.g., return value hints) without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and only a minimal description, the tool lacks information about return values, pagination, sorting, or error handling. It is incomplete for a tool with 4 optional parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with all parameters described in the input schema. The description adds no additional meaning beyond the schema, so a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool fetches trending or popular repositories from GitHub, providing a verb+resource. However, it lacks specificity on what 'trending or popular' means (e.g., by stars or recent activity), preventing a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidelines for when to use this tool versus alternatives are provided, but since no sibling tools are listed, context is minimally adequate. The description implies usage for discovering repositories but lacks exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
With only one tool, there is no possibility of overlap or confusion between tools. The agent will always select the correct tool.
A single tool follows a clear verb_noun pattern (fetch_github_repos), and there are no other tools to introduce inconsistency.
The server has only one tool, which feels insufficient and misaligned with the server name 'book-crawler-mcp'. The tool appears unrelated to book crawling, suggesting poor scoping.
The single tool covers only fetching GitHub repos, leaving any book-crawling domain completely uncovered. The surface is severely incomplete.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Extract papers from ArXiv — titles, abstracts, authors, categories & PDF links. Monitor new AI, phys
Extract PDFs to Markdown, RAG chunks and cited tables; publish tracked Doc Links with read stats.
Hosted MCP server: convert PDFs to clean, LLM-ready Markdown with tables, formulas and OCR.
Scrape, crawl, map & search the web. Open-source, self-hostable Rust crawler & search for AI agents.
Related MCP Servers
- AlicenseNot gradedqualityFmaintenanceA server that allows AI assistants to search for research papers, read their content, and access related code repositories through the PapersWithCode API.26MIT
- -licenseCqualityNot gradedmaintenanceA server that allows fetching web page content using Playwright headless browser with AI-powered capabilities for efficient information extraction.24,9767
- FlicenseAqualityDmaintenanceA server designed for processing PDF documents, enabling text extraction, table data retrieval, and metadata collection from local files. It allows users to scan directories for PDFs and read specific pages, specifically optimized for thesis literature analysis.3
- FlicenseNot gradedqualityDmaintenanceAn intelligent web crawling server that uses Cloudflare's headless browser to render dynamic pages and Workers AI to extract relevant links based on natural language queries. It enables AI assistants to search and filter website content while providing secure access through GitHub OAuth authentication.3
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/LiZhuBin/book-crawler-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server