Huggingface Daily Papers
Provides integration for fetching complete author lists from arXiv papers to supplement HuggingFace daily papers data
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Huggingface Daily Papersshow me today's papers"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
HuggingFace Daily Papers MCP Server
A MCP (Model Context Protocol) server for fetching HuggingFace daily papers.
Features
Fetch today's, yesterday's or specific date HuggingFace papers
Provides paper title, authors, abstract, tags, votes, and submitted by info
Includes paper links and PDF download links
Supports MCP tools and resource interfaces
ArXiv integration for complete author lists
Complete error handling and logging
Comprehensive test coverage
Related MCP server: MCP Hacker News
Installation & Usage
Option 1: Direct execution with uvx (Recommended)
Install and run directly using uvx:
uvx huggingface-daily-paper-mcpThis will automatically install the package and its dependencies, then start the MCP server.
Option 2: Local development
For local development, clone the repository and install dependencies:
git clone https://github.com/huangxinping/huggingface-daily-paper-mcp.git
cd huggingface-daily-paper-mcp
uv syncLocal usage commands
Run as MCP Server (for development):
python main.pyTest Scraper Function:
python scraper.pyRun Tests:
uv run -m pytest test_mcp_server.py -vBuild Package:
uv buildMCP Interface
Tools
get_papers_by_date
Description: Get HuggingFace papers for a specific date
Parameters:
date(YYYY-MM-DD format)
get_today_papers
Description: Get today's HuggingFace papers
Parameters: None
get_yesterday_papers
Description: Get yesterday's HuggingFace papers
Parameters: None
Resources
papers://today
Today's papers JSON data
papers://yesterday
Yesterday's papers JSON data
Project Structure
huggingface-daily-paper-mcp/
├── main.py # MCP server main program
├── scraper.py # HuggingFace papers scraper module
├── test_mcp_server.py # MCP server test cases
├── README.md # Project documentation
├── .gitignore # Git ignore file
├── pyproject.toml # Project configuration file
└── uv.lock # Dependency lock fileTech Stack
Python 3.10+: Programming language
MCP: Model Context Protocol framework
Requests: HTTP request library
BeautifulSoup4: HTML parsing library
pytest: Testing framework
uv: Python package manager
Development Standards
Use uv native commands for package management
Follow Python PEP 8 coding standards
Include type hints and docstrings
Complete error handling and logging
Write unit tests to ensure code quality
Example Output
Single paper data structure:
{
"title": "CMPhysBench: A Benchmark for Evaluating Large Language Models in Condensed Matter Physics",
"authors": ["Weida Wang", "Dongchen Huang", "Jiatong Li", "..."],
"abstract": "CMPhysBench evaluates LLMs in condensed matter physics using calculation problems...",
"url": "https://huggingface.co/papers/2508.18124",
"pdf_url": "https://arxiv.org/pdf/2508.18124.pdf",
"votes": 15,
"submitted_by": "researcher123",
"scraped_at": "2025-08-27T10:30:00.123456"
}MCP Tool output format:
Title: CMPhysBench: A Benchmark for Evaluating Large Language Models in Condensed Matter Physics
Authors: Weida Wang, Dongchen Huang, Jiatong Li, Tengchao Yang, Ziyang Zheng...
Abstract: CMPhysBench evaluates LLMs in condensed matter physics using calculation problems...
URL: https://huggingface.co/papers/2508.18124
PDF: https://arxiv.org/pdf/2508.18124.pdf
Votes: 15
Submitted by: researcher123
--------------------------------------------------AI IDE/CLI Configuration
Claude Code (CLI)
Add to your MCP configuration:
{
"mcpServers": {
"huggingface-papers": {
"command": "uvx",
"args": ["huggingface-daily-paper-mcp"]
}
}
}Cursor IDE
Add to your .cursorrules or MCP settings:
{
"mcp": {
"servers": {
"huggingface-papers": {
"command": "uvx",
"args": ["huggingface-daily-paper-mcp"],
"env": {}
}
}
}
}Windsurf IDE
Add to your Windsurf MCP configuration:
{
"mcpServers": {
"huggingface-papers": {
"command": "uvx",
"args": ["huggingface-daily-paper-mcp"]
}
}
}VS Code with Continue Extension
Add to your continue configuration:
{
"mcp": {
"servers": {
"huggingface-papers": {
"command": "uvx",
"args": ["huggingface-daily-paper-mcp"]
}
}
}
}Other MCP-Compatible Tools
For any MCP-compatible client, use:
# Command
uvx huggingface-daily-paper-mcp
# Or with Python path
python -m mainLicense
MIT License
Available Tools
3 toolsget_papers_by_dateB
Get HuggingFace daily papers for a specific date
| Name | Required | Description | Default |
|---|---|---|---|
| date | Yes | Date in YYYY-MM-DD format |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It states the tool's function but doesn't describe what 'Get' entails (e.g., returns a list, format of papers, pagination, rate limits, or authentication needs). This leaves significant gaps for a tool with no annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with zero waste. It's appropriately sized for a simple tool with one parameter and is front-loaded with the core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description is incomplete. It doesn't explain what the tool returns (e.g., list of papers, metadata format) or behavioral aspects like error handling. For a tool with 1 parameter but missing structured output info, this is inadequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, with the parameter 'date' fully documented in the schema (type, format, pattern). The description adds no additional parameter semantics beyond what the schema provides, so it meets the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Get') and resource ('HuggingFace daily papers') with specific scope ('for a specific date'). It distinguishes from sibling tools by specifying date-based retrieval rather than relative timeframes like 'today' or 'yesterday', though it doesn't explicitly name the siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for date-specific paper retrieval, but doesn't explicitly state when to use this tool versus the sibling tools (get_today_papers, get_yesterday_papers). It provides context about the date parameter but lacks explicit alternatives or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_today_papersB
Get today's HuggingFace daily papers
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the tool fetches papers but doesn't describe traits such as rate limits, authentication needs, data format, or potential errors. This is a significant gap for a tool with zero annotation coverage, making it minimally adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with no wasted words, clearly front-loading the core purpose. It is appropriately sized for a simple tool with no parameters, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity (0 parameters, no output schema, no annotations), the description is complete enough to convey the basic action. However, it lacks details on output format, error handling, or sibling tool differentiation, which are gaps in context for a tool that fetches data, making it minimally viable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters, and the input schema coverage is 100% with an empty object. The description doesn't need to add parameter semantics, so it meets the baseline for tools with no parameters, though it doesn't explicitly state the lack of parameters, which slightly limits clarity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Get') and resource ('today's HuggingFace daily papers'), making the purpose understandable. However, it doesn't explicitly differentiate from sibling tools like 'get_papers_by_date' or 'get_yesterday_papers' beyond the temporal scope implied by 'today's', which prevents a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'get_papers_by_date' or 'get_yesterday_papers'. It implies usage for today's papers but lacks explicit context, exclusions, or prerequisites, leaving the agent to infer based on tool names alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_yesterday_papersB
Get yesterday's HuggingFace daily papers
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden but only states what the tool does without disclosing behavioral traits such as rate limits, authentication needs, or response format. It mentions no constraints or side effects, leaving gaps in understanding how the tool behaves.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without any wasted words. It is appropriately sized and front-loaded, making it easy to grasp immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (0 parameters, no output schema, no annotations), the description is minimally adequate but lacks details on behavioral aspects and sibling differentiation. It covers the basic purpose but doesn't provide enough context for optimal agent use without additional information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters with 100% schema description coverage, so the schema fully documents the absence of inputs. The description adds no parameter information, which is acceptable here as there are no parameters to explain, aligning with the baseline for zero parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Get') and resource ('yesterday's HuggingFace daily papers'), making the purpose immediately understandable. However, it doesn't differentiate from sibling tools like 'get_papers_by_date' or 'get_today_papers' beyond the temporal scope, missing explicit comparison.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'get_papers_by_date' or 'get_today_papers'. It implies usage for yesterday's papers only but lacks explicit when/when-not instructions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
- First observed
get_papers_by_date - First observed
get_today_papers - First observed
get_yesterday_papers
TDQS
Scored across 3 tools
The tools have overlapping purposes with unclear boundaries: get_papers_by_date can retrieve papers for any date, including today and yesterday, making get_today_papers and get_yesterday_papers redundant special cases. This overlap creates ambiguity about which tool to use for date-specific queries, as the descriptions don't clarify when to prefer one over another.
All tool names follow a consistent verb_noun pattern with clear, descriptive naming: get_papers_by_date, get_today_papers, and get_yesterday_papers. The naming convention is uniform throughout, using snake_case and starting with 'get_' followed by the target, making it predictable and readable.
With only 3 tools, the set feels thin and under-scoped for a server named 'Huggingface Daily Papers,' which suggests a broader domain of paper retrieval or analysis. The tools are limited to basic date-based fetching without operations like search, filtering, or metadata access, making the count too low for effective agent use in this context.
There are significant gaps in the tool surface for the implied domain of accessing HuggingFace papers. Missing operations include searching papers by keyword, filtering by categories or authors, retrieving paper details or abstracts, and accessing trends or summaries. This incompleteness will likely cause agent failures when trying to perform common paper-related tasks beyond simple date retrieval.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Agent-native MCP server over the public saagarpatel.dev corpus. Read-only, stateless.
DocBase MCP server for AI agents
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…
Related MCP Servers
- AlicenseBqualityDmaintenanceA Model Context Protocol server that provides Claude and other LLMs with read-only access to Hugging Face Hub APIs, enabling interaction with models, datasets, spaces, papers, and collections through natural language.1072MIT
- AlicenseAqualityDmaintenanceA Model Context Protocol server that enables AI tools like Claude and Cursor to fetch and interact with live Hacker News data (posts, comments, users) via standardized MCP endpoints.114733MIT
- FlicenseNot gradedqualityDmaintenanceA Python implementation of the Model Context Protocol (MCP) server that enables searching and extracting information from arXiv papers, designed to be extensible with additional MCP tools.-
- AlicenseAqualityDmaintenanceMCP server pulling academic publications (arXiv, PubMed, HF Daily Papers), trending code (GitHub, HF Hub), and medical-device regulatory data (FDA 510(k), recalls) into newspaper-style briefings. Per-category round-robin, weighted configuration, sandbox-safe Python launcher.165MIT