Skip to main content
Glama
huangxinping

Huggingface Daily Papers

by huangxinping

HuggingFace Daily Papers MCP Server

A MCP (Model Context Protocol) server for fetching HuggingFace daily papers.

Features

  • Fetch today's, yesterday's or specific date HuggingFace papers

  • Provides paper title, authors, abstract, tags, votes, and submitted by info

  • Includes paper links and PDF download links

  • Supports MCP tools and resource interfaces

  • ArXiv integration for complete author lists

  • Complete error handling and logging

  • Comprehensive test coverage

Related MCP server: MCP Hacker News

Installation & Usage

Install and run directly using uvx:

uvx huggingface-daily-paper-mcp

This will automatically install the package and its dependencies, then start the MCP server.

Option 2: Local development

For local development, clone the repository and install dependencies:

git clone https://github.com/huangxinping/huggingface-daily-paper-mcp.git
cd huggingface-daily-paper-mcp
uv sync

Local usage commands

Run as MCP Server (for development):

python main.py

Test Scraper Function:

python scraper.py

Run Tests:

uv run -m pytest test_mcp_server.py -v

Build Package:

uv build

MCP Interface

Tools

  1. get_papers_by_date

    • Description: Get HuggingFace papers for a specific date

    • Parameters: date (YYYY-MM-DD format)

  2. get_today_papers

    • Description: Get today's HuggingFace papers

    • Parameters: None

  3. get_yesterday_papers

    • Description: Get yesterday's HuggingFace papers

    • Parameters: None

Resources

  1. papers://today

    • Today's papers JSON data

  2. papers://yesterday

    • Yesterday's papers JSON data

Project Structure

huggingface-daily-paper-mcp/
├── main.py                    # MCP server main program
├── scraper.py                 # HuggingFace papers scraper module
├── test_mcp_server.py         # MCP server test cases
├── README.md                  # Project documentation
├── .gitignore                 # Git ignore file
├── pyproject.toml             # Project configuration file
└── uv.lock                    # Dependency lock file

Tech Stack

  • Python 3.10+: Programming language

  • MCP: Model Context Protocol framework

  • Requests: HTTP request library

  • BeautifulSoup4: HTML parsing library

  • pytest: Testing framework

  • uv: Python package manager

Development Standards

  • Use uv native commands for package management

  • Follow Python PEP 8 coding standards

  • Include type hints and docstrings

  • Complete error handling and logging

  • Write unit tests to ensure code quality

Example Output

Single paper data structure:

{
  "title": "CMPhysBench: A Benchmark for Evaluating Large Language Models in Condensed Matter Physics",
  "authors": ["Weida Wang", "Dongchen Huang", "Jiatong Li", "..."],
  "abstract": "CMPhysBench evaluates LLMs in condensed matter physics using calculation problems...",
  "url": "https://huggingface.co/papers/2508.18124",
  "pdf_url": "https://arxiv.org/pdf/2508.18124.pdf",
  "votes": 15,
  "submitted_by": "researcher123",
  "scraped_at": "2025-08-27T10:30:00.123456"
}

MCP Tool output format:

Title: CMPhysBench: A Benchmark for Evaluating Large Language Models in Condensed Matter Physics
Authors: Weida Wang, Dongchen Huang, Jiatong Li, Tengchao Yang, Ziyang Zheng...
Abstract: CMPhysBench evaluates LLMs in condensed matter physics using calculation problems...
URL: https://huggingface.co/papers/2508.18124
PDF: https://arxiv.org/pdf/2508.18124.pdf
Votes: 15
Submitted by: researcher123
--------------------------------------------------

AI IDE/CLI Configuration

Claude Code (CLI)

Add to your MCP configuration:

{
  "mcpServers": {
    "huggingface-papers": {
      "command": "uvx",
      "args": ["huggingface-daily-paper-mcp"]
    }
  }
}

Cursor IDE

Add to your .cursorrules or MCP settings:

{
  "mcp": {
    "servers": {
      "huggingface-papers": {
        "command": "uvx",
        "args": ["huggingface-daily-paper-mcp"],
        "env": {}
      }
    }
  }
}

Windsurf IDE

Add to your Windsurf MCP configuration:

{
  "mcpServers": {
    "huggingface-papers": {
      "command": "uvx",
      "args": ["huggingface-daily-paper-mcp"]
    }
  }
}

VS Code with Continue Extension

Add to your continue configuration:

{
  "mcp": {
    "servers": {
      "huggingface-papers": {
        "command": "uvx",
        "args": ["huggingface-daily-paper-mcp"]
      }
    }
  }
}

Other MCP-Compatible Tools

For any MCP-compatible client, use:

# Command
uvx huggingface-daily-paper-mcp

# Or with Python path
python -m main

License

MIT License

Available Tools

3 tools
get_papers_by_dateB

Get HuggingFace daily papers for a specific date

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYesDate in YYYY-MM-DD format

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden for behavioral disclosure. It states the tool's function but doesn't describe what 'Get' entails (e.g., returns a list, format of papers, pagination, rate limits, or authentication needs). This leaves significant gaps for a tool with no annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with zero waste. It's appropriately sized for a simple tool with one parameter and is front-loaded with the core purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, the description is incomplete. It doesn't explain what the tool returns (e.g., list of papers, metadata format) or behavioral aspects like error handling. For a tool with 1 parameter but missing structured output info, this is inadequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, with the parameter 'date' fully documented in the schema (type, format, pattern). The description adds no additional parameter semantics beyond what the schema provides, so it meets the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get') and resource ('HuggingFace daily papers') with specific scope ('for a specific date'). It distinguishes from sibling tools by specifying date-based retrieval rather than relative timeframes like 'today' or 'yesterday', though it doesn't explicitly name the siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for date-specific paper retrieval, but doesn't explicitly state when to use this tool versus the sibling tools (get_today_papers, get_yesterday_papers). It provides context about the date parameter but lacks explicit alternatives or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_today_papersB

Get today's HuggingFace daily papers

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states the tool fetches papers but doesn't describe traits such as rate limits, authentication needs, data format, or potential errors. This is a significant gap for a tool with zero annotation coverage, making it minimally adequate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with no wasted words, clearly front-loading the core purpose. It is appropriately sized for a simple tool with no parameters, making it highly concise and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's low complexity (0 parameters, no output schema, no annotations), the description is complete enough to convey the basic action. However, it lacks details on output format, error handling, or sibling tool differentiation, which are gaps in context for a tool that fetches data, making it minimally viable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has 0 parameters, and the input schema coverage is 100% with an empty object. The description doesn't need to add parameter semantics, so it meets the baseline for tools with no parameters, though it doesn't explicitly state the lack of parameters, which slightly limits clarity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get') and resource ('today's HuggingFace daily papers'), making the purpose understandable. However, it doesn't explicitly differentiate from sibling tools like 'get_papers_by_date' or 'get_yesterday_papers' beyond the temporal scope implied by 'today's', which prevents a perfect score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like 'get_papers_by_date' or 'get_yesterday_papers'. It implies usage for today's papers but lacks explicit context, exclusions, or prerequisites, leaving the agent to infer based on tool names alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_yesterday_papersB

Get yesterday's HuggingFace daily papers

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden but only states what the tool does without disclosing behavioral traits such as rate limits, authentication needs, or response format. It mentions no constraints or side effects, leaving gaps in understanding how the tool behaves.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that directly states the tool's purpose without any wasted words. It is appropriately sized and front-loaded, making it easy to grasp immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (0 parameters, no output schema, no annotations), the description is minimally adequate but lacks details on behavioral aspects and sibling differentiation. It covers the basic purpose but doesn't provide enough context for optimal agent use without additional information.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has 0 parameters with 100% schema description coverage, so the schema fully documents the absence of inputs. The description adds no parameter information, which is acceptable here as there are no parameters to explain, aligning with the baseline for zero parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get') and resource ('yesterday's HuggingFace daily papers'), making the purpose immediately understandable. However, it doesn't differentiate from sibling tools like 'get_papers_by_date' or 'get_today_papers' beyond the temporal scope, missing explicit comparison.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like 'get_papers_by_date' or 'get_today_papers'. It implies usage for yesterday's papers only but lacks explicit when/when-not instructions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updates
    • First observedget_papers_by_date
    • First observedget_today_papers
    • First observedget_yesterday_papers

TDQS

B3.1/5.0

Scored across 3 tools

Disambiguation2/5

The tools have overlapping purposes with unclear boundaries: get_papers_by_date can retrieve papers for any date, including today and yesterday, making get_today_papers and get_yesterday_papers redundant special cases. This overlap creates ambiguity about which tool to use for date-specific queries, as the descriptions don't clarify when to prefer one over another.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern with clear, descriptive naming: get_papers_by_date, get_today_papers, and get_yesterday_papers. The naming convention is uniform throughout, using snake_case and starting with 'get_' followed by the target, making it predictable and readable.

Tool Count2/5

With only 3 tools, the set feels thin and under-scoped for a server named 'Huggingface Daily Papers,' which suggests a broader domain of paper retrieval or analysis. The tools are limited to basic date-based fetching without operations like search, filtering, or metadata access, making the count too low for effective agent use in this context.

Completeness2/5

There are significant gaps in the tool surface for the implied domain of accessing HuggingFace papers. Missing operations include searching papers by keyword, filtering by categories or authors, retrieving paper details or abstracts, and accessing trends or summaries. This incompleteness will likely cause agent failures when trying to perform common paper-related tasks beyond simple date retrieval.

Maintenance

ActivityInactive
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    D
    maintenance
    A Model Context Protocol server that provides Claude and other LLMs with read-only access to Hugging Face Hub APIs, enabling interaction with models, datasets, spaces, papers, and collections through natural language.
    10
    72
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    A Model Context Protocol server that enables AI tools like Claude and Cursor to fetch and interact with live Hacker News data (posts, comments, users) via standardized MCP endpoints.
    11
    47
    33
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    A Python implementation of the Model Context Protocol (MCP) server that enables searching and extracting information from arXiv papers, designed to be extensible with additional MCP tools.
    -
  • A
    license
    A
    quality
    D
    maintenance
    MCP server pulling academic publications (arXiv, PubMed, HF Daily Papers), trending code (GitHub, HF Hub), and medical-device regulatory data (FDA 510(k), recalls) into newspaper-style briefings. Per-category round-robin, weighted configuration, sandbox-safe Python launcher.
    16
    5
    MIT