Skip to main content
Glama
malkreide

hn-tech-signal-mcp

by malkreide

arxiv_latest

Read-onlyIdempotent

Fetch newly submitted papers from arXiv AI/ML categories to track research frontier before press coverage. Supports cs.AI, cs.LG, cs.CL, cs.CV, cs.NE, and stat.ML.

Instructions

Fetch the most recently submitted papers from arXiv AI/ML categories.

Papers appear hours before press coverage — the fastest signal of what is happening at the AI research frontier.

Categories: cs.AI (Artificial Intelligence), cs.LG (Machine Learning), cs.CL (NLP), cs.CV (Computer Vision), cs.NE (Neural Computing), stat.ML.

All categories are fetched in ONE query (cat:A OR cat:B OR …), so the result is the most recent papers across the requested categories, grouped per category. A paper cross-listed to several requested categories appears under each of them, matching how a single-category query would return it.

Because one window is shared, a small category (cs.NE, stat.ML) can come back with fewer than limit papers even though arXiv holds more. Those categories are named in incomplete_categories — query a single category for full depth.

Args: params (ArxivLatestInput): - categories (List[str]): arXiv category codes - limit (int): Max papers per category (1–20)

Returns: str: JSON with categories, total_papers, distinct_papers, by_category dict, and — only when a category fell short — incomplete_categories plus an explanatory note. Each paper: id, title, abstract (400 chars), authors, published, category (primary), categories (all), url, pdf.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
paramsYes

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
noteNo
categoriesYes
fetched_atYes
by_categoryYes
total_papersYes
distinct_papersYes
incomplete_categoriesNoPresent only when a category returned fewer than limit within the window

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed12 schema fields changedv0.5.0
    • addedOutput schema / $defs
      Added value: +{
      +  "ArxivPaper": {
      +    "additionalProperties": false,
      +    "properties": {
      +      "abstract": {
      +        "description": "First 400 characters",
      +        "title": "Abstract",
      +        "type": "string"
      +      },
      +      "authors": {
      +        "description": "First five",
      +        "items": {
      +          "type": "string"
      +        },
      +        "title": "Authors",
      +        "type": "array"
      +      },
      +      "categories": {
      +        "description": "All categories, primary included",
      +        "items": {
      +          "type": "string"
      +        },
      +        "title": "Categories",
      +        "type": "array"
      +      },
      +      "category": {
      +        "description": "Primary category",
      +        "title": "Category",
      +        "type": "string"
      +      },
      +      "id": {
      +        "title": "Id",
      +        "type": "string"
      +      },
      +      "pdf": {
      +        "title": "Pdf",
      +        "type": "string"
      +      },
      +      "published": {
      +        "title": "Published",
      +        "type": "string"
      +      },
      +      "title": {
      +        "title": "Title",
      +        "type": "string"
      +      },
      +      "url": {
      +        "title": "Url",
      +        "type": "string"
      +      }
      +    },
      +    "required": [
      +      "id",
      +      "title",
      +      "abstract",
      +      "authors",
      +      "published",
      +      "category",
      +      "categories",
      +      "url",
      +      "pdf"
      +    ],
      +    "title": "ArxivPaper",
      +    "type": "object"
      +  }
      +}
    • addedOutput schema / additionalProperties
      Added value: +false
    • addedOutput schema / properties / by_category
      Added value: +{
      +  "additionalProperties": {
      +    "items": {
      +      "$ref": "#/$defs/ArxivPaper"
      +    },
      +    "type": "array"
      +  },
      +  "title": "By Category",
      +  "type": "object"
      +}
    • addedOutput schema / properties / categories
      Added value: +{
      +  "items": {
      +    "type": "string"
      +  },
      +  "title": "Categories",
      +  "type": "array"
      +}
    • addedOutput schema / properties / distinct_papers
      Added value: +{
      +  "title": "Distinct Papers",
      +  "type": "integer"
      +}
    • addedOutput schema / properties / fetched_at
      Added value: +{
      +  "title": "Fetched At",
      +  "type": "string"
      +}
    • addedOutput schema / properties / incomplete_categories
      Added value: +{
      +  "anyOf": [
      +    {
      +      "items": {
      +        "type": "string"
      +      },
      +      "type": "array"
      +    },
      +    {
      +      "type": "null"
      +    }
      +  ],
      +  "default": null,
      +  "description": "Present only when a category returned fewer than limit within the window",
      +  "title": "Incomplete Categories"
      +}
    • addedOutput schema / properties / note
      Added value: +{
      +  "anyOf": [
      +    {
      +      "type": "string"
      +    },
      +    {
      +      "type": "null"
      +    }
      +  ],
      +  "default": null,
      +  "title": "Note"
      +}
    • removedOutput schema / properties / result
      Removed value: -{
      -  "title": "Result",
      -  "type": "string"
      -}
    • addedOutput schema / properties / total_papers
      Added value: +{
      +  "title": "Total Papers",
      +  "type": "integer"
      +}
    • changedOutput schema / required
      Previous value: -[
      -  "result"
      -]New value: +[
      +  "fetched_at",
      +  "categories",
      +  "total_papers",
      +  "distinct_papers",
      +  "by_category"
      +]
    • changedOutput schema / title
      Previous value: -"arxiv_latestOutput"New value: +"ArxivLatestOutput"
  2. First observedv0.2.4

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/openWorld/non-destructive, but the description adds substantial non-obvious behavior: all categories are fetched in ONE OR query, cross-listed papers repeat under each category, a shared window can truncate small categories, and truncation surfaces via incomplete_categories. These are exactly the quirks that would otherwise cause misreads of results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then edge-case behavior, then args/returns. Efficient overall, though the 'Returns' block largely restates the output schema and slightly inflates length against a description that already exists alongside structured return data.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a read-only fetch tool with an output schema present, the description covers everything an agent needs: what is returned, the per-category grouping, the truncation edge case, and the remedy. No meaningful gap remains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is reported as 0%, so the description must carry parameter meaning. It spells out the valid category codes with human-readable names and clarifies that limit is per-category (1–20) rather than a global cap — a distinction the schema alone does not make.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Fetch), resource (recently submitted arXiv papers), and scope (AI/ML categories, latest submissions). 'Most recently submitted' naturally distinguishes it from the sibling arxiv_search, which is a keyword search tool, without needing to name it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear usage context ('Papers appear hours before press coverage — the fastest signal') and an explicit operational rule: query a single category for full depth when a category is truncated. It stops short of naming arxiv_search as the alternative for topical lookups, so the when-not-to-use case is left implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.