AI News MCP
AI News MCP
Local MCP server and ingestion pipeline for AI, tech, business, and research signals. The V2 goal is to:
ingest recent items from newsletters, feeds, blogs, community streams, and research feeds
normalize them into a shared canonical schema
enrich them into editorial post candidates
support recency-first queries for "latest" workflows
expose MCP tools that let a media agent pull ideas on demand
Current V2 Sources
Enabled now:
TLDR AI
Superhuman AI
TheSequence
TechCrunch AI
Lenny's Newsletter
Hugging Face Blog
Hacker News AI query feed
LessWrong
arXiv cs.AI
arXiv cs.LG
Configured but currently second-wave / disabled:
The Neuron
Turing Post
The Batch
The Pragmatic Engineer
a16z AI
Benedict Evans
Quick start
Install dependencies:
py -3 -m pip install -e .Refresh the source database:
py -3 scripts\refresh.pySee top stories:
py -3 scripts\demo_top_stories.pyInspect canonical items:
py -3 -c "from pathlib import Path; from ai_news_mcp.service import AiNewsService; s=AiNewsService(Path('.')); print(s.get_canonical_items(10))"Inspect a full cached article with clean text, markdown, and top images:
py -3 -c "from pathlib import Path; from ai_news_mcp.service import AiNewsService; s=AiNewsService(Path('.')); story=s.get_latest_items(limit=1)[0]; print(s.get_article(article_id=story['article_id'], include_images=True, include_markdown=True))"Run the MCP server in the recommended ChatGPT-compatible mode:
py -3 scripts\run_server.pyDefault runtime behavior:
transport:
streamable-httpbind:
0.0.0.0:8081ChatGPT app URL:
https://<your-public-host>/mcp
Optional overrides:
$env:MCP_TRANSPORT = "sse"
$env:MCP_HOST = "0.0.0.0"
$env:MCP_PORT = "8081"
py -3 scripts\run_server.pyNotes:
For ChatGPT Apps / Developer Mode, prefer a public
/mcpendpoint over Streamable HTTP.Use
/sseonly for older SSE clients or debugging.
Notes
V2 prefers RSS/Atom feeds first, then source-specific HTML adapters where needed.
Storage is local SQLite in
data\ai_news.db.Facts are stored in
canonical_items; editorial outputs are stored inpost_candidates.Rich article cache now stores full HTML, clean text, markdown, top-ranked images, content status, and extraction warnings in
article_cache.refresh.pycurrently rebuilds the source pool on each run so old source generations do not pollute newer results.Query tools now support recency, topic, and
categoryfilters for latest-mode workflows.MCP now exposes
get_article(...)so downstream agents can retrieve the full scraped article body and reference images without doing a separate web fetch.