mcp-research-agent-system
Provides tools for searching arXiv papers, retrieving paper details, and accessing cached summaries via the arXiv public API.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@mcp-research-agent-systemsearch arXiv for recent papers on multi-agent systems"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MCP Research Agent System
A multi-agent research system that combines Model Context Protocol (MCP) servers with LangGraph state machines to automate literature reviews on arXiv. The system decomposes a research goal into sub-queries, searches papers via an MCP-wrapped arXiv API, validates findings through heuristic + LLM-judge review, and synthesizes a structured markdown report.
Architecture Overview
graph TD
A[Research Goal] --> B[Planner Node]
B --> C[Researcher Node]
C --> D[Validator Node]
D -->|valid| E[Next Sub-Query]
D -->|invalid, retries<3| C
D -->|invalid, retries=3| F[Synthesizer Node]
E --> C
F --> G[Final Report]
subgraph "MCP Server (stdio)"
H[search_papers]
I[get_paper_details]
J[get_cached_summary]
end
C -.->|MCP stdio transport| H
C -.->|MCP stdio transport| I
C -.->|MCP stdio transport| JData Flow
Planner — Decomposes the high-level research goal into 3–5 focused sub-queries using an LLM with structured output (
PlannerDecomposition).Researcher — For each sub-query, spawns a fresh MCP subprocess (stdio transport) that wraps the arXiv API. Calls
search_papersandget_cached_summarytools, returningResearchResultobjects.Validator — Heuristic-first validation (zero papers, obvious off-topic) → if ambiguous, falls back to LLM-judge with structured output (
ValidationOutcome). On invalid: revises the sub-query, increments attempt counter, loops back to Researcher (max 3 attempts). On valid: advances to next sub-query or Synthesizer.Synthesizer — Aggregates all
validated_findingsinto a structured markdown report viaSynthesizedReportschema (Overview, Key Themes, Notable Papers, Gaps/Open Questions).
Related MCP server: pubmed-search-mcp
Engineering Decisions & Trade-offs
Decision | Rationale | Trade-off |
Fresh MCP subprocess per sub-query | Guarantees isolation; no cross-query state leakage; matches MCP's designed stateless per-session model. | Higher latency (~100–300ms subprocess spawn) vs. persistent server. Mitigated by SQLite caching. |
Heuristic-first validation + LLM-judge fallback | Avoids LLM cost/latency for obvious failures (empty results, clear off-topic). LLM only invoked on ambiguous cases. | Heuristics can be fooled by adversarial/edge cases; LLM judge adds a safety net. |
Retry loop with query revision (max 3) | Self-correcting: validator suggests a refined query, researcher retries. Prevents dead ends from poor initial decomposition. | Increases total runtime (up to 3× researcher calls per sub-query). Bounded by |
SQLite caching layer | arXiv API has rate limits; repeated queries for same topic across retries/runs benefit from local cache. | Cache invalidation is manual (TTL-based); stale summaries possible if papers updated. |
Provider-agnostic LLM client ( | Works with any OpenAI-compatible endpoint (OmniRoute, vLLM, OpenAI, Gemini, Ollama). Configured via | Requires endpoint to support structured output ( |
Structured JSONL tracing | Every node entry/exit, tool call, and error logged as JSONL to | Log files grow unbounded; no built-in rotation (add logrotate or similar for production). |
Typed state via TypedDict | LangGraph state is fully typed ( | Boilerplate for state updates; |
Pydantic v2 for all schemas | Runtime validation of LLM outputs, tool inputs, and MCP tool results. Fail-fast on schema violations. | Slight overhead vs. raw dicts; worth it for correctness. |
Notable Bugs Found & Fixed
Verbose-mode double-execution bug (commit 91b14e1): In verbose CLI mode (-v), the LangGraph stream (astream with stream_mode="values") was followed by a redundant ainvoke call on the same compiled graph. This caused the entire pipeline to execute twice — doubling LLM API usage, arXiv queries, and trace log entries. The bug was diagnosed by examining logs/trace.jsonl, which showed the planner node re-entering and generating sub-queries a second time after the synthesizer had already completed. The fix captures the final state directly from the stream_mode="values" iterator (the last yielded state contains the complete result), eliminating the second invocation. A regression test in tests/test_cli.py::test_verbose_mode_no_double_execution now asserts that the researcher node is invoked exactly once per sub-query in verbose mode.
Known Limitations
Free-tier LLM API constraints: The default configuration uses Gemini's free tier (20 requests/day), which imposes a hard limit on how often the full pipeline can be run. The planner's fallback-parsing logic (structured output → manual JSON extraction → regex fallback) and the validator's heuristic-first design (zero-paper / off-topic checks before LLM-judge) were built partly to maximize success rate within this quota.
arXiv API rate limits: The built-in 3-second delay between requests and SQLite caching mitigate this, but high-volume runs may still hit limits.
No log rotation: JSONL trace files in
logs/grow unbounded; production deployments should add logrotate or similar.Single-provider structured output assumption: The
get_llmclient assumes the endpoint supportswith_structured_output(OpenAI-compatible function calling). Endpoints without this will fail at runtime.
Quick Start
Prerequisites
Python 3.11+
An OpenAI-compatible LLM endpoint (OmniRoute, vLLM, OpenAI, etc.)
(Optional) Docker for containerized deployment
Installation
# Clone and install in development mode
git clone https://github.com/hemanth2k6/mcp-research-agent-system.git
cd mcp-research-agent-system
pip install -e ".[dev]"
# Configure LLM endpoint (any OpenAI-compatible API)
# Tested default: Gemini
export LLM_BASE_URL="https://generativelanguage.googleapis.com/v1beta/openai/"
export LLM_API_KEY="your-gemini-api-key-here"
export LLM_MODEL="gemini-3.5-flash" # or your model of choiceRun a Research Query
# Basic usage
research-agent "quantum error correction surface codes"
# Verbose mode (streams node progress to stdout)
research-agent "transformer attention mechanisms" -v
# Save report to file
research-agent "graph neural networks for drug discovery" -o report.mdDocker
# Build image
docker build -t mcp-research-agent .
# Run with docker run (pass env vars, or use .env file)
docker run --rm \
--env-file .env \
-v $(pwd)/data:/app/data \
-v $(pwd)/logs:/app/logs \
mcp-research-agent "your research goal"
# Verbose mode
docker run --rm \
--env-file .env \
-v $(pwd)/data:/app/data \
-v $(pwd)/logs:/app/logs \
mcp-research-agent "your research goal" --verbose
# Or with docker compose (loads .env automatically)
docker compose run --rm research-agent "your research goal"
# Verbose with docker compose
docker compose run --rm research-agent "your research goal" --verboseUsage Examples
Example 1: Basic Research Goal (Verified Run)
$ research-agent "recent progress in mixture-of-experts model architectures"Output (trimmed from verified run — August 2026):
Starting research pipeline for: recent progress in mixture-of-experts model architectures
Trace logs will be written to: logs
⟳ researcher: Query 1: ... (attempt 1, 0 papers)
✓ planner: Generated 4 sub-queries: mixture of experts routing mechanisms soft top-k expert choice, sparse mixture of experts load balancing training stability, sparse mixture of experts vision transformers multimodal architectures...
✓ planner: Generated 4 sub-queries: mixture of experts routing mechanisms soft top-k expert choice, sparse mixture of experts load balancing training stability, sparse mixture of experts vision transformers multimodal architectures...
✓ planner: Generated 4 sub-queries: mixture of experts routing mechanisms soft top-k expert choice, sparse mixture of experts load balancing training stability, sparse mixture of experts vision transformers multimodal architectures...
✓ planner: Generated 4 sub-queries: mixture of experts routing mechanisms soft top-k expert choice, sparse mixture of experts load balancing training stability, sparse mixture of experts vision transformers multimodal architectures...
✓ planner: Generated 4 sub-queries: mixture of experts routing mechanisms soft top-k expert choice, sparse mixture of experts load balancing training stability, sparse mixture of experts vision transformers multimodal architectures...
✓ planner: Generated 4 sub-queries: mixture of experts routing mechanisms soft top-k expert choice, sparse mixture of experts load balancing training stability, sparse mixture of experts vision transformers multimodal architectures...
✓ planner: Generated 4 sub-queries: mixture of experts routing mechanisms soft top-k expert choice, sparse mixture of experts load balancing training stability, sparse mixture of experts vision transformers multimodal architectures...
✓ planner: Generated 4 sub-queries: mixture of experts routing mechanisms soft top-k expert choice, sparse mixture of experts load balancing training stability, sparse mixture of experts vision transformers multimodal architectures...
✓ planner: Generated 4 sub-queries: mixture of experts routing mechanisms soft top-k expert choice, sparse mixture of experts load balancing training stability, sparse mixture of experts vision transformers multimodal architectures...
✓ planner: Generated 4 sub-queries: mixture of experts routing mechanisms soft top-k expert choice, sparse mixture of experts load balancing training stability, sparse mixture of experts vision transformers multimodal architectures...
============================================================
Overview
This research report synthesizes recent progress in Mixture-of-Experts (MoE)
model architectures based on validated literature. The surveyed research focuses
on addressing the fundamental challenges of MoE models: routing instability,
load imbalance, high parameter memory footprints, offloading latency during
inference, and cross-domain generalization in vision, graph neural networks, and
recommendation systems.
Key Themes
1. Routing Mechanism Innovations & Load Balancing
Standard top-$k$ routing often suffers from expert under-utilization,
non-differentiability, and routing collapse. Recent work introduces continuous
or soft relaxations, inverse routing paradigms (expert choice),
similarity-preserving router formulations, and copula-based dependence modeling
across tokens to maintain balanced expert load without performance degradation.
2. Inference Efficiency, Offloading & Quantization
Deploying large MoE models on memory-constrained hardware requires efficient
parameter management. Key strategies include predictive expert caching, token
scheduling, dynamic expert quantization, expert pruning/skipping, and
speculative decoding techniques designed to overlap dynamic CPU-to-GPU
offloading latency with compute.
3. Interpretability & Theoretical Foundations
Understanding router behavior and training dynamics is essential for
architectural optimization. Recent studies formulate mathematical bounds for
softmax gating convergence, analyze auxiliary-loss-free load balancing
procedures, and introduce routing signatures to trace task-conditioned expert
activation patterns.
4. Cross-Domain MoE Applications
Beyond text-based Large Language Models (LLMs), MoE principles are increasingly
adapted to vision transformers, graph neural networks facing severe distribution
shifts, and real-time multimodal streaming recommender systems.
Notable Papers
• Mixture-of-Experts with Expert Choice Routing (Yanqi Zhou et al.) — Inverts
token routing by allowing experts to select top-$k$ tokens, ensuring perfect
load balancing (arXiv:2202.09368v2).
• SoftMoE: Soft Differentiable Routing for Mixture-of-Experts in LLMs (Mikołaj
Zasada et al.) — Replaces discrete top-$k$ selection with a soft LapSum
relaxation for fully differentiable routing (arXiv:2606.17952v1).
• Task-Conditioned Routing Signatures in Sparse Mixture-of-Experts Transformers
(Mynampati Sri Ranganadha Avinash) — Formulates routing signatures to study
task-conditioned structural patterns in sparse MoE models
(arXiv:2603.11114v1).
• A Theoretical Framework for Auxiliary-Loss-Free Load Balancing of Sparse
Mixture-of-Experts in Large-Scale AI Models (X. Y. Han, Yuan Zhong) —
Provides rigorous theoretical analysis of auxiliary-loss-free routing
mechanisms (arXiv:2512.03915v3).
• Load Balancing Mixture of Experts with Similarity Preserving Routers (Nabil
Omi et al.) — Develops routers that preserve input similarity to prevent load
imbalance without auxiliary penalties (arXiv:2506.14038v2).
• Hierarchical Copula-Gumbel-Top-K Routing (Richard Yi Da Xu) — Controls joint
routing dependencies across related tokens using exchangeable copula
structures (arXiv:2607.28670v3).
• ExpertFlow: Efficient Mixture-of-Experts Inference via Predictive Expert
Caching and Token Scheduling (Xin He et al.) — Enables single-GPU MoE serving
through predictive caching and offloading (arXiv:2410.17954v2).
• Not All Experts are Equal: Efficient Expert Pruning and Skipping for
Mixture-of-Experts Large Language Models (Xudong Lu et al.) — Proposes
plug-and-play expert pruning and runtime skipping to shrink parameter
footprints (arXiv:2402.14800v2).
• Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference (Kexin
Chu et al.) — Reduces runtime memory consumption using dynamic post-training
quantization on expert parameters (arXiv:2511.15015v3).
• Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with
Speculative Decoding (Zhibin Wang et al.) — Introduces SpecMoEOff to hide
CPU-GPU expert transfer latency during inference (arXiv:2508.21706v2).
• MoE-SpeQ: Speculative Quantized Decoding with Proactive Expert Prefetching
and Offloading for Mixture-of-Experts (Wenfeng Wang et al.) — Combines
speculative quantized execution with proactive prefetching over PCIe
(arXiv:2511.14102v1).
• QuantMoE-Bench: Examining Post-Training Quantization for Mixture-of-Experts
(Pingzhi Li et al.) — Benchmarks diverse post-training quantization schemes
across MoE architectures (arXiv:2406.08155v2).
• Convergence Rates for Softmax Gating Mixture of Experts (Huy Nguyen et al.) —
Establishes theoretical convergence guarantees for softmax gating routers
(arXiv:2503.03213v1).
• Mixture-of-Experts Models in Vision: Routing, Optimization, and
Generalization (Adam Rokah et al.) — Evaluates dense, SoftMoE, and SparseMoE
heads on image classification tasks (arXiv:2601.15021v1).
• GraphMETRO: Mitigating Complex Graph Distribution Shifts via Mixture of
Aligned Experts (Shirley Wu et al.) — Applies MoE architectures to mitigate
out-of-distribution shifts in complex graph data (arXiv:2312.04693v3).
• Efficient Multimodal Streaming Recommendation via Expandable Side
Mixture-of-Experts (Yunke Qu et al.) — Introduces side-MoE modules for
real-time item representation updates (arXiv:2508.05993v3).
• Mixtures of Experts Models (Isobel Claire Gormley, Sylvia
Frühwirth-Schnatter) — Covers foundational statistical frameworks for
covariate-conditioned mixture models (arXiv:1806.08200v1).
Gaps / Open Questions
1 Theoretical Understanding of Complex/Differentiable Routers: While classical
softmax gating has established convergence proofs, full convergence and
stability proofs for continuous relaxations (e.g., SoftMoE, Copula-Gumbel)
under large-scale dynamic pre-training remain open.
2 Hardware-Co-Designed MoE Offloading: Current offloading schemes rely on
speculative decoding and proactive caching to mask PCIe bandwidth
bottlenecks; custom hardware or interconnect-level primitives designed
specifically for irregular expert sparsity are needed.
3 Unified Cross-Modal MoE Standards: Current implementations in vision, graph,
and recommendation domains rely on task-specific heuristics, lacking unified
routing protocols across multimodal representations.
Full trace log written to: logs/trace.jsonlExample 2: Verbose Trace (What You See in logs/trace.jsonl)
{"timestamp": "2026-08-21T15:44:19.971522+00:00", "event_type": "synthesizer_input", "payload": {"research_goal": "quantum error correction", "findings_count": 1}}
{"timestamp": "2026-08-21T15:44:21.574371+00:00", "event_type": "researcher_tool_call", "payload": {"tool": "search_papers", "input": {"query": "surface codes", "max_results": 10}, "sub_query": "surface codes"}}
{"timestamp": "2026-08-21T15:44:22.930660+00:00", "event_type": "tool_result", "payload": {"tool_name": "search_papers", "output_summary": {"paper_count": 10}, "duration_ms": 1352.67}}
{"timestamp": "2026-08-21T15:44:22.937775+00:00", "event_type": "researcher_tool_call", "payload": {"tool": "get_cached_summary", "input": {"topic": "surface codes"}, "sub_query": "surface codes"}}
{"timestamp": "2026-08-21T15:44:22.943917+00:00", "event_type": "tool_result", "payload": {"tool_name": "get_cached_summary", "output_summary": {"match_count": 14}, "duration_ms": 4.30}}
{"timestamp": "2026-08-21T15:44:23.156663+00:00", "event_type": "synthesizer_input", "payload": {"research_goal": "quantum error correction", "findings_count": 10}}
{"timestamp": "2026-08-21T15:44:23.156982+00:00", "event_type": "synthesizer_output", "payload": {"report_length": 2847, "status": "success"}}Note: Each sub-query spawns a fresh MCP subprocess. The trace shows
researcher_tool_call→tool_call→tool_resultpairs for each MCP tool invocation.
Example 3: Retry Loop in Action
When the researcher returns off-topic or empty results, the validator revises the query and retries (up to 3 times):
$ research-agent "impossible query xyz123" -vVerbose output:
[planner] Decomposed into: ['impossible query xyz123']
[researcher] Searching: "impossible query xyz123" → 0 papers
[validator] Heuristic: zero papers → invalid, revised query: "impossible query xyz123 related research"
[researcher] Searching: "impossible query xyz123 related research" → 0 papers
[validator] Heuristic: zero papers → invalid, revised query: "broader topic impossible query"
[researcher] Searching: "broader topic impossible query" → 3 papers
[validator] Heuristic: relevant → valid
[synthesizer] Generating final report...Testing
The test suite (156 tests) is fully offline — no real network calls, arXiv API, or LLM endpoints are invoked. All external dependencies are mocked at the integration layer.
# Run full test suite
pytest -v
# Run with coverage
pytest --cov=src/mcp_research_agent_system --cov-report=term-missing
# Type-check
mypy src/mcp_research_agent_system tests
# Lint
ruff check src testsKey Test Modules
Module | Focus |
| Retry-loop integration tests — Forces researcher to return bad results twice, then good results on 3rd attempt; asserts graph retries correctly and attempts capped at 3. Also tests exhausted attempts path (max retries → synthesizer proceeds). |
| Graph routing: planner → researcher → validator → synthesizer; error handling in each node; router exhaustion logic. |
| CLI argument parsing, pipeline execution, error propagation, verbose streaming. |
| Atom XML parsing, rate limiting, date filtering, error handling. |
| SQLite cache operations, TTL expiry, concurrent access. |
| MCP tool registration, stdio transport, tool input validation. |
| Sub-query decomposition, structured output validation. |
| Researcher sub-agent, MCP subprocess management, error handling. |
| Report synthesis, structured output, fallback handling. |
Retry-Loop Integration Test (from test_validation.py)
async def test_retry_loop_then_success(self):
"""Test graph retries with bad results, then succeeds."""
# 1st call: empty result -> validator says invalid, retry
# 2nd call: off-topic result -> validator says invalid, retry
# 3rd call: good result -> validator says valid
bad_empty = _make_research_result("quantum error correction", [])
bad_offtopic = _make_research_result("quantum error correction", [off_topic_paper])
good_result = _make_research_result("quantum error correction", [valid_paper])
run_research_mock = AsyncMock(side_effect=[bad_empty, bad_offtopic, good_result])
# ... run graph ...
# Researcher called 3 times (initial + 2 retries before success on 3rd)
assert run_research_mock.call_count == 3
assert result["validation_status"] == "valid"This test validates the core self-correcting loop: the system doesn't give up on the first failure — it revises the query and retries, with a hard cap to prevent infinite loops.
Project Structure
mcp-research-agent-system/
├── .github/
│ └── workflows/
│ └── ci.yml # GitHub Actions CI (ruff, mypy, pytest)
├── docs/
│ └── graph_diagram.mmd # Mermaid architecture diagram
├── logs/
│ └── trace.jsonl # JSONL trace output (gitignored)
├── scripts/
│ └── view_trace.py # CLI trace viewer
├── src/
│ └── mcp_research_agent_system/
│ ├── __init__.py
│ ├── agents/
│ │ ├── __init__.py
│ │ ├── graph.py # LangGraph state machine (planner/researcher/validator/synthesizer)
│ │ ├── planner.py # Goal decomposition + validation logic
│ │ ├── researcher.py # MCP subprocess + arXiv search
│ │ ├── synthesizer.py # Report synthesis
│ │ └── state.py # TypedDict ResearchState + factory
│ ├── arxiv_client.py # Async arXiv API client (Atom XML)
│ ├── cache.py # SQLite caching layer
│ ├── cli.py # CLI entrypoint (research-agent)
│ ├── config.py # Pydantic Settings (env vars)
│ ├── errors.py # Custom exceptions
│ ├── logging_utils.py # JSONL structured logging
│ ├── mcp_server.py # MCP stdio server (search_papers, get_paper_details, get_cached_summary)
│ └── trace_viewer.py # Trace log analysis utilities
├── tests/
│ ├── __init__.py
│ ├── conftest.py
│ ├── test_arxiv_client.py
│ ├── test_cache.py
│ ├── test_cli.py
│ ├── test_graph.py
│ ├── test_mcp_server.py
│ ├── test_planner.py
│ ├── test_researcher.py
│ ├── test_synthesizer.py
│ └── test_validation.py # Retry-loop + LLM-judge tests
├── Dockerfile
├── LICENSE
├── pyproject.toml
└── README.mdConfiguration
All settings via environment variables (.env supported via python-dotenv):
Variable | Required | Default | Description |
| Yes | — | OpenAI-compatible API base URL (tested: |
| Yes | — | API key for LLM endpoint (opaque string; no assumed prefix/format) |
| No |
| Model name for all LLM calls |
| No |
| Seconds between arXiv API requests |
| No |
| SQLite database path |
| No |
| Directory for JSONL traces |
Contributing
Fork the repository
Create a feature branch (
git checkout -b feature/amazing-feature)Make changes with tests
Run
ruff check && mypy src tests && pytest— all must passOpen a Pull Request
License
Distributed under the MIT License. See LICENSE for more information.
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Search GitHub, npm, PyPI, StackOverflow, ArXiv from one MCP — built for coding agents.
Central async research commons for persistent AI agents
Research intelligence for AI coding agents. 2M+ CS papers with evidence and tradeoffs.
Persistent memory and knowledge graphs for AI agents. Hybrid search, context checkpoints, and more.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceA LangGraph-powered research agent that performs iterative web searches using Google Search and Gemini models to generate structured reports with citations. It integrates with MCP-compatible clients like Claude and Cursor to enable sophisticated, multi-step AI research workflows.MIT
- AlicenseAqualityBmaintenanceAn intelligent research assistant MCP server for AI agents, providing task-oriented literature search and analysis across multiple academic databases.4626Apache 2.0
- AlicenseNot gradedqualityCmaintenanceEnables deep research tasks using a multi-agent architecture that integrates any LLM and MCP tools. Available via MCP stdio, streamable HTTP, and SSE transports.17MIT
- AlicenseAqualityAmaintenanceA multi-agent Retrieval-Augmented Generation system exposed as an MCP server. Ask a question and a LangGraph pipeline plans the retrieval, pulls evidence from a pgvector knowledge base, optionally augments it with live web research, drafts a cited answer, and then self-critiques it for grounding — revising until the answer is supported by the sources.31MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/hemanth2k6/mcp-research-agent-system'
If you have feedback or need assistance with the MCP directory API, please join our Discord server