Skip to main content
Glama
abhilash2429

agentic-research

by abhilash2429

agentic-research

An agentic web research pipeline. Ask a question in natural language; it searches the web, fetches candidate pages, chunks and embeds them, retrieves only the relevant spans, and returns an answer with citations that point at the exact source passage.

Built to be called by an AI agent as much as by a human. The MCP server exposes it as a deep_research tool, so a coding agent can reach for current, cited information instead of guessing from training data.

Why a retrieval layer

The usual agent approach to the web is: search, fetch the full page, put the whole thing in the context window, answer. That caps you at a handful of sources and charges full token price for pages that turn out to be mostly irrelevant.

typical:   search -> fetch full page -> stuff into context -> answer
           ~5 sources, full token price per page, no span-level citation

here:      search -> fetch N pages -> parent/child chunk -> embed
           -> retrieve only relevant children -> expand parents on demand
           -> compress -> answer
           ~30 sources, lower token cost, structural citations

Related MCP server: Deep Research MCP Server

Status

Work in progress, built in phases. See docs/ for a note on each layer.

Phase

Layer

State

0

Foundation, shared contracts

done

1

Config, role-based LLM factory

done

2

Fetch and extract

done

3

Hierarchical chunking

done

4

Index and rerank

in progress

5

Search providers

done

6

Single researcher loop

todo

7

Full graph: fan-out, clarification, compression

todo

8

Eval harness

todo

9

Service layer and API

todo

10

MCP server

todo

11

Web UI

todo

Setup

python -m venv .venv
.venv/Scripts/python -m pip install -e ".[dev,api,mcp]"
cp .env.example .env      # fill in LiteLLM and search keys
docker compose up -d qdrant
ares doctor

CLI

One command per layer, so you can see what each one does on its own.

ares doctor                        # proxy, models and Qdrant reachable
ares fetch <url> --show            # raw HTML to clean markdown
ares chunk <url>                   # the parent/child tree
ares index <url>                   # embed and store, or report a cache hit
ares search "..." --no-rerank      # retrieval without the reranker
ares search "..." --rerank         # and with it
ares websearch "..."               # provider results
ares research "..."                # the whole loop, cited answer
ares eval                          # scorecard against the committed baseline

License

MIT

Maintenance

ActivityMaintained
ResponsivenessSyncing

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables iterative deep research by integrating AI agents with search engines, web scraping, and large language models for efficient data gathering and comprehensive reporting.
    4
    323
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    An agent-based tool that provides web search and advanced research capabilities including document analysis, image description, and YouTube transcript retrieval.
    17
    Apache 2.0
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enhances LLM applications with deep autonomous web research capabilities, delivering higher quality information than standard search tools by exploring and validating numerous trusted sources.
    368
    MIT
  • A
    license
    B
    quality
    C
    maintenance
    An automated research agent that leverages Google Gemini models and Google Search to perform deep, multi-step web research, generating sophisticated queries and producing citation-rich answers.
    1
    28
    MIT