Skip to main content
Glama
README.md
# ARIA — Autonomous Research & Intelligence Assistant

> An MCP server that autonomously researches any topic: searches the web, scrapes sources, extracts insights, builds a knowledge graph, and synthesizes a structured research brief — in under 90 seconds.

---

## What It Does

Give ARIA a topic → it autonomously:

1. **Searches** the web for relevant sources (Tavily API)
2. **Scrapes** and cleans full page content (httpx + BeautifulSoup)
3. **Extracts** key concepts, claims, and gaps from each source (Claude API)
4. **Builds** a NetworkX knowledge graph of connected concepts
5. **Synthesizes** a final research brief with citations

---

## Setup

### 1. Clone & create virtual environment
```bash
git clone https://github.com/YOUR_USERNAME/aria-mcp.git
cd aria-mcp
python -m venv venv
source venv/bin/activate        # Windows: venv\Scripts\activate
pip install -r requirements.txt
```

### 2. Configure API keys
```bash
cp .env.example .env
# Open .env and fill in your keys
```

Get keys from:
- Anthropic API: https://console.anthropic.com
- Tavily API: https://tavily.com (free tier works)

### 3. Connect to Claude Desktop

Open your Claude Desktop config file:
- **macOS**: `~/Library/Application Support/Claude/claude_desktop_config.json`
- **Windows**: `%APPDATA%\Claude\claude_desktop_config.json`

Add the ARIA server (replace the path with your actual absolute path):
```json
{
  "mcpServers": {
    "aria": {
      "command": "python",
      "args": ["/absolute/path/to/aria-mcp/server/main.py"]
    }
  }
}
```

Restart Claude Desktop. ARIA will appear as an available MCP tool.

### 4. Or use the CLI client
```bash
cd client
python aria_client.py "federated learning in healthcare"
python aria_client.py "transformer architecture" 3
```

---

## Project Structure

```
aria-mcp/
├── server/
│   ├── main.py                  ← MCP server entry point (integration)
│   ├── tools/
│   │   ├── search.py            ← Tavily web search
│   │   ├── scraper.py           ← httpx + BeautifulSoup scraper
│   │   ├── summarizer.py        ← Claude-powered insight extraction
│   │   └── graph.py             ← NetworkX knowledge graph
│   └── utils/
│       └── helpers.py           ← Shared utilities
├── client/
│   └── aria_client.py           ← CLI demo client
├── tests/
│   ├── test_search.py
│   ├── test_scraper.py
│   ├── test_summarizer.py
│   └── test_graph.py
├── output/                      ← Research JSON results (gitignored)
├── .env.example
├── .gitignore
├── claude_desktop_config.json   ← Claude Desktop config snippet
├── requirements.txt
└── README.md
```

---

## Testing Individual Modules

```bash
# From project root, with venv activated
python tests/test_search.py
python tests/test_scraper.py
python tests/test_summarizer.py
python tests/test_graph.py
```

---

## Team Split

| Person | File | Responsibility |
|--------|------|----------------|
| Person 1 | `tools/search.py` | Web search via Tavily |
| Person 2 | `tools/scraper.py` | URL scraping + text extraction |
| Person 3 | `tools/summarizer.py` | Claude-powered summarization + synthesis |
| Person 4 | `tools/graph.py` | Knowledge graph construction |
| All together | `server/main.py` | MCP server integration (Day 2) |

---

## Tech Stack

| Layer | Tool |
|-------|------|
| MCP Framework | `mcp` Python SDK by Anthropic |
| LLM | Claude Sonnet via Anthropic API |
| Web Search | Tavily API |
| Web Scraping | httpx + BeautifulSoup4 |
| Knowledge Graph | NetworkX |
| Language | Python 3.11+ |

---

## Demo

In Claude Desktop, type:

> *"Research the topic: Federated Learning in IoT devices"*

ARIA will autonomously search 5 sources, scrape them, summarize each, build a knowledge graph, and produce a full research brief — all in real time.

---

## License

MIT