Skip to main content
Glama
loksai-dev
by loksai-dev
README.md
# ChronoGraph-Agent

Fair, reproducible benchmark comparing **Standard RAG**, **GraphRAG**, and **bi-temporal ChronoGraph** memory for temporal question answering (Groq LLM, Neo4j, Qdrant).

## Problem

LLM chat memory and naive RAG treat knowledge as static. When facts change, arrive out of order, or are corrected retroactively, vector and property-graph stores often surface **stale or conflated** evidence. That breaks point-in-time questions (“What was true on June 15?”) and audit questions (“What did we believe on July 3?”).

## Core Idea

ChronoGraph stores each fact on **two timelines**:

| Axis | Meaning |
|------|---------|
| **Valid time** (`v_start`, `v_end`) | When the fact was true in the world |
| **Transaction time** (`t_start`, `t_end`) | When the system recorded or superseded the fact |

Retroactive and out-of-order events **version** intervals instead of overwriting history.

## Architecture

```mermaid
flowchart TB
  User --> Agent
  Agent --> MCP
  MCP --> ChronoCore[ChronoGraph Core]
  ChronoCore --> Engine[Temporal Mutation Engine]
  Engine --> Neo4j
  ChronoCore --> Qdrant
```

### RAG (baseline)

```mermaid
flowchart LR
  E[Events] --> C[Chunking]
  C --> Emb[Embeddings]
  Emb --> Qdrant
  Qdrant --> LLM[Groq LLM]
```

### GraphRAG (baseline)

```mermaid
flowchart LR
  E[Events] --> X[Entity/Relation extraction]
  X --> Neo4j
  Neo4j --> T[Traversal]
  T --> LLM[Groq LLM]
```

**GraphRAG in this repo:** non-bi-temporal Neo4j graph; `MERGE` relationships with **latest evidence wins** on the edge. Documented so the comparison is fair, not a strawman.

### ChronoGraph

```mermaid
flowchart LR
  E[Events] --> W[write_fact]
  W --> Engine[Mutation Engine]
  Engine --> Neo4j
  Engine --> Qdrant
  Q[Question] --> R[Temporal retrieval]
  R --> LLM[Groq LLM]
```

## Comparison

| System | Store | Temporal model | Retrieval |
|--------|-------|----------------|-----------|
| RAG | Qdrant chunks | None | Semantic top-K |
| GraphRAG | Neo4j entities/edges | Latest edge only | 1-hop neighborhood |
| ChronoGraph | Neo4j facts + Qdrant | Bi-temporal intervals | `get_state_at`, `get_belief_at`, history |

## Installation

```powershell
git clone <your-repo-url> ChronoGraphAgent
cd ChronoGraphAgent
python -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install -r requirements.txt
copy .env.example .env
# Edit .env with GROQ_API_KEY and Neo4j credentials
```

## Environment Variables

See `.env.example`. Never commit `.env`.

## Docker

```powershell
docker compose up -d
```

- Neo4j Browser: http://localhost:7474 (default `neo4j` / `changeme` from compose)
- Qdrant: http://localhost:6333

## Dataset Generation

```powershell
$env:PYTHONPATH="src"
python scripts/generate_dataset.py
```

Outputs `benchmark/events.jsonl` (138+ events) and `benchmark/questions.jsonl` (210+ questions), seed **42**.

## Ingestion

```powershell
$env:PYTHONPATH="src"
python scripts/ingest_all.py
```

## Running Systems

All three share `MemorySystem`: `ingest`, `query`, `get_history`, `reset`.

## Running Benchmark

```powershell
$env:PYTHONPATH="src"
# Optional: limit cost while testing
$env:BENCHMARK_MAX_QUESTIONS="20"
python benchmark/run_benchmark.py
python scripts/generate_report.py
```

Results: `results/raw_results.json`, `results/results.csv`, `results/summary.json`, `results/plots/`.

Dry run (no DB/LLM calls): `BENCHMARK_DRY_RUN=true`.

## Web UI

```powershell
$env:PYTHONPATH="src"
python -m ui.app
```

Open http://localhost:8080 — compare systems, view metrics, trigger benchmark.

## Demo

```powershell
$env:PYTHONPATH="src"
python demo.py
```

Tom/AWS → Azure → retroactive GCP story with side-by-side answers when Qdrant/Neo4j are up.

## Temporal Example

| Valid interval | Provider |
|----------------|----------|
| Jan 1 → May 1 | AWS |
| May 1 → Jun 1 | GCP (learned later) |
| Jun 1 → Jul 1 | AWS |
| Jul 1 → ∞ | Azure |

**Valid time:** what was true on a calendar date.  
**Transaction time:** what the system knew when you ask belief-as-of questions.

## MCP

```powershell
$env:PYTHONPATH="src"
python -m mcp.server
```

Tools: `store_memory`, `search_memory`, `get_history`, `get_state_at`, `get_belief_at`, `get_changes`.

## Testing

```powershell
$env:PYTHONPATH="src"
python -m pytest tests -q
```

## Benchmark Methodology

- Same `events.jsonl` and `questions.jsonl` for all systems  
- Same Groq model and temperature (`GROQ_MODEL`, `GROQ_TEMPERATURE`)  
- Same answer prompt (`common/llm.py`)  
- Same embedding model for RAG/ChronoGraph semantic paths  
- Documented GraphRAG limitation (no bi-temporal edges)  
- **Temporal accuracy** = point-in-time / retroactive questions vs ground truth from mutation engine  

## Limitations

- Synthetic dataset; extraction heuristics for GraphRAG are simple  
- Python 3.10 works locally; project targets 3.11+  
- Full benchmark requires Docker (Neo4j + Qdrant) and Groq quota  
- LLM still formats final answers; ChronoGraph supplies verified temporal context  
- Neo4j credentials: update `.env` when you share production details  

## Future Work

- Distributed ingestion, richer NER for GraphRAG, learned temporal query planner  
- Production concurrency on mutation engine, larger multi-domain corpora  
- Automated report sections with failure clustering  

## License

MIT — see `LICENSE`.