Skip to main content
Glama
loksai-dev
by loksai-dev

ChronoGraph-Agent

Fair, reproducible benchmark comparing Standard RAG, GraphRAG, and bi-temporal ChronoGraph memory for temporal question answering (Groq LLM, Neo4j, Qdrant).

Problem

LLM chat memory and naive RAG treat knowledge as static. When facts change, arrive out of order, or are corrected retroactively, vector and property-graph stores often surface stale or conflated evidence. That breaks point-in-time questions (“What was true on June 15?”) and audit questions (“What did we believe on July 3?”).

Related MCP server: Hebbrix MCP Server

Core Idea

ChronoGraph stores each fact on two timelines:

Axis

Meaning

Valid time (v_start, v_end)

When the fact was true in the world

Transaction time (t_start, t_end)

When the system recorded or superseded the fact

Retroactive and out-of-order events version intervals instead of overwriting history.

Architecture

flowchart TB
  User --> Agent
  Agent --> MCP
  MCP --> ChronoCore[ChronoGraph Core]
  ChronoCore --> Engine[Temporal Mutation Engine]
  Engine --> Neo4j
  ChronoCore --> Qdrant

RAG (baseline)

flowchart LR
  E[Events] --> C[Chunking]
  C --> Emb[Embeddings]
  Emb --> Qdrant
  Qdrant --> LLM[Groq LLM]

GraphRAG (baseline)

flowchart LR
  E[Events] --> X[Entity/Relation extraction]
  X --> Neo4j
  Neo4j --> T[Traversal]
  T --> LLM[Groq LLM]

GraphRAG in this repo: non-bi-temporal Neo4j graph; MERGE relationships with latest evidence wins on the edge. Documented so the comparison is fair, not a strawman.

ChronoGraph

flowchart LR
  E[Events] --> W[write_fact]
  W --> Engine[Mutation Engine]
  Engine --> Neo4j
  Engine --> Qdrant
  Q[Question] --> R[Temporal retrieval]
  R --> LLM[Groq LLM]

Comparison

System

Store

Temporal model

Retrieval

RAG

Qdrant chunks

None

Semantic top-K

GraphRAG

Neo4j entities/edges

Latest edge only

1-hop neighborhood

ChronoGraph

Neo4j facts + Qdrant

Bi-temporal intervals

get_state_at, get_belief_at, history

Installation

git clone <your-repo-url> ChronoGraphAgent
cd ChronoGraphAgent
python -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install -r requirements.txt
copy .env.example .env
# Edit .env with GROQ_API_KEY and Neo4j credentials

Environment Variables

See .env.example. Never commit .env.

Docker

docker compose up -d

Dataset Generation

$env:PYTHONPATH="src"
python scripts/generate_dataset.py

Outputs benchmark/events.jsonl (138+ events) and benchmark/questions.jsonl (210+ questions), seed 42.

Ingestion

$env:PYTHONPATH="src"
python scripts/ingest_all.py

Running Systems

All three share MemorySystem: ingest, query, get_history, reset.

Running Benchmark

$env:PYTHONPATH="src"
# Optional: limit cost while testing
$env:BENCHMARK_MAX_QUESTIONS="20"
python benchmark/run_benchmark.py
python scripts/generate_report.py

Results: results/raw_results.json, results/results.csv, results/summary.json, results/plots/.

Dry run (no DB/LLM calls): BENCHMARK_DRY_RUN=true.

Web UI

$env:PYTHONPATH="src"
python -m ui.app

Open http://localhost:8080 — compare systems, view metrics, trigger benchmark.

Demo

$env:PYTHONPATH="src"
python demo.py

Tom/AWS → Azure → retroactive GCP story with side-by-side answers when Qdrant/Neo4j are up.

Temporal Example

Valid interval

Provider

Jan 1 → May 1

AWS

May 1 → Jun 1

GCP (learned later)

Jun 1 → Jul 1

AWS

Jul 1 → ∞

Azure

Valid time: what was true on a calendar date.
Transaction time: what the system knew when you ask belief-as-of questions.

MCP

$env:PYTHONPATH="src"
python -m mcp.server

Tools: store_memory, search_memory, get_history, get_state_at, get_belief_at, get_changes.

Testing

$env:PYTHONPATH="src"
python -m pytest tests -q

Benchmark Methodology

  • Same events.jsonl and questions.jsonl for all systems

  • Same Groq model and temperature (GROQ_MODEL, GROQ_TEMPERATURE)

  • Same answer prompt (common/llm.py)

  • Same embedding model for RAG/ChronoGraph semantic paths

  • Documented GraphRAG limitation (no bi-temporal edges)

  • Temporal accuracy = point-in-time / retroactive questions vs ground truth from mutation engine

Limitations

  • Synthetic dataset; extraction heuristics for GraphRAG are simple

  • Python 3.10 works locally; project targets 3.11+

  • Full benchmark requires Docker (Neo4j + Qdrant) and Groq quota

  • LLM still formats final answers; ChronoGraph supplies verified temporal context

  • Neo4j credentials: update .env when you share production details

Future Work

  • Distributed ingestion, richer NER for GraphRAG, learned temporal query planner

  • Production concurrency on mutation engine, larger multi-domain corpora

  • Automated report sections with failure clustering

License

MIT — see LICENSE.

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    Provides long-term memory and a temporal knowledge graph for AI agents, enabling persistent memory and reasoning across sessions.
    33
    30 PyPI
    2
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    A bi-temporal, provenance-carrying memory primitive for AI agents. Enables storing facts, recall, revision, and audit trails via MCP with SQLite storage.
    6
    Apache 2.0
  • A
    license
    A
    quality
    B
    maintenance
    Enables AI agents to record, recall, correct, and forget evidence-backed factual claims with temporal history, while explaining whether remembered information is current, historical, or contested.
    5
    Apache 2.0