Skip to main content
Glama
ininihian

flashback-memory

by ininihian

Flashback Memory

Conversation memory system untuk LLM — simpan histori chat, lalu saat lo tanya, sistem cari memory yang relevan lewat embedding + semantic search (bukan cuma keyword). Dibangun sebagai MCP server supaya bisa dipasang ke klien mana pun (LibreChat, Claude Desktop, dll).

Cara kerja

CHAT ─▶ store_turn (user + AI)
          └─ HotBuffer akumulasi token
               └─ [OFFLOAD] kalau lewat hot window:
                    ├─ [SATPPAM]  LLM (NVIDIA) baca chunk → judul/topik/kategori
                    ├─ [EMBED]    NVIDIA nemotron → vektor 2048-dim
                    └─ [SAVE]     simpan block + vektor ke LanceDB

CHAT ─▶ flashback_memory (query)
          ├─ [EMBED]   query → vektor
          ├─ [SEARCH]  LanceDB cari kandidat mirip
          └─ [RESULT]  balikin memory / klarifikasi kalau ambigu

Setup

1. Prerequisites

  • Python 3.13 (disarankan di proot-distro Debian, karena mcp/lancedb butuh glibc — native Termux Python 3.14 gak kompatibel)

  • venv dengan dependency (lihat requirements.txt)

python3 -m venv /root/.venv_fbn_deb
. /root/.venv_fbn_deb/bin/activate
pip install -r requirements.txt

2. Config & API key

Copy .env.example.env, isi NVIDIA_API_KEY lo:

cp .env.example .env
nano .env   # isi: NVIDIA_API_KEY=***

config.yaml sudah siap pakai. Edit kalau perlu (embedding.dim, hot_window_tokens, dsb).

3. Jalankan

SSE mode (untuk LibreChat / remote client):

bash scripts/start.sh

Server jalan di http://localhost:8001/sse. Log di /root/flashback_sse.log.

Stdio mode (untuk MCP client lokal):

bash scripts/start_stdio.sh

4. Monitoring

Lihat apa yang terjadi di dalam MCP (SATPPAM / EMBED / SEARCH / offload):

bash scripts/logs.sh          # tail -f log
bash scripts/status.sh        # cek server jalan atau tidak

Pasang ke LibreChat

Di librechat.yaml:

mcpServers:
  FlashbackMemory:
    type: sse
    url: http://<IP-DEBIAN-PROOT>:8001/sse

mcpSettings:
  allowedAddresses:
    - <IP-DEBIAN-PROOT>:8001
  allowedDomains: []

⚠️ Pakai IP private (bukan 127.0.0.1) karena LibreChat jalan di container proot berbeda. Cek IP dengan hostname -I di Debian proot.

Tools

  • store_turn(user_msg, model_output, conv_id, conv_key) — simpan 1 turn

  • flashback_memory(query, conv_id, conv_key) — cari memory relevan

Struktur

flashback-memory/
├── README.md
├── .env.example
├── .gitignore
├── config.yaml
├── requirements.txt
├── server.py              # MCP stdio
├── server_sse.py          # MCP SSE (LibreChat)
├── core/                  # store, chunker, retriever, ambiguity
├── models/                # OpenAI-compatible client
├── prompts/               # SATPPAM prompt
├── scripts/               # start / logs / status
└── tests/                 # unit + e2e

Env vars

Var

Default

Keterangan

NVIDIA_API_KEY

(dari .env)

API key NVIDIA (embedding + SATPPAM)

FLASHBACK_PORT

8001

Port SSE server

Catatan

  • DB (/root/flashback_memory_db) dan .env gak di-commit (lihat .gitignore).

  • Embedding: nvidia/nemotron-3-embed-1b (dim 2048).

-
license - not tested
-
quality - not tested
C
maintenance

Maintenance

Maintainers
Response time
Release cycle
Releases (12mo)
Commit activity

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/ininihian/flashback-memory'

If you have feedback or need assistance with the MCP directory API, please join our Discord server