Skip to main content
Glama
noualit

llama-memory

by noualit

llama-memory

MCP-Speicherdienst für llama-server mit persistentem Verlauf und semantischem Speicher über Postgres + PGVector.

Hinweis: Dieses Projekt ist nur für den lokalen/Demo-Gebrauch gedacht. Setzen Sie es nicht ohne zusätzliche Härtung (HTTPS, ordnungsgemäße Authentifizierung, Backups) direkt dem Internet aus.

Was es tut

  • Semantischer Speicher: Erinnerungen nach Bedeutung speichern und abrufen, nicht nur nach Schlüsselwörtern.

  • Gesprächsbrücke: LLM erstellt automatisch Gespräche; Erinnerungen werden verknüpft.

  • Sitzungsübergreifender Abruf: Fragen Sie "Worüber haben wir vorher gesprochen?" und erhalten Sie genaue Antworten.

  • MCP-Protokoll: Funktioniert direkt mit der integrierten MCP-Unterstützung von llama-server.

Related MCP server: engram

Anforderungen

  • Python 3.11 (Miniconda empfohlen)

  • PostgreSQL 16+ mit PGVector-Erweiterung

  • llama-server mit --jinja-Flag (erforderlich für Tool-Aufrufe)

  • nomic-embed-text, das auf llama-server läuft (Standardport 8081)

Installation

# Clone the repo
git clone https://github.com/noualit/llama-memory-local.git
cd llama-memory-local

# Create environment
conda create -n llama-memory python=3.11
conda activate llama-memory

# Install dependencies
pip install -e .

Konfiguration

Kopieren Sie .env.example nach .env und bearbeiten Sie:

cp .env.example .env

Beispiel:

# Database
DATABASE_URL="postgresql://postgres:yourpassword@localhost:5432/llamamem"

# Llama-server (LLM)
LLAMA_SERVER_BASE_URL="http://localhost:8080"

# Embedding model (nomic-embed-text via llama-server)
EMBEDDING_MODEL_URL="http://localhost:8081"

# Embedding model name (default: nomic-embed-text)
EMBEDDING_MODEL_NAME="nomic-embed-text"

# Service port
SERVICE_PORT=9001

Datenbank einrichten

Erstellen Sie die Datenbank und führen Sie Migrationen aus:

psql -U postgres -c "CREATE DATABASE llamamem;"
alembic upgrade head

Die Anwendung stellt beim Start auch das grundlegende Schema sicher, um die Bedienung zu erleichtern.

Dienst ausführen

# Using the script
.\scripts\run_server.ps1

# Or directly
python -m uvicorn app.main:app --host 0.0.0.0 --port 9001

Mit llama-server verbinden

Fügen Sie Ihrer llama-server-MCP-Konfiguration hinzu:

{
  "mcpServers": {
    "llama-memory": {
      "url": "http://YOUR_SERVER_IP:9001/mcp"
    }
  }
}

Der Dienst muss von llama-server aus erreichbar sein. Verwenden Sie die tatsächliche IP, nicht localhost, wenn sie auf verschiedenen Maschinen laufen.

MCP-Tools

Tool

Beschreibung

create_conversation

Neue Gesprächssitzung erstellen

list_conversations

Gespräche mit Speicheranzahl auflisten

get_conversation_history

Alle Erinnerungen in einem Gespräch abrufen

search_memories

Semantische Suche über alle Erinnerungen

save_memory

Wichtige Tatsache oder Entscheidung speichern

System-Prompt

Sie können:

  • Den empfohlenen System-Prompt vom Dienst abrufen:

    • GET /system-prompt → gibt reinen Text zurück.

  • Oder fügen Sie diese minimale Version in llama-server ein:

MEMORY WORKFLOW:
- At the start of each new conversation, call create_conversation with a short title.
- Use the conversation_id from create_conversation when calling save_memory.
- Before answering questions about past topics, call search_memories FIRST.
- When the user shares important information, save it with save_memory.
- If list_conversations has previous chats, check get_conversation_history for context.

Gesundheitscheck

curl http://localhost:9001/health

Gibt DB-Status, Status des Einbettungsdienstes und Tool-Anzahl zurück.

Architektur

Struktur auf hoher Ebene:

  • app/main.py — FastAPI-App, Lebenszyklus, /system-prompt

  • app/settings.py — Pydantic-Einstellungen aus .env

  • app/clients/embeddings.py — Ruft nomic-embed-text für Vektoren auf

  • app/db/engine.py — asyncpg-Verbindungspool (Singleton)

  • app/db/schema.py — Erstellt Tabellen automatisch beim Start

  • app/mcp/endpoint.py — MCP-Protokoll-Handler, Ratenbegrenzer

  • app/mcp/tools/ — Einzelne Tool-Implementierungen

  • migrations/ — Alembic-Datenbankmigrationen

Entwicklung

# Run tests
pytest tests/ -v

# Run with auto-reload
python -m uvicorn app.main:app --host 0.0.0.0 --port 9001 --reload

Siehe CONTRIBUTING.md für Beitragsrichtlinien.

Lizenz

MIT (siehe LICENSE-Datei).

A
license - permissive license
Not graded
quality - not tested
C
maintenance

Maintenance

Maintainers
Response time
Release cycle
Releases (12mo)
Commit activity

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Persistent semantic memory server for AI assistants via MCP, enabling long-term context retention and semantic search across conversations.
    11
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides persistent, local-first AI memory across sessions via MCP tools for storing, searching, and retrieving context from past interactions.
    1
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    Provides persistent memory for AI assistants via MCP, enabling them to store and recall facts, preferences, and tasks across conversations using either local file storage or a cloud backend with semantic search.
    5
    14
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    Provides persistent memory with semantic search for MCP-based AI agents, enabling them to store and recall information across sessions using vector embeddings.
    4
    1
    MIT

View all related MCP servers

Related MCP Connectors

  • Shared, governed long-term memory for AI agents across tools and sessions via MCP and REST.

  • Persistent memory for AI agents. Search, store, and recall across sessions.

  • Cross-AI personal memory. Save once in ChatGPT, recall in Claude, Mistral, Grok, or any MCP client.

View all MCP Connectors

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/noualit/llama-memory-local'

If you have feedback or need assistance with the MCP directory API, please join our Discord server