Skip to main content
Glama

Qwen Memory MCP

Long-term memory for AI agents, powered by Qwen on Alibaba Cloud and exposed over the Model Context Protocol (MCP). Any MCP-capable agent gains durable, cross-session memory that accumulates experience, retrieves what matters within a limited context window, and forgets what is outdated.

Hackathon track: Track 1 - MemoryAgent. License: MIT. Copyright (c) 2026 JHELY GLOBAL SL.

Repository: https://github.com/John-CEO-HQ/qwen-memory-mcp

Demo video: https://youtu.be/ZxXKvVY6iMQ

This project is a learning experiment for the Qwen Cloud Hackathon: a standalone MCP server for agent memory on Alibaba Cloud.

Documentation

Guide

Purpose

docs/TESTING-GUIDE.md

Master index: testing phases and pass criteria

docs/CREDENTIALS-AND-SETUP.md

Accounts, API keys, regions, cost guardrails

docs/PHASE1-REMOTE-INTEGRATION.md

Live Qwen / DashScope tests from your machine

docs/PHASE2-DEPLOYMENT-TESTING.md

Alibaba deploy + verify deployed URL

docs/INSTALL.md

Full install, local run, Alibaba production deploy, troubleshooting

docs/JUDGE-TESTING.md

Instructions for hackathon judges

deploy/README.md

Alibaba ECS / Function Compute quick reference

AGENTS.md

Agent conventions and isolation contract

Related MCP server: mindcore-memory-mcp

Why

Agents feel sharp inside a single conversation and amnesiac across sessions. This server gives an agent a managed memory layer that does four things well:

  • Write - extract a durable memory (preference, fact, commitment, event), with a Qwen-derived summary, tags, importance (salience), and kind.

  • Search - semantic retrieval ranked by similarity + salience + recency + reinforcement.

  • Recall context - pack the most critical memories into a fixed token budget, ready to inject into a small context window.

  • Forget - a maintenance pass that consolidates related memories with Qwen and lets stale, low-value memories decay away.

Architecture

flowchart LR
  agent["Any MCP client / agent"] -->|"MCP: write / search / recall / forget"| server["Qwen Memory MCP server (stdio or HTTP)"]
  server --> service["MemoryService"]
  service -->|"embeddings + reasoning"| qwen["Qwen on Alibaba Cloud Model Studio (DashScope)"]
  service -->|"persist + retrieve"| store["MemoryStore"]
  store --> file["File / in-memory (local + demo)"]
  store --> mysql["Alibaba Cloud RDS / PolarDB for MySQL (production)"]

Memory lifecycle:

flowchart TD
  w["memory_write"] --> analyze["Qwen analyze: summary, tags, salience, kind"]
  analyze --> embed1["Qwen embed (text-embedding-v3)"]
  embed1 --> active["active memory"]
  active --> s["memory_search / memory_recall_context"]
  s --> rank["rank: similarity + salience + recency + reinforcement"]
  rank --> pack["pack into token budget"]
  s -.reinforce.-> active
  active --> f["memory_forget"]
  f --> cluster["cluster by embedding similarity"]
  cluster --> consolidate["Qwen consolidate cluster -> canonical memory"]
  consolidate --> outdated["flag contradicted items -> forgotten"]
  active --> decay["decay score below threshold -> forgotten"]

Quick start

npm install

# Offline demo (no API key needed - deterministic local intelligence):
npm run demo

# Run the test suite:
npm test

# Run as an MCP server over stdio (for MCP Inspector / desktop clients):
npm run build && npm start

To use the real Qwen models, copy .env.example to .env and set QWEN_API_KEY (and optionally QWEN_BASE_URL for your region). Without a key, the server automatically falls back to the offline deterministic intelligence so it always runs.

MCP tools

Tool

Purpose

Key inputs

memory_write

Persist a durable memory

userId, content, sourceSession?, salience?

memory_search

Top-k semantic recall

userId, query, k?

memory_recall_context

Critical memories packed to a token budget

userId, query, tokenBudget

memory_forget

Consolidate + decay maintenance

userId

All memories are namespaced by userId, so one server can serve many agents.

Transports

  • stdio (MCP_TRANSPORT=stdio, default) - launched as a child process by a local MCP client.

  • Streamable HTTP (MCP_TRANSPORT=http) - stateless JSON-RPC at POST /mcp with optional Authorization: Bearer <MCP_AUTH_TOKEN>, plus GET /health. This is the shape used for cloud deployment and remote per-user MCP URLs.

Storage

  • MEMORY_STORE=memory - in-process, ephemeral (tests/demo).

  • MEMORY_STORE=file - single JSON file at MEMORY_FILE_PATH (local default).

  • MEMORY_STORE=mysql - Alibaba Cloud RDS / PolarDB for MySQL (production); schema is created automatically. Vectors are stored as JSON and scored in the app; see src/memory/mysql-store.ts for the AnalyticDB-PG (pgvector) upgrade path.

Alibaba Cloud / Qwen

The only integration points with Alibaba Cloud are src/qwen.ts (DashScope embeddings + chat) and src/memory/mysql-store.ts (RDS/PolarDB). See docs/INSTALL.md for full production setup and deploy/README.md for a short Alibaba quick reference.

Configuration

See .env.example for all variables (Qwen models, store selection, transport, auth token, and forgetting/decay tuning).

Layout

qwen-memory-mcp/
  src/
    qwen.ts               # Alibaba Cloud / Qwen (DashScope) intelligence  [PROOF]
    fake-intelligence.ts  # offline deterministic intelligence (tests/demo)
    intelligence.ts       # picks Qwen vs fake
    config.ts             # env-driven config
    types.ts              # domain types
    memory/
      store.ts            # MemoryStore interface
      file-store.ts       # file / in-memory store
      mysql-store.ts      # Alibaba RDS / PolarDB store               [PROOF]
      create-store.ts     # store factory
      ranking.ts          # retrieval ranking + token-budget packing
      forgetting.ts       # clustering + consolidation + decay
      service.ts          # MemoryService (orchestration)
    server.ts             # MCP server + 4 tools
    transports/
      stdio.ts            # stdio transport
      http.ts             # streamable HTTP transport (stateless)
    index.ts              # entry point
  demo/cli.ts             # multi-session offline demo
  test/                   # vitest suite
  deploy/                 # Alibaba Cloud deployment docs
  Dockerfile

License

MIT License. Copyright (c) 2026 JHELY GLOBAL SL. See LICENSE.

Available Tools

4 tools
memory_forgetConsolidate and forgetA

Runs the maintenance pass: clusters of related memories are merged by Qwen into one canonical memory, contradicted/outdated items are forgotten, and stale low-importance memories decay away. Returns a report of what was consolidated, archived, forgotten, and retained.

ParametersJSON Schema
NameRequiredDescriptionDefault
userIdYes

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It explains the process: merging clusters, forgetting contradicted/outdated items, decaying stale low-importance memories, and returning a report. It could mention idempotency or data loss details, but it is fairly transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core action, and the second sentence explains the output. No extraneous information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simplicity of the tool (1 param, no output schema, no annotations), the description explains the process but lacks parameter semantics and usage guidance. Without an output schema, describing the report structure would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% (no description for the userId parameter). The description does not explain the parameter's meaning or how it affects the operation. With only one required parameter, the description should compensate, but it fails to provide any additional semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool runs a maintenance pass that consolidates, forgets, and decays memories. It uses specific verbs and resources, and distinguishes from siblings (memory_recall_context, memory_search, memory_write) which are for recall, search, and writing, not maintenance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context (maintenance pass) but provides no explicit guidance on when to use this tool versus alternatives or when not to use it. No exclusions or prerequisites are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_recall_contextRecall context within a token budgetB

Returns the most critical memories for a query, greedily packed to fit a token budget, as a ready-to-inject context block. Use this to load long-term memory into a limited context window before answering.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes
userIdYes
tokenBudgetYesApproximate max tokens the returned context may use.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries full burden. Discloses greedy packing and criticality ranking, but does not cover idempotency, state mutations, rate limits, or other behavioral traits. Adds some value beyond the name but lacks comprehensiveness.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences efficiently convey the core function and usage. While concise and front-loaded, the description could benefit from a structured breakdown, but overall it avoids unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of annotations and output schema, the description partially covers purpose and usage but lacks details on parameter semantics, return format, and side effects. Adequate but not fully comprehensive for an agent to use without additional information.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33% (only tokenBudget has a description). The tool description does not explain query or userId parameters, leaving them undefined. Fails to compensate for the low schema coverage, so minimal meaning is added beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Describes returning the most critical memories for a query, greedily packed within a token budget. Clearly states verb and resource, distinguishing from sibling tools by emphasizing the token budget and context-block output, though not explicitly differentiating from memory_search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states to use this tool to load long-term memory into a limited context window before answering. Provides clear context for when to use, but does not explicitly exclude alternatives like memory_search or mention when not to use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_writeWrite memoryA

Persist something worth remembering about a user (a preference, fact, commitment, or event). Qwen derives a short summary, tags, importance (salience), and kind. Call this whenever the user reveals durable information you should recall in future sessions.

ParametersJSON Schema
NameRequiredDescriptionDefault
userIdYesStable id namespacing this user's memories.
contentYesThe information to remember, in plain language.
salienceNoOptional importance override in [0,1]; otherwise Qwen decides.
sourceSessionNoOptional originating session/conversation id.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description reveals that Qwen derives summary, tags, importance, and kind, which is beyond the schema. However, with no annotations, it doesn't cover side effects, idempotency, or limits. Adequate but not thorough.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two efficient sentences: first defines purpose with examples, second provides usage guidance. No wasted words, front-loaded with key information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter write tool with no output schema or annotations, the description covers purpose, usage, and key behavioral nuances (derived fields). Lacks some details like return value or error handling, but sufficient for typical use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds no parameter-specific details beyond the schema, only tying the concept of 'content' to the tool's purpose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool persists user memories, with specific examples (preference, fact, commitment, event). It implicitly distinguishes from siblings (forget, recall_context, search) by focusing on writing new information.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The second sentence explicitly says when to call: 'whenever the user reveals durable information you should recall in future sessions.' No explicit when-not-to-use or alternative tools mentioned, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedmemory_forget
    • First observedmemory_recall_context
    • First observedmemory_search
    • First observedmemory_write

TDQS

A3.9/5.0

Scored across 4 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: writing new memories, searching them, recalling context for a query, and running maintenance. No overlap in functionality.

Naming Consistency5/5

All tool names follow a consistent 'memory_<verb>' pattern, making the action clear and predictable.

Tool Count5/5

Four tools cover the core memory operations (write, search, recall, maintenance) without unnecessary bloat. The scope is appropriate for a memory management server.

Completeness3/5

Missing explicit update and list operations; however, the automatic maintenance and selective retrieval cover common use cases. Gaps exist for direct manipulation of individual memories.

Maintenance

ActivityStale
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    A production-grade long-term memory MCP server that enables AI agents to persist and recall memories across sessions with importance weighting, confidence calibration, and efficient context window management.
    9
    1
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Persistent, searchable memory for AI agents over the Model Context Protocol, enabling memory storage, full-text search with BM25 ranking, and retrieval across sessions.
    6 npm
    MIT
  • F
    license
    Not graded
    quality
    B
    maintenance
    A persistent memory server for AI agents using MCP protocol, enabling semantic storage and retrieval of dialogues, documents, and agent states.
    -