mcp-ollama-qdrant
A persistent vector-memory server (Ollama embeddings + Qdrant) exposing six MCP tools for storing, searching, updating, and deleting memories across collections.
Save one memory —
save_memory(text, metadata, collection): embeds text via Ollama and upserts into Qdrant; metadata accepts a JSON string or object (parsed leniently — invalid JSON stores{}and returns a warning); collection is auto-created if missing.Save many memories —
save_memories(texts, metadata, collection): batch-embeds a list of texts and upserts them together with shared metadata.Semantic search —
search_memory(query, limit=3, filter, collection): similarity search with an optional payload filter (JSON string/object; list values = MatchAny, scalars = exact match, conditions AND-ed); bad filter JSON falls back to unfiltered search plus a warning.Update in place —
update_memory(point_id, text, metadata, collection): re-embeds and overwrites text, replaces metadata when given, keeps existing payload when empty; errors if the point doesn't exist or nothing is supplied to change.Delete —
delete_memory(point_id, collection): hard-deletes a point, erroring if the ID doesn't exist.List collections —
list_collections(): enumerates all Qdrant collections.
Note: this schema is a reduced subset of the README's documented surface — it omits supersede/chunked saves, read/browse tools (get, list_memories, count, values), patch_metadata, archive/unarchive, and search refinements (since, tag, min_score, recency_weight, mmr, include_inactive), and it keeps the old warn-and-continue JSON behavior rather than the README's default hard-reject.
Provides persistent vector memory by using Ollama to generate embeddings for stored memories and search queries, enabling save, batch save, similarity search, and delete operations backed by a Qdrant vector store.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@mcp-ollama-qdrantsave this as a memory: the client prefers email over phone"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
vector-memory
Persistent vector memory for AI agents backed by Ollama (embeddings,
tested with qwen3-embedding:8b) and Qdrant (vector store). Two entry
points over the same core (vector_memory.core):
vector-memory-mcp— MCP stdio server (1:1 with core ops)vector-memory— one-shot CLI (typer, 1:1 with the same core ops)
It exposes these operations:
Tool | Description |
| Embeds |
| Batch version: embeds in |
| Semantic search. Superseded/archived memories are hidden by default ( |
| Re-embeds |
| Metadata-only patch: merge |
| Hard delete by ID (multi-ID supported; missing IDs reported). |
| Archive/unarchive (CLI: |
| Read/browse operations — no embedding call. |
| Lists all existing Qdrant collections. |
| Health checks: Ollama model, Qdrant version/compat, collection, legacy-points report. |
| JSONL backup/restore ( |
| Copy a collection into a new one with the current model (same IDs; source never mutated; resumable). |
| Non-destructive system-field backfill on legacy points (resumable, |
Every data tool takes an optional collection string; an empty value uses
the server-configured collection (--collection / COLLECTION_NAME).
Payload filtering
search_memory accepts an optional filter — a JSON object or a JSON string
built from payload fields. List values become a MatchAny condition (matches if the payload
field contains any of the values), scalar values become exact matches.
Multiple conditions are AND-ed together:
{"tags": ["x"]} // payload.tags contains "x"
{"source": "doc1"} // exact match
{"tags": ["a", "b"], "source": "s"} // AND of MatchAny + matchOn startup the server connects to Ollama and Qdrant and creates the collection
automatically if it does not exist (cosine distance, dimension probed from the
embedding model). If a collection already exists with a different vector
dimension than the current EMBED_MODEL produces — whether the default
collection at startup or an ad-hoc one named in a tool call — the server fails
fast with a clear error instead of silently storing corrupt vectors — fix it
by deleting and recreating the collection, or by switching back to the
original embedding model.
Invalid metadata/filter JSON is rejected by default: the operation fails
with a single stderr line (ArgumentError: metadata is not valid JSON (position N))
and nothing is written. Metadata keys must be lowercase identifiers
(^[a-z][a-z0-9_]{0,63}$, no leading _), and values must be JSON scalars or
lists of scalars. The legacy warn-and-continue behavior is available via
--lenient on save commands or VM_LENIENT=1 in the environment.
Requirements
Python 3.11+ and uv
A reachable Ollama instance (default
http://192.168.X.X:11434)A reachable Qdrant instance (default
http://192.168.X.X:6333)
Related MCP server: Qdrant MCP Server
Installation (CLI)
From the repo directory, install both executables as editable uv tools (on
PATH in ~/.local/bin, edits to the checkout take effect immediately):
uv tool install -e .This installs vector-memory (and the optional vector-memory-mcp server
executable). Verify with vector-memory list-collections.
Configuration
Settings resolve in order: CLI flags > environment variables > defaults.
Setting | CLI flag | Env var | Default |
Ollama base URL |
|
|
|
Qdrant base URL |
|
|
|
Embedding model |
|
|
|
Collection name |
|
|
|
Upgrading from mcp-ollama-qdrant
Upgrading from the old mcp-ollama-qdrant repo/server: collections and
memories carry over unchanged — the package rename does not touch Qdrant.
Register the new entry points (vector-memory CLI, vector-memory-mcp
server) in your client config instead of mcp-ollama-qdrant; the default
collection agent_scenarios is reused as-is.
Running
With uv (recommended — handles the venv and sync automatically):
uv sync
uv run vector-memory-mcp # run the stdio MCP server (or: python mcp_server.py)
uv run vector-memory search "db outage" # one-shot CLI (no daemon)
uv run vector-memory --help # save | save-many | search | update | delete | list-collectionsEvery CLI invocation is one-shot — there is no daemon and no CLI-to-server RPC; the CLI calls the same core functions as the MCP server directly.
Interactive testing / inspection:
uv run mcp dev mcp_server.pyCLI commands
vector-memory save "text" --project p --type decision --tags x [--metadata '{...}'] [--collection C]
vector-memory save "text" --supersedes <old-id> # fact changed: old becomes superseded
vector-memory save-many "text A" "text B" --project p [--metadata '{...}'] [--collection C]
vector-memory search "query" [--limit N] [--filter '{"tags":["x"]}'] [--project p] [--since 7d]
[--tag urgent --tag arch] [--type decision] [--min-score 0.2] [--recency-weight 0.5]
[--mmr 0.7] [--brief] [--format compact] [--include-inactive] [--json] [--collection C]
vector-memory update <point-id> --text "new text" [--metadata '{...}'] [--merge-metadata]
vector-memory patch <point-id> --set '{"tags":["x"]}' [--unset key1 --unset key2] [--collection C]
vector-memory archive <point-id>[,<point-id>...] # hide without deleting
vector-memory get <point-id>[,<point-id>...] [--system] # full text + metadata
vector-memory list [--project p] [--limit 25] [--order-by created]
vector-memory count / stats / values project # inventory, no embedding call
vector-memory migrate [--collection C] [--assume-model M] [--dry-run]
vector-memory reembed --collection SRC --to DST [--resume]
vector-memory export [--collection C] [--with-vectors] > backup.jsonl
vector-memory import backup.jsonl [--collection C] [--reembed] [--on-conflict skip|overwrite]
vector-memory delete <point-id>[,<point-id>...]
vector-memory delete-by-filter --filter '{"project":"p"}' # DRY-RUN by default; --no-dry-run --yes to really delete
vector-memory delete-collection NAME --confirm NAME # whole collection (guarded)
vector-memory list-collections
vector-memory doctor [--json]Semantics worth knowing: identical normalized text is idempotent — saving
it again refreshes the same point (metadata merged) instead of creating a
second one; pass --allow-duplicate to force a new point. Near-duplicates
are reported, never merged (similar <id> <score> <text> lines + a
reminder that similarity does not imply equivalence — check numbers,
versions, negations; --on-similar warn|skip|error, threshold configurable
via VM_DEDUPE_THRESHOLD, calibrated default 0.985 — see
docs/similarity-calibration.md). Superseded/archived memories are hidden
from default search (legacy points without _status stay visible;
--include-inactive shows them annotated). update --metadata replaces the
whole payload metadata (--merge-metadata merges instead); omit --text
for a no-re-embed metadata update. Search hits include ID + score + metadata
text;
--brief/--format compacttruncate (preferringsummary). Invalid metadata/filter JSON fails the command (one stderrArgumentError: ...line, nothing written);--lenient/VM_LENIENT=1restores the old warn-and-continue behavior. Text matching a high-confidence secret rule is rejected (SensitiveContentError; the value is never echoed;--allow-sensitiveoverrides — use only for false positives).
MCP tools vs. CLI (deliberate subset)
The MCP stdio server (vector-memory-mcp) exposes the memory operations an
agent needs at conversation time; operations that are administrative,
destructive-by-design, or one-shot-maintenance stay CLI-only (deliberate
1:1 subset, not an omission):
MCP tool | Notes |
| dedupe idempotency, |
| replaces listed IDs (lifecycle) |
| long-text chunking into |
| batch save |
| full parity with CLI search: |
| read/browse with convenience filters |
| in-place edit; metadata-only patch |
| lifecycle |
| single-ID delete; collection list |
CLI-only (not in MCP): migrate, reembed, delete-by-filter,
delete-collection, multi-ID delete, export/import, consolidate,
stats, doctor — one-shot maintenance/backup/destructive commands that
belong in a shell, not a model-driven tool surface. Any MCP tool failure
raises inside the handler so MCP clients receive a proper error result
(isError), not a successful-looking string.
System payload fields
Every save stamps _-prefixed system fields alongside user metadata:
_created_ts/_updated_ts (epoch floats), _created_at/_updated_at
(ISO-8601), _content_hash (sha256 of the normalized text), _embed_model,
and _agent (when VM_AGENT_ID is set). Lifecycle fields _status,
_supersedes, _superseded_by are reserved for the supersede/archive
lifecycle. Legacy points without these fields stay valid; use
vector-memory migrate to backfill them non-destructively (resumable;
--dry-run reports how many points would change; --assume-model records
the embedding model only if you are certain of what produced the vectors).
Payload indexes (keyword on project/type/tags/source/_status, float
on timestamps) are created idempotently when a collection is opened.
Backups: JSONL export/import and Qdrant snapshots
export/import (JSONL) is a portable, human-inspectable backup of memory
content. For full-fidelity backups (vectors, indexes, collection config,
point versions), use Qdrant's own snapshot API instead — it captures the
collection exactly and restores as-is:
# create a snapshot (full fidelity: vectors, payload indexes, config)
curl -X POST "http://<qdrant-host>:6333/collections/agent_scenarios/snapshots"
# list / download
curl "http://<qdrant-host>:6333/collections/agent_scenarios/snapshots"
# restore into a new collection from a snapshot file
curl -X PUT "http://<qdrant-host>:6333/collections/agent_scenarios_restored?priority=snapshot" \
-H 'Content-Type: application/octet-stream' --data-binary @<snapshot-file>Snapshot files land on the Qdrant server's storage (snapshots/ directory);
schedule the create call with your backup cron. Prefer snapshots for
disaster recovery; prefer export/import when you need the memory content
in a portable format or must re-embed into a different model/collection.
Unarchive semantics
unarchive of a formerly superseded point clears the stale
_superseded_by link together with _status — the point re-enters active
search without a dangling replacement reference.
MCP client config
Add to your client's MCP config (Claude Desktop, Hermes, etc.). The stdio
server entry point is vector-memory-mcp (the vector-memory command is
the CLI, not the server):
{
"mcpServers": {
"vector-memory": {
"command": "uv",
"args": [
"--directory", "/path/to/vector-memory",
"run", "vector-memory-mcp"
],
"env": {
"OLLAMA_URL": "http://192.168.X.X:11434",
"QDRANT_URL": "http://192.168.X.X:6333",
"EMBED_MODEL": "qwen3-embedding:8b",
"COLLECTION_NAME": "agent_scenarios"
}
}
}
}(Env entries are optional if the defaults already point at your instances.)
For Hermes ~/.hermes/config.yaml:
mcp:
servers:
vector-memory:
command: uv
args: ["--directory", "/path/to/vector-memory", "run", "vector-memory-mcp"]Testing
Offline unit tests (Ollama and Qdrant are mocked — no live services needed):
uv run pytest tests/End-to-end smoke test against live Ollama + Qdrant (saves a few memories, searches for them, prints similarity scores):
uv sync
uv run python scripts/live_smoke.pyNotes
All diagnostics are logged to stderr; stdout is reserved for the stdio MCP transport.
Dependency pins:
qdrant-client>=1.15,<2,mcp[cli]<2, numpy 1.x on Python < 3.13 — chosen for compatibility with older x86-64 hardware (pre-x86-64-v2) and the mcp v2 FastMCP rename. Adjust only with reason.
Available Tools
6 toolsdelete_memoryA
Delete a stored memory (point) from the vector DB by ID.
If point_id does not exist, an error is returned (existence is checked
first). If collection is given, the deletion happens there (default: the
server-configured collection).
| Name | Required | Description | Default |
|---|---|---|---|
| point_id | Yes | ||
| collection | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses that a nonexistent point_id raises an error because existence is checked first, and that collection defaults to the server-configured one. It does not state that the deletion is permanent/irreversible, whether it requires elevated permissions, or whether it can remove multiple points.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action, followed by the failure mode and the optional scoping parameter. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and both parameters plus the error path are covered. The one notable omission for a destructive tool is an explicit irreversibility/permission note.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does for both parameters: point_id is the ID of the stored point whose existence is validated, and collection scopes the deletion with a documented default (server-configured collection). No format examples (e.g., ID scheme) are given, but meaning is clear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (delete) and resource (a stored memory/point) in the vector DB, keyed by ID. It is unambiguous what the tool does and clearly distinct from search_memory/update_memory, though it never names the siblings explicitly to route the agent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the obvious CRUD semantics, and the error-on-missing-id note gives some operational context. However, there is no explicit guidance on when to delete versus update_memory or how to handle a memory that should be removed but not destroyed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_collectionsA
List all collections currently present in Qdrant.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry behavioral disclosure. It implies a read-only enumeration with 'currently present', but does not state that it is non-mutating, whether results are paginated or ordered, or whether any auth/scope is required — though for a zero-parameter listing tool there is limited behavior to disclose.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler; the resource and scope appear immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. For a zero-parameter list operation the description is essentially complete, with only minor omissions around ordering/auth that are unlikely to affect correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so there is nothing for the description to disambiguate. Baseline 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('collections'), scoped to 'currently present in Qdrant'. It is unambiguous, and no sibling tool (delete_memory, save_memory, search_memory, update_memory) operates on collections, so no sibling conflict needs resolving.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to call this versus alternatives, nor any prerequisite or follow-up context (e.g., using the result before calling another collection-scoped tool). It only states what the tool does.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_memoriesB
Save multiple documents into the vector DB in one batch.
All texts are embedded in a single Ollama call and upserted together. metadata (JSON string or object) is applied to every document. If collection is given, the memories are stored there (created automatically if missing).
| Name | Required | Description | Default |
|---|---|---|---|
| texts | Yes | ||
| metadata | No | {} | |
| collection | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does add real behavior: embeddings happen in a single Ollama call, writes are upserts, metadata is stamped on every document, and the collection is auto-created if missing. It omits error/failure behavior, overwrite semantics for existing documents, and any size or rate constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences with no filler, front-loaded with what the tool does before the operational details. The information is well ordered, though the parenthetical about JSON string or object slightly duplicates the schema's anyOf.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, and the mutation semantics are partially covered. However, for a no-annotation batch write the description leaves the agent without error handling, overwrite behavior, or batch-size expectations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate; it explains that metadata is applied to every document and that collection defaults to being auto-created. It says nothing about how texts is chunked, per-call limits, or the string-vs-object metadata duality, leaving gaps the schema does not fill.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Save multiple documents into the vector DB') and the scope ('in one batch'), which implicitly but clearly separates it from the singular save_memory sibling. It stops short of naming the sibling it is not, so differentiation relies on the agent inferring from 'multiple'/'batch'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use versus when-not guidance and no acknowledgment of save_memory, search_memory, or update_memory as alternatives. The batch framing is the only signal for choosing this over the single-document tool, which is weak routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_memoryA
Save a new document or scenario outcome into the vector DB.
metadata may be a JSON string (e.g. '{"source": "doc1", "tags": ["a"]}')
or a JSON object — both are accepted. On parse failure an empty dict is
stored instead and a warning is returned alongside the result. If
collection is given, the memory is stored there (created automatically
if missing; default: the server-configured collection).
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| metadata | No | {} | |
| collection | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does real work: it discloses that metadata parse failure stores an empty dict and returns a warning, and that a supplied collection is created automatically if missing. Less positive traits are undisclosed, notably whether saving duplicate text overwrites or appends, and no auth/permission context is given.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose leads, followed by the two non-obvious parameter behaviors, with no filler sentences. The parenthetical JSON example is slightly heavy but earns its place by resolving the string-vs-object ambiguity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description still notes a warning is returned alongside the result. For a 3-parameter mutation tool with no annotations, this is nearly complete; the main omission is duplicate/overwrite behavior and permission requirements.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate and largely does: it explains that metadata accepts either a JSON string or a JSON object, documents the failure fallback, and explains collection semantics including the default. Only 'text', the sole required parameter, is left unelaborated, which is self-evident.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with scope: 'Save a new document or scenario outcome into the vector DB.' An agent knows this is a write-to-vector-store operation, but the description never distinguishes it from the near-identical sibling save_memories (plural) or from update_memory, which is the real risk here.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The word 'new' implies creation as opposed to update_memory, but this is inference rather than guidance. There is no explicit when-to-use, when-not-to-use, or named alternative despite three closely related siblings, so an agent has no stated basis for choosing between save_memory, save_memories, and update_memory.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_memoryA
Search the vector DB for past documents/scenarios semantically similar to a query.
filter is an optional payload filter, as a JSON string or object (e.g.
'{"tags": ["x"]}' — only items whose tags field contains "x").
List values use MatchAny, scalars use exact match, and multiple
conditions are AND-ed. On parse failure the search runs without a filter
and a warning is returned alongside the results.
If collection is given, the search runs there (default: the
server-configured collection).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes | ||
| filter | No | ||
| collection | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does so well: it discloses that a filter parse failure degrades to an unfiltered search with a warning returned alongside results, and that an omitted collection falls back to the server-configured one. It does not state the read-only nature explicitly or discuss result ordering/pagination, but for a search operation the disclosed failure-mode behavior is the valuable part.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose is front-loaded in one sentence, followed by tightly scoped paragraphs on filter and collection. The filter explanation is dense but every clause adds operative detail; nothing reads as filler, though the parenthetical example is slightly verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the description covers the two parameters whose semantics are non-obvious plus the error path. The only meaningful gap is the undocumented 'limit' behavior, which an agent would have to guess from the default of 3.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does an excellent job on 'filter' (JSON string or object, MatchAny for lists, exact match for scalars, AND-ed conditions) and covers the 'collection' default, but 'limit' is never explained and 'query' is left to the obvious. Roughly half the parameters get added meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Search the vector DB for past documents/scenarios semantically similar to a query.' That is unambiguous and clearly distinct from the save/delete/update siblings. It does not explicitly name an alternative for the list_collections style use case, so it stops short of full sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied through the filter and collection mechanics, which tells an agent when those parameters matter, but there is no explicit 'use this instead of X when Y' guidance and no exclusions. An agent must infer that this is the retrieval entry point rather than being told.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_memoryA
Update an existing memory (point) in place under the same ID.
If point_id does not exist, an error is returned (existence is checked
first). When text is given, it is re-embedded and overwrites the stored
text; when text is None the existing text and vector are kept. When
metadata is given (JSON string or object), the payload metadata is
replaced entirely; when left empty ("") the existing metadata is kept.
Passing both text=None and an empty metadata is an error (nothing to
update).
If collection is given, the update happens there.
| Name | Required | Description | Default |
|---|---|---|---|
| text | No | ||
| metadata | No | ||
| point_id | Yes | ||
| collection | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that a missing point_id errors (existence checked first), that text triggers re-embedding and overwrites stored text, that metadata is replaced entirely, that empty string keeps existing metadata, and that text=None plus empty metadata is an error. It omits auth/permission requirements and concurrency/atomicity behavior, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The key fact (in-place update under the same ID) is front-loaded, and the conditional semantics that follow are dense but each sentence carries load-bearing information. Slightly heavy for the topic but not padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be explained, and the description thoroughly covers mutation semantics, error conditions, and validation rules for a no-annotation write tool. The only missing context is permission/authorization and any side effects beyond the target point.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: point_id must exist (checked first), text=None preserves existing text/vector while a value re-embeds and overwrites, metadata accepts a JSON string or object and is replaced wholesale, empty string preserves it, and collection scopes the update. This is more meaning than the bare schema provides for all four parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence gives a specific verb and resource ('Update an existing memory (point) in place under the same ID'), which cleanly separates it from siblings like save_memory and delete_memory. An agent can tell this is an in-place mutation keyed by point_id without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the conditions under which different behaviors occur (text given vs None, metadata given vs empty), which is useful implied guidance. However, it never names alternatives or explicitly states when to prefer update_memory over save_memory/delete_memory, so contextual routing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.1.0- First observed
delete_memory - First observed
list_collections - First observed
save_memories - First observed
save_memory - First observed
search_memory - First observed
update_memory
TDQS
Scored across 6 tools
Each tool targets a distinct operation: save (single), save_memories (batch), search, update, delete, and list_collections. The single-vs-batch split between save_memory and save_memories is explicitly distinguished in the descriptions, so an agent can reliably choose.
All tools follow a consistent snake_case verb_noun pattern (save_memory, search_memory, update_memory, delete_memory, save_memories, list_collections). The only variation, list_collections, reflects a genuinely different resource (collections vs memories), not an inconsistent style.
Six tools is well-scoped for a vector-DB memory server, covering the full point lifecycle plus collection listing without redundancy. Nothing feels padded or missing at the count level.
Core memory lifecycle (create, batch create, search, update, delete) and collection listing are all present, making the surface largely complete. Minor gaps remain: no get_memory-by-ID retrieval and no collection deletion/creation management beyond implicit auto-creation.
Maintenance
Related MCP Connectors
Persistent memory for AI agents. Search, store, and recall across sessions.
Persistent memory and knowledge management for AI agents with semantic search and 50+ tools.
- mem0OAuthio.github.mem0ai
Persistent memory for AI agents: add, search, update, and delete long-term memories.
Memory system for AI agents with semantic search. Store and recall memories with ease.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceProvides intelligent memory management capabilities using Qdrant vector database for semantic search and storage. Supports global, learned, and agent-specific memory types with markdown processing and duplicate detection.-
- AlicenseNot gradedqualityDmaintenanceEnables storing and retrieving information using semantic search with Qdrant vector database. Acts as a memory layer for LLMs to persistently store and semantically search through information and metadata.Apache 2.0
- AlicenseAqualityCmaintenancePersistent semantic memory for AI agents. SQLite-backed, local-first, zero config. Semantic search via Ollama embeddings with keyword fallback. Tools: remember, recall, history, forget, stats.1737 npm1MIT
- FlicenseNot gradedqualityDmaintenanceProvides persistent AI agent memory using a local vector database for long-term semantic storage and short-term session scratchpads. It enables low-latency memory operations including search, storage, and bulk management without external cloud dependencies.-