Skip to main content
Glama

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
KIREO_API_KEYYesBearer token (ki_sk_…). Required.
KIREO_API_URLNoOverride for self-host.https://api.kireo.app
KIREO_LOG_LEVELNodebug / info / warn / error / silent.info
KIREO_PROXY_URLNoHTTP(S) proxy.
KIREO_TELEMETRYNoSet to 0 to disable device-id header.1
KIREO_RETRY_BASE_MSNoExponential backoff base.200
KIREO_ACCEPT_LANGUAGENoLocale for error hints.en
KIREO_REQUEST_TIMEOUT_MSNoPer-request timeout in ms, max 300000 (env alias: KIREO_TIMEOUT_MS).60000
KIREO_RETRY_MAX_ATTEMPTSNo5xx/429 retries (alias: KIREO_RETRY_MAX).3

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
memory_saveA

Persist a long-term memory for the current user.

When to use:

  • The user states a preference, decision, fact, plan, or goal that should outlive this chat.

  • You learn a stable property of the user, their project, or their tooling.

  • The user explicitly says "remember", "记住", "save this".

When NOT to use:

  • Ephemeral chat turns ("ok", "thanks").

  • Sensitive secrets (API keys, passwords) — refuse and warn the user.

  • Information you can re-derive from the codebase on demand.

Returns: { id, created_at, schema_version, embedding_status } — use memory_get for the full record.

memory_searchA

Hybrid semantic + keyword search over the user's memories.

When to use:

  • The user asks "do you remember…" or "what did I say about X".

  • Before answering project-specific questions, search for stored preferences and decisions.

  • You need to ground your answer in user-supplied facts.

When NOT to use:

  • You already have the memory id → use memory_get.

  • You just want the most recent items in a namespace → use memory_recall.

Returns: ranked list of MemoryRecord with a raw RRF relevance score (typically 0.01–0.04).

memory_recallA

Replay recent or important memories from a namespace, without a query.

When to use:

  • At the start of a new chat — to load recent context.

  • The user says "what have we been working on" / "remind me".

  • You want a chronological feed rather than a search match.

When NOT to use:

  • The user has a specific question → use memory_search.

  • You already have the id → use memory_get.

memory_getA

Fetch a single memory by id.

When to use:

  • You already have a memory id from a previous memory_search / memory_recall / memory_save call.

  • The user references "that memory you saved" and you stored the id.

When NOT to use:

  • You only have a vague query — use memory_search or memory_recall instead.

memory_updateA

Patch an existing memory. Only the fields you provide are changed.

When to use:

  • The user corrects a previously stored fact ("actually, I use Tailwind v4 not v3").

  • You need to add tags / entities to an existing memory.

  • The user lowers/raises priority of a memory.

When NOT to use:

  • The memory is wrong AND no longer relevant → use memory_delete.

  • You don't have the id yet → search first.

memory_deleteA

Delete a memory. Defaults to a soft delete (30-day recovery window).

When to use:

  • The user explicitly asks to forget something ("forget that I said …").

  • A memory is clearly wrong AND not worth correcting.

  • GDPR / privacy removal request.

When NOT to use:

  • The memory is just outdated — prefer memory_update.

  • You're unsure — ask the user first.

Note: deletes are always soft — the API has no immediate hard delete. Soft-deleted memories are purged permanently after the 30-day window.

memory_list_namespacesA

List all namespaces the current user has, with per-namespace counts.

When to use:

  • Before deciding which namespace to save to, when the user has multiple projects.

  • To answer "what projects do I have memories for?".

  • For diagnostics.

Returns: array of { name, created_at }.

memory_healthA

Probe the Kireo service and report local + remote health.

When to use:

  • User reports "memory tools aren't working".

  • You hit unexplained errors and need to confirm the API is reachable.

  • First-time setup verification.

Returns: { local: { server_version, node_version, platform }, remote: { status } }.

project_infoA

Resolve the stable identity of the current project and the namespaces its memories live in.

When to use: before every context_save and context_load, and any time the user asks which project or bucket their memories are going to.

ALWAYS call this before context_save or context_load, and ALWAYS show the user the returned key and source — a misidentified project is the one failure the user can catch instantly, and cannot catch at all if hidden.

If warn is non-null the identity is NOT stable across machines; relay it verbatim.

Returns: { key, source, display_name, warn, ctx_namespace, code_namespace, home_namespace }

context_saveA

Persist distilled session context for the current project.

When to use: ONCE per /kireo:save, after you have distilled the session. Always call project_info first and show the user the resolved project.

Each entry needs non-empty evidence — a concrete basis from THIS session (a file path, a command you ran, something the user said). If you cannot cite one, do not write the entry.

Do NOT store things re-derivable from the repo in 60 seconds — that is the code index’s job. Store decisions, constraints, gotchas, open threads, a small map of key files, and stated preferences.

Privacy: content, evidence, and files are pattern-redacted for known credential shapes, but that is NOT a privacy guarantee — it cannot catch business secrets ("customer A’s contract is worth $X") that got paraphrased into evidence/files from things the session read (.env, docker logs, a pasted SQL result). The first run is therefore FORCED into a preview by the tool itself: until a .first-run-acknowledged marker exists next to the outbox directory, every call returns the exact (post-redaction) text and writes/uploads nothing, no matter what dry_run says. Show that text to the user, and only after they confirm there is nothing business-sensitive call again with an explicit dry_run: false — that call performs the real save and records the acknowledgement. From then on dry_run is an ordinary optional preview flag (omitted = real save). A repo with a .kireo/disabled file, or the KIREO_DISABLED env var, disables this tool entirely — it returns immediately and does nothing.

Returns: { stored, deduped, failed, outbox_pending, namespace, dashboard_url }

context_loadA

Load previously saved context for the current project and render it.

When to use: at the start of work on a project — especially on a machine or in a tool where you have not worked on it before.

The first line reports the project identity and the code index freshness. Relay both to the user; never assume the code index is current.

Returns rendered text, grouped constraints → open threads → decisions → gotchas → key files → preferences.

A repo with a .kireo/disabled file, or the KIREO_DISABLED env var, disables this tool entirely — it returns immediately and does nothing: no outbox flush, no request, no CONTEXT.md.

context_archiveA

Save and upload a semantically compressed conversation as .md. When to use: after /kireo:compact has summarized the requested conversation rounds, and the user has asked to upload them. The host model does the compression; this tool redacts, writes the Markdown file, splits it losslessly and uploads it to the memory API. Always pass the conversation directory as cwd. Directory archives use their own namespace (different worktrees and same-named directories stay separate) and never change relay project keys. dry_run:true previews without side effects; otherwise invocation saves AND uploads. A failed or incomplete upload stays queued locally. Run compact again to retry. Returns the project, filename, local_path, namespace, snapshot_id, stored, deduped, outbox_pending, memory_ids and dashboard_url. Never call a queued archive uploaded.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

A4.4/5.0

Scored across 12 tools

Disambiguation5/5

Each tool targets a distinct operation: memory_get/search/recall are clearly separated by id, query, or chronological feed; memory_save/update/delete cover lifecycle edits; context_save/load/archive handle project session context. Even the potentially overlapping memory and context tools are differentiated by their descriptions and usage guidance.

Naming Consistency4/5

Memory tools consistently use memory_<verb> (memory_save, memory_search, memory_delete), and context tools use context_<verb>. The only deviation is project_info, which breaks the verb-first pattern, and memory_list_namespaces which mixes noun phrasing, but overall the convention is predictable and readable.

Tool Count5/5

12 tools is well-scoped for a memory and context persistence server. Each tool earns its place: CRUD for memories, search/recall/health/namespace utilities, and context save/load/archive. No redundant tools or excessive surface area.

Completeness5/5

The tool surface covers the full memory lifecycle: create, read, update, delete, search, recall, and namespace management. Context persistence is complete with save, load, and archive. Project identity resolution and health probing fill the operational gaps, leaving no obvious dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues