kbdb
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@kbdbsearch my notes for how authentication works"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.

@dikolab/kbdb
A file-based knowledge base with ranked keyword and semantic (hybrid) search -- learn your documents, then recall the relevant knowledge. No external server. Runs as a CLI and MCP server.
๐ Documentation ยท MCP Setup ยท CLI Reference
GitLab | NPM | JSR | License: AGPL-3.0
Runs on Node.js 20+ or Deno 2.6+. No database server, no cloud account -- just files on disk.
What is kbdb?
kbdb gives AI agents a persistent, searchable second brain. Point it at your Markdown docs and it indexes them into a file-based knowledge base -- then agents (and you) recall the most relevant knowledge by ranked keyword and semantic search, not exact-key lookup. It is a living store: agents learn new facts, update them, and recall them across sessions.
No external server to install, no cloud account -- just files on disk. It runs anywhere Node.js or Deno runs, and works as an MCP server, so agents like Claude can plug it in as a memory tool.
How search works: kbdb uses keyword search by default -- synonyms are expanded, terms are ranked by relevance, and headings carry 2ร weight in scoring. When an exact query finds nothing, kbdb automatically loosens the match so you still get the best available results.
Want smarter results? Use --algo hybrid to blend
keyword matching with similarity search -- finding
results even when different words describe the same
concept. The default TF-IDF embedding provider
works offline with zero setup. Swap it for a
third-party provider (local ONNX model or remote
API) in worker.toml when you need richer
embeddings.
Knowledge stays fresh: Re-learn a file and kbdb
replaces the old version automatically.
Near-duplicate detection warns you when you are
learning something you already have -- by embedding
similarity, so it catches the same fact reworded, not
just the same bytes. kbdb contradictions reports
sections that cover the same ground so you can read
them together. Integrity checks verify checksums,
orphans and references. Confidence scores help agents
tell strong matches from weak ones.
Related MCP server: Librarian
Getting Started
What You Need
One of these (pick whichever you already have):
Node.js version 20 or newer -- Download
Deno version 2.6 or newer -- Download (2.6 is the floor: the storage engine loads its WebAssembly through source-phase imports, which is what lets it run offline after one
deno install. Older Deno fails with a misleadingModule not foundnaming a.wasmfile that is present.)
That's it. No database server. No extra tools.
Install
Using Node.js:
CLI build hosted on NPM.
npm install -g @dikolab/kbdbUsing Deno:
CLI build hosted on JSR.
deno install -Agf jsr:@dikolab/kbdb/cliSee the CLI Installation Guide for prerequisites and verification steps.
Try It Out
1. Create a knowledge base
kbdb db init --db ./my-kbThis creates a .kbdb folder that holds all your
data.
2. Feed it your docs
kbdb learn ./docsPoint it at a folder of Markdown files. kbdb reads
them, breaks them into sections, and builds a
search index. Add --tags design,v2 to tag
sections for scoping, --replace to update
existing sections from the same source, or
--level 2 to set the hierarchical depth
(1 = broadest, 6 = narrowest). When learning a
directory, level is auto-detected from folder
depth.
3. Search
kbdb search "how does auth work"Results are ranked by relevance with snippets
showing where your terms matched. Output defaults
to --format rec (recfile: one field: value per
line) for easy grepping. Other formats: json
(machine-readable), text (numbered list), and
mcp (JSON-RPC 2.0 envelope). Use --offset to
page through large result sets.
To try hybrid search (keyword + AI similarity):
kbdb search "how does auth work" --algo hybridTip:
--dbis optional for the CLI. kbdb walks up from your working directory to the nearest.kbdbfolder, so commands just work anywhere inside a project. Point at a specific base with--db <dir>(the parent of.kbdb), or setKBDB_DB_DIR. Only themcpserver requires an explicit--db-- it never searches the working directory.
Search across bases: enrich results with
read-only knowledge from other databases using
--other-db <dir> (repeatable), or add --cascade
to also pull from .kbdb folders in parent
directories:
kbdb search "how does auth work" \
--other-db ~/shared-kb --cascadeEvery result carries a source_db field -- the
database root it came from -- which you can paste
straight back into --db or --other-db.
Scripting: Add
--format jsonto get structured JSON output for parsing. Use--non-interactiveor setKBDB_NON_INTERACTIVE=1to suppress prompts in CI pipelines.
4. Recall context
kbdb recall <kbid> --depth 1Start with a search result's kbid and expand context progressively: depth 0 gives the section content, depth 1 adds parent documents and back-references, depth 2 adds siblings and forward references, depth 3 includes full text of referenced sections.
Knowledge Base
Build, search, and maintain your knowledge store.
Import Markdown and plain text files with tags and source tracking
Smart updates -- re-learning a file supersedes the old version instead of duplicating it
History -- a superseded section is retired, not deleted:
kbdb historywalks the chain from either end, and an old kb-id still resolvesSearch with three algorithms: keyword (default), AI similarity, or hybrid (both)
Auto-fallback -- if your exact query finds nothing, kbdb loosens the match automatically
Recall sections with progressive context -- from a quick summary to full related content, or as deep as a
--max-tokensbudget allowsMeasure whether retrieval is actually any good --
kbdb evalscores Recall@k, MRR and nDCG@k against your own dataset, and exits non-zero when a change makes ranking worseNeighbourhood --
kbdb neighbourhoodsays what relates to a section and how: eight typed edges, seven of them recorded facts and one inferredConsolidate --
kbdb consolidateproposes groups of sections that could become one. It proposes only; you write the merge and apply it yourselfExport -- snapshot your knowledge base for backup
Verify database integrity and clean up stale data
Rebuild indexes if anything goes wrong
See the Knowledge Base Guide for the full walkthrough, including export and backup.
Agent Tooling
Integrate kbdb with AI agents and custom tools.
MCP quick-start (Claude CLI):
claude mcp add kbdb -- \
npx @dikolab/kbdb mcp --db /path/to/projectSee the MCP Installation Guide for Claude Code, VS Code, and Claude Desktop config files, plus troubleshooting.
MCP server with 30 tools -- search, recall, learn, revise, gaps, contradictions, export, skill/agent search, and more
Skills -- store reusable prompt templates with fill-in-the-blank arguments
Agents -- create AI agent profiles that combine a persona with skills
Capture policy -- the server tells the agent what to store during the MCP handshake itself, so it needs no per-host configuration. Two of its six clauses are about what not to store: chat summaries, guesses, secrets, and anything the code already says. kbdb delivers the policy; it cannot make an agent follow it
Auto-capture -- can ask the host's own model to pick out knowledge worth storing. It needs the MCP
samplingcapability, and Claude Code does not advertise it, so auto-capture is inert there. Every other feature in this list is unaffected -- see Host SupportDaemon resilience -- configurable request timeout and automatic retry with daemon respawn
Worker daemon lifecycle management -- stop and restart the background process
Granular Deno permissions -- the daemon runs with scoped permissions instead of
--allow-allPath confinement -- the daemon rejects path traversal (
..) in export/import
What the server tells an agent. The initialize
response carries an instructions string -- the one
channel every compliant MCP host receives without any
setup. kbdb spends it on capture policy: search before
answering, treat an unanswered verdict as a gap to
investigate rather than guess at, store decisions and
corrections that cost real effort to find, and do not
store what the code already says. The same sentences
are quoted in the learn, revise and search tool
descriptions rather than paraphrased, so there is one
source for all of them.
See the Agent Tooling Guide for MCP setup, skills, agents, and the library API, and Capture Policy for the six clauses in full and why they are written once.
For Developers
Library API
Use kbdb programmatically in your Node.js or Deno project:
import { createWorkerClient } from '@dikolab/kbdb';
// Spawns a background worker if not already running
const client = await createWorkerClient({
contextPath: '/path/to/.kbdb',
requestTimeoutMs: 30_000,
});
const results = await client.search({
query: 'authentication',
limit: 10,
offset: 0,
});
console.log(results.items);
client.disconnect();Pass contextPath (the .kbdb directory itself)
or dbPath (the parent directory -- kbdb discovers
.kbdb inside it).
See the Library API Reference for the full API.
Development Setup
git clone https://gitlab.com/diko316/knowledge-base-db.git
cd knowledge-base-db
npm install
npm testDocker
Two Dockerfiles, and they are not interchangeable.
Dockerfile at the repository root builds the MCP
server -- that is the one MCP directories build, and
the one to use if you want kbdb in a container. See
Install MCP Server
for the host configuration and why it needs a named
volume rather than a bind mount.
Dockerfile.tooling builds the development
toolchain (Node.js and Deno), which every make
target uses through docker-compose.yaml:
HOST_UMASK=$(umask) docker compose run --rm tool shRun make benchmark to measure search and rebuild
latency at scale -- results are written to
docs/benchmark/benchmark.md
automatically.
See the Makefile for all available build targets.
Contributing
Fork the repository
Create a feature branch
Make your changes and add tests
Run
npm testandnpm run lintOpen a merge request
Documentation
Second Brain with Claude Code -- the canonical setup guide: workspace layout, correct commands, MCP wiring
CLI Installation Guide -- prerequisites, npm/JSR install, verification
Knowledge Base Guide -- importing, searching, recall, export
Agent Tooling Guide -- MCP, skills, agents, library API
CLI Reference -- full command list with examples
MCP Installation Guide -- Claude CLI, Claude Code, VS Code, Claude Desktop
MCP Server Guide -- setup, tools, environment config, and what the server tells an agent at
initializeHost Support -- which MCP hosts deliver the capture policy and advertise sampling, measured rather than assumed
Capture Policy -- what kbdb tells an agent to store, and why it is stated once
Search and Ranking -- how search works under the hood
Storage Architecture -- file formats, directory layout, and what retained history costs
Deno Permissions -- the permission flags kbdb needs, and why
Benchmark Results -- search and rebuild latency at scale
Learn scaling -- how learn cost grows with corpus size
Contradiction signals -- calibrating the near-duplicate and contradiction thresholds over 1128 labelled pairs
The search engine
Storage, indexing and ranking come from
@dikolab/vdb,
kbdb's sibling project by the same author. Its
documentation covers the retrieval side in depth:
vdb Overview -- storage model, partitions, BM25F, vector and hybrid search
vdb Examples -- worked queries and ranking behaviour
Support
kbdb is free, AGPL-licensed software. If it earns a place in your workflow, you can support ongoing development via PayPal.
License
This project is dual-licensed:
Open source under the GNU Affero General Public License v3.0 (
AGPL-3.0-only)Commercial license available for closed-source or SaaS use
Versions <= 0.5.0 remain under the ISC license.
See LICENSING.md for details and contact information.
Maintenance
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables intelligent ingestion and querying of PDF, Markdown, and text files using hybrid search that combines keyword matching and semantic embeddings with citations.2
- AlicenseNot gradedqualityAmaintenanceProvides AI agents with persistent knowledge storage, enabling them to store, search, and retrieve text, documents, and files using semantic and keyword search via MCP tools.31Apache 2.0
- AlicenseNot gradedqualityDmaintenanceEnables AI assistants to perform semantic, hybrid, and filtered search on indexed local documentation with RAG capabilities.2MIT
- AlicenseAqualityAmaintenanceProvides persistent, searchable memory for AI agents, enabling them to retain, recall, and reflect on information across conversations.191MIT
Related MCP Connectors
Persistent memory and knowledge management for AI agents with semantic search and 50+ tools.
Persistent memory for AI agents. Search, store, and recall across sessions.
Universal memory for AI agents and tools. Save, organize and search context anywhere.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/diko316/knowledge-base-db'
If you have feedback or need assistance with the MCP directory API, please join our Discord server