knowledgebased
# knowledgebased
A reusable [Model Context Protocol](https://modelcontextprotocol.io) server that provides semantic search and a tag-based knowledge graph for any project. Auto-discovers a knowledge directory from cwd; silently disables when absent.
Written in TypeScript. Uses local sentence-transformer embeddings (`Xenova/multilingual-e5-small`) ā no API keys, no network calls after the first model download.
## Features
- š **Semantic search** ā embedding-based natural language queries (multilingual)
- š¤ **RAG search** ā tiered results with automatic LLM summarization via MCP sampling
- š·ļø **Tag search with graph traversal** ā follow `related:` links across fragments
- š **Markdown fragments with YAML frontmatter** ā human-readable, git-friendly
- š **Zero overhead when unused** ā exits silently if no knowledge is present
- š§ **Flexible auto-discovery** ā co-located, hidden, sibling, or user-global
## Quick Start
### Install
```bash
npm install -g knowledgebased
# or run on demand:
npx -y knowledgebased setup
```
`setup` registers the server in `~/.copilot/mcp-config.json` (or you can configure any MCP client manually). It will:
- **Auto-activate** in any project where knowledge is discovered
- **Stay disabled** (zero overhead) elsewhere
### Per-repo install (any MCP client)
Add to your `.mcp.json` / client config:
```json
{
"mcpServers": {
"knowledge": {
"type": "stdio",
"command": "npx",
"args": ["-y", "knowledgebased"]
}
}
}
```
## Knowledge Discovery
The server discovers knowledge from two independent phases, then **unions** all results.
Given `cwd = ~/workspace/my-project/`, here is every location the server checks:
```
~/
āāā .knowledgebased.json ā Phase 2: user-global config (always read)
āāā notes/ ā Phase 2: external KB (declared in bases)
ā āāā *.md
ā
āāā workspace/
āāā my-project.knowledge/ ā Phase 1 ā£: sibling folder
ā āāā *.md
ā
āāā my-project/ ā cwd
āāā .knowledge.json ā Phase 1 ā : config pointer (highest pri)
āāā knowledge/ ā Phase 1 ā”: co-located, visible
ā āāā *.md
āāā .knowledge/ ā Phase 1 ā¢: co-located, hidden
ā āāā *.md
āāā src/
```
### Phase 1 ā project source
Walks up from cwd. At **each** ancestor directory, tries four patterns in order ā **first match stops the entire walk**:
| Priority | Pattern | Within git root | Beyond git root |
|----------|---------|:-:|:-:|
| ā | `.knowledge.json` | ā
| ā
(explicit intent) |
| ā” | `knowledge/` | ā
| ā (too generic) |
| ⢠| `.knowledge/` | ā
| ā (too generic) |
| ⣠| `../<project>.knowledge/` | ā
| ā
(explicit naming) |
Beyond the git root, only explicitly-intentioned patterns (ā config pointer and ⣠sibling) are checked. If **no git root is found at all**, generic patterns are never used ā only ā and ⣠apply. This prevents accidental matches with unrelated `knowledge/` directories outside a project context.
Result: 0 or 1 **project source** (alias: `repo`, refs validated against cwd).
### Phase 2 ā external knowledge bases
Always runs (even if Phase 1 found a project source). Reads `~/.knowledgebased.json` and matches cwd against `repos` entries.
Result: 0āN **external sources** (alias: base ID, refs unscoped). Both phases are unioned and deduped by canonical directory hash.
### User-global config (`~/.knowledgebased.json`)
Defines named knowledge bases and binds them to repos:
```json
{
"bases": {
"personal": "~/notes",
"team": { "knowledge": "~/team/conventions", "cacheDir": "~/.cache/team" }
},
"repos": {
"*": ["personal"],
"~/workspace/my-project": ["team"]
}
}
```
| Field | Description |
|-------|-------------|
| `bases.<id>` | A string path (shorthand) or `{ "knowledge": "...", "cacheDir": "..." }`. Paths support `~` expansion. |
| `repos."*"` | Wildcard ā these bases are active in **every** project. |
| `repos.<path>` | Array of base IDs to activate when cwd is inside this path. **Longest-prefix match** wins (segment-boundary, case-insensitive on Windows). |
In the example above:
- `personal` is available everywhere (wildcard `"*"`)
- `team` is only available when working inside `~/workspace/my-project`
- Fragments from external sources are prefixed with their alias: `personal@notes/foo.md`
### Per-project config (`.knowledge.json`)
Points to a knowledge directory that lives elsewhere:
```json
{ "knowledge": "../shared-kb", "cacheDir": "./.cache/embeddings" }
```
| Field | Required | Description |
|-------|----------|-------------|
| `knowledge` | optional | Path to the knowledge directory. Resolved relative to the config file. Defaults to `./knowledge`. |
| `cacheDir` | optional | Override for the embedding cache. Defaults to `~/.cache/knowledgebased/<hash>`. |
### Validation rules
These conditions cause a **loud startup error**:
- `repos` references a non-existent base ID
- Base ID is `"*"`, or contains `@`, `/`, or spaces
- Two bases resolve to the same canonical directory
## Knowledge Fragments
Markdown files with YAML frontmatter:
```markdown
---
tags: [workflow, git]
related: [workflow/branch-naming]
source: session/2026-04-21
verified: false
refs: [src/utils.ts::parseArgs]
---
# Fragment Title
Content goes here...
```
## MCP Tools
| Tool | Description |
|------|-------------|
| `search_knowledge` | Tag-based search with graph traversal |
| `search_semantic` | Embedding-based semantic search with similarity scores |
| `search_rag` | Semantic search with automatic LLM summarization via MCP sampling |
| `list_tags` | List all tags with counts |
| `list_sources` | List loaded knowledge sources |
| `add_knowledge` | Create a new fragment |
| `update_knowledge` | Update an existing fragment |
| `delete_knowledge` | Delete a fragment permanently |
| `audit_knowledge` | Validate refs and related links |
| `reload_sources` | Re-discover sources from config |
### Which search tool to use?
```
User question
ā
āā "What topics does the KB cover?" ā search_semantic (explore)
ā Low threshold, scan fragment titles and scores.
ā
āā "How does X work?" ā search_rag (answer)
ā Returns concise summary + references.
ā If key details are missing, follow up with search_knowledge.
ā
āā "Give me everything about Y" ā search_knowledge (enumerate)
tags=["Y"], returns full unabridged content.
```
### search_rag ā RAG-style search
`search_rag` combines semantic search with MCP client [sampling](https://modelcontextprotocol.io/specification/2025-03-26/server/sampling) to deliver concise, query-aware results. Results are split into tiers:
| Tier | Score | Behavior |
|------|-------|----------|
| **direct** | ā„ `directThreshold` (0.85) | Full content returned verbatim |
| **related** | One-hop graph neighbors of direct hits | Summarized via LLM sampling |
| **summarized** | ā„ `threshold` (0.80), < `directThreshold` | Summarized via LLM sampling |
Every response includes a **references table** listing all used fragments with their similarity score, tier, and reason for inclusion.
When the MCP client doesn't support sampling, summarized/related fragments fall back to metadata-only output (title, tags, and a content preview).
**Parameters:**
| Parameter | Default | Description |
|-----------|---------|-------------|
| `query` | ā | Natural language search query |
| `threshold` | 0.80 | Minimum similarity score for inclusion |
| `directThreshold` | 0.85 | Score above which fragments are returned verbatim |
| `maxTokens` | 500 | Max tokens for the LLM summary |
## CLI Commands
```bash
knowledgebased setup # Register globally in ~/.copilot/mcp-config.json
knowledgebased init # Create knowledge/ in cwd
knowledgebased init --knowledge ../other/kb # Create .knowledge.json pointing elsewhere
```
## Development
```bash
npm install
npm run build # compile TS ā dist/
npm test # run unit tests via node:test + tsx
npm start # run from compiled output
npm run watch # incremental rebuild
```
## License
MIT
TDQS
Scored across 10 tools
The three search tools (search_knowledge, search_semantic, search_rag) have distinct purposes but search_semantic and search_rag both perform semantic search, differing only in summarization. The descriptions provide clear guidance, but the overlap between these two could cause misselection.
All tool names follow a consistent verb_noun pattern with lowercase and underscores (e.g., search_knowledge, add_knowledge, update_knowledge, list_tags). Verbs are clear and actions are predictable, making the pattern uniform throughout.
With 10 tools, the server is well-scoped for knowledge base management. Each tool serves a clear purpose covering CRUD operations, searching, tagging, auditing, and source management, with no redundant or excessive additions.
The tool set provides solid CRUD coverage (add, update, delete, search) plus list_tags, list_sources, audit, and reload. A minor gap is the lack of a direct 'get by ID' tool, but the search variants effectively cover retrieval needs.