local-corporate-kb
The server provides a read-only RAG interface for a corporate knowledge base, offering the following tools:
kb_search – Semantic or hash-based search with natural language queries and optional filters (domain, status, service, authority, source_type, document_type), plus top_k and min_score controls.
kb_get_document – Retrieve a complete normalized document by ID, typically after a search.
kb_list_documents – List document metadata with optional filtering and pagination.
kb_stats – Return index statistics including counts, model identity, timestamps, and resolved local directories. The index is built from Markdown, HTML, TXT, and similar formats, converted in-memory for efficient retrieval. Connections are possible locally via stdio or remotely via Streamable HTTP (with Bearer token security). The server is strictly read-only; it cannot modify documents or execute commands, and all processing is local (no external API calls).
Allows indexing and searching Confluence pages exported as HTML or Markdown, providing a local knowledge base for retrieval-augmented generation.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@local-corporate-kbКак рассчитывается дневной лимит?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Локальная корпоративная база знаний для Qwen Code
Это локальный MVP корпоративного RAG: документы индексируются Python-процессом, embeddings
сохраняются в проверяемый файловый кэш, а при поиске целиком находятся в RAM. Qwen Code остаётся
единственной генеративной моделью и получает найденные фрагменты через read-only MCP tools по
локальному stdio или удалённому Streamable HTTP. MCP-сервер не формулирует финальные ответы,
не исполняет shell-команды и не изменяет документы.
Проект рассчитан на Python 3.12 и standalone FastMCP==3.4.4. Запуск после установки не зависит
от uv: все runtime-скрипты вызывают Python из .venv напрямую. Для разработки доступна
воспроизводимая установка через uv.lock, а для корпоративных машин — отдельная установка через
обычный pip.
Отдельные пошаговые инструкции:
Архитектура
Confluence export
↓
knowledge/*.md, *.html, *.txt
↓
loader + normalizer + structural chunker
↓
local feature hashing (по умолчанию) или локальная embedding-модель
↓
NumPy matrix in RAM
↓
MCP stdio / Streamable HTTP
↓
Qwen Code CLIВ удалённом режиме тот же индекс один раз загружается в память серверного процесса, после чего к нему одновременно подключаются Qwen Code CLI с разных машин:
Qwen CLI ─┐
Qwen CLI ─┼─ HTTPS / Bearer token ─ MCP Streamable HTTP ─ in-memory index
Qwen CLI ─┘DocumentLoader безопасно обходит только KB_KNOWLEDGE_DIR, нормализует Markdown/TXT и переводит
экспортированный HTML в Markdown-подобный текст. StructuralChunker сохраняет путь заголовков,
списки, таблицы и code fences. KnowledgeService координирует кэш и работает только через
интерфейс KnowledgeStore; MCP-слой не знает о NumPy.
В RAM находятся документы, чанки, отображение chunk_id -> index и нормализованная NumPy-матрица
[chunk_count, embedding_dimension]. Cosine similarity считается как matrix @ query_vector.
Экономия контекста Qwen
Поиск не передаёт модели все найденные тексты. Сервер сначала находит до 12 кандидатов внутри индекса, затем отдаёт максимум 3 наиболее релевантные выдержки из разных документов: до 260 условных токенов на выдержку и до 1000 на один ответ инструмента. Выдержка выбирается по словам вопроса и сохраняет ссылку на источник. Это ограничивает расход контекста, даже если в базе десятки тысяч страниц.
Если выбранный результат требует деталей, Qwen вызывает kb_get_chunk с chunk_id; полный текст
не загружается автоматически. kb_get_document также возвращает ограниченный извлекаемый фрагмент.
Лимиты настраиваются через KB_SEARCH_* и KB_DOCUMENT_CONTEXT_TOKENS в .env.example.
Защищённый kb_run_context_benchmark сравнивает прежние top-5 полных чанков с текущими top-3
выдержками: Hit@K, оценку токенов, точные JSON bytes, процент сжатия и latency. Инструмент требует
отдельный KB_BENCHMARK_PASSWORD, который Qwen запрашивает перед каждым вызовом.
В HTTP-режиме по /admin доступна встроенная парольная панель: usage-счётчики, живая системная
нагрузка, память и uptime процесса, список документов, загрузка Markdown/HTML/TXT и фоновая
инкрементальная переиндексация. Декларативные MCP search-tools по-прежнему доступны через
защищённый API, но технический редактор schemas из основной панели убран.
На диске в .cache/kb/ находятся только:
manifest.json— версии схемы, идентичность модели, chunking config и knowledge hash;documents.json— нормализованные документы и metadata;chunks.json— чанки без отдельной копии embedding;embeddings.npy— матрица без pickle.
Это не Vector DB: нет отдельного сервиса хранения, индекса ANN или SQL. Удалённый HTTP — только read-only MCP-фасад; поиск выполняется полным cosine scan по NumPy-матрице в памяти, а диск используется для ускорения старта.
Related MCP server: rag-retriever-mcp
Первый запуск
Подключение сотрудника к удалённой базе
RAG, индекс и документы находятся только на сервере. Для старых версий Qwen сотруднику
передаётся один файл clients/corporate_kb_stdio_proxy.py.
Qwen запускает его локально через uv как stdio MCP, а Python-процесс ходит к удалённому
серверу через обычные HTTP GET-запросы. Node.js, npx, mcp-remote, Nginx, копия проекта,
.venv, документы и индекс на клиенте не нужны.
Готовый settings находится в
examples/qwen-uv-stdio-settings.example.json, а полная
инструкция для сотрудника — в README.client.md.
Новые версии Qwen также могут подключаться к /mcp напряму через Streamable HTTP; скрипт
install.sh оставлен как опциональный способ для таких клиентов.
Скопируйте из неё mcpServers в ~/.qwen/settings.json и замените четыре placeholder:
REPLACE_WITH_ABSOLUTE_UV_PATH— результатwhich uvилиwhere uv(в Windowsuv.exe);REPLACE_WITH_ABSOLUTE_PATH— каталог, в котором сотрудник сохранил единственный.py-файл;REPLACE_WITH_SERVER_IP_OR_DOMAIN— адрес удалённого сервера;REPLACE_WITH_SERVER_TOKEN— Bearer-токен.
Qwen запускает этот файл как локальный MCP по stdio командой uv run. Скрипт содержит inline
dependency на FastMCP, поэтому uv сам создаёт изолированное кэшированное окружение. Локальный MCP
обращается к удалённому RAG только через обычные авторизованные JSON GET endpoints /api/v1/*.
node, npx, mcp-remote, локальная копия документов и локальный индекс не нужны.
Установка серверной части
Убедитесь, что доступен Python 3.12. На корпоративной машине рекомендуется pip-вариант: он не
читает uv.lock, не запускает uv и не зависит от установленной в системе версии uv.
Hugging Face, PyTorch и sentence-transformers в базовую установку не входят:
./scripts/setup-pip.sh
source ./scripts/activate-venv.shДля разработки с точными версиями из lock-файла остаётся вариант:
./scripts/setup-venv.sh
source ./scripts/activate-venv.shОба варианта создают одинаковую .venv; runtime-команды используют только .venv/bin/python.
Полностью локальный режим по умолчанию использует hash provider и не требует модели или сети:
./scripts/dev.sh index-hash
./scripts/dev.sh search-hashHash provider строит локальные lexical vectors из слов и символьных триграмм. Он пригоден для полностью автономного поиска по совпадающей терминологии, но не понимает смысл и синонимы так же хорошо, как semantic embedding model.
Для качественного semantic search сначала положите заранее полученные и одобренные model files в локальный каталог. Этот проект не скачивает их. Например:
models/Qwen3-Embedding-0.6B/После этого активируйте окружение, укажите только локальный путь и постройте индекс:
./scripts/dev.sh install-pip-semantic
source ./scripts/activate-venv.sh
export KB_EMBEDDING_PROVIDER=sentence_transformers
export KB_EMBEDDING_MODEL="$KB_PROJECT_ROOT/models/Qwen3-Embedding-0.6B"
export KB_EMBEDDING_LOCAL_FILES_ONLY=true
./scripts/dev.sh index-semantic
./scripts/dev.sh search "Какой сервис владеет дневными лимитами?"local_files_only=true, HF_HUB_OFFLINE=1 и TRANSFORMERS_OFFLINE=1 запрещают обращения к
Hugging Face. Если model files отсутствуют, индексирование завершится понятной ошибкой без попытки
скачивания. По умолчанию выбирается CUDA, затем MPS, затем CPU.
scripts/start-mcp.sh по умолчанию запускает MCP с KB_EMBEDDING_PROVIDER=hash, поэтому обычное
подключение Qwen полностью offline. Для локальной semantic-модели явно передайте provider и путь в
environment Qwen-конфигурации. Все runtime wrappers вызывают Python из готовой .venv напрямую:
после установки они не обращаются к package registry и не меняют окружение.
CLI
./.venv/bin/python -m corporate_kb.cli index
./.venv/bin/python -m corporate_kb.cli index --force
./.venv/bin/python -m corporate_kb.cli search "Как рассчитывается дневной лимит?" --top-k 5
./.venv/bin/python -m corporate_kb.cli search "Как рассчитывается дневной лимит?" --service limits-service
./.venv/bin/python -m corporate_kb.cli documents
./.venv/bin/python -m corporate_kb.cli stats
./.venv/bin/python -m corporate_kb.cli eval --top-k 5У search, documents, stats и eval есть --json. В этом режиме stdout содержит только JSON,
а логи остаются в stderr.
Если кэша нет или он несовместим, обычный поиск при KB_AUTO_INDEX=false завершится практичным
сообщением Run: ./scripts/dev.sh index. Это предотвращает неожиданную сетевую активность во время
MCP discovery.
Подключение к Qwen Code
Если MCP-серверы хранятся в отдельном каталоге, установите туда автономную runtime-копию. Скрипт
создаёт подкаталог corporate-kb, копирует только необходимые файлы, создаёт собственный .venv,
ставит locked runtime dependencies без dev-пакетов, строит hash-индекс и печатает готовый server
entry для Qwen:
./scripts/install-mcp-server.sh /absolute/path/to/mcp-serversЕсли версия uv на целевой машине отличается или uv запрещён политиками, установите ту же
runtime-копию через pip:
./scripts/install-mcp-server.sh /absolute/path/to/mcp-servers --pipДля закрытого окружения можно сначала только скопировать файлы, затем настроить корпоративный Python package registry и завершить установку командами, которые напечатает скрипт:
./scripts/install-mcp-server.sh /absolute/path/to/mcp-servers --copy-onlyСкопируйте examples/qwen-settings.example.json в .qwen/settings.json проекта и замените все
/ABSOLUTE/PATH/... реальными абсолютными путями. Не рассчитывайте на раскрытие ${PROJECT_ROOT}
в JSON. В command указан абсолютный путь к .venv/bin/python, а в args — запуск модуля
corporate_kb.mcp.server. Поэтому Qwen не зависит от глобальных python, uv, PATH, shell
activation или wrapper-скрипта.
Минимальная форма server entry:
{
"command": "/absolute/path/to/repository/.venv/bin/python",
"args": ["-m", "corporate_kb.mcp.server"],
"cwd": "/absolute/path/to/repository",
"env": {
"PYTHONPATH": "/absolute/path/to/repository/src"
}
}Альтернатива через CLI (выполняйте из корня этого репозитория, подставив абсолютные пути):
qwen mcp add \
--scope project \
--timeout 120000 \
-e KB_KNOWLEDGE_DIR=/absolute/path/to/repository/knowledge \
-e KB_CACHE_DIR=/absolute/path/to/repository/.cache/kb \
-e KB_EMBEDDING_PROVIDER=hash \
-e KB_EMBEDDING_LOCAL_FILES_ONLY=true \
-e HF_HUB_OFFLINE=1 \
-e TRANSFORMERS_OFFLINE=1 \
-e PYTHONUNBUFFERED=1 \
-e PYTHONNOUSERSITE=1 \
-e PYTHONPATH=/absolute/path/to/repository/src \
-e KB_AUTO_INDEX=false \
local-corporate-kb \
/absolute/path/to/repository/.venv/bin/python \
-m corporate_kb.mcp.serverstdio — транспорт по умолчанию, поэтому --transport http здесь не нужен. Синтаксис команды
сверен с официальной документацией Qwen Code,
но в среде разработки этого репозитория qwen не был установлен, и команда локально не выполнялась.
JSON-конфигурация также задаёт cwd и trust: false; статического фильтра tools в ней нет.
Проверка подключения:
qwen
/mcpТестовый запрос:
Используй corporate knowledge MCP.
Найди, какой сервис владеет дневными лимитами,
объясни правило и обязательно укажи использованные источники.Встроенные tools сервера:
kb_search— поиск сtop_k,min_score, metadata filters и компактными выдержками;kb_get_chunk— лениво загружает один ограниченный фрагмент поchunk_id;kb_run_context_benchmark— защищённый паролем read-only замер качества и сжатия;kb_get_document— ограниченный извлекаемый фрагмент документа поdocument_id;kb_list_documents— metadata документов без embeddings;kb_stats— состояние индекса и абсолютные пути.
Не задавайте статический includeTools в Qwen settings, если используете управляемые tools из UI:
клиентский allowlist скроет новые схемы от LLM. Прямой HTTP-клиент обновляет /mcp discovery, а
однофайловый stdio proxy перечитывает каталог после перезапуска Qwen.
Ручной запуск stdio server:
KB_LOG_LEVEL=DEBUG ./.venv/bin/python -m corporate_kb.mcp.serverstdout зарезервирован для MCP-протокола; все application logs направляются в stderr.
Удалённый MCP по HTTP
Удалённый режим заранее загружает готовый индекс и только после этого открывает порт. Поэтому все подключённые Qwen CLI используют один прогретый процесс и не строят embeddings при каждом запросе. Endpoint реализует рекомендованный для удалённых MCP-серверов Streamable HTTP, а не устаревший SSE.
1. Подготовить сервер
Скопируйте репозиторий и документы на сервер, установите runtime и один раз постройте индекс:
cd /opt/corporate-kb
./scripts/setup-pip.sh --no-dev
./scripts/dev.sh index-hashСгенерируйте отдельный секрет длиной не менее 32 символов:
openssl rand -hex 32Для прямого запуска внутри доверенной сети или VPN задайте секрет и запустите listener на всех сетевых интерфейсах:
export KB_MCP_HTTP_BEARER_TOKEN='PASTE_GENERATED_TOKEN'
export KB_MCP_HTTP_HOST='0.0.0.0'
export KB_MCP_HTTP_PORT='8000'
export KB_AUTO_INDEX='false'
./scripts/start-mcp-http.shПубличный health check не раскрывает тексты документов:
curl http://10.0.0.5:8000/healthДля stdio-клиента сервер также предоставляет защищённый read-only JSON API. Например, проверка поиска использует обычный GET и тот же Bearer-токен:
curl -G 'http://10.0.0.5:8000/api/v1/search' \
-H 'Authorization: Bearer PASTE_GENERATED_TOKEN' \
--data-urlencode 'query=какой сервис владеет дневными лимитами' \
--data-urlencode 'top_k=3'Доступны /api/v1/search, /api/v1/document, /api/v1/chunk, защищённый
/api/v1/admin/context-benchmark, /api/v1/documents и /api/v1/stats. Они используют
тот же прогретый индекс, что и MCP tools, не строят embeddings на клиенте и не изменяют документы.
Сам /mcp требует заголовок Authorization: Bearer .... Ограничения по Host, Origin, домену или IP
нет: сервер принимает клиента с любого адреса, если передан правильный токен. KB_AUTO_INDEX=false
гарантирует, что удалённый процесс не начнёт неожиданную переиндексацию.
Проверяйте с клиентской машины не только /health, но и настоящий MCP initialize:
curl -i --max-time 15 \
'http://10.0.0.5:8000/mcp' \
-H 'Authorization: Bearer PASTE_GENERATED_TOKEN' \
-H 'Content-Type: application/json' \
-H 'Accept: application/json, text/event-stream' \
--data '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"curl-test","version":"1.0"}}}'Ожидается HTTP/1.1 200. 401 означает неверный токен, 404 — неверный путь, а 421 — что
запущена старая сборка с Host allowlist. Прямой FastMCP listener, запущенный через Python или uv,
использует обычный HTTP. https:// указывайте только при наличии TLS reverse proxy; иначе Qwen
обычно сообщает TypeError: fetch failed.
2. Подключить Qwen CLI
Перед раздачей впишите в корневой install.sh публичный HTTPS-адрес сервера и созданный токен:
default_mcp_url="https://kb.company.example/mcp"
default_mcp_token="THE_SERVER_TOKEN"Сотруднику передаётся только этот один файл. В любом каталоге он выполняет:
bash install.shСкрипт не скачивает репозиторий, документы или Python-зависимости и не создаёт каталог RAG. Он
только добавляет подключение corporate-kb в пользовательскую конфигурацию уже установленного
Qwen Code. После запуска сотрудник перезапускает qwen и проверяет соединение через /mcp.
Bearer-токен внутри готового скрипта является секретом: раздавайте файл через защищённый корпоративный канал. Для отзыва доступа замените токен на сервере и выпустите новый скрипт.
3. Доступ через интернет
Не передавайте Bearer-токен по открытому интернету через обычный HTTP. Оставьте backend на
127.0.0.1:8000, а наружу опубликуйте его как HTTPS через Nginx, Caddy, ingress или корпоративный
API gateway:
export KB_MCP_HTTP_BEARER_TOKEN='PASTE_GENERATED_TOKEN'
export KB_MCP_HTTP_HOST='127.0.0.1'
./scripts/start-mcp-http.shМинимальные существенные параметры location для Nginx:
location /mcp {
proxy_pass http://127.0.0.1:8000;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_buffering off;
proxy_read_timeout 3600s;
}После этого клиент подключается к https://kb.example.com/mcp. TLS-сертификат и сетевой доступ
настраиваются на reverse proxy; порт 8000 не должен быть открыт наружу.
Когда документы изменились, выполните ./scripts/dev.sh index-hash и перезапустите HTTP-процесс.
Уже работающий процесс намеренно продолжает обслуживать согласованную старую версию индекса до
рестарта.
Добавление Confluence-страницы
Экспортируйте страницу в HTML либо сохраните её как Markdown и положите внутрь knowledge/.
Поддерживаются .md, .markdown, .html, .htm, .txt. Скрытые каталоги, .git, .cache,
__pycache__, node_modules, бинарные и неподдерживаемые файлы игнорируются. После изменения
перестройте индекс; при обычном запуске несовпадение knowledge_hash также инвалидирует кэш.
Пример front matter:
---
document_type: service
service: limits-service
domain: payments
status: current
authority: confluence
authority_priority: 80
owner: limits-team
source_id: "confluence-12345"
source_url: "https://confluence.example.com/pages/12345"
last_reviewed: "2026-07-20"
custom_field: "неизвестные поля тоже сохраняются"
---
# Limits ServiceБез front matter заголовок берётся из первого H1 или имени файла, source_id — из относительного
пути, status=current, authority=local_file, authority_priority=50.
Кэш и конфигурация
Пересобрать кэш:
./scripts/dev.sh indexПолностью удалить его можно командой rm -rf .cache/kb, после чего снова выполнить kb index.
Запись каждого файла атомарна, а manifest.json заменяется последним. Повреждение JSON/NumPy,
несовпадение схемы, модели, dimension, query instruction, chunking config или knowledge hash приводит
к понятной invalidation, а не к неясной NumPy-ошибке.
Все параметры перечислены в .env.example. Основные:
KB_EMBEDDING_PROVIDER=sentence_transformers|hash;KB_EMBEDDING_MODEL=./models/Qwen3-Embedding-0.6B— локальный каталог model files;KB_EMBEDDING_LOCAL_FILES_ONLY=true— fail-closed запрет сетевой загрузки модели;KB_EMBEDDING_DEVICE=auto|cpu|mps|cuda;KB_EMBEDDING_DIMENSION=1024;KB_CHUNK_SIZE_TOKENS=700,KB_CHUNK_HARD_MAX_TOKENS=900,KB_CHUNK_OVERLAP_TOKENS=80;KB_AUTO_INDEX=false.KB_MCP_HTTP_HOST,KB_MCP_HTTP_PORT,KB_MCP_HTTP_PATH;KB_MCP_HTTP_BEARER_TOKEN— обязательный секрет для HTTP-режима;
Относительные пути разрешаются относительно текущего project working directory; kb stats
показывает итоговые абсолютные пути.
Проверки
./scripts/dev.sh lint
./scripts/dev.sh typecheck
./scripts/dev.sh test
./scripts/dev.sh checkОбычный shell-скрипт scripts/dev.sh также объединяет повседневные команды:
./scripts/dev.sh install
./scripts/dev.sh install-pip
./scripts/dev.sh install-semantic
./scripts/dev.sh install-pip-semantic
./scripts/dev.sh test
./scripts/dev.sh lint
./scripts/dev.sh typecheck
./scripts/dev.sh index-hash
./scripts/dev.sh search-hash
./scripts/dev.sh index
./scripts/dev.sh search
./scripts/dev.sh index-semantic
./scripts/dev.sh eval
./scripts/dev.sh serve
./scripts/dev.sh serve-httpТесты всегда инжектируют hash provider и не требуют интернета, Hugging Face, GPU, Qwen Code, Docker или внешней БД. Интеграционные тесты проверяют как in-memory MCP transport, так и HTTP handshake через ASGI без открытия сетевого порта.
Ограничения MVP и развитие
Полный brute-force cosine scan подходит для небольшой локальной базы, но не для миллионов чанков.
При изменении документов неизменившиеся chunks и их embeddings переиспользуются из предыдущего кэша; полный пересчёт нужен только для новых или изменившихся chunks.
Нет Confluence REST API, OAuth, фоновой синхронизации и HTML-адаптеров под каждый вариант экспорта.
Нет reranker, hybrid/BM25 retrieval и отдельной оценки authority при ранжировании.
Точный token counter реальной модели не используется для предварительного chunking: интерфейс
TokenCounterотделён, поэтому его можно подключить без связи chunker с SentenceTransformer.Статический Bearer-токен даёт всем клиентам одинаковые права; для персональных учётных записей, отзыва сессий и аудита нужен внешний gateway/IdP либо полноценный OAuth.
Для перехода на настоящую Vector DB нужно реализовать PostgresKnowledgeStore или
QdrantKnowledgeStore с тем же контрактом KnowledgeStore, выбрать реализацию при сборке
KnowledgeService и сохранить API сервиса/MCP без изменений. Следующим этапом стоит добавить
инкрементальный cache manifest, batch upsert, hybrid retrieval и production evaluation corpus.
Available Tools
4 toolskb_get_documentA
Return one complete normalized document after kb_search identifies its document_id.
| Name | Required | Description | Default |
|---|---|---|---|
| document_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It lacks details on permissions, rate limits, side effects, or what 'normalized' means. As a retrieval tool, it likely has no destructive side effects, but this is not confirmed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. Every part adds value, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with one parameter and an output schema, the description covers purpose and usage context. However, it lacks behavioral transparency and parameter details, leaving gaps in completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description mentions document_id comes from kb_search but does not explain its format, constraints, or allowed values. With 0% schema description coverage, more detail is needed to compensate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('return'), the resource ('one complete normalized document'), and the context ('after kb_search identifies its document_id'), effectively distinguishing it from sibling tools like kb_search and kb_list_documents.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implicitly tells when to use this tool (after kb_search) and implies it is not for searching or listing. However, it does not explicitly state when not to use or mention alternatives beyond the sibling names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kb_list_documentsB
List filtered document metadata without document bodies or embeddings.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| domain | No | ||
| status | No | ||
| service | No | ||
| document_type | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without annotations, the description carries the burden of behavioral disclosure. It clarifies that bodies and embeddings are not included, which is a key behavioral trait. However, it does not mention other aspects like pagination, ordering, or whether filtering is exact or partial. It is adequate but not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that immediately conveys the core purpose and scope. Every word contributes value, and there is no redundancy or filler. It is perfectly concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 5 parameters (0 required) and a tool with multiple filter options, the description is too brief. It does not explain how filters combine, the meaning of each field, or what the output schema contains. An agent cannot fully judge whether this tool meets a specific filtering need without additional inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description provides no information about the 5 parameters (limit, domain, status, service, document_type). The description must compensate due to low coverage but does not, leaving the agent to rely solely on parameter names and types, which may be insufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it lists filtered document metadata and explicitly excludes bodies and embeddings. This distinguishes it from siblings like kb_search (full text) and kb_get_document (full document). The verb 'list' combined with the resource 'document metadata' is specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives such as kb_search or kb_get_document. The description implies use for metadata, but does not mention exclusions or specific scenarios. An agent would have to infer usage from the purpose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kb_searchA
Search corporate knowledge before architectural analysis or changes spanning multiple services. Use it for business rules, ADRs, APIs, events, and runbooks. Cite source_path or source_url in the final answer. A retrieved fragment is evidence, not the only source of truth; call kb_get_document when the complete document is needed.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| top_k | No | ||
| domain | No | ||
| status | No | current | |
| service | No | ||
| authority | No | ||
| min_score | No | ||
| source_type | No | ||
| document_type | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. It discloses that results are fragments, not complete documents, and mentions citing sources. However, it does not specify behavior like how results are ranked, pagination, or any side effects. Still, for a search tool, this is reasonably transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each serving a distinct purpose: what it does, how to use it, and when to use an alternative. No wasted words, well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (9 parameters, no schema descriptions, no annotations), the description covers high-level purpose and guidance but fails to document parameters. Output schema exists but does not compensate for missing parameter semantics. Completeness is moderate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has 9 parameters with 0% description coverage. Description only implicitly mentions 'query' as the search term and does not explain any other parameters (top_k, domain, status, service, authority, min_score, source_type, document_type). This leaves the agent guessing about their meaning and usage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states it searches corporate knowledge, lists concrete use cases (architectural analysis, changes across services), and mentions specific content types (business rules, ADRs, APIs, events, runbooks). It also distinguishes from sibling tool kb_get_document for complete documents.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear instructions: use before architectural analysis, cite source_path or source_url, and notes that a fragment is evidence, not the only source of truth. Explicitly suggests using kb_get_document when the complete document is needed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kb_statsB
Return index counts, identity, timestamps, and resolved local directories.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so the description must cover behavioral traits. It discloses the output elements but does not mention performance, required permissions, or side effects. As a stat retrieval tool, it is likely safe, but lacks explicit reassurance.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, front-loaded with key outputs. Could be improved by structuring or clarifying terms like 'identity', but overall efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero parameters and an output schema, the description is nearly complete for its simplicity. However, it does not address how this tool relates to sibling tools, leaving some contextual gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (no parameters). Baseline score of 3 applies as the description adds no parameter info, but none is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it returns index counts, identity, timestamps, and resolved local directories, indicating a read-only stats tool. However, it does not differentiate from siblings like kb_search, kb_get_document, kb_list_documents, and 'identity' is vague.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. It is implied for getting overall knowledge base stats, but no explicit when-to-use or when-not-to-use criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
kb_get_document - First observed
kb_list_documents - First observed
kb_search - First observed
kb_stats
TDQS
Scored across 4 tools
Each tool has a clearly distinct purpose: kb_search for searching, kb_get_document for retrieving full documents, kb_list_documents for listing metadata, and kb_stats for index statistics. No overlap or ambiguity.
All tools follow a consistent 'kb_verb' pattern with snake_case (e.g., kb_search, kb_get_document). Naming is predictable and uniform.
With 4 tools, the server is well-scoped for a corporate knowledge base: search, retrieval, listing, and statistics. Each tool adds value without redundancy.
The tool surface covers the core operations for a read-only knowledge base: search, get full document, list metadata, and view stats. No obvious gaps given the intended use case.
Maintenance
Related MCP Connectors
Cloud or self-hosted knowledge for AI agents: hybrid search, reranking, GraphRAG, scoped MCP tools.
Personal knowledge base MCP server with semantic search, auto-categorization, metadata extraction
Ingest, manage, and retrieve documents for RAG-powered AI applications
DocBase MCP server for AI agents
Related MCP Servers
- AlicenseAqualityDmaintenanceLocal-first RAG indexing and semantic search MCP server. Enables document retrieval and context-aware queries using local embedding models.36 npmMIT
- FlicenseAqualityDmaintenanceA local-first document retrieval engine that mounts as an MCP tool for agents to index files, search for relevant passages, and let the agent's own LLM answer.4-
- FlicenseNot gradedqualityDmaintenanceA local RAG knowledge base MCP server that exposes semantic document search as tools using zvec for vector storage and Qwen3-Embedding for text embedding.-
- AlicenseAqualityBmaintenanceIndexes local Markdown/text files into a SQLite database with vector embeddings and provides MCP tools for semantic search without cloud dependencies.3AGPL 3.0