Open WebUI Knowledge MCP
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Open WebUI Knowledge MCPsearch my knowledge base for Qdrant setup steps"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Open WebUI Knowledge MCP
Компактный MCP для Goose: управление знаниями и поиск в Qdrant через Open WebUI.
Goose / MCP client → этот MCP → Open WebUI → Qdrant
↓
настроенные embedding / rerankingMCP использует только Open WebUI HTTP API. Open WebUI отвечает за загрузку, chunking, embeddings, хранение оригиналов, векторов и поиск. MCP возвращает фрагменты, а не генерирует итоговый ответ. Прямых клиентов Qdrant/Ollama и ML-библиотек в зависимостях MCP нет.
Установка
Нужны Python 3.10+, uv и уже запущенный Open WebUI, настроенный на Qdrant. В каталоге репозитория:
uv venv --python 3.12
uv pip install --python .venv/bin/python -e '.[dev]'
cp .env.example .envУкажите в .env URL и API key Open WebUI. Запуск из терминала:
set -a
. ./.env
set +a
.venv/bin/openwebui-rag-mcpСервер ожидает MCP-сообщения на stdin; произвольного вывода в stdout нет.
Сам MCP не читает .env: пример выше экспортирует его значения в окружение.
Для обычной установки без инструментов разработки замените '.[dev]' на ..
Related MCP server: Open WebUI Knowledge Base MCP Server
Настройка Open WebUI → Qdrant
Эти переменные задаются процессу/контейнеру Open WebUI, а не MCP:
VECTOR_DB=qdrant
QDRANT_URI=http://qdrant:6333
QDRANT_API_KEY=
RAG_EMBEDDING_ENGINE=ollama
RAG_OLLAMA_BASE_URL=http://ollama:11434
RAG_EMBEDDING_MODEL=your-installed-embedding-modelqdrant и ollama в примере — имена сервисов в одной контейнерной сети; замените
адреса на доступные Open WebUI. Проверьте сохранённые настройки Documents в Admin
Panel: часть параметров Open WebUI хранится в его БД. Модель выбирает администратор.
Это настройка backend для баз, которыми управляет Open WebUI; существующие
произвольные коллекции Qdrant автоматически базами Open WebUI не становятся.
Описание переменных Open WebUI.
Hybrid search и reranker настраиваются в Open WebUI. MCP сохраняет выбранный режим
и передаёт k/k_reranker для конкретного поиска. Совместимость reranker с Ollama
зависит от возможностей Open WebUI и выбранного адаптера; MCP не эмулирует /rerank.
Для первого запуска можно использовать обычный vector search без reranker.
В Open WebUI разрешите API keys и создайте ключ пользователя с нужными правами
на Knowledge/Files/Retrieval API. Если включены ограничения endpoints, разрешите
используемые ниже пути. rag_health проверяет только авторизованный Knowledge API,
а не фактическую доступность Qdrant или моделей.
Окружение MCP
Переменная | По умолчанию | Значение |
|
| Базовый URL; поддерживается path prefix |
| обязательно | Bearer token Open WebUI |
|
| Timeout каждой HTTP-операции, секунды |
|
| Проверять TLS, строго |
|
| Предел пагинации; превышение возвращает ошибку |
|
| Максимум фрагментов, 1–100 |
|
| Лимит текста фрагмента в выдаче |
|
| Лимит извлечённого текста файла в выдаче |
|
| Лимит входного текста/query в символах |
Усечение всегда обозначается truncated; у файла есть total_chars.
RAG_SCORE_THRESHOLD старого прототипа удалён: MCP не может одинаково трактовать
оценки всех режимов retrieval. Threshold задаётся в Open WebUI.
Невалидная конфигурация завершает запуск с кодом 2 и сообщением в stderr.
Недоступный Open WebUI при старте логируется; MCP остаётся запущен и может
восстановиться при следующем вызове.
Goose
Добавьте stdio extension через интерфейс Goose или объедините этот блок со своим
~/.config/goose/config.yaml. Замените путь и ключ своими значениями:
extensions:
openwebui_knowledge:
name: openwebui_knowledge
type: stdio
enabled: true
cmd: /absolute/path/to/GooseMCPopenwebui/.venv/bin/openwebui-rag-mcp
args: []
timeout: 600
envs:
OPENWEBUI_URL: http://localhost:3000
OPENWEBUI_API_KEY: replace-with-your-key
OPENWEBUI_TIMEOUT: "120"
OPENWEBUI_VERIFY_TLS: "true"
RAG_TOP_K: "8"Timeout Goose учитывает, что запись состоит из нескольких HTTP-запросов. Используйте абсолютный путь: запуск уже установленного пакета не требует PyPI. Формат конфигурации Goose.
Tools
knowledge_id — ID базы Open WebUI; file_id — ID документа. Это разные сущности.
Tool | Назначение |
| Проверить доступ к Knowledge API |
| Все доступные базы, кратко и с пагинацией |
| Создать базу |
| Сведения о базе и краткий список её файлов |
| Загрузить UTF-8 |
| Изменить текст общего файла, подтвердить чтением, обновить индекс базы |
| Отсоединить файл от базы с |
| Извлечённый текст файла |
| Найти релевантные фрагменты |
Старые имена rag_list_knowledge, rag_search, rag_get_file сохранены как aliases
с теми же входными параметрами. Формат результатов обновлён; это не полная обратная
совместимость старого прототипа.
Пример последовательности arguments:
{"name":"Рабочие заметки","description":"Решения команды"}Из ответа knowledge_create возьмите knowledge_id:
{"knowledge_id":"<id-базы>","text":"Согласовали выпуск в пятницу.","title":"Решение","metadata":{"project":"demo"}}Из ответа knowledge_add возьмите file_id для чтения, обновления или удаления.
Поиск:
{"query":"Когда выпуск?","knowledge_ids":["<id-базы>"],"top_k":5}knowledge_ids=null ищет во всех доступных базах, [] — ни в одной. MCP отправляет
один retrieval-запрос для выбранного набора и сохраняет порядок Open WebUI.
Каждый результат содержит text, truncated, chunk_id, file_id, knowledge_id,
title, source, metadata, score, distance. Отсутствующие upstream поля — null.
Score/distance сохраняются без преобразований: поле distances в разных режимах
Open WebUI может содержать разные типы оценок. Отдельные vector/reranker scores
и достоверное число chunks API не гарантирует. Embeddings из ответов исключаются.
Контракт API и ограничения MVP
Контракт сверялся с официальным исходным кодом Open WebUI main 22.09.2026:
Knowledge API,
Files API,
Retrieval API.
23.09.2026 дополнительно проверены установленный Open WebUI 0.11.3 и
контракт Retrieval API этого тега.
Операция | HTTP API |
Базы |
|
База / файлы |
|
Загрузка |
|
Обработка |
|
Привязка / переиндексация / удаление связи |
|
Извлечённый текст |
|
Изменение текста |
|
Поиск по одной базе |
|
Поиск по нескольким базам |
|
Список баз поддерживает items/total, data/total и старый плоский массив.
Файлы читаются из вложенного files старых ответов либо из отдельного paginated API.
Retrieval поддерживает одну вложенную строку результатов и плоский массив.
Это совместимость форматов, а не обещание поддержки любого релиза. Для мутаций
нужны перечисленные endpoints и поддержка delete_file=false установленной версией.
Причина HTTP 400 на проверенном Open WebUI 0.11.3
Ошибка устранена на проверенном экземпляре: пользователь переключил Open WebUI на встроенную модель reranker и восстановил доступ API key. Итоговая live-проверка 23.09.2026 через MCP прошла без 400/403 и без предупреждений о пропущенных оценках:
Сценарий | top_k | Получено фрагментов | Результат |
«НПА России» | 8 | 8 | У всех есть числовая оценка в |
«Тестовая база знаний», alias | 1 | 1 | Контрольная фраза найдена |
Обе указанные базы, общий запрос | 3 | 3 | У всех есть числовая оценка в |
Все 3 доступные базы, | 3 | 3 | Контрольная фраза найдена |
Оценки возвращены в upstream-поле distances; отдельное поле score отсутствует
и сохраняется как null. Это не отсутствие оценок. Порядок выдачи сохраняется,
лимиты соблюдаются. Проверка подтверждает выполнение поиска, но не является
оценкой качества релевантности на эталонном наборе. Новые документы при этом
не загружались; использован ранее созданный проверочный файл.
Предшествующая диагностика установила следующую цепочку сбоя:
Open WebUI не мог соединиться с настроенным внешним reranker на
host.docker.internal:11435/v1/rerank. В логах есть ошибка установления соединения; на хосте порт 11435 не слушается.ExternalReranker.predict в 0.11.3 перехватывает исключение и возвращает
Noneвместо оценок.RerankCompressor в 0.11.3 при этом возвращает исходные документы. У части документов нет
score.merge_and_sort_query_resultsсравниваетNoneсfloatи вызывает TypeError; retrieval handler преобразует его в HTTP 400.
Запрос MCP соответствует QueryCollectionsForm: collection_names: list[str],
query: str, k: int, k_reranker: int. Необязательные hybrid, r,
hybrid_bm25_weight, enable_enriched_texts не передаются, поэтому настройки выбирает
Open WebUI. Контракт проверен по
официальному router 0.11.3;
общая документация API
не заменяет контракт конкретного релиза.
Первоначальный обход через /query/doc подтвердил получение фрагментов, но не
исправность reranking: отсутствие общего этапа сортировки позволяет вернуть 200
с отсутствующими оценками. MCP сохраняет этот штатный endpoint для одной базы,
но теперь добавляет warnings, если у возвращаемого фрагмента отсутствуют и score,
и distance. Это предупреждение о непроверяемом ранжировании, а не доказательство
ошибки reranker для любого ответа без метрик. Оценки и порядок не изменяются.
Несколько баз используют один глобальный запрос; при HTTP 400 возвращаются
isError=true, http_status, endpoint и безопасная подсказка проверить зависимости.
Если используется внешний reranker, первопричина устраняется администратором Open WebUI:
В Admin Panel → Settings → Documents проверить external reranker URL/model/key и доступность из контейнера Open WebUI.
host.docker.internalобозначает хост; сервис должен быть запущен и принимать соединения из Docker.Восстановить существующий reranker на нужном порту либо указать его действующий адрес. Не подставлять URL embeddings или chat API: нужен совместимый rerank endpoint.
Согласно официальному адаптеру, Open WebUI отправляет POST с
model,query,documents(массив строк),top_n. Ответ должен содержатьresultsсindexи числовымrelevance_scoreдля документов.Повторить
knowledge_searchс двумяknowledge_idsи сknowledge_ids=null. Успешное чтение файлов илиrag_healthне проверяет reranker.
Во время диагностики API key получал 401 на административный GET /api/v1/retrieval/config.
MCP не обходит это ограничение, не меняет глобальные настройки, не выключает
hybrid/reranking автоматически и не запускает модели. В этой установке пользователь
выбрал встроенный reranker; восстановление прежнего внешнего сервиса больше
не является условием работы поиска.
Изменение настроек reranker само по себе не требует переиндексации документов.
До переключения reranker live-проверка через MCP при top_k=8 вернула восемь фрагментов,
пять без обеих метрик; MCP добавил предупреждение. Запрос по двум базам вернул
явную ошибку 400 с новой диагностикой. Локально проходят 86 тестов, включая
восстановление поиска через MCP SDK после ошибки, и ruff check/format.
Фактически проверено: список 35 документов «НПА России», получение полного текста
одного файла (24 768 символов), два поисковых фрагмента через knowledge_search;
загрузка mcp-check-20260923.txt в «Тестовая база знаний», чтение 156 символов обратно
и нахождение контрольной фразы через alias rag_search. Проверочный файл сохранён
в тестовой базе, его file_id: cd8a48cf-03d3-478a-aaf0-a4397d5acdf1.
Live update/delete и запуск из Goose в этой проверке не выполнялись.
Записи неатомарны. Ошибка возвращается с MCP
isError=true,ok=false,stage, известнымfile_idиoutcome=partial_or_unknown. При timeout запрос мог завершиться: сначала проверьте Open WebUI, затем решайте, повторять ли операцию. Автоповторов нет.При сбое привязки загруженный файл сохраняется в Open WebUI; его можно проверить и прикрепить через UI. MCP не удаляет его автоматически.
Update меняет общий файл. Другие базы, использующие его, могут измениться; поведение обновления их индексов зависит от версии Open WebUI. Индекс указанной базы обновляется отдельным вызовом. Исходный скачиваемый файл и извлечённый текст могут отличаться после редактирования — MCP читает именно извлечённый текст.
Delete удаляет связь и поручает удаление векторов Open WebUI. Сам файл и другие базы сохраняются. Некоторые ошибки очистки/поиска Open WebUI скрывает внутри успешного HTTP-ответа; MCP не может независимо подтвердить состояние Qdrant.
Source/title/metadata передаются как metadata загрузки (Open WebUI сохраняет их в
file.meta.data); попадание произвольных полей в retrieval metadata зависит от backend. Изменение metadata, metadata-filter, удаление всей базы и внешние read-only Knowledge Sources не входят в CRUD MVP.Конкурирующие записи сериализуются внутри одного MCP-процесса. Транзакций между несколькими MCP или пользователями Open WebUI нет.
Разработка и проверки
.venv/bin/pytest -q
.venv/bin/ruff check src/openwebui_rag_mcp tests
.venv/bin/ruff format --check src/openwebui_rag_mcp testsВсе HTTP-вызовы тестов mock'аются через httpx. Отдельный тест запускает реальный stdio subprocess, выполняет MCP initialize/list_tools/call_tool и проверяет восстановление после ошибки tool. Qdrant/Ollama/Open WebUI для unit-тестов не нужны. План и оценка исходного кода — TODO.md, инструкции — AGENTS.md.
Available Tools
12 toolsknowledge_addC
Upload text, wait for processing, then attach to the given knowledge base.
A failed write may be partial: inspect stage and file_id before retrying.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| title | No | knowledge.txt | |
| source | No | ||
| metadata | No | ||
| knowledge_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses a key behavioral trait: a failed write may be partial and recommends inspecting stage and file_id before retrying. This goes beyond the annotations, which only provide basic hints (readOnlyHint, idempotentHint, destructiveHint), adding practical failure-handling context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the primary action and a useful warning. Efficient and to the point, with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is insufficient for an agent to correctly use the tool: it doesn't explain parameter semantics, doesn't specify the return structure (though no output schema exists), and lacks usage guidance relative to siblings. It only partially addresses failure handling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description provides no information about any of the five parameters, despite the schema having zero description coverage. An agent would not know the purpose of title, source, metadata, or even that knowledge_id identifies the target knowledge base. This is a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: upload text, wait for processing, then attach to a knowledge base. It distinguishes itself from sibling tools by focusing on adding content, though it doesn't explicitly name alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus alternatives like knowledge_create or knowledge_update. The description implies usage for adding text to a knowledge base but doesn't mention any exclusions or alternative conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
knowledge_createB
Create an Open WebUI knowledge base; returns its knowledge_id.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| description | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the read-only/idempotent/destructive profile, so the description only needs to add return behavior and side effects. It does add the key behavioral fact that a new knowledge base is created and its knowledge_id is returned, which is essential for the agent to use the result.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one concise sentence with no filler. It front-loads the action and resource, then specifies the return value, making every word useful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple and annotations cover the mutation profile, but the description omits usage guidance and parameter semantics, both of which are relevant for correct invocation. The return value is stated, which partially compensates for the lack of an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it never explains the 'name' or 'description' parameters beyond what the schema already shows. The only parameter information comes from field titles and a default value, which is insufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Create') and a specific resource ('Open WebUI knowledge base'), and adds the useful return value ('returns its knowledge_id'). This clearly differentiates it from sibling read/update/delete/search tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives, especially the similarly named 'knowledge_add'. There are no stated prerequisites, exclusions, or conditions that would route an agent to a different sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
knowledge_deleteBDestructive
Detach a file from this base. Retain the shared file; do not delete the whole base.
| Name | Required | Description | Default |
|---|---|---|---|
| file_id | Yes | ||
| knowledge_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already include destructiveHint=true, so the description's job is to add context. It does clarify that the operation preserves the shared file and doesn't delete the base, which is valuable context about scope. However, it doesn't explain consequences for references or permissions, nor does it detail idempotency, which is not covered by annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with zero waste. It front-loads the core action ('Detach a file from this base') and adds the retention constraint immediately. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (2 params, no nested objects, no output schema) and annotations covering destructive behavior, the description is mostly adequate. However, the complete lack of parameter explanation and lack of any mention of return value or side effects beyond base retention leaves some gaps for an agent unfamiliar with 'knowledge' concepts.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain what knowledge_id and file_id represent or how they relate. It only says 'file from this base' which hints at file_id being the file and knowledge_id the base, but that's minimal. With zero coverage, the description provides insufficient compensation for understanding the parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('Detach') and resource ('file from this base'), distinguishing it from destructive deletion. It differentiates from siblings by implying it's not a whole-base deletion, but it doesn't explicitly name a sibling like knowledge_get_file or knowledge_list for contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for detaching a file while retaining the shared file, and the exclusion of deleting the whole base is a when-not. However, it doesn't explicitly state when to use this tool versus alternatives like knowledge_update or rag_get_file, and provides no conditions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
knowledge_getCRead-only
Get a knowledge base and brief file metadata; file_id identifies each document.
| Name | Required | Description | Default |
|---|---|---|---|
| knowledge_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds that the result includes brief file metadata and mentions file_id as an identifier, but it does not clarify the output structure or whether pagination or other limits apply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with the core action front-loaded. However, the trailing clause about file_id is ambiguous and slightly detracts from clarity despite the brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 1-parameter read-only retrieval tool, the description is minimal but lacks an explanation of knowledge_id, the relationship between knowledge_id and file_id, or when this tool should be used over knowledge_get_file. The missing output schema makes this gap more significant.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain the required knowledge_id parameter beyond the schema title. It introduces file_id, which is not an input parameter, potentially confusing agents rather than helping them understand what knowledge_id refers to.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Get') and resource ('a knowledge base'), and adds 'brief file metadata' to suggest its scope. However, it does not explicitly differentiate from the sibling knowledge_get_file, and the phrase 'brief file metadata' is somewhat ambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus siblings like knowledge_list, knowledge_get_file, or rag_get_file. No alternatives or conditions are mentioned, leaving the agent to infer usage from the tool name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
knowledge_get_fileARead-only
Get extracted file text with total_chars and explicit truncated flag.
| Name | Required | Description | Default |
|---|---|---|---|
| file_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=true, destructiveHint=false), and the description adds genuinely useful behavior: it returns a character count and an explicit truncated flag, implying long content may be cut off and that the caller can detect this. Nothing here contradicts the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence with no filler; the verb comes first and every phrase ('extracted file text', 'total_chars', 'explicit truncated flag') adds meaning a caller benefits from. Nothing is redundant or missing in terms of economy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read-only tool with annotations covering safety, this is nearly complete: the return characteristics are summarized despite there being no output schema. However, it leaves unclear how a valid file_id is obtained and which of the many knowledge/rag siblings is the intended choice for a given retrieval need.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it says nothing about file_id—its format, provenance, or how to obtain it. The single parameter is self-descriptive by name and required, which limits the harm, but the description adds no parameter-level meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ('Get extracted file text') and names the return fields (total_chars, truncated flag), so an agent can understand what the tool does. It does not explicitly distinguish this from siblings like knowledge_get or rag_get_file, leaving some ambiguity about which getter to reach for.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus knowledge_get, rag_get_file, knowledge_search, or the other siblings. No context for when-not-to-use or which alternative fits a different intent is given, so an agent must rely on naming conventions alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
knowledge_listARead-only
List all accessible Open WebUI knowledge bases, without documents or vectors.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is known. The description adds meaningful behavioral context: results are filtered by 'accessible' (permission scoping) and exclude documents/vectors, clarifying the output's size and nature. It doesn't detail output format, but for a simple list that's acceptable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence of 11 words, front-loaded with the verb and resource, and every phrase earns its place by clarifying scope ('all accessible') and filtering ('without documents or vectors'). No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless, read-only listing tool, the description is complete. It states what it lists, the scope of the list, and what it excludes. With no output schema, the description still gives enough for an agent to call it correctly and set expectations without additional context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and schema coverage is 100%, so the baseline is 4. The description adds nothing about parameters, but none are needed. No further semantic value is required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('List'), a specific resource ('Open WebUI knowledge bases'), and explicitly scopes the result ('all accessible', 'without documents or vectors'). This distinguishes it from siblings that fetch full contents or specific items, such as knowledge_get or rag_get_file.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: it lists accessible bases only and excludes documents/vectors, which implies when to choose it over heavier retrieval tools. However, it does not explicitly name alternatives or state when not to use it, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
knowledge_searchARead-only
Retrieve fragments in Open WebUI rank order. No answer generation.
Omitted knowledge_ids searches all accessible bases; [] searches none. Scores/distances are raw upstream values, not normalized relevance.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| top_k | No | ||
| knowledge_ids | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds valuable behavioral details: 'No answer generation' clarifies the output is raw retrieval only, and 'raw upstream values, not normalized relevance' warns that scores are not calibrated. This goes beyond the annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact—two sentences—with the core purpose and ordering front-loaded. Each sentence conveys necessary information without padding. The key behavioral caveats ('No answer generation' and raw scores) are placed immediately after the purpose, making it easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 3 parameters, no output schema, and a non-trivial search tool, the description covers purpose, behavior, and one parameter nuance. It omits the meaning of top_k, default count limits, and the exact structure of returned fragments. While it mentions scores/distances, it does not fully specify the output shape. This is adequate but not comprehensive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the meaning. It precisely explains the semantics of knowledge_ids (omitted vs empty), which is non-obvious and valuable. However, it does not clarify the query parameter (obvious) or top_k (what it controls, default behavior). Since it covers only one of three parameters in depth, it partially compensates but leaves gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Retrieve' and the resource 'fragments', and specifies the ordering ('rank order'). It also differentiates itself from answer-generation tools via 'No answer generation', which helps distinguish it from siblings like rag_search. However, it does not explicitly name any sibling tool, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides concrete guidance for the knowledge_ids parameter (omitted vs empty array), which tells the agent how to restrict the search scope. It does not, however, offer explicit guidance on when to prefer this tool over alternatives like rag_search or knowledge_get. The differentiation is implied ('No answer generation') but not stated as a usage rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
knowledge_updateADestructive
Replace shared file text and reindex this base. Other bases sharing this file may change.
Writes are not atomic. Inspect stage/file_id on failure before retrying.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| file_id | Yes | ||
| knowledge_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (destructiveHint=true, idempotentHint=false), the description adds crucial behavioral context: writes are not atomic, other bases may change, reindexing occurs, and failed calls should prompt inspection of stage/file_id before retrying. This is exactly the kind of non-obvious behavior an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each earning its place: the core action, the side effect, and the retry/failure caveat. Critical information is front-loaded and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the most important operational risks for a destructive, non-idempotent mutation with no output schema: side effects, non-atomicity, and retry guidance. It falls slightly short on explaining the knowledge_id parameter and what exactly 'stage' refers to in the failure inspection step.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only implicitly ties 'text' to file text and mentions file_id in the failure guidance. The knowledge_id parameter is left entirely unexplained, and no parameter-level semantics are provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action: replace shared file text and reindex the base. It also discloses an important scoping fact ('Other bases sharing this file may change'), which clearly separates this update tool from sibling create/add/delete/search tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is implied by the purpose statement, and the warning about other bases sharing the file suggests caution, but the description does not explicitly state when to choose this tool over alternatives or when not to use it. No sibling tool is named as a fallback.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rag_get_fileDRead-only
Alias of knowledge_get_file.
| Name | Required | Description | Default |
|---|---|---|---|
| file_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotated readOnlyHint=true and destructiveHint=false, the description adds no behavioral context. It doesn't describe what retrieving a file entails, error scenarios, or return format. The alias reference is purely structural, not behavioral.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one short sentence, which is concise in size, but it under-specifies nearly everything. The sentence's only content is a cross-reference, not a meaningful description, so it fails to earn its place as a functional explanation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and a single parameter, the description must carry the full burden, but it leaves the tool's behavior, input semantics, and return value entirely to inference from a sibling. This is incomplete and unusable as a standalone definition.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description provides no information about file_id. The agent only sees 'File Id' as a string with no explanation of its source, format, or usage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Alias of knowledge_get_file' does not state the tool's own purpose; it only redirects to a sibling tool. An agent reading this in isolation must look up knowledge_get_file to learn the function, which is a tautological cross-reference rather than a clear verb+resource statement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus knowledge_get_file or any other sibling. The alias remark implies interchangeability but does not state conditions, exclusions, or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rag_healthARead-only
Check authenticated Open WebUI Knowledge API; does not verify Qdrant or models.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds the behavioral scope: it only checks the authenticated API and does not verify Qdrant or models. This is useful context beyond the annotations, though it doesn't describe the response format or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that front-loads the main action and immediately clarifies scope with the exclusion. Every word earns its place; no fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter health-check tool with readOnlyHint=true and destructiveHint=false, the description is nearly complete. It clearly states what is checked and what is not. The only minor gap is the lack of detail about the return value (e.g., status codes or response shape), but with no output schema and a simple health check, this is a minor omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters, so the schema is trivially complete. The description doesn't need to explain parameters, and the baseline for 0 params is 4. The description adds no parameter info, but none is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Check') and resource ('authenticated Open WebUI Knowledge API'), and distinguishes itself from siblings by explicitly noting what it does not verify ('does not verify Qdrant or models'). It is clear and concise, though it could be slightly more explicit about the health-check nature.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool: to check the health of the Knowledge API specifically, not Qdrant or models. It doesn't explicitly name alternatives, but the sibling list and the exclusion of Qdrant/models provide clear context for when this tool is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rag_list_knowledgeBRead-only
Alias of knowledge_list.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description only adds that the tool is an alias and provides no additional behavioral context such as return format, pagination, or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single four-word sentence with no filler or redundant information. It is appropriately sized for a simple alias tool and front-loads the only meaningful fact.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only alias, the description is minimally viable, and no output schema exists to explain return values. It would be more complete if it briefly stated what the alias returns or explicitly said 'use knowledge_list' as the canonical implementation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so the schema covers everything and the description has no parameter burden. This matches the baseline for parameter-free tools.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as an alias of knowledge_list, which distinguishes it from siblings and points to the intended operation. It does not explicitly restate the underlying verb/resource, but the alias relationship makes the purpose unambiguous if knowledge_list is understood.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this alias versus knowledge_list or the other sibling tools. The alias statement implies equivalence, but there is no explicit when-to-use, when-not-to-use, or alternative-selector information.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rag_searchDRead-only
Alias of knowledge_search.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| top_k | No | ||
| knowledge_ids | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, but the description adds no behavioral context beyond pointing to knowledge_search. It does not disclose return format, pagination, ordering, or any side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely short, but this is under-specification rather than effective conciseness. It spends only four words and omits substance that should be present in a standalone tool definition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter tool with no output schema and no schema descriptions, the description is far from complete. It relies entirely on knowledge_search being understood elsewhere and provides no fallback context for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description provides no parameter information. The query, top_k, and knowledge_ids parameters are left entirely undocumented by the description, so the agent gains no additional meaning beyond the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description only says 'Alias of knowledge_search', so it does not state any concrete verb or resource. It delegates meaning to a sibling tool's description rather than explaining what rag_search actually does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to use this tool or when to prefer alternatives. The alias relationship implies interchangeability with knowledge_search, but it does not explain contexts, exclusions, or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
12 tool updates
v0.2.0- First observed
knowledge_add - First observed
knowledge_create - First observed
knowledge_delete - First observed
knowledge_get - First observed
knowledge_get_file - First observed
knowledge_list - First observed
knowledge_search - First observed
knowledge_update - First observed
rag_get_file - First observed
rag_health - First observed
rag_list_knowledge - First observed
rag_search
TDQS
Scored across 12 tools
Several tools are exact aliases (rag_get_file/knowledge_get_file, rag_list_knowledge/knowledge_list, rag_search/knowledge_search), so an agent must arbitrarily choose between duplicate entry points. The remaining tools are mostly distinct, but the duplicated surface creates real selection ambiguity.
The knowledge_* tools follow a mostly consistent snake_case verb_noun pattern, but rag_* aliases introduce a second prefix convention. Names like knowledge_get and knowledge_get_file are also close enough to be mildly confusing.
Twelve tools is within a reasonable range, but three are exact aliases padding the count. The effective surface is about nine unique operations, which is well-scoped for this domain.
File operations are well covered with add, get, update, delete, and search, and knowledge bases support list, create, and get. However, knowledge bases lack update and delete operations, leaving an obvious lifecycle gap.
Maintenance
Related MCP Connectors
Knowledge base MCP for AI agents on iknow.dev. Search, read, and maintain via OAuth.
Cloud or self-hosted knowledge for AI agents: hybrid search, reranking, GraphRAG, scoped MCP tools.
DocBase MCP server for AI agents
Your memory, everywhere AI goes. Build knowledge once, access it via MCP anywhere.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceMCP server for Open WebUI Knowledge Bases – search and access your knowledge bases from Cursor, Claude Desktop, and other MCP clients.41 npm11MIT
- AlicenseNot gradedqualityDmaintenanceMCP server that exposes Open WebUI Knowledge Bases as tools and resources, enabling AI assistants to search and access knowledge bases.4MIT
- AlicenseAqualityCmaintenanceEnables indexing and semantic search of codebases and documents via MCP, using Ollama embeddings and Qdrant vector store.5Apache 2.0
- AlicenseNot gradedqualityDmaintenanceA knowledge base MCP server backed by Qdrant vector database with local embeddings for semantic search and document management.2 npm1ISC