Skip to main content
Glama
cmacdonald0514

glean-company-docs

Чат-бот по документации Halcyon

Ответы на вопросы на основе локального корпуса документов, построенные на Indexing, Search и Chat API от Glean и представленные как единый MCP-инструмент.

Веб-интерфейса нет. Интерфейс чата — это MCP-клиент: Cursor, Claude Desktop или любой другой.

Как это работает

data/Halcyon Shared Drive/  ->  Indexing API  ->  Search API  ->  Chat API  ->  {answer, sources, diagnostics}

Search выполняется до Chat, и Chat никогда не извлекает данные: отрывки извлекаются явно и передаются в него. Если ничего не проходит порог релевантности (пересечение терминов с вопросом), ответ — честное «проиндексированный контент не найден», и Chat не вызывается.

Related MCP server: faq-rag

Настройка

Требуются Poetry и Python 3.12+.

poetry config virtualenvs.in-project true   # keeps the venv at ./.venv
poetry install
cp .env.example .env      # then fill in the tokens

.env находится в .gitignore. Никогда не коммитьте реальные токены.

Переменная

Используется

Примечания

GLEAN_INSTANCE

оба

SDK строит https://{instance}-be.glean.com

GLEAN_INDEXING_TOKEN

только индексация

никогда не загружается на пути запроса

GLEAN_CLIENT_TOKEN

search + chat

область Chat/Search, тип Global

GLEAN_DATASOURCE

оба

общая песочница, поэтому это разделяет ID документов

GLEAN_DOCS_ROOT

только индексация

корень корпуса (data/Halcyon Shared Drive)

GLEAN_ACT_AS

search + chat

email для выполнения; требуется для Global-токенов

Необязательные, со значениями по умолчанию: GLEAN_DOC_ID_PREFIX (halcyon), GLEAN_TOP_K (5), GLEAN_MAX_SNIPPET_SIZE (2000), GLEAN_MIN_TERM_OVERLAP (0.30), GLEAN_CHAT_TIMEOUT_MS (60000).

Два токена структурно разделены: Settings.for_indexing() читает GLEAN_INDEXING_TOKEN, Settings.for_query() читает GLEAN_CLIENT_TOKEN и никогда не касается переменной индексации.

Использование

poetry run python -m glean_chat_bot     # the MCP server, on stdio
poetry run glean-index --dry-run        # extract and report, send nothing
poetry run glean-index                  # extract and bulk-push
poetry run glean-index --process-now    # ask Glean to process immediately (1 per 3h)
poetry run pytest                       # contract tests: no network, no tokens
poetry run pytest -m live               # the eval set, against real Glean

Добавьте -v для отладочного журналирования.

Индексация асинхронна: glean-index возвращает управление, как только Glean принял документы, за несколько минут до того, как они станут доступны для поиска. Подтвердите полное покрытие в консоли администратора Glean, прежде чем полагаться на ответы.

Конфигурация MCP-клиента

{
  "mcpServers": {
    "glean-company-docs": {
      "command": "/absolute/path/to/glean-chat-bot/.venv/bin/python",
      "args": ["-m", "glean_chat_bot"],
      "env": {
        "GLEAN_INSTANCE": "support-lab",
        "GLEAN_CLIENT_TOKEN": "...",
        "GLEAN_ACT_AS": "you@example.com",
        "GLEAN_DATASOURCE": "interviewds3"
      }
    }
  }
}

Один инструмент, ask_company_docs(question, top_k=None, include_citations=True) -> dict, возвращающий {answer, sources, diagnostics}. diagnostics сообщает, что было искано и что вернулось, чтобы вызывающая модель могла отличить «ничего не совпало» от «моя формулировка не подошла» и повторить попытку соответственно.

Структура

glean_chat_bot/
  __main__.py      `python -m glean_chat_bot`, the MCP server: one tool over query.ask.ask()
  client.py        indexing and query client factories, the ActAs header
  extraction.py    one adapter per file type, path signals, walk
  indexing.py      the whole write path, behind the `glean-index` command
  models.py        Passage, Source, Answer, ExtractedDoc (pydantic)
  query/           search.py (search -> Passage, the relevance floor)
                   chat.py (chat -> answer + resolved citations)
                   ask.py (ask() — the single orchestration function)
  utils/           config.py (env loading, one Settings, two constructors)
                   logging.py (log format, timing wrapper on every Glean call)
data/              the corpus
docs/              extraction notes
tests/             test_contract.py (the invariants, offline)
                   eval_cases.py + test_eval_live.py (the eval set, `-m live`)

Poetry для зависимостей и упаковки, Ruff для линтера и форматирования. Запустите poetry run ruff check . и poetry run ruff format . перед коммитом.

Ещё не реализовано

Права групп и пользователей, фильтрация по отделам, аннотации свежести, манифест хешей контента, переписывание запросов и адаптивные повторные попытки, потоковая передача, память разговора, повторные попытки и экспоненциальная задержка, Docker, CI.

F
license - not found
Not graded
quality - not tested
C
maintenance

Maintenance

Maintainers
Response time
Release cycle
Releases (12mo)
Commit activity

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables answering natural-language questions from FAQ documents using vector search and LLM generation via an MCP tool.
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables hybrid document search (BM25 and dense) over a configurable corpus via MCP tools, returning passages and sources for AI agents to cite in answers.
    MIT

View all related MCP servers

Related MCP Connectors

  • Query any docs site via MCP. Submit a URL, ask questions, get cited answers.

  • Google AI Overview answers and cited sources via the Apify Google AI Overview API, hosted MCP.

  • Your company's brain for AI agents. Cited, permission-aware knowledge across every system.

View all MCP Connectors

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/cmacdonald0514/glean-chat-bot'

If you have feedback or need assistance with the MCP directory API, please join our Discord server