Skip to main content
Glama
cmacdonald0514

glean-company-docs

Halcyon Docs Chatbot

Beantwortung von Fragen auf Basis eines lokalen Dokumentkorpus, aufgebaut auf Gleans Indexing-, Search- und Chat-APIs und bereitgestellt als einzelnes MCP-Tool.

Es gibt keine Weboberfläche. Die Chat-Oberfläche ist ein MCP-Client – Cursor, Claude Desktop oder ein anderer.

So funktioniert es

data/Halcyon Shared Drive/  ->  Indexing API  ->  Search API  ->  Chat API  ->  {answer, sources, diagnostics}

Search läuft vor Chat, und Chat ruft nie selbst ab: Passagen werden explizit abgerufen und übergeben. Wenn nichts die Relevanzschwelle überschreitet (Begriffsüberlappung mit der Frage), lautet die Antwort ehrlich „keine indizierten Inhalte gefunden“ und Chat wird nie aufgerufen.

Related MCP server: faq-rag

Einrichtung

Erfordert Poetry und Python 3.12+.

poetry config virtualenvs.in-project true   # keeps the venv at ./.venv
poetry install
cp .env.example .env      # then fill in the tokens

.env ist in gitignored. Committe niemals echte Tokens.

Variable

Verwendet von

Hinweise

GLEAN_INSTANCE

beide

SDK baut https://{instance}-be.glean.com

GLEAN_INDEXING_TOKEN

nur Indexing

wird auf dem Query-Pfad nie geladen

GLEAN_CLIENT_TOKEN

Search + Chat

Scope Chat/Search, Typ Global

GLEAN_DATASOURCE

beide

gemeinsame Sandbox, daher namespaced dies Doc-IDs

GLEAN_DOCS_ROOT

nur Indexing

Korpus-Root (data/Halcyon Shared Drive)

GLEAN_ACT_AS

Search + Chat

E-Mail, als die agiert wird; für Global-Tokens erforderlich

Optional, mit Standardwerten: GLEAN_DOC_ID_PREFIX (halcyon), GLEAN_TOP_K (5), GLEAN_MAX_SNIPPET_SIZE (2000), GLEAN_MIN_TERM_OVERLAP (0.30), GLEAN_CHAT_TIMEOUT_MS (60000).

Die beiden Tokens sind strukturell getrennt: Settings.for_indexing() liest GLEAN_INDEXING_TOKEN, Settings.for_query() liest GLEAN_CLIENT_TOKEN und fasst die Indexierungsvariable nie an.

Verwendung

poetry run python -m glean_chat_bot     # the MCP server, on stdio
poetry run glean-index --dry-run        # extract and report, send nothing
poetry run glean-index                  # extract and bulk-push
poetry run glean-index --process-now    # ask Glean to process immediately (1 per 3h)
poetry run pytest                       # contract tests: no network, no tokens
poetry run pytest -m live               # the eval set, against real Glean

Füge -v für Debug-Logging hinzu.

Indexierung ist asynchron: glean-index kehrt zurück, sobald Glean die Dokumente akzeptiert hat, Minuten bevor sie durchsuchbar werden. Bestätige die vollständige Abdeckung in der Glean-Admin-Konsole, bevor du dich auf die Antworten verlässt.

MCP-Client-Konfiguration

{
  "mcpServers": {
    "glean-company-docs": {
      "command": "/absolute/path/to/glean-chat-bot/.venv/bin/python",
      "args": ["-m", "glean_chat_bot"],
      "env": {
        "GLEAN_INSTANCE": "support-lab",
        "GLEAN_CLIENT_TOKEN": "...",
        "GLEAN_ACT_AS": "you@example.com",
        "GLEAN_DATASOURCE": "interviewds3"
      }
    }
  }
}

Ein Tool, ask_company_docs(question, top_k=None, include_citations=True) -> dict, das {answer, sources, diagnostics} zurückgibt. diagnostics berichtet, was durchsucht wurde und was zurückkam, damit das aufrufende Modell „nichts gefunden“ von „meine Formulierung hat gefehlt“ unterscheiden und entsprechend erneut versuchen kann.

Aufbau

glean_chat_bot/
  __main__.py      `python -m glean_chat_bot`, the MCP server: one tool over query.ask.ask()
  client.py        indexing and query client factories, the ActAs header
  extraction.py    one adapter per file type, path signals, walk
  indexing.py      the whole write path, behind the `glean-index` command
  models.py        Passage, Source, Answer, ExtractedDoc (pydantic)
  query/           search.py (search -> Passage, the relevance floor)
                   chat.py (chat -> answer + resolved citations)
                   ask.py (ask() — the single orchestration function)
  utils/           config.py (env loading, one Settings, two constructors)
                   logging.py (log format, timing wrapper on every Glean call)
data/              the corpus
docs/              extraction notes
tests/             test_contract.py (the invariants, offline)
                   eval_cases.py + test_eval_live.py (the eval set, `-m live`)

Poetry für Abhängigkeiten und Paketierung, Ruff für Linting und Formatierung. Führe poetry run ruff check . und poetry run ruff format . vor dem Committen aus.

Noch nicht gebaut

Gruppen- und Benutzerberechtigungen, Abteilungsfilter, Frische-Anmerkungen, ein Inhalts-Hash-Manifest, Query-Umschreibung und adaptives Wiederholen, Streaming, Konversationsspeicher, Wiederholung und Backoff, Docker, CI.

F
license - not found
Not graded
quality - not tested
C
maintenance

Maintenance

Maintainers
Response time
Release cycle
Releases (12mo)
Commit activity

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables answering natural-language questions from FAQ documents using vector search and LLM generation via an MCP tool.
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables hybrid document search (BM25 and dense) over a configurable corpus via MCP tools, returning passages and sources for AI agents to cite in answers.
    MIT

View all related MCP servers

Related MCP Connectors

  • Query any docs site via MCP. Submit a URL, ask questions, get cited answers.

  • Google AI Overview answers and cited sources via the Apify Google AI Overview API, hosted MCP.

  • Your company's brain for AI agents. Cited, permission-aware knowledge across every system.

View all MCP Connectors

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/cmacdonald0514/glean-chat-bot'

If you have feedback or need assistance with the MCP directory API, please join our Discord server