Skip to main content
Glama

Official-Docs Ops Agent (Gemini + MCP + Vertex AI Search)

An infrastructure-operations agent that only answers from official upstream documentation and cites the exact source for every step. Built on Gemini, the Model Context Protocol (MCP), and Vertex AI Search.

▶️ Demo video (2 min, English narration, EN/ZH subtitles): https://youtu.be/l67ilza4aGk · download MP4 · subtitles: EN / 中文 / EN+中文

✍️ Author's blog / 作者博客: https://api-cloud.cc

The problem

AI assistants are now part of day-to-day ops work (Kubernetes, nftables, fail2ban, vLLM, systemd...). Their most dangerous failure is not "I don't know"; it's a confident answer built from an outdated blog post or a half-remembered flag. On production firewalls and clusters, one wrong line can lock you out of a server.

Related MCP server: insight-mcp

What this does

 Question ──> Gemini agent ──(MCP tools/call)──> search_official_docs
                 │                                     │
                 │                       ┌─────────────┼──────────────┐
                 │                     demo          vertex          ssh
                 │                 (bundled      (Vertex AI       (existing `ask`
                 │                  sample)       Search engine)   endpoint)
                 ▼
 Answer with [n] citations → official doc / GitHub source URLs
 or "Not found in the official docs I have access to."
  • MCP server (docs_agent.server): exposes one tool, search_official_docs(query, max_results). Any MCP client can use it: this Gemini agent, Gemini CLI, Claude Code, and others.

  • Gemini agent (docs_agent.agent): google-genai with the MCP session passed straight in as a tool. A strict system prompt makes it search first, answer only from the passages, cite every step, and refuse when nothing relevant comes back.

  • Pluggable backends:

    • demo: a small bundled corpus, so anyone can run it in 2 minutes with only a Gemini API key.

    • vertex: a Vertex AI Search engine over your own docs corpus. In production we index 30 upstream repos (8,530 documents).

    • ssh: calls an existing locked-down ask --json endpoint. This is how our production knowledge base is consumed today.

Quick start

git clone <this repo> && cd gemini-mcp-docs-agent
python -m venv .venv && . .venv/bin/activate
pip install -e .
export GEMINI_API_KEY=...          # free key from https://aistudio.google.com/apikey
export DOCS_BACKEND=demo
docs-agent "How do I drain a Kubernetes node safely?"
docs-agent "How do I configure Redis Sentinel?"   # out of scope -> refuses instead of guessing

Test only the MCP server (no Gemini key needed):

DOCS_BACKEND=demo python tests_smoke.py "nftables set with CIDR ranges"

Use it from any MCP client (example mcpServers entry):

{ "official-docs": { "command": "docs-mcp-server", "env": { "DOCS_BACKEND": "demo" } } }

Vertex AI Search backend:

gcloud auth application-default login
export DOCS_BACKEND=vertex VERTEX_PROJECT=your-project VERTEX_ENGINE=your-engine-id

Development environment

Layer

What we used

Workstation

Windows 11 laptop + WSL2 (Ubuntu, Python 3.12, uv)

Server

One small GCP Spot VM (2 vCPU / 6 GB, no GPU) in the Jakarta region, hardened with nftables + fail2ban + CrowdSec + PortSentry

Knowledge base

Vertex AI Search (Agent Search, Enterprise + LLM features) + Cloud Storage bucket, paid from GenAI App Builder credits

Corpus

30 official upstream repos: AI platforms (vLLM, LiteLLM, Ollama, Open WebUI, Dify, LangGraph, KServe, Ray, HF Transformers, HF TGI, MCP spec), containers (Kubernetes, K3s, Docker), network/security (nftables, fail2ban, CrowdSec, PortSentry, Hysteria2, sing-box, CoreDNS, Caddy, nginx, acme.sh, Cloudflare), systems (systemd, Prometheus), languages (Python, Bash, Git)

Ops

Daily incremental sync via cron, Telegram alert after ≥2 consecutive failures, private GitHub repo for server code, per-run IDs in all logs

This demo

mcp 2.x Python SDK (MCPServer), google-genai, gemini-2.5-flash

AI collaborators

This project was built by one human operator working with several AI agents, each with a clear role and boundary:

Who

Role

Human operator

Owns every decision: approved each production change, chose the architecture pivots, and set the "no AI acts without assignment" rule

Claude Code (WSL)

Main builder: server setup, sync pipeline, ask CLI, SSH gate, this MCP + Gemini demo, plus daily journal and rollback checkpoints

Codex

Independent read-only reviewer. Its acceptance review found 4 real issues (evidence-folder collisions under concurrency, partial-import handling, token-format validation, out-of-scope detection), all fixed

Gemini / Grok

Consumers: they query the knowledge base with their own restricted key instead of living on the server

The point of this setup: stability over capability. An agent that drifts, guesses, or changes defaults on its own is not usable on shared infrastructure, even if it's smart. Grounding answers in official docs is the same idea applied to knowledge.

How the design evolved and why

Full story in Chinese: docs/JOURNEY.zh-CN.md. Summary:

Initial plan

What we changed

Why

A shared "AI workbench" VM: 4 resident AI user accounts (Claude/Codex/Gemini/Grok), each with its own repo folder; Codex CLI installed on the server

No AI lives on the server. One restricted kb SSH account (restrict + forced-command gate) that can only run ask; the 4 AI users and the server-side Codex CLI were removed

Each resident AI needed its own Google identity, token refresh, file permissions and repo hooks. Every "it works" had to be re-proven per user, and one missed credential cost hours. The AIs already run, logged in, on the workstation; the server only has to answer questions

Google-managed remote MCP for Agent Search, called with each AI's credentials

A thin ask CLI on the server using Agent Search's own grounded answer generation; MCP is offered on the client side (this repo)

The remote MCP tools/call needed extra IAM roles per identity and failed with 403 errors that took a while to trace. Server-side generation is paid by Search credits and uses no AI subscription quota. Answers come back in 4–6 s

One test source (vLLM, 75 docs, random IDs)

30 sources, 8,530 docs, stable document IDs = hash(source:path) with daily incremental sync (unchanged run takes 7 s)

Random IDs create duplicates on every re-import. Stable IDs let upstream edits overwrite in place and upstream deletions get removed

Long-lived service-account key

Root cron impersonates the SA every 30 min to mint a 1-hour token

The org policy forbids SA key creation, and short-lived tokens are safer anyway

Parallel imports

Fetch/upload stay parallel; imports are serialized per data store

A data store accepts only one import at a time (HTTP 409)

Trust answered=true

Add a grounding score; "found something" ≠ "covers the product you asked about"

A PostgreSQL replication question was "answered" with "no direct guide". The refusal path needs a confidence signal

Local model / Docker

Neither

2 vCPU / 6 GB with no GPU can't serve a useful model; Docker's iptables rules conflict with the hand-built nftables baseline

Honest limitations

  • The demo corpus is a handful of short paraphrased passages with links to the real docs, enough to show the flow but not a knowledge base.

  • The model can still misread a correct passage. Citations make that checkable but don't prevent it.

  • The ssh backend returns one grounded summary plus citations, not raw passages.

License

Apache-2.0


More write-ups on AI-native infrastructure operations: https://api-cloud.cc

Available Tools

1 tool
search_official_docsB

Search official upstream documentation and return passages with source URLs.

Args: query: A focused technical question or keywords, in English. max_results: Number of passages to return (1-10).

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes
max_resultsNo

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does disclose the return shape (passages plus source URLs), but says nothing about read-only status, authentication needs, rate limits, coverage of upstream sources, or behavior when no results match.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core statement of what the tool does is front-loaded in the first sentence, followed by a compact Args block. Every line carries information; only minor polish is missing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read tool with no output schema, the description covers the essential semantics of both inputs and hints at the return format. It falls short on freshness, source coverage, and failure behavior, which an agent cannot infer elsewhere.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: it explains that query should be a focused technical question or keywords in English, and that max_results ranges from 1-10 with a default of 5. Both parameters gain meaning beyond the bare JSON schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (search) and resource (official upstream documentation) and says what comes back (passages with source URLs). There are no sibling tools to distinguish from, so the only missing element of a 5 is sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies a search is for technical questions in English, but never says when to use this tool versus other retrieval tools, nor any preconditions or exclusions. No when/when-not guidance is present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.1.0
    • First observedsearch_official_docs

TDQS

A3.5/5.0

Scored across 1 tool

Disambiguation5/5

With only one tool, there is no possibility of overlap or misselection. The single tool's purpose is unambiguous.

Naming Consistency5/5

The lone tool follows a clear snake_case verb_noun pattern (search_official_docs). No inconsistency exists because there is only one name.

Tool Count3/5

One tool is borderline thin for a documentation server. While a search-only interface is functional, the scope could reasonably support additional tools for fetching full documents or listing sources.

Completeness4/5

The server covers core documentation retrieval with source URLs, but lacks obvious operations like fetching a full document, listing available doc sets, or handling versions. These are minor gaps that agents can work around.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides semantic search over markdown documentation using RAG, allowing natural language queries and integration with MCP clients.
    1
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables hybrid document search (BM25 and dense) over a configurable corpus via MCP tools, returning passages and sources for AI agents to cite in answers.
    MIT
  • F
    license
    Not graded
    quality
    B
    maintenance
    Provides MCP tools to search engineering runbooks and historical incidents using semantic retrieval, supporting evidence-grounded incident investigation.
    -
  • A
    license
    A
    quality
    B
    maintenance
    Enables any MCP host to search and answer over your own documents with hybrid BM25+dense retrieval, cross-encoder reranking, and grounded, cited responses that refuse when no evidence is found.
    5
    MIT