rbac-rag-assistant
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@rbac-rag-assistantWhat's the university policy on academic probation?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
rbac-rag-assistant
Related: drivescore-cloud, a Go service on Kubernetes with Terraform and an incident log.
An internal knowledge assistant for a university: staff and students ask questions in plain language, and get answers grounded in policy documents they are cleared to read. Retrieval-augmented, tool-calling, role-aware, and exposed over MCP so any client can use the same tools under the same rules.
This is Layer 0 of a five-layer system. Layers above it add the delivery pipeline, reliability engineering, evaluation, security testing and observability, each on top of this application rather than beside it.
Run it
Python 3.13. Note that python on this machine resolves to the msys2 build,
which has none of these packages, so call the venv interpreter directly.
py -3.13 -m venv .venv
.\.venv\Scripts\python.exe -m pip install -r requirements.txt.\.venv\Scripts\python.exe test_rag.py
.\.venv\Scripts\python.exe test_tools.pyBoth print ok and need no API key: vectors.json is committed and covers the
corpus and every question in the golden set. Retrieval and the access rules stay
testable offline, which is the point of keeping them out of the model.
Ask it something (needs GEMINI_API_KEY):
.\.venv\Scripts\python.exe agent.py student "I have a 1.8 GPA, what happens?"
.\.venv\Scripts\python.exe agent.py staff "the wifi is down in the east wing, log it"Retrieval on its own, no key required:
.\.venv\Scripts\python.exe rag.py admin "adjunct pay per credit hour"Over HTTP. KB_TOKENS is mandatory, because a server that cannot identify a
caller would have to refuse every request; ADR-13 covers why the role is never
read from a header the caller writes:
$env:KB_TOKENS="devstudent:student,devstaff:staff,devadmin:admin"
.\.venv\Scripts\python.exe serve.pycurl -H "Authorization: Bearer devstudent" "localhost:8080/search?q=how+late+can+I+enroll"Route | Auth | Needs a model |
| no | no |
| bearer token | no |
| bearer token | yes, else |
Related MCP server: agentic-ops-builder
The pipeline
.github/workflows/ci.yml runs on every push and can refuse to publish. Three
things make it a gate rather than a report:
Gate | Fails the build when |
| a golden case that used to pass stops passing, or anything leaks |
| p95 on |
Trivy | a critical CVE with an available fix is in the image |
Both scanners caught something real on their first run. Trivy refused to
publish over three criticals in perl-base — a heap overflow compiling regular
expressions, a path traversal in Archive::Tar — all with a patched version
already released, inherited from a python:3.12-slim base rebuilt on somebody
else's schedule. semgrep found SHA1 in the embedding cache key, which was
changed to SHA256 rather than suppressed (ADR-17).
And the load test found a bug nothing else could see. The first k6 run
reported p95 40.94ms, p90 40.93ms, median 40.91ms, min 1.75ms. A distribution
with no spread is not a workload, it is a constant, and ~40ms is the Linux
delayed-ACK timer: http.server flushes headers in one write and the body in
another, so Nagle held the second segment waiting for an acknowledgement the
client would not send for 40ms because it was waiting for more data. Retrieval
itself takes under 2ms, so nearly all of the measured latency was two TCP timers
arguing with each other.
before | after | |
p95 | 40.94 ms | 9.01 ms |
median | 40.91 ms | 3.71 ms |
throughput | 242 req/s | 2,287 req/s |
The budget is set at 30ms rather than a round 250ms for that reason: it sits below the delayed-ACK constant, so reintroducing that bug fails the build instead of being absorbed by generous headroom.
Where it runs
A managed Kubernetes control plane costs about $73/month whether anyone visits or not. Both halves of what a cluster would prove are available for nothing, so the pipeline does both (ADR-19).
A real cluster, for two minutes. kind builds one inside the CI runner,
deploys deploy/k8s.yaml to it, and asserts three things by watching them
happen:
deleting pod/kb-5cd8dc8b86-xzjpx → back to 2/2 ready, service healthy
error: deployment "kb" exceeded its progress deadline
rollout refused, as it should be → service stayed up throughout
deployment.apps/kb rolled back → serve: okThe rollback step inverts the exit code of kubectl rollout status on purpose:
a rollback test that never observes a failed rollout proves nothing. It also
gives Layer 2 somewhere to run chaos experiments — "kill pods mid-request" is
not a sentence that means anything on a serverless host.
A URL, on Cloud Run. The container runs only while somebody is asking it
something and sleeps at zero otherwise, inside a permanent free allowance of two
million requests a month. A new revision deploys carrying no traffic, the
end-to-end suite runs against it on its own tagged URL, and only then does
traffic move — so a broken revision is never in front of a user and there is
nothing to roll back from. Setup is deploy/GOOGLE_CLOUD_SETUP.md; the job is
skipped entirely until GCP_PROJECT_ID exists, so this repository stays green
for anyone who clones it without a cloud account.
Authentication is Workload Identity Federation — GitHub signs a statement naming the repository and the run, Google is configured to accept exactly that, and no long-lived key is stored anywhere.
The quality gate compares cases, not percentages, because a percentage cannot see a swap: one case fixed, one broken, score unchanged (ADR-14). It was verified by breaking it on purpose — raising the retrieval floor from 0.61 to 0.68 made it exit 1, name fifteen regressed cases, and report four improvements that a percentage would have netted off.
Everything on that path runs without an API key, which is what makes it
affordable per push: the free tier allows 20 model requests per day per
model, so a live evaluation in CI would be exhausted before lunch. The live
run is .github/workflows/eval.yml, started by hand.
.\.venv\Scripts\python.exe eval\retrieval.py --gate # what CI runs
.\.venv\Scripts\python.exe eval\retrieval.py --accept # deliberately move the bar
.\.venv\Scripts\python.exe test_serve.py # end to end over HTTPHow it fits together
File | What it owns |
| Parsing, chunking, the embedding index, and the clearance filter |
| The three tools, their access rules, and argument validation |
| The tool-calling loop, and the only file that knows which model vendor is used |
| The same tools over MCP. No implementation of its own |
| The corpus. Front matter carries each document's |
Three decisions worth defending
The caller's role is not a tool parameter. It is bound out of band before
the turn starts. Had search_docs(query, role) existed, the model would hold
its own clearance, and a sentence buried in a retrieved document telling it to
"use role=admin" would be an escalation path. There is nothing to pass, so
there is nothing to talk it into.
The clearance filter runs on the candidate set, not the results. A chunk the caller may not see never enters the ranking, so its existence cannot be inferred from a gap in the results or from a score that moved.
Retrieval by meaning, and the number that justified it. This started as word matching on purpose, as the baseline the evaluation layer existed to beat. Swapping before the golden set existed would have been a guess. Swapping after produced a table: 83.3% to 91.7% overall, paraphrase matching 50% to 75%, and no category got worse. ADR-3.
MCP
mcp_server.py registers the functions from tools.py and adds nothing else,
which is the argument for MCP in one file: one implementation, one set of access
rules, many clients. Add to a client's config:
{
"mcpServers": {
"rbac-rag-assistant": {
"command": "C:\\Users\\Shrey\\Documents\\CLAUDE CODE\\rbac-rag-assistant\\.venv\\Scripts\\python.exe",
"args": ["C:\\Users\\Shrey\\Documents\\CLAUDE CODE\\rbac-rag-assistant\\mcp_server.py"],
"env": { "KB_ROLE": "staff" }
}
}
}An MCP client has no notion of a university role, so KB_ROLE is read once at
startup and applies to the whole connection. The role belongs to the server
process rather than to the conversation, which is what keeps the model from
choosing it.
Behaviour when access is denied
A student asking about staff-only material is told the material exists and that they are not cleared for it, rather than being told nothing exists (ADR-7). Verified live:
Guidance on what to do if you suspect your account has been compromised exists in the university policy documentation, but accessing it requires staff clearance. Please contact the office that owns this policy.
What crosses the boundary is a classification label and nothing else. No passage, no title, no filename. A question no document covers gets a different answer, which is the whole point of the distinction.
What surprised me
One. faculty compensation bands returned nothing for an administrator,
while professor salary returned the same document at 0.29. Those three words
appear only in the document's title, and parse() lifts the title out of the
body into metadata before anything is indexed. Users search with a document's
title far more often than its wording, so the most natural query for the most
sensitive document was the one guaranteed to fail. One line to fix: index
title + text, store text.
Worth recording because of where it was found. The access-control tests passed throughout, including the one asserting a student never sees confidential content, which was vacuously true when nobody could retrieve the document at all. A test that passes because nothing was returned looks identical to a test that passes because the rule works.
Two, found while implementing ADR-7. The "restricted material exists" notice fires on innocent questions. "Who is a full professor here" scores 0.28 against the compensation bands, which is higher than a legitimate staff match at 0.19. No threshold separates them, so this is not tunable; it is lexical matching being unable to tell a shared word from a shared meaning. Nothing leaks, and the answer is still wrong. First thing for Layer 3 to measure.
Three. The model fills small gaps the documents do not cover. Asked about a compromised account it suggested contacting "the IT Service Desk or IT Security team" when the tool had only said "the office that owns it". Plausible, harmless here, and not grounded in anything retrieved. That is exactly the behaviour a groundedness scorer exists to catch, and it appeared within the first five live questions.
Environment
Variable | Default | |
| Required by | |
|
| Pinned, not an alias. ADR-8 |
|
| Retrieval. Vectors cached in |
|
| MCP server only |
|
| Corpus directory |
|
| Ticket log |
| Required by | |
|
|
|
|
| Seconds between live evaluation questions, to stay under the rate limit |
This server cannot be deployed
Maintenance
Related MCP Connectors
- docs2mcpOAuthcom.docs2mcp
Query your own PDFs and documents from any MCP client. Every answer cites the page it came from.
Query any docs site via MCP. Submit a URL, ask questions, get cited answers.
Query your org's data in natural language — read-only MCP access to SQL, NoSQL, files & warehouses.
Read-only hosted MCP over CanonicAI's cited Answers corpus on canonicai.com.
Related MCP Servers
- FlicenseNot gradedqualityCmaintenanceEnables querying company knowledge base using RAG, providing accurate answers from internal documents via MCP.-
- FlicenseNot gradedqualityBmaintenanceEnables natural-language Q&A, human-approved actions, and dashboard generation over a data ontology via MCP.-
- FlicenseNot gradedqualityCmaintenanceProvides read-only, citation-backed semantic search and retrieval-augmented generation over enterprise documents via standardized MCP tools, with local embeddings for privacy.-
- AlicenseNot gradedqualityCmaintenanceEnables MCP clients to query a modular RAG knowledge hub with hybrid dense/sparse retrieval, reranking, multimodal image support, and traceable observability, all through natural language tools.MIT