retrieval_mcp_server
Provides tools for managing IT access operations through Webex, including handling tickets and requests.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@retrieval_mcp_serverRetrieve the current employee handbook and summarize the remote work policy."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
NovaOps Company Brain
An AI assistant for internal company operations, with evidence, permissions, and human approval built into its workflows.
NovaOps helps answer onboarding questions, coordinate IT access requests, extract vendor records, and process renewals. It uses synthetic company data so the workflows can be inspected and replayed without a real employer's systems.
Final verified release: 145/145 tests · 33/33 required-plus-optional evaluation records · 4/4 real process-death recovery stages.
The local showcase presents the recorded release evidence as a five-step demo; use the evaluation commands below to reproduce the deterministic checks.
Start here
Try it: five-minute local demo, including a deterministic run with no model API credentials.
Understand it: architecture and implementation evidence.
Inspect the results: evaluation improvement record and submission evidence.
Next direction: Incident Investigator proposal — planned, not implemented.
Related MCP server: enterprise-agent-lab
What it does
Workflow | Example | Boundary |
Maya — onboarding | Answer a policy question with cited evidence | Caller permissions filter retrieval before ranking |
Webex — IT access | Track an access request through approval | Model text cannot authorize a write |
Vendor — extraction | Turn a supplied document into a validated CRM record | Missing facts remain missing |
Renewal — operations | Resume a renewal after an approval event | Persistent state and replay protection govern effects |
Stack: Python · AWS Bedrock · MCP · SQLite · Langfuse · FastAPI · Docker
Course foundation and implementation focus
This is my AI Engineering course capstone, built around the supplied NovaOps brief, synthetic dataset, and evaluation scenarios. The repository documents the implemented application, its controls, and the evidence used to validate it.
The implementation focuses on permission-aware retrieval, durable approval and renewal state, tool-boundary enforcement, binding evaluation checks, trace instrumentation, and adversarial retrieval testing. The improvement record explains specific defects found and corrected. Course requirements and synthetic scenarios are credited as the foundation; evaluation results are scoped project evidence, not production customer outcomes.
What is implemented
Claim | Repository evidence |
One entry point classifies and scopes Maya and Webex requests |
|
Actual classify → scope → execute → tools → answer work is traced |
|
MCP carries caller identity, scope, and required evidence |
|
Manager-only sources are filtered before ranking |
|
Manager access derives from reporting relationships, not a claimed group |
|
Operational state and pending approvals survive restart |
|
Model text cannot directly authorize a write |
|
Existing Webex ticket |
|
Missing employee-to-seat data is reported rather than inferred |
|
Binding golden facts, sources, permissions, and tool-use rules are scored |
|
All 27 required measured turns have a trace index |
|
Vendor extraction uses one forced-tool call and validates the supplied schema locally |
|
Renewal requires scoped approval and clearance, activates on the effective date, survives process death, and replays every event exactly once |
|
Adversarial permission, approval, and provider proposals fail closed |
|
A typed input guard runs before planning and can be tested independently of server authorization |
|
Retrieved content can be semantically screened or quarantined by exact ID/source/SHA-256 before the answer model |
|
A project-specific poison reaches a real OpenSearch-backed Bedrock answer, then is contained, deleted, and cleared from thread memory |
|
API, MCP, and worker ship as separate non-root containers, with a cost-gated AWS/EFS deployment path |
|
Architecture
flowchart LR
U[Caller request] --> A[Company Brain router]
A --> T[MCP / local tool gateway]
T --> R[Permission-filtered evidence]
T --> S[SQLite operational state]
A --> M[Grounded answer]
S --> H[Recorded human approval]
H --> W[Controlled write / renewal worker]
A --> E[Traces and evaluation]CompanyBrainAgent owns routing and the shared conversational result contract. Maya and Webex remain focused internal scopes. Vendor is a synchronous document-in/record-out function; Renewal begins from a schedule and resumes on persisted inbound events. All conversational retrieval and operational calls cross a ToolGateway, which can run in-process for deterministic tests or against the FastMCP server. SQLite is the source of durable operational, approval, renewal, outbox, and idempotency state; the default database is .state/novaops.sqlite3 and can be overridden with NOVAOPS_DB_PATH.
The dataset deliberately has no employee-to-Webex-seat relationship. Role entitlement is not proof of assignment, so inspect_software_seat_assignments returns that limitation explicitly.
Install and verify
Python 3.11 or newer is required.
git clone https://github.com/eddieiskl/novaops-company-brain.git
cd novaops-company-brain
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[dev]'
python -m pytest -qRun the committed deterministic evaluations:
python evals/run_maya_s2.py
python evals/run_maya_s9.py
python evals/run_webex_s8.py
python evals/run_submission.py
python evals/run_submission.py --include-optional
python evals/run_guardrail_attacks.py
python evals/run_lesson13_offline.py
python evals/run_lesson15.pyWith the synthetic AWS/OpenSearch environment configured, the Lesson 13 live evidence commands are:
python evals/run_lesson13_live.py
python evals/run_lesson13_capstone_poison.pyThe submission runner maps every required measured input to binding expectations from spec/GOLDEN-DATASETS.json or an explicit Webex check. It exits non-zero when a required fact/source, permission rule, tool-use rule, or generic safety invariant fails.
MCP and live model mode
Start the MCP server in one terminal:
source .venv/bin/activate
python retrieval_mcp_server.pySmoke-test it from another:
source .venv/bin/activate
NOVAOPS_TOOL_MODE=mcp python scripts/smoke_mcp.pyFor a traced Bedrock run, create a local .env or export credentials for AWS and Langfuse, then run:
NOVAOPS_TOOL_MODE=mcp NOVAOPS_ANSWER_MODE=bedrock NOVAOPS_GUARD_MODE=bedrock \
python evals/run_submission.py --include-optional --trace --update-submissionThis sends synthetic evaluation prompts and retrieved synthetic NovaOps evidence to the configured AWS Bedrock and Langfuse projects. Never commit .env; it is ignored.
Evaluation and observability
Each measured turn creates one Langfuse trace—27 for required scope and 33 with both optional workflows. Multi-turn trace inputs include prior conversation, declared evaluation criteria, and the evidence or verified operation state used by the answer. Large source documents stay in observation inputs rather than propagated metadata. The phase and tool observations wrap live execution, and tool observations record real arguments, results, completion state, and errors. Bedrock answer and structured-extraction calls appear as generation observations. Deterministic score comments explain every result rather than reporting an unexplained aggregate pass.
The bounded improvement record is in evals/IMPROVEMENT_REPORT.md. The final trace IDs and reviewed commit are recorded in SUBMISSION.md.
GitHub Actions runs pytest, the optional-inclusive 33-case promotion gate, the focused guardrail attacks, Compose validation, and a build of all three container images.
Completion status
Stage | Status |
Required Maya workflow | Complete |
Required Webex workflow | Complete |
Lesson 11 observability and evaluation | Complete; instructor membership remains an external submission step |
Lesson 12 eval-loop engineering | Complete for the required scope; before/after gate is documented |
Vendor workflow | Complete; three schema-valid 0/1/2-gap extractions |
Renewal workflow | Complete; three scoped, restart-safe, replay-safe outcomes and four process-death recovery checks |
Lesson 13 security homework | Complete; 16-case live semantic guard, independent enforcement proof, project-specific OpenSearch poisoning/recovery evidence, and application-owned retrieval provenance |
Lesson 14 packaging/deployment | Historical deployment stage verified on ECS/LiteLLM/EFS: 108 tests, 22 local checks, 16 cloud checks, task-replacement persistence and automatic stop; cleanup verified. See |
Lesson 15 system design and renewal hardening | Complete; PRD review, HLD, as-built record, scoped authority, Finance clearance, effective-date activation, and 145-test acceptance suite. See |
Security and data handling
Secrets and runtime databases are ignored by Git.
Regular employees never receive manager-only chunks.
Direct write tools are absent from Maya’s model-visible loadouts.
A recorded, assigned human approval is required before a gated write can be released.
Retrieved chunks must match the application-owned ID/source/SHA-256 corpus manifest; semantic inspection and exact quarantine provide additional containment.
Quarantined evidence is explicitly invalidated from in-process conversation memory after corpus cleanup.
Answers distinguish observed facts, actions taken, recommendations, and blockers.
See SUBMISSION.md for the deliverable index and spec/PROJECT-DESCRIPTION.md for the supplied project brief.
The Lesson 13 boundary map, evidence limitations, offline commands, and live-resumption checklist are in docs/lesson13-security-homework.md. The isolated showcase includes a Security view with live blocked, allowed, and human-review probes that report the tool path and durable-state delta.
Containers
IMAGE_TAG="$(git rev-parse HEAD)" docker compose build
IMAGE_TAG="$(git rev-parse HEAD)" docker compose up -d
curl http://127.0.0.1:18080/health/live
curl http://127.0.0.1:18080/health/readySee docs/deployment.md for entry points, durability, credentials, cost, cleanup, and cloud-deployment boundaries. docs/lesson14-cloud-evidence.md is the reviewer-facing live evidence record.
The Lesson 14 vendor extension added bounded provider retries, one schema repair, and an SQS consumer sharing the HTTP extraction function. That historical stage passed 118 tests and all three live 0/1/2 missing-field fixtures.
The final Lesson 15 as-built design hardens Renewal with scoped authority, explicit Finance and security clearance, source-conflict handling, future-agreement persistence, effective-date activation, and transactional notification intent. The current release passes 145 tests, all 33 required-plus-optional evaluation records, and four real process-death recovery checks. The original course evaluation predates the PRD v1.3 scope fields; the local supplemental fixtures are explicitly labeled under evals/fixtures/lesson15/.
This server cannot be deployed
Maintenance
Related MCP Connectors
Shared, permission-aware company context for AI agents, with provenance, approvals and audit.
Your company's brain for AI agents. Cited, permission-aware knowledge across every system.
- StackOneOAuthcom.stackone
Give AI agents 30,000+ safe, token-optimized actions across Workday, SAP, Oracle + hundreds more.
Operate Linux, macOS and Windows from your LLM. Every action runs through an auditable allowlist.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables AI agents to search HR policies, create IT support tickets, and send external emails with configurable security levels and human-in-the-loop validation.-
- FlicenseNot gradedqualityCmaintenanceEnables controlled AI-agent access to enterprise-shaped tools with a deny-by-default gated write path, human approval, dry-run execution, and append-only audit logging.1-
- FlicenseNot gradedqualityCmaintenanceEnables AI agents to securely call enterprise MCP tools with tenant-scoped RBAC, human approvals, audit logging, and multi-tool workflows across customer, order, document, and ticket data.-
- AlicenseBqualityCmaintenanceEnables AI agents to act on live business objects under enforceable per-call identity, per-tool grants, and mandatory human approval for irreversible actions, with connectors isolated from core logic.8MIT