gov-mcp
Provides tools for auditing AI workloads in a Google Cloud project, including discovering AI agents (even unregistered ones), running deterministic governance checks, and producing evidence-backed compliance reports.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@gov-mcpRun a governance audit on my GCP project"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Agent Governance Auditor
An ADK agent that audits other AI agents for compliance — finds the AI workloads running in a GCP project, including the ones nobody registered, runs 11 deterministic governance checks against a versioned policy pack, and produces an auditor-ready evidence package mapped to the EU AI Act (Art. 6, 9, 11, 12, 13, 15, 50) and SOC 2 Trust Services Criteria. Every finding carries an immutable evidence citation (GCS + SHA-256). Every write action is human-approved.
Each policy states why it maps to the articles it claims — and two mappings were removed on review for over-reaching, because over-mapping is what discredits a compliance report fastest. The tool audits for the absence of required declarations, never for inferred legal violations: "no documented risk classification" is checkable and is what an auditor writes up; "this agent is high-risk under Annex III" is a legal judgement it is not entitled to make.
Built for the Gemini Enterprise hackathon — Stream 2 (High-Code: ADK + custom MCP server + Agent Runtime on the Gemini Enterprise Agent Platform).
Why
The EU AI Act timeline moved, and that makes the problem more urgent rather than less. Regulation (EU) 2026/1744 (the "Digital Omnibus on AI", in force 27 July 2026) deferred Annex III high-risk obligations to 2 December 2027. Deferred, not cancelled — and meanwhile Article 50 transparency duties, Article 5 prohibited practices and the GPAI provider obligations are all live today, with penalties reaching €35M or 7% of global turnover.
So enterprises have roughly sixteen months to build an evidence trail for high-risk systems, while already carrying obligations on the agents they run right now. Both depend on answering one question auditors always ask — "what AI systems do you have, and can you prove they're governed?" — to which the honest answer is usually "we're not entirely sure."
This tool answers it, with evidence.
Related MCP server: EU AI Act Compliance MCP Server
How it works
The diagram above shows what exists. Core design rules (full architecture):
The LLM never decides compliance — code does. Pass/fail is computed in MCP tool code; the agent orchestrates, prioritizes and narrates. A mutation guard aborts any run where triage tries to flip a status.
Grounding by construction. A finding without an
evidence_refcannot exist in the data model — the schema rejects it.Read-only by default. The only write path (a remediation order) requires a human-minted, single-use approval token, enforced server-side so the server never trusts the agent's claim that a human approved.
The auditor audits itself — it deploys under its own Agent Identity and is discovered by its own sweep, passing 6 of its own 11 checks.
One audit, end to end
Two things in there are the design, not decoration. Triage calls no MCP tool at all — the one step doing real reasoning is denied any reach outside the process. And the run suspends before anything is written: no token exists yet, so the server would refuse a write even if the agent attempted one.
Discovery is behavioural, not name-matching
The question this product lives or dies on is "would you find an agent that isn't called agent-something?" Name matching answers "no" — it misses customer-insights-api and flags an nginx called agent-proxy. So a workload is classified on five signals, with the confidence and the reasons reported on every candidate:
Signal | What it observes | Strength |
| the service account appears in Cloud Audit Logs calling a model API | confirmed |
| deployed on Agent Runtime — an agent by construction | confirmed |
| carries | declared |
| environment references a model or agent framework | likely |
| the name looks agent-ish — kept, but demoted to weakest | possible |
The first one is the point: a workload that talks to a model cannot hide behind a boring name. The demo fleet includes customer-insights-api — a real ADK agent with no agent-ish name and no labels — precisely so this claim is testable rather than asserted.
What it covers, and what it doesn't
Discovery reaches further than auditing, and the report says which is which:
Audited in full | Cloud Run · Agent Runtime — all 11 checks read their configuration |
Detected, not yet audited | Cloud Functions · GKE · Compute — found, but their config is not readable the same way |
Reported as an open question | any identity running inference that matches no workload we found |
Genuinely blind | models called outside Google Cloud, a model running locally on a VM, cross-project calls, or audit logging switched off |
The last row is the honest one. A workload calling an external provider never
touches Google's logs, and a compliance auditor does not catch that — it checks
whether the controls that would have stopped it are switched on, and reports
when they are not. GOV-NET-007 is exactly that check.
The floor: using a model requires authenticating, and authenticating is logged. So the worst case is "here is one we cannot attribute — go look", never silence.
Quickstart
All development happens inside a container — Ubuntu 26.04 LTS with gcloud, Terraform, Node and a pinned Python 3.12. Nothing is installed on your machine, and the environment is identical on macOS, Windows (Docker Desktop or WSL2) and Linux.
Prerequisites: Docker running (OrbStack, Docker Desktop, or WSL2) and this repo cloned. Nothing else.
1. Build the image and get a shell
Open a terminal in the repo root — the folder holding this README and the gov_mcp/ auditor/ infra/ folders:
cd path/to/agent-governance-auditor # wherever you cloned it
# Build. First time ~3-5 min; afterwards it's instant (layer cache), so it's
# safe to just always run it.
docker build -t agv-dev docker/
# Start a shell inside the container.
docker run -it --rm \
-v "$PWD":/workspace \
-v agv-gcloud:/home/ubuntu/.config/gcloud \
-v agv-venv:/opt/venv \
-p 8080:8080 -p 8000:8000 -p 6274:6274 -p 6277:6277 \
agv-dev bashdocker run -it --rm `
-v "${PWD}:/workspace" `
-v agv-gcloud:/home/ubuntu/.config/gcloud `
-v agv-venv:/opt/venv `
-p 8080:8080 -p 8000:8000 -p 6274:6274 -p 6277:6277 `
agv-dev bashWhat those flags do:
Flag | Why |
| Live-mounts your repo folder. Edit files on your machine in any editor; the container sees changes instantly. Nothing is copied. |
| Keeps your gcloud login in a Docker volume, so you log in once — not every session — and nothing is written to your host. |
| Keeps installed Python packages between sessions (and off the slow bind mount). |
| Publishes ports so a browser on your Mac can reach services running in the container: |
| Deletes the container on exit. Safe — everything worth keeping is in the two volumes above. |
Your prompt becomes ubuntu@…:/workspace$. You're in.
Everything from here runs inside the container.
2. One-time setup and authentication
bash docker/post-create.sh # creates the python env, installs deps, runs the tests
# BOTH logins are required and they are NOT interchangeable:
# the first authenticates the gcloud CLI
# the second writes Application Default Credentials, which Terraform and
# every google-cloud-* python client read instead
gcloud auth login --no-launch-browser
gcloud auth application-default login --no-launch-browser
gcloud auth application-default print-access-token >/dev/null && echo "ADC OK"Each login prints a URL to open in your browser and asks you to paste a code back. On the second consent screen, tick every permission box ("Select all") — partial consent fails with a confusing Scope has changed crash, and you need the cloud-platform scope for anything to work. Don't move on until ADC OK prints.
→ if the consent screen errors: gotcha 0b
3. Create your sandbox GCP project
export PROJECT_ID="agent-gov-auditor-$(date +%y%m%d)" # must be globally unique
gcloud projects create "$PROJECT_ID" --name="agent-governance-auditor"
gcloud config set project "$PROJECT_ID"
gcloud billing accounts list # copy your account id
gcloud billing projects link "$PROJECT_ID" --billing-account=XXXXXX-XXXXXX-XXXXXX
# REQUIRED: attribute ADC API calls to your project. User credentials carry no
# project of their own, so without this Terraform gets a 403 SERVICE_DISABLED
# blaming Google's shared ADC client project (764086051850).
gcloud auth application-default set-quota-project "$PROJECT_ID"
gcloud config set run/region us-central1Billing must be linked before Terraform runs — enabling APIs requires it.
→ if you hit a 403 naming project 764086051850: gotcha 0c
4. Provision the infrastructure
Creates the enabled APIs, service accounts (including the deliberately over-permissioned rogue one), evidence bucket, budget, audit-log sink and Artifact Registry repo.
cd infra
cp terraform.tfvars.example terraform.tfvars
# edit terraform.tfvars: project_id, billing_account_id, region
terraform init
terraform plan
terraform applyIf apply fails on the billing budget, that's expected on some accounts — budgets need billing-account-level permission, not project-level. Create it once in the Console and terraform import, or comment the resource out. Don't lose an evening to it.
→ if the first apply fails with a wall of SERVICE_DISABLED: gotcha 0e — usually just re-run it
5. Deploy everything
One command builds and deploys the whole estate on top of the Terraform baseline, in dependency order, printing elapsed time per stage. Measured: 6m 57s for the full four-agent fleet.
./scripts/deploy-all.shThen open the Gemini Enterprise app it registered (Agents → 3-dot → Preview) and send:
Run a governance audit of this project.
The run stops at the approval gate. Reply APPROVE to authorise remediation, or APPROVE <finding id> for a subset, or DECLINE. (GE and the Agent Runtime Playground render no confirm button for ADK's experimental confirmation primitive, which is why the gate also accepts a typed reply.)
Prefer the terminal, or want to drive it without a browser:
python scripts/query_agent_runtime.py # multi-turn chat against the deployed agent→ if the deployed agent 401s or stalls at the gate: gotchas 0q and 0r
6. Tear down, rebuild, repeat
A soft teardown removes everything a deploy script created and keeps everything Terraform owns — so you can exercise the entire deploy path in minutes without a 20-minute project bootstrap. It is also the best rehearsal for the demo, because it is the same sequence.
./scripts/teardown-workloads.sh # prompts first; --yes to skip
terraform -chdir=infra plan # expect NO changes — proves the split is clean
./scripts/deploy-all.sh # back up in ~6 minutesRemoved | Kept |
GE app + agent registration | project, enabled APIs |
Agent Runtime deployment | service accounts and their IAM |
the stale IAM binding naming the deleted agent | evidence + staging buckets |
| budget, audit-log sink |
Artifact Registry and its images, so the rebuild is fast |
Evidence is deliberately not deleted: the bucket carries a 30-day retention policy and will refuse. That is the immutability the design claims — and watching a delete get rejected demonstrates it better than any assertion.
The deploy chain
Useful when something fails and you need to know which link to re-run:
scripts/deploy-all.sh
├─ 1. auditee fleet
│ auditees/deploy-{compliant,legacy,rogue,insights}.sh
│ └─ each sources auditees/common.sh → build_image()
│ └─ gcloud builds submit (Dockerfile + main.py + requirements.txt)
│ └─ Artifact Registry
│ then gcloud run deploy, with posture set by FLAGS only
├─ 2. gov_mcp/deploy.sh → Cloud Build → Cloud Run (MCP server)
├─ 3. auditor/deploy.sh → Agent Runtime + its two IAM bindings
└─ 4. scripts/setup-gemini-enterprise.sh → GE app + agent registration + sharingThree things worth knowing about that chain:
common.shis a sourced library, not a script. It definesPROJECT_ID,REGION,IMAGEandbuild_image(). Running it directly does nothing.One image, four deployments. All four auditees run the same container; their governance posture lives entirely in the
gcloud run deployflags — labels, service account, env vars — which is exactly what the auditor inspects. So the first script builds and the rest reuse.build_image()skips the build when the image already exists. After editingauditees/main.py, a plain redeploy ships the old image and your change silently doesn't land. Force it once:FORCE_BUILD=1 ./auditees/deploy-compliant.sh.
Steps 2 and 3 also work standalone (./auditor/deploy.sh re-deploys just the agent), and every script is idempotent — re-running is safe.
Leaving and returning: exit ends the session and removes the container. Re-run the same docker run … command to come back — your gcloud login and installed packages are still there, because they live in the agv-gcloud and agv-venv volumes rather than in the container. To wipe everything and start clean: docker volume rm agv-gcloud agv-venv.
When something breaks, check gotchas & sharp edges before debugging — it covers the failures we already hit, including the Scope has changed auth crash, the mcp.shared.session import error, the gcloud virtualenv/VPN failure, and why several checks legitimately report SKIPPED on a personal project.
The steps above are the happy path. docs/build-plan.md carries the rest: the operating rules, every gotcha we actually hit, and a decision log explaining why things are the way they are.
Repository
Path | What |
Architecture & component plan | |
Operating rules, gotchas, decision log | |
Demo script, rubric mapping, prepared answers, coverage limits | |
What the EU AI Act actually requires today, with sources | |
Versioned policy pack (YAML — the governance rules, reviewable & git-versioned) | |
| |
ADK app — SequentialAgent pipeline, typed session state, Agent Runtime | |
Demo fleet with deliberate postures | |
Terraform for the whole sandbox | |
Deterministic + agent-layer evaluation harness — 171 offline tests, 8 live | |
The development container |
scripts/
Script | What |
Build and deploy every workload, in order, timed | |
Soft teardown — removes workloads, keeps the Terraform baseline | |
Create the GE app and register the agent — entirely by API, no console clicks and no OAuth client needed | |
Multi-turn chat with the deployed agent from a terminal | |
Re-hash a report and every piece of evidence it cites, without trusting the auditor | |
Browse and pretty-print the evidence store | |
Build the OAuth authorization URI, if a GE integration ever needs one | |
Run the pipeline locally and print wall clock + per-step token cost — seconds per iteration instead of a 3-minute redeploy | |
The live rogue-agent demo sequence |
Testing
pytest evals/ -q # 171 offline, no GCP needed, free
pytest evals/agent -m live # 8 end-to-end agent evals (~100s, ~$0.02, needs ADC)The live suite drives the real pipeline and asserts what unit tests structurally cannot: that the run reaches the gate, that nothing is remediated before a human approves, that the approval token never reaches the chat, and that every failing finding carries a valid evidence hash.
This server cannot be deployed
Maintenance
Related MCP Connectors
Register every AI agent, log every action, prove it. EU AI Act compliance built in.
Threat modeling, code/cloud/pipeline scanning, shadow-AI discovery, compliance checks and fixes.
Runtime AI governance: decision gates, human approval, hash-chained audit, compliance mapping.
EU AI Act sovereignty scanning. Provider residency, registration status, audit trail support.
Related MCP Servers
- AlicenseAqualityBmaintenanceProvides cryptographic signing and verification for AI decisions to generate verifiable, Ed25519-signed receipts for compliance and auditing. It automatically maps AI actions to regulatory frameworks like HIPAA and SOX with high-performance, sub-3ms signing.4MIT
- AlicenseAqualityCmaintenanceProvides automated EU AI Act compliance tools, including risk classification, role determination, transparency disclosures, content watermarking, deepfake labeling, and security threat detection.1633Apache 2.0
- AlicenseNot gradedqualityDmaintenanceEnables EU AI Act compliance for AI agent systems by providing risk classification, audit trails, gap analysis, and evidence package generation.60 npmMIT
- AlicenseNot gradedqualityBmaintenanceProvides 62 AI governance tools for compliance with regulations like the EU AI Act, enabling risk management, transparency, bias detection, and more.MIT