MCP Virtual User
Supports testing Google's Gemini Live voice conversations by injecting audio into the virtual microphone and capturing/transcribing the spoken responses.
Enables testing OpenAI's ChatGPT voice mode by injecting synthetic speech into the virtual microphone and transcribing the assistant's spoken response for assertions.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MCP Virtual UserOpen ChatGPT and ask it for the weather in Tokyo, then transcribe the spoken reply."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MCP Virtual User
A synthetic human in a Docker container. Has ears, a mouth, eyes, and hands.
What is this?
A fully self-contained environment that simulates a real user interacting with web apps via voice and browser. It has:
Virtual microphone — MCP can inject audio (TTS or raw) that any app thinks is coming from a hardware mic
Virtual speakers — MCP can capture and transcribe whatever audio the OS plays back
Real browser — Chromium with Playwright control, persistent login sessions
Real display — Xvfb + VNC for debugging (watch what's happening live)
MCP interface — Everything exposed as tools via Streamable HTTP
Related MCP server: Lotus MCP
Use Cases
Test ChatGPT voice mode — inject "What's the weather?" into the mic, capture ChatGPT's spoken response, transcribe it, assert on it
Test Gemini Live — same flow against Google's voice AI
Test our Mobile Mesh UI — full end-to-end voice conversation testing against our own app
Any voice-enabled web app — if it uses the browser's mic/speaker, we can test it
Architecture
┌─────────────────────────────────────────────────────────────┐
│ Docker Container │
│ │
│ ┌─────────────────────────────────────────────────────┐ │
│ │ MCP Server (port 8360) │ │
│ │ │ │
│ │ Audio Tools: │ │
│ │ tts_to_mic(text) → Piper TTS → virtual mic │ │
│ │ transcribe_speakers() → parec → Whisper STT │ │
│ │ inject_audio(b64) → raw audio → virtual mic │ │
│ │ capture_audio(secs) → raw audio from speakers │ │
│ │ wait_for_speech() → detect + transcribe │ │
│ │ │ │
│ │ Browser Tools: │ │
│ │ browser_navigate, click, type, screenshot, etc. │ │
│ │ browser_grant_mic_permission(origin) │ │
│ │ │ │
│ │ Screen Tools: │ │
│ │ screen_screenshot, screen_size, vnc_url │ │
│ └──────────┬──────────────────────────┬───────────────┘ │
│ │ │ │
│ ┌──────────▼──────────┐ ┌───────────▼───────────────┐ │
│ │ PulseAudio │ │ Playwright + Chromium │ │
│ │ │ │ │ │
│ │ virtual_mic ◀──────│──│── browser reads as mic │ │
│ │ (pipe-source) │ │ │ │
│ │ │ │ browser plays audio ──▶ │ │
│ │ virtual_speaker ───│──│── captured via .monitor │ │
│ │ (null-sink) │ │ │ │
│ └─────────────────────┘ └───────────────────────────┘ │
│ │
│ ┌─────────────────────┐ ┌───────────────────────────┐ │
│ │ Xvfb :99 │ │ x11vnc + noVNC │ │
│ │ 1920x1080x24 │ │ port 5900 / 6080 │ │
│ └─────────────────────┘ └───────────────────────────┘ │
└─────────────────────────────────────────────────────────────┘Quick Start
# Build and start
docker compose up -d --build
# Watch the virtual display (open in your browser)
open http://localhost:6080/vnc.html?autoconnect=true
# Test the MCP server
curl http://localhost:8360/mcp -X POST \
-H "Content-Type: application/json" \
-H "Accept: application/json, text/event-stream" \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"health","arguments":{}}}'
# Run smoke tests
bash test-smoke.shFirst-Time Setup (One-Time)
# 1. Build and start
docker compose up -d --build
# 2. Open VNC in your browser to see the desktop
open http://localhost:6080/vnc.html?autoconnect=true
# 3. In the VNC window, Chromium is running. Log into:
# - https://chat.openai.com (ChatGPT)
# - https://gemini.google.com (Gemini)
# 4. Export cookies to DragonsKeep so they persist
bash scripts/export-browser-cookies.sh
# 5. Done! Future container starts auto-inject the stored sessions.Running Tests
# Infrastructure tests (no login required)
cd tests && pip install -r requirements.txt
pytest test_audio_roundtrip.py test_browser.py -v
# ChatGPT conversation tests (requires login)
pytest test_conversations.py -v -m chatgpt --timeout=120
# Gemini conversation tests
pytest test_conversations.py -v -m gemini --timeout=120
# Our Mesh UI conversation tests
pytest test_conversations.py -v -m meshui --timeout=120
# All conversation tests
pytest test_conversations.py -v --timeout=120Example: Test ChatGPT Voice Mode
# From any MCP client (Kiro, our mesh agent, etc.)
# 1. Navigate to ChatGPT
browser_navigate("https://chat.openai.com")
# 2. Grant mic permission
browser_grant_mic_permission("https://chat.openai.com")
# 3. Click the voice mode button
browser_click("[data-testid='voice-mode-button']")
# 4. Speak a question (injected into the virtual mic)
tts_to_mic("What is the capital of France?")
# 5. Wait for ChatGPT to respond and transcribe what it says
response = wait_for_speech(timeout_seconds=15)
# response == "The capital of France is Paris."
# 6. Assert
assert "Paris" in responseSession Management
Login sessions persist in ./data/browser-profile/ (mounted volume).
For ChatGPT/Gemini auth tokens, use DragonsKeep:
Store cookies in DragonsKeep as
chatgpt-session/gemini-sessionOn container start, inject them into the browser profile
Or: log in manually once via VNC (http://localhost:6080), session persists.
Ports
Port | Service |
8360 | MCP Server (Streamable HTTP) |
6080 | noVNC (browser-based VNC viewer) |
5900 | VNC direct |
MCP Tools
Audio
Tool | Description |
| Synthesize text → inject as microphone input |
| Capture speaker output → transcribe to text |
| Push raw audio bytes into the virtual mic |
| Record raw audio from the virtual speakers |
| Detect speech on speakers, wait for it to finish, transcribe |
Browser
Tool | Description |
| Go to a URL |
| Click an element |
| Type into an input |
| Press a keyboard key |
| Take a page screenshot |
| Get page text content |
| Execute JavaScript |
| Wait for text to appear |
| Get current URL |
| Allow mic access for an origin |
Screen
Tool | Description |
| Full desktop screenshot |
| Get display resolution |
| Get the live VNC viewer URL |
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- FlicenseAqualityDmaintenanceEnables AI assistants to perform automated web testing by controlling a Chrome browser to navigate, interact with pages, capture screenshots, extract console logs, and simulate mobile devices for responsive design testing.7
- FlicenseNot gradedqualityDmaintenanceEnables creation of reusable browser automation skills through demonstration by recording user actions in a browser while narrating, then converting those workflows into executable skills that can be invoked through natural language.
- FlicenseNot gradedqualityCmaintenanceEnables automated end-to-end testing and verification of web applications through natural language, with self-healing selectors and dual-mode execution.15
- AlicenseNot gradedqualityAmaintenanceEnables plain-English browser automation via an MCP server, allowing agents to run objectives or test suites in a real browser without selectors or scripts.105Apache 2.0
Related MCP Connectors
Simulation, evaluation and monitoring for voice agents.
Hosted real Google Chrome MCP with per-user persistent state. Navigate, click, type, screenshot.
AI QA tester — real browsers scan sites for bugs, SEO, perf, and accessibility issues via chat.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/macbeth76/mcp-virtual-user'
If you have feedback or need assistance with the MCP directory API, please join our Discord server