MCP Virtual User
MCP Virtual User
Синтетический человек в Docker-контейнере. У него есть уши, рот, глаза и руки.
Что это?
Полностью автономная среда, имитирующая реального пользователя, который взаимодействует с веб-приложениями через голос и браузер. В ней есть:
Виртуальный микрофон — MCP может внедрять аудио (TTS или сырое), которое любое приложение считает идущим с аппаратного микрофона
Виртуальные динамики — MCP может захватывать и транскрибировать аудио, которое воспроизводит ОС
Настоящий браузер — Chromium с управлением через Playwright, постоянные сессии входа
Настоящий дисплей — Xvfb + VNC для отладки (смотрите, что происходит в реальном времени)
Интерфейс MCP — всё доступно через инструменты Streamable HTTP
Related MCP server: Lotus MCP
Варианты использования
Тест голосового режима ChatGPT — внедрить «Какой прогноз погоды?» в микрофон, захватить голосовой ответ ChatGPT, транскрибировать его и выполнить проверку
Тест Gemini Live — тот же процесс для голосового ИИ от Google
Тест нашего Mobile Mesh UI — полное сквозное тестирование голосовых разговоров в нашем приложении
Любое веб-приложение с голосовыми функциями — если оно использует микрофон/динамики браузера, мы можем его протестировать
Архитектура
┌─────────────────────────────────────────────────────────────┐
│ Docker Container │
│ │
│ ┌─────────────────────────────────────────────────────┐ │
│ │ MCP Server (port 8360) │ │
│ │ │ │
│ │ Audio Tools: │ │
│ │ tts_to_mic(text) → Piper TTS → virtual mic │ │
│ │ transcribe_speakers() → parec → Whisper STT │ │
│ │ inject_audio(b64) → raw audio → virtual mic │ │
│ │ capture_audio(secs) → raw audio from speakers │ │
│ │ wait_for_speech() → detect + transcribe │ │
│ │ │ │
│ │ Browser Tools: │ │
│ │ browser_navigate, click, type, screenshot, etc. │ │
│ │ browser_grant_mic_permission(origin) │ │
│ │ │ │
│ │ Screen Tools: │ │
│ │ screen_screenshot, screen_size, vnc_url │ │
│ └──────────┬──────────────────────────┬───────────────┘ │
│ │ │ │
│ ┌──────────▼──────────┐ ┌───────────▼───────────────┐ │
│ │ PulseAudio │ │ Playwright + Chromium │ │
│ │ │ │ │ │
│ │ virtual_mic ◀──────│──│── browser reads as mic │ │
│ │ (pipe-source) │ │ │ │
│ │ │ │ browser plays audio ──▶ │ │
│ │ virtual_speaker ───│──│── captured via .monitor │ │
│ │ (null-sink) │ │ │ │
│ └─────────────────────┘ └───────────────────────────┘ │
│ │
│ ┌─────────────────────┐ ┌───────────────────────────┐ │
│ │ Xvfb :99 │ │ x11vnc + noVNC │ │
│ │ 1920x1080x24 │ │ port 5900 / 6080 │ │
│ └─────────────────────┘ └───────────────────────────┘ │
└─────────────────────────────────────────────────────────────┘Быстрый старт
# Build and start
docker compose up -d --build
# Watch the virtual display (open in your browser)
open http://localhost:6080/vnc.html?autoconnect=true
# Test the MCP server
curl http://localhost:8360/mcp -X POST \
-H "Content-Type: application/json" \
-H "Accept: application/json, text/event-stream" \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"health","arguments":{}}}'
# Run smoke tests
bash test-smoke.shПервоначальная настройка (однократно)
# 1. Build and start
docker compose up -d --build
# 2. Open VNC in your browser to see the desktop
open http://localhost:6080/vnc.html?autoconnect=true
# 3. In the VNC window, Chromium is running. Log into:
# - https://chat.openai.com (ChatGPT)
# - https://gemini.google.com (Gemini)
# 4. Export cookies to DragonsKeep so they persist
bash scripts/export-browser-cookies.sh
# 5. Done! Future container starts auto-inject the stored sessions.Запуск тестов
# Infrastructure tests (no login required)
cd tests && pip install -r requirements.txt
pytest test_audio_roundtrip.py test_browser.py -v
# ChatGPT conversation tests (requires login)
pytest test_conversations.py -v -m chatgpt --timeout=120
# Gemini conversation tests
pytest test_conversations.py -v -m gemini --timeout=120
# Our Mesh UI conversation tests
pytest test_conversations.py -v -m meshui --timeout=120
# All conversation tests
pytest test_conversations.py -v --timeout=120Пример: тестирование голосового режима ChatGPT
# From any MCP client (Kiro, our mesh agent, etc.)
# 1. Navigate to ChatGPT
browser_navigate("https://chat.openai.com")
# 2. Grant mic permission
browser_grant_mic_permission("https://chat.openai.com")
# 3. Click the voice mode button
browser_click("[data-testid='voice-mode-button']")
# 4. Speak a question (injected into the virtual mic)
tts_to_mic("What is the capital of France?")
# 5. Wait for ChatGPT to respond and transcribe what it says
response = wait_for_speech(timeout_seconds=15)
# response == "The capital of France is Paris."
# 6. Assert
assert "Paris" in responseУправление сессиями
Сессии входа сохраняются в ./data/browser-profile/ (смонтированный том).
Для токенов аутентификации ChatGPT/Gemini используйте DragonsKeep:
Сохраните куки в DragonsKeep как
chatgpt-session/gemini-sessionПри запуске контейнера внедрите их в профиль браузера
Или: войдите вручную один раз через VNC (http://localhost:6080) — сессия сохранится.
Порты
Порт | Сервис |
8360 | MCP Server (Streamable HTTP) |
6080 | noVNC (просмотр VNC в браузере) |
5900 | VNC direct |
Инструменты MCP
Аудио
Инструмент | Описание |
| Синтезировать текст → внедрить как вход микрофона |
| Захватить вывод динамиков → транскрибировать в текст |
| Передать сырые аудио-байты в виртуальный микрофон |
| Записать сырое аудио с виртуальных динамиков |
| Обнаружить речь на динамиках, дождаться окончания и транскрибировать |
Браузер
Инструмент | Описание |
| Перейти на URL |
| Кликнуть по элементу |
| Ввести текст в поле ввода |
| Нажать клавишу клавиатуры |
| Сделать скриншот страницы |
| Получить текстовое содержимое страницы |
| Выполнить JavaScript |
| Подождать появления текста |
| Получить текущий URL |
| Разрешить доступ к микрофону для origin |
Экран
Инструмент | Описание |
| Скриншот всего рабочего стола |
| Получить разрешение дисплея |
| Получить URL живого VNC-просмотра |
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- FlicenseAqualityDmaintenanceEnables AI assistants to perform automated web testing by controlling a Chrome browser to navigate, interact with pages, capture screenshots, extract console logs, and simulate mobile devices for responsive design testing.7
- FlicenseNot gradedqualityDmaintenanceEnables creation of reusable browser automation skills through demonstration by recording user actions in a browser while narrating, then converting those workflows into executable skills that can be invoked through natural language.
- FlicenseNot gradedqualityCmaintenanceEnables automated end-to-end testing and verification of web applications through natural language, with self-healing selectors and dual-mode execution.15
- AlicenseNot gradedqualityAmaintenanceEnables plain-English browser automation via an MCP server, allowing agents to run objectives or test suites in a real browser without selectors or scripts.105Apache 2.0
Related MCP Connectors
Simulation, evaluation and monitoring for voice agents.
Hosted real Google Chrome MCP with per-user persistent state. Navigate, click, type, screenshot.
AI QA tester — real browsers scan sites for bugs, SEO, perf, and accessibility issues via chat.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/macbeth76/mcp-virtual-user'
If you have feedback or need assistance with the MCP directory API, please join our Discord server