Skip to main content
Glama

MCP Virtual User

Docker コンテナ内で動作する仮想的な人間です。耳、口、目、手を備えています。

これは何ですか?

音声とブラウザーを介して Web アプリを操作する実ユーザーをシミュレートする、完全自己完結型の環境です。特徴は次のとおりです。

  • 仮想マイク — MCP が音声(TTS または raw)を注入し、どのアプリからもハードウェアマイクからの入力として認識されます

  • 仮想スピーカー — MCP が OS で再生される音声をキャプチャして文字起こしできます

  • 実際のブラウザー — Playwright で操作できる Chromium と、永続化されるログインセッション

  • 実際のディスプレイ — Xvfb + VNC でデバッグ(実際の動作をライブで視聴可能)

  • MCP インターフェース — すべての機能を Streamable HTTP 経由のツールとして公開

Related MCP server: Lotus MCP

ユースケース

  • ChatGPT の音声モードをテスト — マイクに "What's the weather?" を注入し、ChatGPT の音声応答をキャプチャして文字起こし、検証します

  • Gemini Live をテスト — Google の音声 AI に対して同じフローを実行します

  • 自社の Mobile Mesh UI をテスト — 自社アプリに対するエンドツーエンドの音声会話テストを行います

  • 音声対応の Web アプリ全般 — ブラウザーのマイク/スピーカーを使用するアプリであれば、テスト可能です

アーキテクチャ

┌─────────────────────────────────────────────────────────────┐
│  Docker Container                                            │
│                                                             │
│  ┌─────────────────────────────────────────────────────┐    │
│  │  MCP Server (port 8360)                              │    │
│  │                                                     │    │
│  │  Audio Tools:                                       │    │
│  │    tts_to_mic(text)     → Piper TTS → virtual mic   │    │
│  │    transcribe_speakers() → parec → Whisper STT      │    │
│  │    inject_audio(b64)    → raw audio → virtual mic   │    │
│  │    capture_audio(secs)  → raw audio from speakers   │    │
│  │    wait_for_speech()    → detect + transcribe       │    │
│  │                                                     │    │
│  │  Browser Tools:                                     │    │
│  │    browser_navigate, click, type, screenshot, etc.  │    │
│  │    browser_grant_mic_permission(origin)             │    │
│  │                                                     │    │
│  │  Screen Tools:                                      │    │
│  │    screen_screenshot, screen_size, vnc_url          │    │
│  └──────────┬──────────────────────────┬───────────────┘    │
│             │                          │                     │
│  ┌──────────▼──────────┐  ┌───────────▼───────────────┐    │
│  │  PulseAudio         │  │  Playwright + Chromium    │    │
│  │                     │  │                           │    │
│  │  virtual_mic ◀──────│──│── browser reads as mic    │    │
│  │  (pipe-source)      │  │                           │    │
│  │                     │  │  browser plays audio ──▶  │    │
│  │  virtual_speaker ───│──│── captured via .monitor   │    │
│  │  (null-sink)        │  │                           │    │
│  └─────────────────────┘  └───────────────────────────┘    │
│                                                             │
│  ┌─────────────────────┐  ┌───────────────────────────┐    │
│  │  Xvfb :99           │  │  x11vnc + noVNC          │    │
│  │  1920x1080x24       │  │  port 5900 / 6080        │    │
│  └─────────────────────┘  └───────────────────────────┘    │
└─────────────────────────────────────────────────────────────┘

クイックスタート

# Build and start
docker compose up -d --build

# Watch the virtual display (open in your browser)
open http://localhost:6080/vnc.html?autoconnect=true

# Test the MCP server
curl http://localhost:8360/mcp -X POST \
  -H "Content-Type: application/json" \
  -H "Accept: application/json, text/event-stream" \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"health","arguments":{}}}'

# Run smoke tests
bash test-smoke.sh

初回セットアップ(一度だけ)

# 1. Build and start
docker compose up -d --build

# 2. Open VNC in your browser to see the desktop
open http://localhost:6080/vnc.html?autoconnect=true

# 3. In the VNC window, Chromium is running. Log into:
#    - https://chat.openai.com (ChatGPT)
#    - https://gemini.google.com (Gemini)

# 4. Export cookies to DragonsKeep so they persist
bash scripts/export-browser-cookies.sh

# 5. Done! Future container starts auto-inject the stored sessions.

テストの実行

# Infrastructure tests (no login required)
cd tests && pip install -r requirements.txt
pytest test_audio_roundtrip.py test_browser.py -v

# ChatGPT conversation tests (requires login)
pytest test_conversations.py -v -m chatgpt --timeout=120

# Gemini conversation tests
pytest test_conversations.py -v -m gemini --timeout=120

# Our Mesh UI conversation tests
pytest test_conversations.py -v -m meshui --timeout=120

# All conversation tests
pytest test_conversations.py -v --timeout=120

例:ChatGPT の音声モードをテスト

# From any MCP client (Kiro, our mesh agent, etc.)

# 1. Navigate to ChatGPT
browser_navigate("https://chat.openai.com")

# 2. Grant mic permission
browser_grant_mic_permission("https://chat.openai.com")

# 3. Click the voice mode button
browser_click("[data-testid='voice-mode-button']")

# 4. Speak a question (injected into the virtual mic)
tts_to_mic("What is the capital of France?")

# 5. Wait for ChatGPT to respond and transcribe what it says
response = wait_for_speech(timeout_seconds=15)
# response == "The capital of France is Paris."

# 6. Assert
assert "Paris" in response

セッション管理

ログインセッションは ./data/browser-profile/(マウント済みボリューム)に保持されます。

ChatGPT/Gemini の認証トークンには DragonsKeep を使用してください。

  1. Cookies を chatgpt-session / gemini-session として DragonsKeep に保存する

  2. コンテナ起動時にそれらをブラウザープロファイルに注入する

または、VNC(http://localhost:6080)で一度手動ログインすれば、セッションは永続化されます。

ポート

ポート

サービス

8360

MCP Server (Streamable HTTP)

6080

noVNC (ブラウザーベースの VNC ビューアー)

5900

VNC 直接接続

MCP ツール

音声

ツール

説明

tts_to_mic

テキストを音声合成し、マイク入力として注入します

transcribe_speakers

スピーカー出力をキャプチャしてテキストに文字起こしします

inject_audio

生のオーディオバイト列を仮想マイクにプッシュします

capture_audio

仮想スピーカーから生のオーディオを録音します

wait_for_speech

スピーカーで音声を検出し、終了するまで待ってから文字起こしします

ブラウザー

ツール

説明

browser_navigate

URL に移動します

browser_click

要素をクリックします

browser_type

入力欄にテキストを入力します

browser_press_key

キーボードのキーを押します

browser_screenshot

ページのスクリーンショットを撮影します

browser_get_text

ページのテキスト内容を取得します

browser_evaluate

JavaScript を実行します

browser_wait_for_text

指定したテキストが表示されるまで待機します

browser_url

現在の URL を取得します

browser_grant_mic_permission

オリジンに対してマイクアクセスを許可します

画面

ツール

説明

screen_screenshot

デスクトップ全体のスクリーンショットを取得します

screen_size

ディスプレイ解像度を取得します

vnc_url

ライブ VNC ビューアーの URL を取得します

F
license - not found
Not graded
quality - not tested
C
maintenance

Maintenance

Maintainers
Response time
Release cycle
Releases (12mo)
Commit activity

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Servers

  • F
    license
    A
    quality
    D
    maintenance
    Enables AI assistants to perform automated web testing by controlling a Chrome browser to navigate, interact with pages, capture screenshots, extract console logs, and simulate mobile devices for responsive design testing.
    7
  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables creation of reusable browser automation skills through demonstration by recording user actions in a browser while narrating, then converting those workflows into executable skills that can be invoked through natural language.
  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables automated end-to-end testing and verification of web applications through natural language, with self-healing selectors and dual-mode execution.
    15
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables plain-English browser automation via an MCP server, allowing agents to run objectives or test suites in a real browser without selectors or scripts.
    105
    Apache 2.0

View all related MCP servers

Related MCP Connectors

  • Simulation, evaluation and monitoring for voice agents.

  • Hosted real Google Chrome MCP with per-user persistent state. Navigate, click, type, screenshot.

  • AI QA tester — real browsers scan sites for bugs, SEO, perf, and accessibility issues via chat.

View all MCP Connectors

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/macbeth76/mcp-virtual-user'

If you have feedback or need assistance with the MCP directory API, please join our Discord server