Skip to main content
Glama
nhodges
by nhodges

mcp-vroid

VRoid Studio의 GUI를 구동하는 MCP 서버입니다. 모든 MCP 클라이언트 (Claude Code 또는 프로토콜을 지원하는 다른 무엇이든)에게 앱을 실행하고, 화면을 보고, 이미지에서 위젯을 찾고, 클릭·타이핑하고, 매개변수를 설정하고, .vrm을 내보내는 도구 모음을 제공합니다 — Arch + Hyprland (Wayland) 환경에서, VRoid Studio는 Steam/Proton으로 실행됩니다.

VRoid Studio에는 스크립팅 API가 없으므로, 이것을 가능한 유일한 방식으로 동작합니다: 창을 스크린샷하고, OCR과 색상 매칭으로 물체를 찾고, 실제 포인터와 키보드 이벤트를 주입합니다.

   grim ──► PNG ──► tesseract / cv2 ──► (x, y) ──► virtual pointer / XTEST
    ▲                                                        │
    └────────────────────  screenshot again  ◄───────────────┘

서버의 엔진은 제 arrakis 프로젝트의 tools/vroid-driver 스파이크를 이 프로젝트에 mcp_vroid.driver로 벤더링한 것입니다 — 동일한 코드를 재패키징하여 MCP 클라이언트가 설치하고 시작할 수 있게 했습니다.


요구사항

항목

이유

Hyprland (>= 0.55, Lua dispatch API)

창 탐색, 포커스, 작업 공간

VRoid Studio via Steam/Proton (appid 1486350)

구동 대상 앱

grim

스크린샷

tesseract + eng traineddata

OCR

gcc, wayland-scanner, libwayland-client

포인터 헬퍼를 빌드하기 위한 컴파일러

Xwayland (DISPLAY)

키보드와 휠은 X11 XTEST를 통해 전달됨

Python 3.11+, uv

서버 자체

Python 의존성(uv sync가 설치 요구): mcp, pillow, numpy, opencv-python-headless, pytesseract, python-xlib.

Related MCP server: Persona Motion Studio

설치

git clone https://github.com/nhodges/mcp-vroid
cd mcp-vroid
uv sync                 # virtualenv + dependencies
bash native/build.sh    # builds native/vpointer  <-- REQUIRED, not optional

native/build.sh는 zwlr_virtual_pointer_unstable_v1용 약 150줄 짜리 C 클라이언트를 컴파일합니다(프로토콜 XML은 native/protocols/ 안에 벤더링되어 있음). 이것이 없으면 모든 포인터 도구가 native/vpointer missing 메시지와 함께 실패합니다. vroid_status는 존재 여부를 보고합니다.

왜 C 헬퍼인가: 기준 머신에 ydotool이 설치되어 있지 않고 /dev/uinput이 0600 root:root이므로 evdev 주입은 sudo 또는 udev 규칙이 필요합니다. Wayland 가상 포인터 프로토콜은 둘 다 필요하지 않으며, 실제 컴포지터 커서를 움직이고 어떤 창에서도 동작합니다.

클라이언트에 등록하기

Claude Code:

claude mcp add vroid -- uv run --directory /path/to/mcp-vroid mcp-vroid

표준 표준 mcpServers JSON:

{
  "mcpServers": {
    "vroid": {
      "command": "uv",
      "args": ["run", "--directory", "/path/to/mcp-vroid", "mcp-vroid"]
    }
  }
}

클라이언트는 종종 정화된 환경으로 서버를 실행합니다. 이 서버는 시작 시(src/mcp_vroid/session_env.py)에서 XDG_RUNTIME_DIR, WAYLAND_DISPLAY, HYPRLAND_INSTANCE_SIGNATURE, DISPLAY를 런타임 디렉터리에서 복구하므로 hyprctl/grim/XTEST가 그대로 작동합니다. vroid_status는 무엇을 채워 넣었는지 보여줍니다. 이미 환경에 존재하는 값이 우선합니다.

선택적 환경 변수:

변수

기본값

설명

MCP_VROID_CAPTURES

$XDG_STATE_HOME/mcp-vroid/captures

스크린샷이 저장되는 경로

MCP_VROID_OUT

$XDG_STATE_HOME/mcp-vroid/out

내보내기/저장 기본 경로

MCP_VROID_VPOINTER

<checkout>/native/vpointer

포인터 헬퍼 경로

MCP_VROID_MAX_IMAGE_PX

1600

클라이언트로 보내는 이미지의 긴 쪽 최대 픽셀 (0 = 축소 안 함)

도구

생명주기

도구

동작

vroid_launch(restart=false, timeout=240)

필요 시 Steam을 통해 VRoid를 시작하고, Hyprland 작업 공간 9에 배치하고, 원래 자신이 있던 작업 공간을 변수에 저장하고, 포커스 + 전체 화면으로 전환합니다. restart=true면 먼저 실행 중인 인스턴스를 종료합니다 — 저장되지 않은 작업은 유실됩니다.

vroid_status()

창 존재/포커스/제목/지오메트리, 활성 작업 공간, 캡처 디렉터리, 그리고 vpointer/grim/tesseract/hyprctl 접근 가능 여부. 읽기 전용, OCR 없음.

vroid_release()

사용자가 있던 작업 공간으로 다시 전환합니다. VRoid는 작업 공간 9에서 계속 실행됩니다.

보기

도구

동작

vroid_screenshot(region?, tag?, whole_screen?, full_resolution?)

창(또는 Wine 저장 대화상자용 전체 출력)을 캡처하고, 캡처 디렉터리에 저장하며, MCP 이미지 콘텐츠로 반환해 클라이언트의 모델이 볼 수 있게 합니다. 원본 이미지 크기와 전송 시 적용된 downscale 배율을 보고합니다.

vroid_find_text(query, region?, exact?, limit?)

새 캡처 + tesseract; 일치하는 단어 박스와 중심점을 이미지 픽셀 단위로 반환합니다. region을 전달하세요 — 전체 프레임 OCR은 약 10초, 패널은 약 2초입니다.

vroid_find_button(color='primary'|'disabled', label?, region?)

VRoid의 단색 #0096FA 약(파란 pill)을 색으로 찾습니다. tesseract는 흰색에 파란 라벨을 잘 인식하지 못하기 때문입니다. 회색 pill는 비활성을 뜻합니다.

vroid_current_screen()

start / editor / export_vrm / hair_editor / unknown을 반환합니다.

동작 (원시 입력)

도구

동작

vroid_click(x, y, space='image', button='left', double=false)

포인터를 몇 단계로 부드럽게 이동(호버 상태 발생)시킨 후 클릭합니다.

vroid_drag(x1, y1, x2, y2, space='image', button='left')

누르고 → 24단계 이동 → 놓기. 우클릭 드래그는 카메라를 궤도 이동, 휠 클릭 드래그는 팬합니다.

vroid_scroll(dy, dx=0, x=?, y=?, space='image')

마우스 휠을 X11 버튼 4/5(가로는 6/7)로 전달합니다. 포인터를 스크롤하려는 패널 위에 올려놓으세요.

vroid_type(text, clear_first=false)

포커스된 위젯에 XTEST를 통해 입력합니다.

vroid_key(combo, times=1)

Return, Escape, ctrl+s, ctrl+shift+s, … 등을 실행합니다.

동작 (플로우)

도구

동작

vroid_new_character(base='Fem'|'Masc')

시작 화면 → 새로 만들기 → 베이스 → 에디터로 이동합니다.

vroid_open_tab(name)

이름 / 헤어스타일 / 바디 / 의상 / 엑세서리 / 룩 탭입니다.

vroid_set_slider(label, value)

파라미터 패널을 해당 행까지 스크롤하고 숫자 상자에 정확한 값을 입력합니다.

vroid_set_color(label, hex)

#RRGGBB 색상 상자에 대해 동일하게 작동합니다.

vroid_export_vrm(path, avatar_name, creator, version='1.0')

Export-as-VRM 전체 과정, VRM 설정 메타데이터 모달과 Wine의 저장 대화상자를 포함합니다. version은 VRM1.0 또는 VRM0.0을 선택합니다.

vroid_save_project(name?)

명시된 .vroid 경로로 Ctrl+S 또는 Ctrl+Shift+S 저장(인자가 없으면 그냥 저장).

모든 동작 도구는 VRoid를 먼저 포커스하고, 포커스된 창이 VRoid Studio가 아니면 온전히 거부합니다.

구동 방법

주로: 스크린샷 → 확인 → 찾기 → 동작 → 다시 스크린샷

  1. vroid_launch()

  2. vroid_screenshot() 후 이미지를 살펴보세요

  3. vroid_find_text("Export") (또는 vroid_find_button())로 좌표 획득

  4. vroid_click(x, y) — 반드시 멀리 캡처에서 얻은 좌표를 사용

  5. vroid_screenshot()하여 실제로 어떤 일이 일어났는지 확인

원래 스파이크에서 어렵게 배운 경험 법칙:

  • 일부만 읽지 말고 전체 프레임을 읽으세요. "Close Hairstyle Editor" 확인 모달이 화면 중간에 있는 데도 상단 60픽셀만 검사해서 클릭이 6번 실패했습니다.

  • 3D 뷰포트의 변화로 판단하지 마세요. VRoid는 모든 프레임을 디더링하므로, 아무것도 없어도 전체 창 diff는 0.98로 읽힙니다. UI 형태를 관찰하세요.

  • 슬라이더 드래그보다 숫자 상자를 우선하세요. vroid_set_slider는 정확한 값을 입력하고, 드래그는 숫자 상자가 없는 제어에 사용됩니다.

  • 기본 버튼은 색상으로 찾습니다. 파란색 예상 위치의 회색 약이 보이면 앱이 필수 필드가 비어 있다고 알리는 것입니다.

  • 전체 2560×1440 프레임의 OCR은 약 10초가 걸립니다. 반드시 region을 전달하세요.

좌표 공간

소스 머신에서는 세 가지 공간이 활성화되어 있고 서로 다릅니다:

공간

기준머신 크기

사용처

Hyprland 레이아웃 (논리적)

2048 × 1152

hyprctl, 가상 포인터

캡처 이미지의 픽셀

2560 × 1440

tesseract, cv2, 사용자가 이미지에서 보는 것

X11 픽셀 (Xwayland)

2560 × 1440

XTEST

도구는 기본적으로 이미지 픽셀(space="image")을 받고 변환하므로, vroid_find_text 결과를 바로 vroid_click에 전달하세요. MCP_VROID_MAX_IMAGE_PX로 이미지가 다운스케일 된 경우 이미지에서 직접 좌표를 읽어 사용하려면 보고된 downscale 배율의 역수를 먼저 곱하거나, vroid_find_text가 항상 원본 픽셀을 반환하므로 그냥 사용하세요.

UI 맵 (VRoid Studio 2.14.0, 영문)

전체 화면 창 2560×1440 크기의 캡처 이미지를 기준으로 좌표를 제공합니다. 이것들은 힌트일 뿐이며 — 도구는 OCR로 위치를 찾습니다.

시작 화면 - Create New + 카드가 ≈ (118, 218)에 있고, 캡션은 (118, 328)에 있습니다. New / Open은 오른쪽 상단 (2439, 99) / (2495, 100)에 있으며, 아래에 Sample Models 그리드가 있습니다. Create New를 열면 모달 *"Select a base to start with"*이 표시되고 Fem (1199, 862) 및 Masc (1359, 862) 캡션이 있습니다 — 캡션 위의 ~100px 지점에 있는 썸네일을 클릭하세요.

편집기 — 탭 줄은 y ≈ 23에 있습니다: Face 97 · Hairstyle 198 · Body 302 · Outfit 392 · Accessories 509 · Look 622. 햄버거 메뉴 ☰ (29, 23) → Save (Ctrl+S), Save As… (Ctrl+Shift+S), import/bulk export, undo/redo, 모델 선택 화면으로 돌아가기 — Escape는 이 메뉴가 닫히지 않으므로, 다른 곳을 클릭하세요. 오른쪽 상단 툴바: 카메라 (2415, 23), 공유/export (2464, 23), 케밥 메뉴 ⋮ (2512, 23). 왼쪽의 아이콘 레일(x ≈ 24, 첫 번째 아이콘 y ≈ 77, 이후 ~48px 간격) = 현재 탭의 하위입니다. 왼쪽 패널 = 프리셋 그리드에 Presets/Custom이 y ≈ 120에 있으며; 오른쪽 패널 = Customize, 그 다음 Parameters.

오른쪽 패널 컨트롤

컨트롤

조작 방법

slider

x ≈ 2505의 숫자 상자 (vroid_set_slider); 트랙은 x ≈ 2278 → 2516이며 0.0이 중앙에 있습니다

colour

x ≈ 2450의 #RRGGBB 상자 (vroid_set_color)

checkbox / radio

사각형/원형 클릭

accordion

캡션 클릭(예: > Reduce Polygons)

dropdown

네이티브 Wine 대화 상자에서만; 클릭한 후 화살표 키 사용.

Body 파라미터는 Model's Height : 161.2 cm로 시작하며, Fem Height, Masc Height, Body Size, Head Size, Head Width, Head Tip (Y), Neck Length/Thickness/Width, Soften Collarbone 등이 뒤따릅니다. Face 파라미터: Eye Size X/Y, Eyes Position (X/Y), Rotate Eye Socket, Inner/Outer Eye Slant/, Iris Size X/Y, Gaze (Y)` … (약 40개 행; 도구가 스크롤을 대신 처리합니다).

헤어 편집기 — Hairstyle 탭 → 왼쪽 레일의 파트 아이콘 → Custom 하위 탭 → + Create New → 오른쪽 패널 Edit Hairstyle. 내부: Add Freehand Hair Guides / Add Procedural Hair Guides, Hair Groups 목록, 툴 팔레트가 (330 / 365 / 398 / 432, 83)에 있고, 취소/다시 실행이 (76, 23) / (133, 23)에 있습니다. 편집기를 나가면 먼저 확인을 요청: (23, 23)의 ✕를 클릭하면 Close Hairstyle Editor 모달이 미리 열립니다ol "하려면 Save as new item / Overwrite / Close without saving.

VRM으로 내보내기 — 공유 아이콘 (2464, 23) → Export as VRM → 전체 화면으로 전환됩니다 내보내기 페이지에 파란색 Export 알약이 ≈ (2412, 197)에 있습니다 → VRM Settings 모달(중앙 정렬, ~x 1000~1560, 스크롤 가능): Export Format 라디오 VRM1.0 / VRM0.0, Avatar Name 필수, Version, Creators 필수, 저작권/연락처/참조/사용 관련 체크박스가 있습니다. 내보내기 알약은 두 필수 필드가 모두 채워질 때까지 잿빛으로 비활성 상태를 유지합니다 → Wine 저장 대화 상자 (별도 창이며 제목 Export): File name: 필드가 포커스되고 선택된 상태에서 열리므로 Windows 경로를 입력하면 그 필드가 교체되고 Return 키가 기본선택을 누릅니다. Proton prefix가 Z:\를 /에 매핑하므로, /home/nuri/x는 Z:\home\nuri\x가 됩니다. OCR을 사용해 찾아낸 Save라는 메시지는 클릭하지 않으세요. Save in:레이블이 동일한 needle에 매칭될 수 있습니다.

Brittle한 것들

  • OCR이 위치를 파악하는 유일한 요소입니다. 작거나, 간격이 넓거나, 어두운 배경의 레이블은 분리되거나 버려질 수 있습니다(Export → E+xport). 아이콘은 텍스트가 전혀 없습니다 — 그 닻들은 창의 크기에 하드코딩된 비율이며, pixiv가 UI를 플로시키면 이동합니다.

  • 고정 앵커는 2560×1440에서 배율 1.25 기준으로 보정된 분수입니다. 다른 것 모니터는 다시 측정해야 할 수도 있습니다.

  • 모달은 서치 영역에 나타나지 않고 클릭을 조용히 삼킵니다.

  • 타이밍. 3D 뷰포트는 베이스가 선택된 후 ~5초 후에 나타납니다. export는 5~30초 소요됩니다(무거운 모델은 더 오래 걸립니다).

  • Wine 대화 상자는 별도의 창이며 자체 클래스와 렌더링 영역이 있습니다. 그 안에서는 vroid_screenshot(whole_screen=true)를 사용하세요.

  • Language. 이 기준은 영어 UI를 가정합니다. VRoid가" 창가 메뉴 ⋮ → Settings → Language에서 전환하세요.

  • 대기 화면 보호기가 실행 중 세션을 가져갈 수 있습니다. The guard is 화면 보호기에 입력을 거부하고, 작업 전에 해당 창 하나(그 창만)를 닫습니다.

보안 참고

이 서버는 실제 마우스와 키보드 이벤트를 현재 데스크톱 세션에 주입하고 화면을 캡처합니다. 그것이 이 서버의 존재 이유이며, 위험이기도 합니다:

  • 스크린샷에는 화면에 입력되는 모든 정보가 포함될 수 있고, whole_screen=true를 사용하면 모든 것이 캡처되며, 캡처 파일은 암호화되지 않은 디스크에 저장됩니다.

  • 키 입력은 포커스를 가진 대상으로 전달됩니다. VRoid Studio가 포커스된 창일 때만 실행되지만, 손상되었거나 부주의한 프롬프트는 여전히 VRoid 내부의 아무 위치에서도 클릭할 수 있습니다.

  • vroid_launch(restart=true)는 VRoid Studio를 죽이고 저장되지 않은 작업을 잃습니다.

  • 여기에는 어떤 것도 샌드박스가 없으며 확인 단계도 없습니다.

보는 사람이 있는 곳에서 실행하세요. 자신이 보고 있는 세션에서만 실행하고, 에이전트가 감독 없이 운전하지 않도록 하십시오. vroid_release()를 호출하면 완료 후 데스크톱을 반환합니다.

Development

uv run python scripts/smoke_test.py             # start the server, list tools, call vroid_status
uv run python scripts/smoke_test.py --screenshot # + one passive capture if VRoid is open
uv run vroid-driver shot                        # the original driver CLI, still here

vroid-driver( mcp_vroid.driver.cli)는 스파이크의 셸( shell ) 인터페이스입니다 — launch, shot, find, click, tab, slider, export, cam, apply-params, … — MCP 클라이언트 없이 디버깅에 편리합니다.

Credits and licence

드라이버(src/mcp_vroid/driver/, native/)는 원래 제 arrakis 프로젝트의 tools/vroid-driver 스파이크에서 시작되었으며, 여기에는 MCP 서버가 그 주위에 감싸인 채 포함되어 있습니다.

MIT — LICENSE 참조.

Available Tools

18 tools
vroid_clickA

Click a point in the VRoid window with the real compositor cursor.

Refuses unless VRoid Studio is the focused window; it focuses the window itself first (workspace 9, fullscreen) and raises rather than clicking into somebody else's app.

The pointer glides to the target in a few steps so hover states fire, then clicks and settles ~0.35 s. Coordinates must come from a CURRENT capture - take a fresh vroid_screenshot or vroid_find_text right before clicking, because panels reflow and modals move.

ParametersJSON Schema
NameRequiredDescriptionDefault
xYesX coordinate in `space`.
yYesY coordinate in `space`.
spaceNo'image' = pixels of a window capture (what vroid_screenshot / vroid_find_* report - the default, and almost always what you want); 'window' = Hyprland layout units relative to the window's top-left; 'layout' = absolute Hyprland layout units of the whole output.image
buttonNoMouse button. right/middle also orbit/pan the 3D viewport when dragged.left
doubleNoSend two clicks (selects a word in a text box).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without any annotations, the description carries the full burden and does so richly: it discloses the focus/raise behavior, refusal condition, pointer glide timing (~0.35 s), and the requirement for fresh captures. This goes well beyond what the schema conveys and gives the agent a clear model of side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three paragraphs, with the core action stated first, followed by essential behavioral caveats. It is reasonably tight—each sentence contributes (focus behavior, pointer glide, coordinate freshness). No redundant fluff, though it could be compressed slightly. The structure is logical and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists and parameter schema covers everything, the description supplies needed behavioral context: when it refuses, how it moves the cursor, and the freshness requirement. It does not mention return values, but the output schema handles that. It is complete enough for an agent to invoke correctly, leaving no major gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already describes all 5 parameters with 100% coverage. The description adds only context about coordinate freshness (stale coordinates cause misses), but does not elaborate on 'space', 'button', or 'double'. Since the schema is exhaustive, the description adds little extra but is not deficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States precisely that it clicks a point in the VRoid window using the real compositor cursor, distinguishing it from drag, scroll, and other sibling actions. The verb, resource, and mechanism are explicit, so an agent can clearly understand the tool's role.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit guidance that coordinates must come from a current capture, and that it will focus/raise the VRoid window itself, refusing to click into other apps. This conveys when to use it (for interacting with VRoid) and warns against stale coordinates. However, it does not explicitly contrast with sibling click-like tools (though drag/scroll are self-evident).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vroid_current_screenA

Identify which VRoid screen is on top: start, editor, export_vrm, hair_editor, or unknown.

Cheap-ish (OCRs only the top strip, ~2 s) and worth calling whenever a flow tool fails - most failures are "you are not on the screen this tool expects". 'unknown' usually means a modal is up; screenshot the whole window and look.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that the tool is 'cheap-ish' (only OCRs the top strip, ~2 s) and that 'unknown' indicates a modal, which are useful behavioral traits. However, it does not explicitly state that the tool is read-only or has no side effects. For a diagnostic tool, this is a minor omission, so a 4 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-structured: the first sentence states the primary purpose, and the second provides usage guidance and troubleshooting. Every sentence contributes meaningful information without redundancy. The 'unknown' tip is an efficient way to add depth.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter diagnostic tool, the description is complete. It covers purpose, possible outputs, cost, common failure interpretation, and a recommended next step. The output schema likely provides the exact return type, but the description already gives enough to the agent. No significant gaps remain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline is 4. The description adds value by enumerating the possible output values (the five screen states), which helps the agent interpret the result. No parameter details are needed, so this dimension is well-served.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: to identify which VRoid screen is on top, listing the possible outputs (start, editor, export_vrm, hair_editor, unknown). The verb 'identify' and the resource 'screen state' are specific, and the explicit list of possible values distinguishes it from sibling tools that perform actions rather than diagnostics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use it: 'worth calling whenever a flow tool fails' and explains the common failure mode. It also provides guidance on interpreting the 'unknown' result (modal is up) and suggests a follow-up action (take a full-window screenshot). This gives clear decision-making context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vroid_dragA

Press, glide and release - slider handles, the 3D camera, hair guides.

For Parameters sliders prefer vroid_set_slider, which types an exact value into the row's numeric box; dragging is only for controls that have no numeric box. The glide is 24 interpolated steps, which is what the app needs to register a drag rather than a click.

ParametersJSON Schema
NameRequiredDescriptionDefault
x1YesPress point X in `space`.
x2YesRelease point X in `space`.
y1YesPress point Y in `space`.
y2YesRelease point Y in `space`.
spaceNo'image' = pixels of a window capture (what vroid_screenshot / vroid_find_* report - the default, and almost always what you want); 'window' = Hyprland layout units relative to the window's top-left; 'layout' = absolute Hyprland layout units of the whole output.image
buttonNoleft = slider handles and drawing; right = orbit the camera (~400 image px is 90 deg of yaw); middle = pan the model.left

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It adds key details: the glide is 24 interpolated steps to register a drag rather than a click, and the kinds of operations it performs (slider, camera orbit, pan). While it doesn't discuss failure modes or side effects, it covers the essential behavioral traits an agent would need to know before calling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short paragraphs: the first delivers the primary purpose in a compact phrase, and the second adds usage guidance and a behavioral detail. Every sentence earns its place, and the key scoping information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description gives enough context for correct invocation: it names target resources, the condition for using it, and the interpolation detail. An output schema exists, so return values are covered. It stops short of describing error handling or coordinate system conversion, but for a drag tool that's acceptable given the schema richness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides 100% coverage with detailed descriptions for x1, y1, x2, y2, space, and button, including enums and defaults. The description does not add parameter-specific explanation beyond what the schema offers, so a baseline of 3 is appropriate because the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb (press, glide, release) and the specific resources it acts on (slider handles, 3D camera, hair guides), and it distinguishes itself from sibling vroid_set_slider by noting dragging is only for controls without a numeric box. This leaves no ambiguity about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use this tool versus vroid_set_slider, with a concrete rule: 'prefer vroid_set_slider for parameters sliders' and 'dragging is only for controls that have no numeric box.' It also mentions camera and hair guides, giving clear contexts for alternative use. No inference is needed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vroid_export_vrmA

Walk the entire Export-as-VRM flow and write a .vrm file.

Editor toolbar share icon -> 'Export as VRM' -> the blue Export pill -> the VRM Settings modal (fills Avatar Name and Creators, picks the export format, scrolls to the bottom and clicks Export) -> Wine's save dialog (types a Z:\ path and presses Return) -> waits for the file size to stop growing.

Must be started from the EDITOR screen with a model loaded. Takes 30 s to a few minutes depending on the model. Returns the written path and its size; raises with the path of a diagnostic screenshot if any step fails - read that screenshot before retrying, since a half-finished flow usually leaves a modal open that the next attempt will trip over.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesWhere to write the .vrm, as a normal Linux path. Translated to the Proton prefix's Z:\ mapping for the Wine save dialog.
creatorYesVRM metadata 'Creators' - also REQUIRED.
timeoutNoSeconds to wait for the file to finish being written.
versionNoVRM spec version to export: '1.0' (VRoid's default) or '0.0' for the legacy VRM0.0 format.1.0
avatar_nameYesVRM metadata 'Avatar Name' - REQUIRED by VRoid; the Export button stays grey until it is filled.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden, and it excels. It discloses the side effects (opens modals, uses Wine save dialog), the time cost (30s to minutes), the failure behavior (raises with a diagnostic screenshot path), and the recommended retry strategy (read the screenshot, beware of leftover modals). This is unusually rich and actionable, covering the obvious 'what happens to the system' and 'what to expect' questions an agent would have.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than average, but it front-loads the core action ('Walk the entire Export-as-VRM flow and write a .vrm file') and structures the rest as a concise step chain, a prerequisite, a timing note, and a failure/retry note. Each sentence earns its place; there is no filler. It could arguably be trimmed slightly (e.g., the detailed UI path), but the richness contributes to transparency, so a 4 is fair.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having an output schema, the description explicitly states what it returns (the written path and its size) and how it signals failure (a diagnostic screenshot path), which complements the schema. It also covers prerequisites, duration, and the retry hazard. For a complex multi-step GUI automation tool, this is fully complete; an agent has everything needed to invoke it correctly and handle outcomes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, so all five parameters already have explanations (e.g., path is a Linux path translated to Z:\, avatar_name is required for the Export button to activate). The tool description does not add any parameter-specific semantics beyond what the schema provides. Per the rubric, with full schema coverage, the baseline is 3, and the description adds no extra value here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool walks the entire 'Export as VRM' flow and writes a .vrm file, with a verb (walk/export), a specific resource (VRM file), and an explicit step sequence. It distinguishes itself from low-level siblings like vroid_click or vroid_type by being a composite workflow, and from vroid_save_project (which saves a project, not exports a VRM). The first sentence alone conveys the exact purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear prerequisite: 'Must be started from the EDITOR screen with a model loaded.' It also implies when to use it (when you need a .vrm export) and orients the user by narrating the UI steps. It does not explicitly name an alternative tool or describe when NOT to use it, but given the sibling set (most are atomic actions), the usage context is sufficient. A minor gap is the lack of explicit 'use this instead of manual steps' guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vroid_find_buttonA

Find VRoid's primary action buttons by their colour, not their text.

VRoid's confirm buttons ('Export', 'OK', 'Create') are solid #0096FA pills whose white labels tesseract regularly loses, so this matches the button chrome. Blobs come back biggest-first, in image px, ready for vroid_click.

A grey pill (color='disabled') where you expected a blue one means the action is disabled - on the VRM Settings modal that means Avatar Name or Creators is still empty.

ParametersJSON Schema
NameRequiredDescriptionDefault
colorNo'primary' finds the enabled blue #0096FA pill; 'disabled' finds the grey pill, which is the app telling you a required field is still empty.primary
labelNoOptional label to disambiguate when several pills are visible; the button interior is OCR'd at high upscale to check it.
limitNoMax blobs to return.
regionNoRestrict the search to this rectangle (image px).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description fully carries the burden of behavioral disclosure. It reveals the matching mechanism (button chrome), the output ordering (biggest-first, image px), and the semantic meaning of a disabled grey pill, including the specific validation context on the VRM Settings modal. This goes well beyond the schema and gives agents critical operational knowledge.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a compact three sentences, each earning its place: purpose and differentiation (sentence 1), technical rationale and output format (sentence 2), and practical interpretation of a key signal (sentence 3). It is front-loaded with the core purpose and contains no filler, making it highly efficient for agent parsing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is moderately complex with 4 parameters and an output schema, but the description covers the essential context: why this tool exists (OCR failure), how results are returned (blobs, ordered), and what a specific result indicates (disabled state). Combined with the exhaustive schema descriptions, nothing critical for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds little beyond what the schema already explains (e.g., the color enum semantics are already in the schema). It does reinforce the color meanings but does not add new parameter-level detail for label, limit, or region. The description therefore does not elevate above the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: finding VRoid buttons by color rather than text. It specifies the resource (primary action buttons), the method (color matching), and explicitly differentiates from text-based approaches. The mention of 'ready for vroid_click' also clarifies the intended downstream use, making the tool's role unambiguous relative to siblings like vroid_find_text.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides strong usage context by explaining why color-based matching is necessary (tesseract loses white labels), implicitly steering agents away from text-based tools. It also gives guidance on interpreting results (grey pill means disabled) and notes the typical next step (click). However, it does not explicitly name alternative tools or enumerate when NOT to use this tool, which keeps it from a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vroid_find_textA

OCR the current VRoid window and return where query appears.

Takes its own fresh screenshot, so coordinates are current. Results are image px, ordered by OCR confidence, ready to pass straight to vroid_click (which defaults to space='image').

Known OCR weaknesses in this UI: small, letter-spaced or light-on-dark labels get split or dropped ('Export' -> 'E' + 'xport'), and white text on VRoid's blue primary buttons often disappears entirely - use vroid_find_button for those. Icons (toolbar, left rail) have no text at all; the README's UI map has their fractional positions.

If nothing is found, that is information: the screen may not be the one you think it is, or a modal is covering it. Take a screenshot and look.

ParametersJSON Schema
NameRequiredDescriptionDefault
exactNoRequire the whole word to match, not a substring. Use for captions that share a prefix with a heading (e.g. 'Name' vs 'Avatar Name').
limitNoMax matches to return.
queryYesLabel to look for, e.g. 'Export', 'Hairstyle', 'Avatar Name'. Matching is case- and punctuation-insensitive substring by default.
regionNoRestrict OCR to this rectangle (image px). Strongly recommended: OCR of a full 2560x1440 frame takes ~10 s, a panel-sized region under 2 s.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description fully discloses behavior: it takes a fresh screenshot, returns coordinates in image pixels ordered by OCR confidence, and is designed for direct use with vroid_click. It candidly lists known OCR weaknesses ('Export' splits to 'E'+'xport', white text on blue buttons disappears) and performance implications (~10s full frame vs ~2s region). This is exceptionally transparent for a tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose in the first sentence, then flows logically through currentness, output format, limitations, alternatives, and failure handling. Every sentence carries substantive information—no filler. It is long but appropriately dense for a tool with this many behavioral caveats and integration points.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers everything an agent needs to invoke this tool correctly: what it does, output format and ordering, how to feed results into vroid_click, known failure modes (OCR splitting, missing white-on-blue text), performance guidance via region, alternative tool for buttons, and what a null result implies. Since an output schema exists, return details are not required in the description. For a computer-vision OCR tool with multiple integration touchpoints, this is complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and each parameter (query, exact, limit, region) already has detailed descriptions including defaults and examples. The tool description adds little beyond schema—it only mentions that results are ready for vroid_click, which is about output usage rather than parameter meaning. Baseline 3 is appropriate because the schema fully documents parameters and the description does not enhance it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise action and resource: 'OCR the current VRoid window and return where `query` appears.' It immediately distinguishes itself from the sibling vroid_find_button by noting it handles text labels, and even mentions specific limitations for button text. This makes the tool's purpose unmistakable relative to its siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use this tool versus alternatives: use vroid_find_button for blue primary buttons, and it notes that icons have no text so users should consult the UI map. It also advises when no result is found (screen may be wrong or modal covering) and recommends taking a screenshot. This is clear conditional guidance with exclusions and alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vroid_keyA

Press a key combination in the focused VRoid widget.

Useful ones: Return confirms a value box or a Wine dialog's default button; ctrl+s saves the project; ctrl+shift+s is Save As; ctrl+z/ctrl+y undo/redo in the editor.

Note Escape does NOT close VRoid's hamburger menu - click elsewhere to dismiss it.

ParametersJSON Schema
NameRequiredDescriptionDefault
comboYesA key, optionally with modifiers, e.g. 'Return', 'Escape', 'Tab', 'BackSpace', 'ctrl+s', 'ctrl+shift+s', 'ctrl+z'. Names follow X keysyms; 'enter', 'esc', 'space', arrows are aliased.
timesNoRepeat count.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses a notable behavioral nuance: Escape does not close the hamburger menu, and that key presses go to the focused widget. This is sufficient for a key-press tool, though it doesn't describe potential side effects or focus requirements beyond what's implied.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured: the primary objective is stated first, followed by a list of useful combos and a caveat. Every sentence contributes value, and there is no fluff. The layout is easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple key-press tool, the description covers what keys to press, gives practical examples, and highlights a relevant gotcha. It doesn't mention return values, but the output schema likely indicates success/failure, and for this tool it's not critical. It is sufficiently complete for an agent to call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already describes both parameters (combo and times) with examples and aliases, so the baseline is 3. The description adds value by providing concrete, useful key combinations (e.g., 'Return confirms a value box', 'ctrl+s saves'), which goes beyond the schema's generic format explanation and clarifies likely inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'Press a key combination in the focused VRoid widget.' It uses a specific verb and resource, and distinguishes itself from siblings like vroid_click and vroid_type by focusing on key combinations. The context of 'focused widget' adds clarity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides practical guidance by listing useful key combinations (Return, ctrl+s, ctrl+z, etc.) and their effects, which helps an agent decide when to use this tool. It also gives a specific caution about Escape not closing the hamburger menu. However, it does not explicitly mention alternatives or when not to use it, though the examples imply its use for shortcuts and confirmations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vroid_launchA

Start VRoid Studio (Steam appid 1486350, Proton) and take control of it.

Idempotent: if the window already exists it is reused, not relaunched. Then the window is parked on Hyprland workspace 9, the workspace the user was on is remembered (vroid_release puts them back), and the window is focused and fullscreened so its geometry - and therefore every coordinate you will read off a screenshot - is stable.

Cold start over Proton takes 30-90 s; the call blocks until the window is up. It does NOT wait for the start screen to finish drawing, so take a vroid_screenshot and look before clicking anything.

ParametersJSON Schema
NameRequiredDescriptionDefault
restartNoKill a running VRoid Studio first (UNSAVED WORK IS LOST) and start a clean instance.
timeoutNoSeconds to wait for the window to appear.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description fully carries behavioral disclosure. It details idempotency, window parking on workspace 9, focusing and fullscreening, blocking behavior, cold start timing, and the caveat that the start screen may not be ready. This is comprehensive for a launch tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and logically sequenced: purpose first, then idempotency, workspace details, and timing caveat. Each sentence adds necessary information without fluff. Slightly longer than minimal but justified by the behavioral details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and an output schema present, the description covers essential launch behavior, blocking, and post-launch guidance. It does not explicitly state return values, but the output schema likely handles that. It also does not mention error handling, but that is not critical for a launcher.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both parameters have descriptive schema entries (restart warns about unsaved work, timeout defines wait). The description adds context around restart via idempotency but does not materially extend parameter meaning beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Start VRoid Studio (Steam appid 1486350, Proton) and take control of it.' It distinguishes itself from siblings by describing idempotent behavior and workspace parking, making it clear this is the launcher, not any other action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit context: it can be reused if window exists, blocks for 30-90s on cold start, and instructs to take a screenshot before clicking. It names vroid_release as complementary, implying when not to use (though it does not explicitly exclude other tools). The guidance is clear and actionable, though slightly implicit about alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vroid_new_characterA

From the start screen: Create New -> pick a base -> land in the editor.

Locates the 'Create New' tile by text (so it survives the Recently Edited grid growing), clicks the card above its caption, picks the base thumbnail, then waits up to a minute for the editor - the 3D viewport takes several seconds to appear after a base is chosen.

ParametersJSON Schema
NameRequiredDescriptionDefault
baseNoWhich base model the 'Select a base to start with' modal offers.Fem

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It discloses that it locates the 'Create New' tile by text (robust to the Recently Edited grid growing), clicks the card, picks the base thumbnail, and waits up to a minute for the editor to appear, acknowledging the 3D viewport delays. This goes beyond the short intent, but it does not mention potential failure modes or side effects (e.g., what happens if the editor doesn't load). Still, it is substantially transparent for a creation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a compact block of text that clearly front-loads the purpose and then details the steps. Each sentence earns its place: the text-locating technique, the click sequence, and the wait time are all useful for execution. It could be slightly shorter, but it avoids redundancy and remains focused on actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (multi-step navigation with timing considerations) and the presence of an output schema (so return values need not be described), the description covers the essential context: the start screen prerequisite, the steps, and the wait behavior. It does not specify what happens if the start screen is not present or if the base selection fails, but these are edge cases that are not typically required for standard usage. Overall, it is complete for an agent to call the tool correctly under normal conditions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The parameter 'base' is fully described in the schema with an enum and a clear description ('Which base model the 'Select a base to start with' modal offers'). The tool description does not add any additional meaning beyond the schema, which is already explicit. Since schema coverage is 100%, a baseline score of 3 is appropriate; the description does not compensate with extra semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: navigating from the start screen to create a new character by clicking 'Create New', selecting a base, and landing in the editor. It specifies the exact workflow and distinguishes itself from sibling tools like vroid_launch (which likely starts the app) or vroid_open_tab (which opens tabs). The verb-resource pair is specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description sets clear context by starting with 'From the start screen', implying the app must already be launched and on that screen. It does not explicitly mention alternatives or when not to use, but the procedural nature and the prerequisite are evident. It lacks explicit exclusions, but the context is clear enough for an agent to infer appropriate usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vroid_open_tabA

Switch the editor to a top-level tab, located by OCR of the tab strip.

Only works on the editor screen (not the start screen, the export screen or the hair editor). Waits ~2 s for the panels to redraw and returns the path of a verification screenshot - take a vroid_screenshot if you want to see the result.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesOne of Face, Hairstyle, Body, Outfit, Accessories, Look (the editor's top tab strip).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses the ~2 second wait for redraw, the return of a verification screenshot path, and the suggestion to take a screenshot to view results. It does not specify failure modes (e.g., what happens if OCR cannot find the tab), but the disclosed timing and return behavior are useful. No contradictions with annotations (since none exist).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise, with the core purpose in the first sentence and supporting details in a short second paragraph. It front-loads the function and then adds necessary constraints and behavior without any redundant or irrelevant content. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the availability of an output schema (though not shown) and high parameter schema coverage, the description provides sufficient context for an agent to call the tool correctly. It covers the environment limitation, the wait time, and the return value. Missing details include error handling when OCR fails, but for a tool that is part of a larger suite with similar patterns, this is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description covers 100% of the single parameter 'name' with clear allowed values (Face, Hairstyle, Body, Outfit, Accessories, Look). The description does not add additional meaning about the parameter beyond what the schema already provides, which is acceptable given the high schema coverage. The baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Switch'), a clear resource ('top-level tab'), and the method ('located by OCR of the tab strip'). It clearly distinguishes this from sibling tools like vroid_click (which clicks generic elements) or vroid_find_text (which locates text), since it is specialized for tab switching. The screen restriction further refines its purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states the tool only works on the editor screen, excluding start, export, and hair editor screens, which tells the agent when not to use it. However, it does not directly mention alternatives or contrast with sibling tools like vroid_click when a tab is already visible. The guidance is clear on the environment constraint but lacks explicit 'use this instead of X' phrasing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vroid_releaseA

Hand the desktop back: switch to the workspace the user was on before.

Leaves VRoid running on workspace 9. Call this when you are done with a session, or before handing control back to the human.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It discloses the workspace switch and that VRoid remains running, which is the primary behavior. It does not mention potential side effects like saving state or closing dialogs, but these are likely irrelevant for a simple release action. Slightly more detail would earn a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action and usage context. Every word earns its place, with no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with an output schema present, the description covers the purpose, the behavioral outcome, and the appropriate invocation time. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline is 4. No parameter descriptions are needed, and the description adds sufficient context for how the tool behaves regardless of inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise action: 'switch to the workspace the user was on before' and clarifies that VRoid remains running on workspace 9. This clearly distinguishes it from siblings like vroid_launch or vroid_open_tab, and leaves no ambiguity about the tool's core function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to call: 'Call this when you are done with a session, or before handing control back to the human.' This gives clear situational context and implies it should not be used mid-task, providing strong usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vroid_save_projectA

Save the .vroid project - plain Save, or Save As to an explicit path.

With name: presses Ctrl+Shift+S and drives Wine's save dialog the same way the VRM export does, then waits for the file. Without name: opens the hamburger menu and clicks Save, which overwrites the project's existing file and opens the Wine dialog only if the project has never been saved (in that case call this again WITH a name).

Worth doing before any risky experiment: nothing else in this server persists your work, and vroid_launch(restart=true) discards it.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoSave As target: a bare name (written into the server's out dir as <name>.vroid) or an absolute Linux path. Omit to do a plain Save, which silently overwrites the project's existing file.
timeoutNoSeconds to wait for the file to be written.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description bears full responsibility. It discloses that plain Save silently overwrites the existing file, that Save As opens Wine's save dialog, and that without a name the dialog only appears if never saved. It also mentions it waits for the file and explains the persistence context. This is transparent and goes beyond mere operation to side effects and prerequisites.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: it states the core purpose first, then explains the two modes, then the critical warning. Each sentence earns its place, with no redundant phrasing or filler. The structure makes it easy for an agent to parse and act on.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with two modes and edge cases (unsaved project), the description covers everything needed: mode selection, what happens with and without a name, the dialog behavior, and the importance relative to other tools. It also addresses persistence and the destructive nature of vroid_launch(restart=true). No critical information is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already covers both parameters fully (100% coverage) with descriptions. The tool description adds value by explaining the behavioral difference between supplying `name` (Save As) vs omitting it (plain Save), and clarifies that omitting it opens the dialog only if never saved. This adds operational meaning beyond the schema's straightforward parameter descriptions, so it exceeds the baseline of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description starts with 'Save the .vroid project - plain Save, or Save As to an explicit path' which clearly states the tool's specific action and distinguishes the two modes. It is distinct from siblings (no other save tool) and gives a concrete resource (the .vroid project).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly provides usage context: 'Worth doing before any risky experiment...' and explains that nothing else persists work and vroid_launch(restart=true) discards it. It also details when to use name (Save As) vs plain Save, including the caveat about the dialog appearing only when never saved and the instruction to retry with a name. This is clear, actionable guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vroid_screenshotA

Screenshot VRoid and return the image, plus the path it was saved to.

LOOK at the returned image before you decide anything - this is the only way to see the app. Every capture is also written to the captures dir so it can be re-read later.

Coordinate space: the reported image_size is the native capture size (2560x1440 on the reference machine) and that is the space every other tool means by space="image". The transported image may be downscaled (downscale in the text block says by how much); if you read a coordinate off the picture by eye, divide it by that factor before clicking. Better: get coordinates from vroid_find_text / vroid_find_button, which always report native image px.

Gotchas carried over from the driver: capture the WHOLE window, not a crop, when checking "did that work" - modals appear in the middle of the screen and a top-strip-only check will miss them. And do not judge change by the 3D viewport, which dithers every frame; watch a UI strip instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagNoShort label used in the saved filename.
regionNoOptional crop in image px of the window capture. Omit for the whole window.
whole_screenNoCapture the whole output instead of just the VRoid window - needed for the Wine save/export dialog, which is a separate window.
full_resolutionNoReturn the image at native resolution instead of downscaling it for transport. Large.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full behavioral burden and does so thoroughly. It discloses that every capture is saved to a captures directory, returned images may be downscaled, coordinates use native image pixel space, and there are driver quirks around modals and viewport dithering. This is unusually transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every section earns its place: purpose, return behavior, coordinate system, and practical gotchas. It is front-loaded with the core action and then layers important operational details so an agent can use the tool correctly without skimming irrelevant prose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a screenshot tool with no output schema and no annotations, this is complete: the agent learns what is returned, where it is written, how coordinates should be scaled, what mode to use for verification, and what UI regions are reliable to observe. I see no important calling context missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already describes all four parameters, so the baseline is 3; the description adds meaningful coordinate-space and downscale context that directly affects how region and full_resolution should be interpreted. It does not separately expand on tag or whole_screen, but the schema handles those sufficiently.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Screenshot VRoid and return the image, plus the path it was saved to.' This clearly states the tool's purpose and main outputs, and the later coordinate-space discussion differentiates it from related perception tools like vroid_find_text and vroid_find_button.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit usage context: look at the returned image before deciding anything, capture the whole window when verifying changes, avoid relying on the dithered 3D viewport, and prefer vroid_find_text / vroid_find_button for coordinates. This goes far beyond a generic description and routes the agent to the right behavior and alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vroid_scrollA

Wheel-scroll a panel or zoom the 3D viewport.

Sent as X11 button 4/5 (6/7 horizontally): VRoid is an XWayland client and ignores Wayland virtual-pointer axis events, so this is the only wheel that works on it.

Long Parameters lists need this - a label that vroid_find_text cannot see is usually just below the fold. vroid_set_slider scrolls to its own row automatically, so you rarely need to do it by hand.

ParametersJSON Schema
NameRequiredDescriptionDefault
xNoPark the pointer here first - the wheel goes to whatever is under the cursor, so this decides WHICH panel scrolls. ~(0.93w, 0.60h) is the right-hand Parameters panel, the centre is the 3D viewport.
yNoSee `x`.
dxNoHorizontal wheel notches, positive = right. Rarely useful; VRoid's panels scroll vertically only.
dyNoVertical wheel notches. POSITIVE scrolls DOWN (further into a panel); over the 3D viewport, positive zooms OUT.
spaceNo'image' = pixels of a window capture (what vroid_screenshot / vroid_find_* report - the default, and almost always what you want); 'window' = Hyprland layout units relative to the window's top-left; 'layout' = absolute Hyprland layout units of the whole output.image

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It explains that the wheel is sent as X11 button 4/5 (6/7 horizontally) because VRoid is an XWayland client that ignores Wayland virtual-pointer axis events. It also reveals the direction semantics: 'POSITIVE scrolls DOWN' and 'over the 3D viewport, positive zooms OUT.' This is rich, honest behavioral context beyond a simple 'scroll' statement, with no contradiction to any annotations (none exist).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences with zero waste. It front-loads the purpose, then explains the platform-specific mechanism, then gives a concrete use case with a reference to alternatives. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity, the rich input schema (with parameter descriptions), and an output schema, the description is fully sufficient. It explains why the tool exists (XWayland limitation), when to use it (scrolling long panels), and how it contrasts with siblings. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so per the rubric the baseline is 3. The description does not add parameter-specific meaning beyond the schema; it focuses on when to use the tool and the underlying mechanism. The schema already contains detailed descriptions for each parameter (e.g., x: 'Park the pointer here first...'), so the description's lack of direct parameter elaboration is acceptable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource combination: 'Wheel-scroll a panel or zoom the 3D viewport.' It immediately differentiates from siblings by noting that vroid_set_slider scrolls automatically and vroid_find_text cannot see labels below the fold. This clearly tells an agent what the tool does and how it differs from related tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit when-to-use context: 'Long Parameters lists need this - a label that vroid_find_text cannot see is usually just below the fold.' It also explicitly states when not to use it: 'vroid_set_slider scrolls to its own row automatically, so you rarely need to do it by hand.' This gives clear guidance on selecting this tool over alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vroid_set_colorA

Set a colour swatch by typing a hex code into its #RRGGBB box.

Scrolls the right-hand panel to the labelled row, clicks the hex field just below the label, replaces its contents and presses Return. Same caveat as sliders: the row must belong to the currently open tab and sub-category (left icon rail).

ParametersJSON Schema
NameRequiredDescriptionDefault
hexYesColour as '#RRGGBB' or 'RRGGBB'.
labelYesThe colour row's label, e.g. 'Main Color', 'Highlight Color', 'Base Color'.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full responsibility of behavioral disclosure. It transparently describes the step-by-step mechanism: scrolling the right-hand panel, clicking the hex field, replacing contents, and pressing Return. It also discloses the limiting condition about the currently open tab. This is a solid level of transparency for a UI automation tool, though it does not mention potential failure modes (e.g., label not found) or side effects beyond the described actions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is exceptionally concise—two sentences—and front-loads the purpose in the first sentence. The second sentence packs in the operational steps and the caveat with no wasted words. Every element contributes to understanding, and nothing is redundant with the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the essential aspects of the tool: its purpose, the detailed interaction sequence, and a critical precondition (correct tab and sub-category). It does not explain error handling or what happens if the label is missing, but given that an output schema exists (though not shown) and the tool is relatively simple, the description is sufficiently complete for an agent to use it effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Both parameters have schema descriptions with examples, so the baseline is 3. The description adds extra meaning by explaining how each parameter is used in the operation: 'label' identifies the row to scroll to, and 'hex' is the value typed into the field. This contextualizes the parameters beyond their simple data type definitions, making the tool easier to invoke correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear, specific action: 'Set a colour swatch by typing a hex code into its #RRGGBB box.' It identifies the resource (colour swatch) and the method (typing hex code), and the focus on color differentiates it from sibling tools like vroid_set_slider. The purpose is unambiguous and actionable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a caveat ('Same caveat as sliders: the row must belong to the currently open tab and sub-category') which gives some usage context, but it does not explicitly state when to use this tool versus alternatives like vroid_set_slider or vroid_click. There is no direct 'use this when...' or 'do not use when...' guidance, leaving the agent to infer the appropriate scenario from the tool's purpose.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vroid_set_sliderA

Set a Parameters slider exactly, by typing into its numeric box.

Scrolls the right-hand panel from the top until the labelled row is visible (so it works for parameters below the fold), clicks the numeric box at the end of that row, clears it, types the value and presses Return.

This is far more reliable than dragging the handle - use vroid_drag only for controls with no numeric box. Open the right tab first (vroid_open_tab): each tab has its own parameter list.

ParametersJSON Schema
NameRequiredDescriptionDefault
labelYesThe parameter's label as printed in the right-hand panel, e.g. 'Fem Height', 'Head Size', 'Eye Size X'.
valueYesValue to type into the row's numeric box. Most VRoid parameters run -1.0..1.0 with 0 centred; the app clamps out-of-range values.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses the exact sequence of actions (scrolls, clicks, clears, types, presses Return) and notes that it works for parameters below the fold. It does not cover failure behavior, but the described mechanics are transparent and complete for normal usage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight paragraphs. The first sentence states the purpose directly. The second details the method without fluff. The third gives usage distinction and prerequisite. Every sentence earns its place; no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a UI automation tool with two parameters and an output schema, the description covers the core usage: what it does, how it works, when to use it, and the prerequisite tab. It omits error handling and return-value details, but those are not critical for a correct first call. Minor gap, but overall sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%—both parameters have descriptive comments. The description does not add extra meaning beyond what the schema already provides (e.g., 'Value to type into the row's numeric box' matches). Baseline of 3 is appropriate; no further compensation needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Set') and resource ('Parameters slider') with the method ('typing into its numeric box'). It explicitly distinguishes from sibling vroid_drag by saying 'use vroid_drag only for controls with no numeric box', making it clear which tool to select.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear when-to-use ('This is far more reliable than dragging the handle'), when-not-to-use (use vroid_drag for controls without numeric box), and a prerequisite ('Open the right tab first (vroid_open_tab)'). No ambiguity about context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vroid_statusA

Report whether VRoid Studio is running, focused, and where it sits.

Cheap and side-effect free (no OCR, no input) - use it before anything else, and after anything that might have lost focus. focused false means every acting tool will refuse until vroid_launch (or any acting tool's own focus step) takes the window back.

Also reports whether the external helpers this server needs are present: the vpointer binary (built by native/build.sh), grim, tesseract, hyprctl.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It explicitly states that the tool is 'Cheap and side-effect free (no OCR, no input)' and describes the side effects of the `focused` flag on other tools, as well as the external helper dependencies it reports. This is transparent and thorough, even anticipating edge cases.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured into three clear parts: primary purpose, usage guidance, and helper reporting. It front-loads the most critical information (status report) and continues with actionable context. Every sentence adds value—'Cheap and side-effect free' justifies frequent use, and the helper list tells the agent what to expect. No fluff or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with an output schema (which covers return format), this description provides all necessary context: what it reports, when to use it, side effects on other tools, and external dependencies. It even explains how `focused` false affects subsequent acting tools and how to recover. The tool is simple, and the description fully covers its use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline for this dimension is 4. The description appropriately focuses on behavior and usage rather than parameter details. No parameter explanations are needed, and the description does not introduce any parameter ambiguity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states exactly what the tool does: 'Report whether VRoid Studio is running, focused, and where it sits.' The verb 'Report' is specific, the resource (VRoid Studio) is clear, and the scope is distinct from the acting sibling tools (vroid_launch, vroid_click, etc.), which perform actions rather than status checks.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit guidance on when to use the tool: 'use it before anything else, and after anything that might have lost focus.' It also explains the consequence of `focused` false and names the alternative (vroid_launch) that can restore focus. This is clear, actionable usage guidance that distinguishes it from other tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vroid_typeA

Type into whatever widget currently has keyboard focus.

Click the field first (vroid_click) - this tool has no idea where the caret is. Keystrokes go through X11 XTEST because the Wayland virtual keyboard is mis-read by this Proton client (a whole string arrives as a single character).

Always screenshot afterwards to confirm the text landed in the field you meant: VRoid's forms have several boxes with near-identical captions, and typing into the wrong one leaves the primary button greyed out with no other symptom.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesLiteral text to type. '\n' presses Return.
clear_firstNoSelect-all + backspace before typing, so the field is replaced rather than appended to.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It discloses the XTEST mechanism and the Proton mis-read issue, directly warns about caret independence, and explains the symptom of wrong-field typing. It also notes the effect of clear_first semantically. This is exemplary transparency for an input tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-organized: purpose statement, usage direction, and verification step. It front-loads the critical warning about clicking first. Slightly verbose but every sentence adds operational value, so it earns a high score rather than a 3.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (keyboard input on a misbehaving client), the description covers prerequisites, mechanism, verification, and failure modes. Output schema exists, so return value details aren't needed. All information an agent needs to call it correctly and detect errors is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both parameters have descriptions. The text parameter's escape sequence ('\n' for Return) is already in the schema. The description adds context about literal text and the clear_first behavior, which is beyond the schema's simple description. A high score because the description reinforces the meaning and clarifies usage in context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (type), resource (keyboard focus), and clearly distinguishes from siblings (vroid_click, vroid_key). The purpose is unmistakable: type text into the currently focused widget. The caveat about having no idea where the caret is clarifies its scope versus click.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs to click the field first (naming vroid_click as prerequisite), warns about XTEST versus Wayland issue, and mandates a screenshot afterwards to verify. It also identifies a specific failure mode (grayed-out primary button) and provides the verification step. This is complete when-to-use guidance with clear sequencing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 18 tool updatesv0.1.0
    • First observedvroid_click
    • First observedvroid_current_screen
    • First observedvroid_drag
    • First observedvroid_export_vrm
    • First observedvroid_find_button
    • First observedvroid_find_text
    • First observedvroid_key
    • First observedvroid_launch
    • First observedvroid_new_character
    • First observedvroid_open_tab
    • First observedvroid_release
    • First observedvroid_save_project
    • First observedvroid_screenshot
    • First observedvroid_scroll
    • First observedvroid_set_color
    • First observedvroid_set_slider
    • First observedvroid_status
    • First observedvroid_type

TDQS

A4.4/5.0

Scored across 18 tools

Disambiguation5/5

Each tool addresses a distinct capability: launching, screen identification, OCR, clicking, typing, saving, exporting, etc. Even the 'seeing' tools (screenshot, find_text, find_button) have clearly separated purposes: one captures, one finds text, one finds buttons. The interaction primitives (click, drag, scroll, type, key) are mutually exclusive and well-defined. No two tools could be confused for the same action.

Naming Consistency4/5

All tools share the 'vroid_' prefix and mostly follow a verb_noun pattern (open_tab, export_vrm, set_slider), but a few are single verbs or nouns (status, release, screenshot, click, type). The naming is predictable and readable, with only minor deviations like 'status' and 'release' being not strictly verb_noun. Overall consistent enough for an agent to infer actions.

Tool Count4/5

18 tools is on the higher side but justified for a GUI automation server that needs primitives for every interaction type plus higher-level workflows like export and save. The count feels well-scoped for the domain—each tool earns its place since there's no redundant functionality. Slightly heavy but still reasonable.

Completeness5/5

The surface covers the full lifecycle: launch, status, screen detection, navigation (open_tab, scroll), interaction (click, drag, type, key, sliders, colors), inspection (screenshot, OCR), persistence (save_project), export (export_vrm), and creation (new_character). No obvious gaps—even edge cases like disabled buttons and modal detection are addressed. The server appears fully equipped for its stated purpose of automating VRoid Studio.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI agents to control the Blockout previs desktop app for AI filmmaking, allowing staging of 3D worlds, character animation, camera framing, timeline control, and viewport screenshotting through MCP tools.
    6
    Apache 2.0
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI assistants to control a desktop virtual character (VRM) by playing animations, showing/hiding the character, and checking runtime status through the MCP protocol.
    562,949 npm
    2
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables Windows MCP clients to author MikuMikuDance scenes by controlling characters, cameras, lighting, physics, timeline editing, MME effect assignments, and AVI video output through natural language or scripted tool calls.
    30
    MIT No Attribution