Skip to main content
Glama

Studio Live

Real-time Roblox Studio for AI agents. One MCP server, nine tools, an agent runtime that lives inside Studio.

CI MIT License Node ≥ 20 MCP

Install · Quick start · Tools · Works with · Docs · Security

Studio Live connects an AI coding agent to a running Roblox Studio over a single WebSocket. Instead of one tool call per property, the agent ships whole Luau programs against a resident API, gets errors and assertions pushed into its context, keeps playtests alive while it hot-patches code, reaches Open Cloud for the open place, and asks a vision sidecar visual questions that come back as text.

  • Fast where it matters — 16 ms round trip (one Heartbeat frame), a hot-patch in a live playtest in ~17 ms, screenshots in ~30 ms even while Studio is minimised. Every number is measured; see Why it feels fast.

  • One undo step per program — an edit-DataModel run is one Ctrl+Z for the human and rolls back on error.

  • Push, not poll — errors, assertions, milestones, playtest, peer and job events arrive as they happen.

  • Playtests stay up — install controllers that run at Heartbeat in-engine, hot-patch scripts, push instances, run_until a predicate — no stop/start per change.

  • Multi-agent, multi-Studio — several agents share one Studio with FIFO writes; several Studio windows (server + N clients) form a multiplayer playtest.

  • Geometry is enforced — overlapping parts and "parts in parts" are detected on every edit and can be rejected outright.

  • Any MCP client — Claude Code, Codex CLI, Cursor, Claude Desktop, or a plain shell.

The design record is in docs/architecture.md, the wire contract in docs/protocol.md, the engine measurements it rests on in docs/live-measurements.md. Agents should read docs/agent-guide.md; the cloud and look tools have their own pages, docs/cloud.md and docs/vision.md.

Works with any MCP client

It is a plain stdio MCP server. Nothing in the plugin, the protocol or the tools is tied to one client; only the way pushed events reach the model differs.

Client

Tool calls

Push events

look (vision)

Claude Code

stdio MCP (claude mcp add …)

its Monitor tool subscribes to ws://127.0.0.1:47800/events

Anthropic API credential, or your Claude Code login

Codex CLI · Cursor · Claude Desktop · any MCP client

stdio MCP

the events tool (long-poll), or subscribe to /events yourself

Anthropic API credential, or a claude CLI on PATH

Shell, scripts, other agents

studio-live call <tool> or POST /rpc

/events WebSocket

same

Several clients can run at once — the first bridge owns Studio, later ones proxy through it automatically. look is the one tool with a model behind it and today it only speaks to Anthropic models (the backend is the VisionProvider interface in bridge/src/vision/types.ts). Everything else works with no AI-vendor credential at all.

Related MCP server: roblox-studio-mcp

Why it feels fast

Every number below was measured on one machine (Studio 0.738, Windows 11); see docs/live-measurements.md.

Cost

Existing connectors

Studio Live

Measured

Transport wait per hop

community plugin polls every 0.5 s: mean 250 ms, worst 500 ms, 2 req/s dispatch

one persistent WebSocket, push both ways

echo RTT p50 16.5 ms, max 17.7 ms (= one Heartbeat frame); 200 requests fired in one frame all answered in the next

Changing code in a playtest

stop → edit → start ≈ 2.7 s plus two LLM turns

the playtest stays up; Script.Source written + Disabled toggled

hot-injected Script/LocalScript ran within 16.6 / 16.9 ms

Screenshot

in-engine CaptureService: 649 / 592 ms (≈1.5 fps), never fires while minimized

Win32 PrintWindow on the Studio HWND from the bridge

28–44 ms for 1734×1399, works while occluded; ≈100 ms round trip including bicubic resize to 1024 px and JPEG

Structured observation

one tree read or screenshot per LLM turn

pushed events + diff since a cursor

full 2 381-instance snapshot: walk 2.6 ms + JSON 0.55 ms = 114 KB

LLM turns per task

2–10 s per micro-tool, ten tools for "walk to the door and try it"

one turn ships a Luau program or a controller that runs at Heartbeat in-engine

the turn itself cannot be removed, only the count

The turn is the real cost, so the design (1) moves closed loops into Studio (controllers, run_until predicates, in-engine assertions), (2) makes each turn do more (programs against a resident S API, saved skills), and (3) streams observations to the agent instead of making it ask.

Requirements

  • Windows 10/11. The bridge and the plugin are platform-neutral Node and Luau, but the screenshot worker is Win32 (PrintWindow) and install / twin look up Windows paths, so macOS is not supported yet.

  • Node.js ≥ 20.

  • Roblox Studio with Load User Plugins In Run Modes on (File → Studio Settings → Studio).

  • Optional: an Anthropic credential or a Claude Code login for look; a Roblox Open Cloud API key for cloud.

Install

Full walkthrough with the reasons behind each step: docs/install.md.

git clone https://github.com/gurmyd/roblox-studio-live.git
cd roblox-studio-live
npm install
npm run build            # tsc + copy worker.ps1 into dist + pack plugin/bootstrap.luau -> dist/StudioLive.rbxmx
npm run install:studio   # copies the plugin into %LOCALAPPDATA%\Roblox\Plugins and prints the next steps

Then, once: restart Roblox Studio. The edit DataModel loads plugin files only at start; after that the runtime is pushed by the bridge on every connect, so updating this package never needs another restart.

Claude Code

The install step prints the exact line with the absolute path:

claude mcp add studio -- node "<absolute path>\dist\bridge\cli.js" serve

Recommended .claude/settings.json, so tool calls and the push monitor run without prompts:

{ "permissions": { "allow": ["mcp__studio__*", "Monitor"] } }

Codex CLI

~/.codex/config.toml:

[mcp_servers.studio]
command = "node"
args = ["C:\\path\\to\\roblox-studio-live\\dist\\bridge\\cli.js", "serve"]
# tool calls longer than Codex's tool timeout return a job handle; use job { action: "wait" }

Cursor / Claude Desktop / other JSON-configured clients

{ "mcpServers": { "studio": { "command": "node", "args": ["C:\\path\\to\\roblox-studio-live\\dist\\bridge\\cli.js", "serve"] } } }

Quick start

In a Claude Code session with Studio open on a place:

  1. Arm push once — events then arrive in context without asking: Monitor({ ws: { url: 'ws://127.0.0.1:47800/events' }, persistent: true }) (other clients: call events { since: 0, timeout_ms: 25000 } when you want a batch)

  2. observe { what: "status" } — confirms the hub is connected, which DataModels are alive, capabilities.

  3. run { dm: "edit", undo_label: "agent: first part", code: "return S.path(S.part{ Name='Hello', Size=Vector3.new(4,1,4), CFrame=CFrame.new(0,3,0) })" } — one undo step the human can Ctrl+Z.

  4. playtest { action: "start", mode: "play" }, then playtest { action: "install", dm: "client:1", name: "walker", code: "<controller>" } — the controller plays and reports assertions as events while you keep editing.

Worked examples for every acceptance test (T1–T4) are in docs/agent-guide.md. A record of six agents building a playable game concurrently through the bridge is in docs/multi-agent-build-report.md.

Tools

Tool

Purpose

run

Execute a Luau program against the resident S API in edit (one undo step, rolled back on error) or server/client:N (ephemeral). code or code_file (the bridge reads the file — no shell/JSON escaping touches Luau). Returns value, captured output, change counts, duration, plus warnings (a :Destroy( call in an edit-DM program is not undoable) and unwritable (properties the plugin VM cannot set).

observe

Read-only: status, tree (an explicit root is always returned, n: 0 when empty), props, find, diff (since cursor), logs (dm: "all" merges every DataModel; play-DM startup prints have real seqs), script (path, from, to line ranges), stats, selection, player, screenshot, windows, geometry, selftest.

playtest

start (play / run / multiplayer with players client Studios) / stop / status / add_players; run_until (predicate evaluated in-engine; predicate_file); install/uninstall/list controllers (code_file; persist: true is stored by the bridge on disk and re-installed on every peer hello — survives runtime, bridge and Studio restarts; an unsaved place, placeId 0, is memory-only); hotpatch a script's source in a live DataModel (source_file); push edit-DM instances into the live test (replace default true: a same-named, same-class sibling at the target is removed first, replaced: n; Terrain, the camera and player characters are never touched, skipped).

input

Human-like input sequences in the play client (keys, clicks, move, look, text, wait, focus; no scroll — an engine limit) through UserInputService:CreateVirtualInput().

events

Long-poll backfill of the event journal (since, kinds, timeout_ms ≤ 50000) — the portable fallback to Monitor push.

skills

Luau program library on disk: list/get/save (source_file)/delete/run; eight read-only builtins ship with the bridge (settle_physics, device_sim, profile_scripts, bulk_attributes, insert_asset, lighting_preset, list_scripts, remote_map), overridable by name.

job

status/cancel/wait/list for long operations that returned a handle.

cloud

Roblox Open Cloud for the open place: datastore / ordered / memory / message / info / publish / asset_upload / asset / luau / instance / restriction / notify. publish makes a saved place file the live version; luau and instance act on that published place. info what:"key" reads what the API key may do — writes included — from Roblox's key introspection. Universe, place and creator ids default from the connected Studio session; the API key is read per call from ROBLOX_OPEN_CLOUD_KEY or <STUDIO_LIVE_HOME>/opencloud.json / .key (docs/cloud.md).

look

Vision sidecar: {question} screenshots Studio and answers in text through a Claude vision model, so no image enters the agent's context; {watch: {question, interval_s, stop_when}} keeps capturing, skips unchanged frames and streams vision events to /events; list / stop (docs/vision.md).

observe, events and look are annotated read-only so Claude Code runs them in parallel with other calls; job is not (cancel rolls back an edit-DM recording). cloud and look are the only tools that leave the machine (openWorldHint: Open Cloud, the Claude API). With two Studio windows connected, write tools require session and reads carry a session_note.

look reaches its vision model one of two ways. With an Anthropic API credential in the bridge's environment (ANTHROPIC_API_KEY / ANTHROPIC_AUTH_TOKEN, or an ant auth login profile) it calls the Claude API directly — about 2 s per look. Without one, a logged-in Claude Code install is enough: when claude is on the bridge's PATH the sidecar runs claude -p --output-format json per frame on your subscription (~10–15 s per look, watch interval at least 15 s, usage counted against your plan's limits rather than an API bill). STUDIO_LIVE_VISION_PROVIDER = auto | api | claude-cli chooses (default auto: the API when a credential resolves, else the CLI), and every answer and vision event names the provider that served it — see docs/vision.md.

Programs in play DataModels get the same S API plus S.ensure(path, class, props?) (create-if-missing for roots shared between agents), S.script.get(path, {from, to}) line ranges, and in controllers ctx.pathTo(pos, timeout) (PathfindingService) next to the straight-line ctx.moveTo. Several agents can share one Studio: edit-DM writes queue FIFO (no busy until 500 are waiting; measured live with 30 and 100 concurrent programs), and docs/agent-guide.md has the multi-agent contract they follow.

Geometry is enforced ("parts in parts"): every edit-DM run checks the parts it added or moved through S for intersections with other parts (touching faces are fine) and for Parts parented under Parts, and reports them as geometry + warnings; geometry_policy: "reject" (or STUDIO_LIVE_GEOMETRY_POLICY=reject) rolls such a run back with geometry_violation, observe { what: "geometry" } audits any subtree, and S.placeOn / S.fits / S.overlaps make correct placement one line — see the agent guide's "Geometry rules (enforced)".

Command line

Everything the tools do is also reachable from a shell through the running bridge (POST /rpc):

studio-live serve                                   # the bridge + MCP server on stdio (what Claude Code launches)
studio-live install                                 # install the bootstrap plugin (STUDIO_LIVE_PORT baked in)
studio-live status                                  # GET /status of the running bridge (sessions, journal, jobs)
studio-live call <tool> [json | -] [--raw] [--port N] [--timeout ms]
studio-live call run --code-file .\build.luau --args-file .\build.args.json
studio-live call playtest '{"action":"hotpatch","dm":"server","path":"ServerScriptService.Main"}' --source-file .\Main.server.luau
studio-live call playtest '{"action":"run_until","dm":"server"}' --predicate-file .\ready.luau
studio-live sync <dir> [--pull] [--once] [--no-hotpatch] [--port N]
studio-live twin <place.rbxl> [--port N] [--timeout ms] [--exe C:\...\RobloxStudioBeta.exe]
  • call performs one tool call and prints the tool's text (--raw: the whole MCP result JSON); exit code 1 when the result isError or no bridge answers. --args-file <json> supplies the whole argument object; --code-file, --source-file, --predicate-file read a Luau file (UTF-8, BOM stripped) into code / source / predicate — the supported way to hand Luau to the tools from scripts and subagents, because shell heredocs turn \n inside Luau strings into real newlines (the resulting Malformed string error says so).

  • sync mirrors a folder of .luau files into the open place (or the place into the folder with --pull) — docs/sync.md.

  • twin launches a second Roblox Studio on a local place file (the newest RobloxStudioBeta.exe under %LOCALAPPDATA%\Roblox\Versions\version-*\, then a per-machine install under %ProgramFiles(x86)%\Roblox\Versions; --exe <path> or STUDIO_LIVE_STUDIO_EXE skips the scan; spawned detached with the .rbxl as its argument), waits up to 90 s (--timeout ms) for the new session to show up in GET /status, and prints its session id and place. A Studio that cannot start, or exits before connecting, is reported immediately (exit 1). **Multiple sessions:** each Studio process is its own session; while more than one is connected, pass session: "<guid or unique prefix>" on tool calls — reads default to the active session and say so (session_note), writes (run, playtest, input, skills run) refuse to guess. Naming a session once makes it the active one.

Troubleshooting

Screenshot fails with minimized. PrintWindow cannot render a minimized window (Studio suspends its render loop). By default the bridge un-minimizes it without taking focus (ShowWindow(SW_SHOWNOACTIVATE) and a SetWindowPos to the bottom of the Z-order) and waits 400 ms; pass restore: false to refuse instead. A window that was maximized before being minimized comes back at its normal size — that is a Win32 limit of non-activating restores. Occluded windows capture fine; you do not need Studio in front.

Screenshot fails with no_window. No visible top-level window of a RobloxStudioBeta process has "Roblox Studio" in its title: Studio is not running, is still on the splash screen, or title_match did not match (the message lists the titles it saw). With several Studio processes the foreground one wins, else the topmost in Z-order; within a process the largest window wins (the main window over floating docks). observe { what: "windows" } lists them; pass hwnd from that list to pin one. The first screenshot after a bridge start can take a few seconds while the PowerShell worker compiles its Win32 shim; the bridge warms it up in the background and does not count that time against the 10 s capture timeout.

playtest start says no_peer. The play server/client DataModels never said hello. Open File → Studio Settings → Studio and turn Load User Plugins In Run Modes on (Faster Play Solo disables user plugins in test mode by default), and check that StudioLive.rbxmx is installed and Studio was restarted after installing it. observe status shows peers once they connect.

Port 47800 is in use. The bridge listens on 127.0.0.1:47800 (STUDIO_LIVE_PORT overrides; Studio connects to the literal 127.0.0.1). A second MCP process on the same machine must not fight for the socket: it detects the primary and runs in proxy mode, forwarding tool calls over POST /rpc to the primary, which owns the single Studio connection; when the primary exits it first waits for jobs the proxy started, and the proxy then takes the port over on its next call. If the port is held by something else entirely, either free it or pick another: the plugin reads no settings, so set STUDIO_LIVE_PORT in the environment, run npm run install:studio again (it bakes the port into StudioLive.rbxmx), restart Studio once, and register the MCP server with --env STUDIO_LIVE_PORT=<port> (the install output prints the exact line).

syntax_error: Malformed string on code that is valid Luau. The program reached Studio with a real newline inside a quoted string: a shell heredoc or a JSON layer collapsed \\n to \n before the bridge saw it (the bridge itself is escape-clean — its selftest compiles "\n" and "[^\n]+" inside Luau strings). The error message appends the hint "your transport turned \n into a newline — pass code from a file (code_file)" when that is what happened. Write the program to a file and pass code_file / source_file / predicate_file (or studio-live call --code-file); never work around it with string.char(10).

Persisted controllers vanished after a restart / playtest list shows none. Persistence lives in the bridge (<STUDIO_LIVE_HOME>/persist/<placeId>.json, default home ~/.studio-live), not in Studio: plugin:GetSetting returns nil for every key on Studio 0.738. The bridge re-sends the list to the hub on every hello, so entries return when the bridge that stored them (or one sharing its home directory) is running; an entry is removed by uninstall or by a later install of the same name without persist. An unsaved place (placeId 0) has no file: its entries are memory-only and gone after a bridge restart. persist_note instead of persist_file on an install means the bridge could not write the file (read-only home, disk full) — the entry lives in memory for that bridge process only. Two Studios on the same placeId (a twin) share the file, which holds the union of both sessions' lists.

Monitor is unavailable (Claude Code). Monitor is gated off when DISABLE_TELEMETRY / CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC is set, on Bedrock/Vertex/Foundry, in non-interactive runs, and in some clients. Use the events tool instead: events { since: <last seq>, timeout_ms: 25000 } long-polls the same journal, and a PostToolBatch hook can attach "events since cursor" to every tool round trip. Also call events whenever a pushed frame shows dropped > 0 or a seq gap.

Nothing happens after editing plugin/bootstrap.luau. Studio does not load a newly written plugin file while running. Only the bootstrap lives on disk; iterate on runtime code through the bridge (pushed on connect) and restart Studio only when the bootstrap itself changes.

Development

npm run typecheck                 # tsc --noEmit
npm test                          # vitest (tests/**/*.test.ts); capture tests use a fake worker, no Studio needed
npm run build                     # dist/bridge/*.js + worker.ps1 + dist/StudioLive.rbxmx + dist/StudioLive.lua
npm run selftest                  # offline end-to-end: real bridge on a random port + scripts/fake-hub.mjs, PASS/FAIL per protocol check
npm run luau:check                # parse/type-check plugin/**/*.luau with luau-lsp (fetches Roblox definitions once)
npm run fake-hub -- --port 47800  # pretend to be Studio against a running bridge (handshake, events, canned answers)
node scripts/capture-smoke.mjs    # lists Studio windows and captures two frames with timings (Studio must be open)

Layout: bridge/src (Node bridge: MCP server, sessions, journal, jobs, skills, capture/ worker, cloud/, vision/, sync/), plugin/bootstrap.luau (the only code installed into Studio) and plugin/runtime/** (the hub/agent runtime the bridge pushes on connect; see docs/luau-runtime.md), skills/builtin (Luau programs shipped with the bridge), scripts/ (build, packaging, selftest, smoke), docs/, tests/.

The bridge's stdout is the MCP stdio transport; every log line goes to stderr. Screenshots land in %TEMP%\studio-live\frames (newest 200 kept). The rules that keep the live system working (the bootstrap is frozen, tests never bind 47800, keys are never printed) are in AGENTS.md.

Documentation

Page

What it covers

docs/agent-guide.md

How an agent should work with the tools: working style, worked examples, the multi-agent contract, geometry rules, look vs observe

docs/architecture.md

The design record: what was built and why

docs/protocol.md

The wire contract between bridge and plugin: bootstrap, ops, events, the S API, value serialization

docs/luau-runtime.md

The in-Studio runtime that the bridge pushes on every connect

docs/install.md

Install walkthrough with the reason behind each step

docs/cloud.md

The cloud tool: Open Cloud actions, permissions, key setup

docs/vision.md

The look tool: providers, watch mode, costs

docs/sync.md

studio-live sync: mirroring .luau files into and out of a place

docs/live-measurements.md

Engine measurements on Studio 0.738 the design rests on

docs/live-test-results.md

Acceptance-test results against a live Studio, defects found and fixed

docs/multi-agent-build-report.md

Six agents building a playable game concurrently through the bridge

docs/research-brief.md

The architecture brief that preceded the build: what a connector can and cannot make fast

Security

The bridge listens on 127.0.0.1 only and has no authentication: any process on your machine can drive Studio through it, and the plugin runs whatever bundle the bridge sends it. Only look (a screenshot to the Anthropic API) and cloud (your request to Open Cloud) leave the machine. API keys are read from the environment or ~/.studio-live, never returned by a tool, and masked in logs. Nothing saves or publishes a place on its own: cloud publish makes a place file the live version and cloud restriction ban bans players, but only when an agent calls them, and only with a key granted those permissions. Give an agent's key no more than its work needs — cloud { action: "info", what: "key" } shows what it has.

License

MIT. Not affiliated with Roblox Corporation.

Available Tools

9 tools
cloudRoblox Open CloudA
Destructive

Roblox Open Cloud for the place open in Studio. IDs default to the connected session (universe = game.GameId, place = game.PlaceId, creator = the place owner); pass universe_id / place_id / id / creator only to override. Needs an API key: env ROBLOX_OPEN_CLOUD_KEY or /opencloud.json {"key":"…"}, re-read every call and never returned. A 403 names the exact Creator Hub permission to add — or call info what:"key" first: it probes the key and reports allowed | denied | unknown per action. action (ops in parens; arg help is on each field):

  • datastore (list_stores|list_entries|get|set|delete|increment), ordered (list|get|set|delete|increment): persistent entries; etag makes a set conditional.

  • memory (map_list|map_get|map_set|map_delete|queue_add|queue_read|queue_discard): MemoryStore sorted maps + queues, fast cross-server state with a ttl.

  • message: topic + message (≤1 KB) → MessagingService in live servers, not playtests.

  • info (universe|place|group|user|me|key|memberships|roles|inventory|subscription): reads. key = the capability probe.

  • publish: uploads a local .rbxl/.rbxlx as the place's new live version. Do this before luau / instance, which read the PUBLISHED place, never the Studio session.

  • asset_upload: file + asset_type (Model|Decal|Audio|Video|Animation|Mesh|Image) → asset_id + moderation. asset (get|update|versions|rollback|archive|restore): the rest of the lifecycle; update puts a new .fbx behind an existing Model asset_id.

  • luau: runs a script in a fresh server copy of the published place; returns results + logs.

  • instance (get|update|children): read/edit instances of the published place ("root" = the DataModel). For the OPEN place use the run tool.

  • restriction (list|get|ban|unban|logs): ban a user from the experience or one place.

  • notify: send an experience notification to a user (message_id = a Creator Hub template). Results are JSON ≤ 20 KB; long calls return pending: true + a re-poll handle after timeout_ms.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoinfo group/user/inventory/subscription: id; restriction: the player’s user id; notify: the user to notify
opNodatastore: list_stores|list_entries|get|set|delete|increment; ordered: list|get|set|delete|increment; memory: map_list|map_get|map_set|map_delete|queue_add|queue_read|queue_discard; asset: get|update|versions|rollback|archive|restore; restriction: list|get|ban|unban|logs; instance: get|update|children
keyNodatastore/ordered/memory map: entry or item key
deepNoinfo key: also probe writes that a delete against a reserved name can settle (nothing real is deleted)
etagNodatastore set: write only if the entry still has this etag (from get)
fileNoasset_upload / asset update (.fbx Models only) / publish: absolute path of the file
nameNoasset_upload / asset update: display name
taskNoluau: re-poll an earlier task path instead of creating one
whatNoinfo: universe | place | group | user | me (owner of the open place) | key (what this API key may do) | memberships | roles | inventory | subscription
countNomemory queue_read: how many items to read (default 1, max 200)
levelNorestriction: "universe" (default, the whole experience) or "place" (one place only)
scopeNodatastore: scope (omit = global); ordered: scope (default global)
storeNodatastore/ordered: data store name; memory: the sorted map or queue name (created by its first write)
topicNomessage: MessagingService topic (≤ 80 chars)
ttl_sNomemory map_set / queue_add: how long the item lives, in seconds
usersNodatastore set/increment: user ids the value belongs to
valueNoset ops: JSON value (ordered: integer; memory queue_add: the item payload)
actionYes
amountNoincrement: integer amount to add (negative allowed)
filterNolist_entries: id.startsWith("prefix"); ordered list: entry >= 10 && entry <= 30; memory map_list: sortKey > 100 (NOT numericSortKey); restriction logs: user == "users/156"; info memberships/inventory: see the reference
reasonNorestriction ban: the private moderation note (not shown to the player)
scriptNoluau: Luau source run server-side in a fresh copy of the published place
creatorNoasset_upload: owner (default: owner of the open place)
messageNomessage: string or JSON (≤ 1 KB)
read_idNomemory queue_discard: the read_id from queue_read — a whole read batch is acknowledged at once
versionNoasset rollback: the version number to restore (from asset versions)
asset_idNoasset ops: the asset to act on
order_byNoordered list: "value desc" (default ascending); memory map_list: only "id" / "id desc"
place_idNooverride: place (default game.PlaceId of the open place)
priorityNomemory queue_add: higher runs closer to the front (equal priorities keep insertion order)
sort_keyNomemory map_set: sort key — a number sorts numerically, a string alphabetically
page_sizeNolist ops: page size (memory and restrictions cap at 100, asset versions at 50)
read_maskNoasset get: comma-separated extra metadata fields to return
asset_typeNoasset_upload: Model | Decal | Audio | Video | Animation | Mesh (also Image)
attributesNodatastore set/increment: metadata object
class_nameNoinstance update: Folder | Script | LocalScript | ModuleScript — the only classes this API can write
duration_sNorestriction ban: length in seconds; omit entirely for a permanent ban
message_idNonotify: the notification string id from Creator Hub
page_tokenNolist ops: nextPageToken from the previous page
parametersNonotify: values for the {placeholders} in the notification string (string or integer each)
product_idNoinfo subscription: the subscription product id
propertiesNoinstance update: PascalCase properties, e.g. {"Source":"print(1)"} or {"Enabled":false}
timeout_msNolong calls: how long to wait (default 60000, max 300000)
descriptionNoasset_upload / asset update: description
instance_idNoinstance: the instance to act on; "root" (default) is the DataModel
launch_dataNonotify: launch data carried into the experience when the player taps (≤ 200 bytes)
universe_idNooverride: universe (default game.GameId of the open place)
exclude_altsNorestriction ban: do not extend the ban to detected alt accounts
operation_idNoasset_upload: re-poll an earlier upload operation instead of uploading
show_deletedNodatastore list ops: include deleted
version_typeNopublish: "Published" (default, goes live) or "Saved" (stored as a version without publishing)
all_or_nothingNomemory queue_read: return nothing (404) unless the full count is available
display_reasonNorestriction ban: the message the banned player sees
invisibility_sNomemory queue_read: seconds the read items stay hidden from other readers before reappearing
analytics_categoryNonotify: analytics category for this notification

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark destructive/write behavior to no surprise, and the description adds substantial operational context: API key resolution from env or file, re-read every call and never returned, 403 error messages naming the required permission, a built-in key probe (info what:'key'), return size limits, pending async handles, and the critical distinction that luau/instance read the published place, not the Studio session. These details go far beyond what annotations convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but appropriately dense for a tool with 55 parameters across 12 actions. It front-loads the most critical constraints (defaults, authentication, error behavior) before listing actions, and each action group is compactly summarized. A few sentences could be tightened, but overall the structure aids scanning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the enormous scope and absence of an output schema, the description covers essentials: authentication setup, default scoping, override mechanism, error probing, result size limits, async polling pattern, publishing workflow, and where the tool does not apply (open place → run tool). No critical usage detail appears to be missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 98%, so the baseline is 3; the description adds meaningful cross-cutting semantics: default universe/place/creator resolution with overrides, etag-based conditional writes, memory store TTL behavior, message size limits, and the action-specific op groupings. It also refers users to field-level help ('arg help is on each field'), acknowledging the schema's completeness while clarifying global defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear resource and scope: 'Roblox Open Cloud for the place open in Studio.' It enumerates a dozen distinct actions, making the tool's breadth apparent. However, it does not explicitly contrast with sibling tools like observe/run/playtest, though the name and subtitle imply a cloud API access role.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit routing guidance: 'For the OPEN place use the run tool,' distinguishing when not to use this tool for instance operations. It also explains sequencing ('Do this before luau / instance, which read the PUBLISHED place, never the Studio session') and environmental caveats ('message … in live servers, not playtests'). While not covering every possible alternative, it provides clear context for major decision points.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

eventsEvent journalA
Read-onlyIdempotent

Backfill from the bridge's event journal (ring of 10,000 per session). Returns events with seq > since, oldest first → {cursor, events[], dropped, truncated, latest_seq, hub_dropped}. cursor is the seq of the last event actually returned: pass it as the next since. truncated means more events wait after cursor; dropped > 0 means the ring evicted events you never saw. kinds filters by event type: log, error, assert, milestone, custom, playtest, peer, selection, job, controller, change (default all); levels filters log events (print|info|warn|error). timeout_ms > 0 long-polls: returns as soon as a matching event arrives, or empty when the timeout (≤ 50000) elapses. This is the portable fallback to push. Prefer Monitor on ws://127.0.0.1:/events: batched frames ≤ 4 KB carrying seq and dropped; default kinds error, assert, milestone, custom, playtest, peer, controller, job plus warn/error logs (?kinds=…&levels=… per socket). After any dropped > 0 or seq gap on the socket, call events with since = the last seq you saw.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindsNoEvent types to include (default all)
limitNoMax events (default 500)
sinceNoReturn events with seq > since (default 0)
levelsNolog levels to include: print|info|warn|error (default all)
sessionNoStudio session GUID or unique prefix (default: the active hub; required for writes when several Studios are connected)
timeout_msNoLong-poll up to this long for a matching event (default 0 = return now)

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds substantial behavioral context beyond that: the ring buffer eviction behavior (dropped > 0 means events were evicted), the long-poll timeout semantics, the cursor contract, and the relationship to the push socket. It doesn't describe every edge case, but it discloses the non-obvious behaviors that matter for correct use.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-organized: it front-loads the core backfill behavior and return contract, then filters, then long-polling, then the fallback relationship. Every sentence carries information. It is longer than the typical description, but the complexity of the tool (cursor semantics, ring eviction, long-polling, push fallback) justifies the length. A slight deduction for the dense run-on style in the first sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, idempotent, non-destructive tool with 100% schema coverage and no output schema, the description is remarkably complete. It explains the return shape, the cursor contract, the eviction failure mode, the long-poll behavior, and the relationship to the push alternative. An agent has everything it needs to call this tool correctly, including how to recover from dropped events.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all six parameters. The description adds meaning for kinds and levels by enumerating the accepted values, and for timeout_ms by explaining the long-poll behavior. However, since the schema already covers the basics, the description's added value is moderate rather than transformative. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Backfill from the bridge's event journal', and immediately distinguishes itself from the push alternative. It names the exact return shape and the semantics of each field, so an agent can tell exactly what this tool does and how it differs from siblings like observe or look.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use this tool: as the portable fallback to push, and after any dropped > 0 or seq gap on the socket. It also names the preferred alternative (Monitor on ws://127.0.0.1:<port>/events) and explains the conditions that select between them. This is explicit when/when-not guidance with a named alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inputVirtual inputA

Human-like input in a play client through the real input pipeline (dm 'client' = lowest-numbered client, or 'client:N'). actions run in order:

  • {type:'key', key:'W', hold_ms:600 (≤ 60000)} press, hold, release; {type:'key', key:'Space', down:true|false} single edge

  • {type:'click', x, y, button:'left'|'right'|'middle'} down, 50 ms, up; {type:'move', x, y} absolute cursor; {type:'look', dx, dy} relative camera look — best effort: the step fails unless the camera actually rotated (the default camera script ignores virtual deltas; drive workspace.CurrentCamera from a run instead)

  • {type:'text', text}; {type:'focus', path:'PlayerGui.Hud.Input'} TextBox:CaptureFocus(); {type:'wait', ms (≤ 60000)} x,y default to GUI space (gui: true): the coordinates a GuiObject reports as AbsolutePosition; the runtime adds GuiService:GetGuiInset(). gui: false sends raw viewport pixels (inset included). Screenshot pixels are NOT viewport pixels: the screenshot is the whole Studio window scaled by scale; the 3D viewport sits inside it at an offset you must calibrate (see the agent guide). Input only reaches the game while the Studio window renders (not minimized). Result: {steps:[{i, ok, error?}], elapsed_ms}. A failing step does not abort the sequence unless abort_on_error=true; keys still held when a sequence aborts or is cancelled are released. No scroll action: virtual input produces no MouseWheel events. For reactive or long-running input, install a controller with playtest.

ParametersJSON Schema
NameRequiredDescriptionDefault
dmNoTarget DataModel: 'edit' (default) | 'server' | 'client' (lowest-numbered) | 'client:N'
actionsYes
sessionNoStudio session GUID or unique prefix (default: the active hub; required for writes when several Studios are connected)
wait_msNoWait this long for completion before returning a {job_id,status:"running"} handle (default 25000)
abort_on_errorNoStop at the first failing step (default false)

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations are all false and provide no safety profile, so the description carries the full burden. It thoroughly discloses action execution order, hold/release semantics, 50 ms click timing, coordinate-space behavior with GUI inset, the fact that screenshot pixels are not viewport pixels, absence of MouseWheel events, key release on abort/cancel, per-step result format, and non-aborting failure behavior unless abort_on_error is set.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense and every sentence contributes necessary operational detail, with no filler. It is front-loaded with the core purpose and action list, but the heavy use of parentheticals and run-on sentence structure makes it harder to parse than a more bulleted layout would be. Still, it is appropriately sized for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity and lack of an output schema, the description is complete: it covers action order, hard limits, coordinate-system calibration, result shape, abort and error semantics, key-release behavior, no-scroll limitation, window rendering prerequisite, and when to delegate to playtest or run. An agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

While the schema provides descriptions for dm, session, wait_ms, and abort_on_error, the complex `actions` array properties lack schema descriptions. The description compensates by defining each action type and its fields (key, down, hold_ms, x/y, button, dx/dy, text, path, ms, gui) with constraints and coordinate defaults, adding substantial meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states exactly what the tool does: it provides human-like input through the real input pipeline in a play client, with a detailed breakdown of supported action types (key, click, move, look, text, focus, wait). This clearly distinguishes it from siblings like observe, run, and playtest, and goes far beyond a restatement of the title.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit routing guidance: for reactive or long-running input, install a controller with `playtest`; for camera rotation that the default script ignores, drive workspace.CurrentCamera from a `run` instead. It also states a key prerequisite—input only reaches the game while the Studio window renders (not minimized)—so an agent knows when the tool will not work.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jobJob handleA
Destructive

Track long operations that returned {job_id, status:'running'}. status {job_id} → {status: running|done|error, op, dm, elapsed_ms, progress, notes, hub_connected, result?, error?}. wait {job_id, wait_ms ≤ 50000} blocks until the job finishes or the wait elapses, then returns the same snapshot. list → {jobs:[…]} running first, then recent ones (find an id you lost). cancel {job_id} asks Studio to stop the program at its next S.yield()/slice boundary and rolls back an edit-DM recording; the job then finishes with error.code 'cancelled'. Jobs survive a Studio reconnect (hub_connected false while it is away) and end at their deadline. Finished jobs expire 10 minutes after completion.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYes
job_idNoThe job handle returned by a running tool call (status/cancel/wait)
wait_msNowait: max time to block (default 25000, max 50000)

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate destructiveHint=true)Skiputations; the description thoroughly details cancel's behavior: stops at next yield/slice, rolls back edit-DM, job finishes with error.code 'cancelled'. It also explains survival across reconnects and expiration. This goes well beyond annotations, providing comprehensive behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but organized by subcommands, front-loading the core purpose. Each sentence carries meaningful info (status format, wait limits, cancel rollback, expiration). Minor redundancy: repeating {job_id, status:'running'} and action details could be trimmed, but overall efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description defines the exact output snapshot for status/wait (fields like status, op, dm, elapsed_ms). It covers all actions, edge cases (lost IDs, reconnect, expiration), and parameter constraints. An agent has full info to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 67% (job_id and wait_ms are described, but action enum values are not). The description compensates by explaining each action's meaning and input expectations (e.g., wait_ms ≤ 50000). But it doesn't add additional parameter-level details beyond what schema provides, so a 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states that this tool tracks long operations and demonstrates the subcommands (status, wait, list, cancel) with their exact output formats. It clearly identifies the resource (job handles) and the actions available. However, it doesn't explicitly name sibling tools to differentiate from, relying on the subcommands to convey purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use each action: status to check progress, wait to block until finish, list to find lost IDs, cancel to stop and rollback. It also notes jobs survive reconnects and expire after 10 minutes, giving clear context. It doesn't explicitly say when NOT to use this tool versus alternatives, but the subcommands are self-contained.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lookLook at StudioA
Read-only

Look at the Roblox Studio window through a vision model and get a short TEXT answer instead of an image — the screenshot never enters your context (a 768 px frame is ≈ 500 image tokens on the sidecar, a few hundred characters back to you). One-shot: {question, max_width? (768), region? {x,y,w,h} window px, model?} → {answer, model, provider, captured_ms, model_ms, usage, frame_path}. Ask concrete visual questions: 'is there a red error in the Output panel? quote it', 'where is the Play button (region or px)?', 'is the character standing on the platform or falling?'. The model sees only the screenshot and answers 'not visible' when it cannot tell. Watch: {watch: {question, interval_s (5, min 2; 15 on claude-cli), max_frames (60), stop_when? (substring, or /regex/i), diff_only? (true)}} → {watch_id}. Frames are analysed in the background, one model call at a time; each answer arrives as a Monitor event {type:'vision', watch_id, frame, answer, changed, provider}. With diff_only, frames whose bytes differ < 2% from the last analysed frame are skipped (no model call, no event). The watch ends on max_frames, when the answer matches stop_when, on {stop: watch_id | 'all'}, or after 3 consecutive failures; the last event has done:true and reason. {list: true} shows running watches with counts and the last answer. Prefer observe tree|props|find|player for state — they are exact and free. look is for what only pixels can tell: rendering, layout, UI text, visual glitches, what a playtest looks like. Works with an Anthropic API key (ANTHROPIC_API_KEY / ant auth login, ~2 s per look) OR a logged-in Claude Code install (claude on PATH, ~10–15 s per look on the subscription); STUDIO_LIVE_VISION_PROVIDER = auto (default) | api | claude-cli. Models: STUDIO_LIVE_VISION_MODEL (look; default claude-opus-5 / sonnet on the CLI) and STUDIO_LIVE_WATCH_MODEL (watch; default claude-sonnet-5 / haiku).

ParametersJSON Schema
NameRequiredDescriptionDefault
listNoList running watches
stopNoStop a watch by watch_id, or 'all'
modelNoModel override: an id (default claude-opus-5 for look, claude-sonnet-5 for watch) or, on the claude-cli provider, an alias (sonnet / haiku / opus)
watchNoStart a background watch: frames → text events
regionNoCrop in window pixels before scaling (coordinates as in observe screenshot / windows)
questionNoOne-shot: what to look for in the Studio window
max_widthNoFrame width in px (default 768; image tokens ≈ w×h/750)

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, idempotentHint=false, and openWorldHint=true, covering the safety and side-effect profile. The description adds valuable behavioral context: the screenshot never enters the user context, the tool returns a text answer, it can return 'not visible' when inconclusive, and it details watch termination conditions and failure handling. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is lengthy but information-dense, covering modes, parameters, providers, and examples. It front-loads the core purpose and the key benefit (text answer, no image tokens) before diving into details. Some redundancy, such as repeating model defaults, but each section serves a purpose. Not excessively verbose for the complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (one-shot, watch mode, region cropping, model overrides, provider differences), the description provides comprehensive coverage: it explains the output fields, event structure, diff_only behavior, stop conditions, and failure handling. It also references sibling tools for alternatives and lists concrete example questions. The schema and annotations cover the rest, so nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% as all parameters are described in the schema. The description adds some context beyond schema, such as default values for interval_s (5, min 2; 15 on claude-cli) and max_frames (60), and notes diff_only default true. However, this is mostly redundant with the schema descriptions and adds limited new meaning, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool performs visual inspection of the Roblox Studio window via a vision model, returning text answers. It specifies the verb 'Look at' and resource 'the Roblox Studio window', and distinguishes from siblings by stating it is for pixel-level information, while exact state tools like observe are preferred for other queries.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use this tool versus alternatives: 'Prefer observe tree|props|find|player for state — they are exact and free. look is for what only pixels can tell.' It also explains when to use one-shot versus watch mode, and details the difference between providers and models.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

observeObserve StudioA
Read-onlyIdempotent

Read-only view of Studio. what:

  • status: session, place{placeId,placeName,universeId,creatorType,creatorId}, peers, playtest, capabilities, fps, journal seq.

  • tree: root (game), depth (2, max 6), max (500), classes, fields (extra props) → {nodes:[{path,class,name,n,props?}], truncated} (explicit root always returned).

  • props: paths[], props[] → {items:[{path,class,props}], missing} (unreadable props omitted).

  • find: root, name (substring), class (IsA), tag, attr{name,value}, prop{name,value} → {items, total}.

  • diff: since (seq) → {added[], removed[]} instance paths changed since a cursor.

  • logs: since, level, filter, tail (100), dm ('all' = merged journal, lines carry src) → {items:[{seq,t,level,msg,src}], next}; startup prints of play DMs included.

  • script: path, from, to (1-based lines) → {path, class, lines, total_lines, text} (not cut at 8 KB). stats: fps, frameMs, heartbeatMs, physicsMs, instances, memoryMB. selection: {paths} (edit only).

  • geometry: root (Workspace), max (5000; sampled beyond), tolerance (0.05), include_nested (true) → {overlaps:[{a,b,depth}], nested:[{path,parent}], checked, sampled, ms, totals}: parts intersecting parts, BaseParts under BaseParts. Audit a build; fix all listed.

  • player: dm (client) → name, position, velocity, state, health, walkSpeed, cameraCFrame, floorMaterial, seated.

  • screenshot: the Studio window captured by the bridge (works occluded, un-minimizes without focus). max_width (1024; 0 = none), format jpeg|png, quality (70), region{x,y,w,h} window px, hwnd | title_match → image + {path,width,height,source_width,source_height,scale,windowTitle,hwnd,captured_ms}; image px × scale = window px. windows: lists Studio windows.

  • selftest: the hub's runtime self-check → {ok, checks[]}; use after install or when results look wrong. Prefer structured reads (exact, cheap); screenshots only for visual checks (look answers them in text). dm routes the read into the playtest. response_format 'detailed': ~200 KB cap, full CFrames.

ParametersJSON Schema
NameRequiredDescriptionDefault
dmNoTarget DataModel: 'edit' (default) | 'server' | 'client' | 'client:N'; logs also accepts 'all' (merged journal)
toNoscript: last line inclusive (default: end)
maxNotree: default 500; find: default 100; geometry: parts checked (default 5000, sampled beyond)
tagNofind: CollectionService tag
attrNofind: attribute match
fromNoscript: first line (1-based, default 1)
hwndNoscreenshot: pin one Studio window (from observe windows)
nameNofind: name substring, case-insensitive
pathNoscript: path of a Script/LocalScript/ModuleScript
propNofind: property match
rootNotree/find: root path (default game); geometry: default Workspace
tailNologs: default 100
whatYes
classNofind: IsA class
depthNotree: default 2, max 6
levelNologs: minimum level
pathsNoprops: instance paths
propsNoprops: property names (default: notable props, ≤ 60)
sinceNodiff/logs: seq cursor
fieldsNotree: extra props per node, e.g. ["Position","Size"]
filterNologs: substring
formatNoscreenshot: default jpeg
regionNoscreenshot: crop in window pixels
classesNotree: keep only these classes
qualityNoscreenshot: jpeg quality, default 70
restoreNoscreenshot: un-minimize without stealing focus (default true)
sessionNoStudio session GUID or unique prefix (default: the active hub; required for writes when several Studios are connected)
wait_msNoWait this long for completion before returning a {job_id,status:"running"} handle (default 25000)
max_widthNoscreenshot: default 1024; 0 disables scaling
toleranceNogeometry: studs of penetration ignored so touching faces are not overlaps (default 0.05)
title_matchNoscreenshot: choose the Studio window whose title contains this (case-insensitive)
include_nestedNogeometry: also list BaseParts parented under BaseParts (default true)
response_formatNoconcise ≈ 20 KB (default) | detailed ≈ 200 KB

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate readOnlyHint=true, idempotentHint=true, and destructiveHint=false, but the description adds significant detail beyond these: it explains that 'screenshot' works occluded and un-minimizes without focus (with a 'restore' parameter), that 'geometry' samples beyond max and has tolerance for touching faces, that 'props' omits unreadable props, and that 'detailed' response has a ~200 KB cap. These behaviors are not implied by the annotations, so the description adds value. However, some behaviors like rate limits or auth are not mentioned, but not contradictory.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely long, listing every mode in bullet points with detailed sub-options. While it is well-structured and front-loaded with the core purpose, it is far from concise. Every sentence does add information, but the sheer volume makes it hard to scan. The formatting helps, but it could be trimmed by relying more on the schema for parameter details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (33 parameters, 14 modes), the description is remarkably complete. It covers all major modes, explains output shapes, provides defaults, and gives usage tips. It even addresses edge cases like sample thresholds and response size limits. There is no output schema, so the description is the primary source for return value structure, and it handles this comprehensively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 97%, so the schema already documents most parameters well. However, the description adds critical context for parameters like 'max' (different defaults per mode), 'dm' (routing), 'wait_ms' (job handle behavior), and 'response_format' (size caps). It also explains that 'script' is not cut at 8 KB, which is not in the schema. Since coverage is high, the baseline is 3, but the extra semantic details justify a 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Purpose is clearly stated as a 'Read-only view of Studio' with a comprehensive list of read operations (status, tree, props, find, diff, logs, script, stats, selection, geometry, player, screenshot, windows, selftest). It distinguishes itself from siblings by emphasizing it is read-only, while siblings like 'run' and 'playtest' likely perform actions. The description is specific and resource-oriented, making it easy to identify when to use this tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit guidance on when to use different modes, e.g., 'Prefer structured reads (exact, cheap); screenshots only for visual checks (`look` answers them in text).' It also explains when to use 'dm' to route reads into playtest, and mentions 'selftest: use after install or when results look wrong.' While it doesn't name sibling alternatives explicitly, it references the sibling 'look' for visual checks, and the structured reads vs. screenshots trade-off is clearly stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

playtestPlaytest controlA
Destructive

Control the live playtest and the in-engine runtime. Keep ONE playtest alive; a restart costs ~3 s and loses play-DM state.

  • start {mode: play|run|multiplayer, players}: waits for the runtimes → {running, mode, players, peers, started_ms}. multiplayer spawns a server DM plus players (1-8, default 2) client Studios (client:1..N; slow → job handle). no_peer: turn on "Load User Plugins In Run Modes" / install the plugin (test still running: Stop in Studio).

  • stop → {stopped_ms}. status → {running, mode, players, peers, elapsed_s, controllers}. add_players {count} (multiplayer).

  • run_until {dm, predicate | predicate_file, timeout_ms ≤ 120000, interval_ms, args}: Luau predicate evaluated in-engine each Heartbeat until truthy → {result: true|'timeout', value, elapsed_ms, checks}.

  • install {dm, name, code | code_file, persist}: code returns { load = function(ctx) … end, unload = function() … end }. ctx: S, assert(name,cond,detail), milestone, emit, log, onHeartbeat, onEvent, every, after, player, character(), input.*, moveTo (straight line), pathTo (pathfinding), state(), storage. Same name replaces; asserts/milestones → /events. persist: kept by the bridge, re-installed whenever that DM appears in a later playtest.

  • uninstall {dm, name} (drops the persisted entry). list → controllers on all peers + persisted.

  • hotpatch {dm, path, source | source_file, restart}: writes a script Source in the live DM and restarts it, playtest kept (ModuleScript: re-require needed).

  • push {paths, dm, parent, replace}: copies edit-DM instances into the live DM at their own paths (or under parent) → {dm, paths, count, bytes, replaced, skipped?, replicated}; replace (default true) removes a same-name, same-class sibling first (never Terrain, the camera, a character); ephemeral, not undoable. dm: 'server' | 'client' | 'client:N'. *_file = absolute path the bridge reads (never heredoc Luau). Slow actions return {job_id, status:'running'} after wait_ms; use job.

ParametersJSON Schema
NameRequiredDescriptionDefault
dmNoTarget DataModel: 'edit' (default) | 'server' | 'client' (lowest-numbered) | 'client:N'
argsNoAvailable to the program as ARGS
codeNoinstall: Luau returning { load = function(ctx) … end, unload = function() … end }
modeNostart: default play
nameNoinstall/uninstall: controller name
pathNohotpatch: script path, e.g. ServerScriptService.Main
countNoadd_players: clients to add to the running multiplayer test, 1-8 (default 1)
pathsNopush: edit-DM instance paths to serialize into the live DM
actionYes
parentNopush: parent path in the target DM (default: each instance lands at its own edit-DM path)
sourceNohotpatch: full new source
persistNoinstall: keep it in the bridge and re-install whenever that DM appears in a later playtest
playersNostart (multiplayer): client Studio processes to spawn, 1-8 (default 2)
replaceNopush: destroy a same-named sibling at the target parent first (default true); false keeps both
restartNohotpatch: toggle Disabled to restart the script (default true)
sessionNoStudio session GUID or unique prefix (default: the active hub; required for writes when several Studios are connected)
wait_msNoWait this long for completion before returning a {job_id,status:"running"} handle (default 25000)
code_fileNoinstall: instead of code: absolute path of a file holding the Luau (read by the bridge; UTF-8, BOM ok, ≤ 4 MB). No shell/JSON escaping touches it
predicateNorun_until: Luau expression or chunk; truthy ends the wait
timeout_msNorun_until/push: default 30000, max 120000
interval_msNorun_until: 0 = every Heartbeat (default)
source_fileNohotpatch: instead of source: absolute path of a file holding the Luau (read by the bridge; UTF-8, BOM ok, ≤ 4 MB). No shell/JSON escaping touches it
predicate_fileNorun_until: instead of predicate: absolute path of a file holding the Luau (read by the bridge; UTF-8, BOM ok, ≤ 4 MB). No shell/JSON escaping touches it

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true, and the description elaborates on this: 'push ... ephemeral, not undoable' and 'replace (default true) removes a same-name, same-class sibling first (never Terrain, the camera, a character)'. It also discloses the restart cost (~3 s), state loss ('loses play-DM state'), persistence semantics ('persist: kept by the bridge, re-installed whenever that DM appears in a later playtest'), and async behavior ('Slow actions return {job_id, status:"running"} after wait_ms; use `job`'). This goes well beyond the annotations, which only say destructiveHint=true. The only minor gap is that it doesn't explicitly state that start/stop are also destructive in the sense of ending a running playtest, but the 'Keep ONE playtest alive' warning covers that.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-structured: it opens with the most important constraint ('Keep ONE playtest alive'), then groups actions with their parameters in a compact notation. Every sentence carries information: the restart cost, the no_peer workaround, the predicate semantics, the persist behavior, the push replacement rules, and the async job handle note. It loses one point because the density makes it somewhat hard to parse at a glance – the action-parameter groupings are packed into a single paragraph without line breaks, and an agent might need to re-read to extract the exact parameter list for a given action. Still, it is far from verbose and front-loads the critical warning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 23 parameters, no output schema, and a complex multi-action surface, the description covers the essential behavioral context: what each action does, which parameters apply, what the return values look like (e.g., {running, mode, players, peers, started_ms}), and the async job handle pattern. It also covers edge cases like no_peer, persist, and replace. It doesn't fully document every return shape for every action (e.g., status returns {running, mode, players, peers, elapsed_s, controllers} but stop only shows {stopped_ms}), and it doesn't explain the `args` parameter's role in run_until beyond 'available to the program as ARGS'. Given the tool's complexity, these are minor gaps; the description is remarkably complete for a 23-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 96%, so the schema already documents most parameters. The description adds meaning by grouping parameters with their actions (e.g., 'start {mode: play|run|multiplayer, players}', 'run_until {dm, predicate | predicate_file, timeout_ms ≤ 120000, interval_ms, args}'), which clarifies which parameters apply to which action. It also adds constraints not fully in the schema, such as 'players (1-8, default 2)' and 'timeout_ms ≤ 120000'. The description doesn't repeat every schema description, but it adds the action-parameter mapping that the flat schema lacks. A 4 is appropriate because the schema does most of the work, but the description adds meaningful grouping and defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear verb and resource: 'Control the live playtest and the in-engine runtime.' It then enumerates each action (start, stop, status, run_until, install, uninstall, list, hotpatch, push) with its specific purpose, which distinguishes it from sibling tools like observe, look, run, input, events, skills, job, and cloud. The scope is unambiguous: this is the tool for managing playtests and the live DM runtime, not for observing or running code in other contexts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage guidance: 'Keep ONE playtest alive; a restart costs ~3 s and loses play-DM state.' It also gives conditional guidance for multiplayer ('no_peer: turn on "Load User Plugins In Run Modes" / install the plugin'), for run_until ('Luau predicate evaluated in-engine each Heartbeat until truthy'), and for push ('replace (default true) removes a same-name, same-class sibling first'). It names alternatives implicitly by listing all actions and their parameters, and the sibling list shows this is the only playtest control tool. The description also warns about slow actions returning job handles, which tells the agent when to expect async behavior.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

runRun Luau in StudioA
Destructive

Run a Luau program in Roblox Studio against the resident S API. Ship whole programs, not micro-calls: build, query and verify in one call and return one JSON-safe summary. Globals: S, ARGS (= args), print/warn (captured), game, workspace, task. dm 'edit' (default): ONE ChangeHistory recording — one undo step, rolled back on error/timeout. 'server' | 'client' | 'client:N': inside the live playtest, ephemeral (lost on stop, not undoable). code or code_file (absolute path the bridge reads). Never heredoc Luau: \n in a string becomes a real newline → syntax_error "Malformed string". S: get, ensure(path,class,props), find{root,name,class,tag,attr,max}, tree, props, new, set (unwritable props → result.unwritable), batchSet, clone, destroy (undoable), part, model, grid, placeOn(part,target,{align,gap}) sets a part on top of target, fits(cframe,size) → ok, blockers, overlaps(parts), script.get(path,{from,to})/set/patch/restart/create, emit, log, yield() (in long loops), wait, raycast, distance, remaining(). Geometry is enforced: parts intersecting parts (touching faces are fine) and BaseParts parented under BaseParts (use Models/Folders) are reported. geometry_policy warn (default) → result.geometry {overlaps[{a,b,depth}], nested[{path,parent}], checked, totals} + warnings; reject → rolled back, error geometry_violation with the report; off. Fix before moving on. Result: {value, output[], duration_ms, changes{added,removed,paths}, undo: committed|cancelled|unavailable|n/a, ephemeral, dm, warnings?, detached?, unwritable?, geometry?}. response_format 'detailed' adds 12-component CFrames. Errors: {error:{code: luau_error|syntax_error|timeout|no_peer|cancelled|busy|geometry_violation|…, message, stack, output}}. Still running after wait_ms (default 25000)? {job_id, status:'running'} — the program keeps running; use job. Edit-DM writes queue FIFO. dry_run (edit only) rolls back; a raw :Destroy() is neither rolled back nor undoable (use S.destroy). Several Studios connected: pass session.

ParametersJSON Schema
NameRequiredDescriptionDefault
dmNoTarget DataModel: 'edit' (default) | 'server' | 'client' (lowest-numbered) | 'client:N'
argsNoAvailable to the program as ARGS
codeNoLuau program. Globals: S (resident API), ARGS, print/warn (captured), game, workspace, task. A top-level `return` is captured.
dry_runNoedit DM only: run, then roll back
sessionNoStudio session GUID or unique prefix (default: the active hub; required for writes when several Studios are connected)
wait_msNoWait this long for completion before returning a {job_id,status:"running"} handle (default 25000)
code_fileNoInstead of code: absolute path of a file holding the Luau (read by the bridge; UTF-8, BOM ok, ≤ 4 MB). No shell/JSON escaping touches it
timeout_msNoExecutor deadline in ms (default 30000)
undo_labelNoChangeHistory waypoint name (edit DM)
geometry_policyNowarn (default; env STUDIO_LIVE_GEOMETRY_POLICY): overlapping/nested parts → result.geometry + warnings; reject: such a run is rolled back (error geometry_violation); off: no check. Play DMs check only when set
response_formatNodetailed: full 12-component CFrames ('c') in the returned value (default concise)

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations carry only destructiveHint=true and no readOnly/idempotent flags, so the description must carry the behavioral burden—and it does. It discloses rollback semantics per dm mode ('ONE ChangeHistory recording… rolled back on error/timeout'), ephemerality of play DMs, geometry enforcement with policy outcomes, the running-job escape hatch, and the FIFO queue. None of this contradicts the annotations; it substantially extends them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is long, but the tool is complex and the density is justified. The text is organized into scannable blocks (dm modes, code input, S API, geometry, results/errors, long-running) and front-loads the core contract before enumerating API details. Every section earns its place, even if a future pass could trim the S-method listing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 11 params, no output schema, and a mutation-heavy tool, the description is unusually complete: it defines the success result shape, error envelope with codes, geometry report structure, long-running job handle, and undo/dry-run caveats. It even covers edge cases like a raw `:Destroy()` not being rolled back. Nothing essential is missing for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3; the description adds real value on top by explaining `code_file`'s bypass of shell/JSON escaping, `geometry_policy` outcomes in the result, `response_format`'s effect on CFrames, `dry_run` rollback limits, and `session` requirements. Slight gap: `args`/`timeout_ms` rely on the schema alone, but that is acceptable at full coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Run a Luau program in Roblox Studio against the resident S API.' It goes beyond the name by defining the execution model ('Ship whole programs, not micro-calls') and enumerating provided globals, distinguishing it from siblings like playtest, observe, and job.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit direction on when to use the tool: 'Ship whole programs, not micro-calls' and 'build, query and verify in one call,' and it points to the sibling `job` for long-running programs ('use `job`'). It also flags what not to do ('Never heredoc Luau') and when `session` is required. It doesn't contrast every sibling explicitly, but the context is clear enough for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

skillsSkill libraryA
Destructive

Luau program library on disk (/skills/.luau with a --[[ studio-live skill … ]] header holding name, description and params). Save programs you will run again; run them with args (ARGS), same environment and result shape as run. action: list → [{name, description, params, builtin}]; get {name} → {source, …}; save {name, source | source_file (absolute path), description, params}; delete {name}; run {name, args, dm, undo_label, dry_run, geometry_policy, timeout_ms, response_format, wait_ms, session} → the run result (geometry checked exactly as for run). Builtins ship read-only (builtin: true): settle_physics, device_sim, profile_scripts, bulk_attributes, insert_asset, lighting_preset, list_scripts, remote_map — get one to read its params. Saving a skill with a builtin name overrides it; deleting the override restores it; builtins themselves cannot be deleted.

ParametersJSON Schema
NameRequiredDescriptionDefault
dmNoTarget DataModel: 'edit' (default) | 'server' | 'client' (lowest-numbered) | 'client:N'
argsNoAvailable to the program as ARGS
nameNoSkill name: letters, digits, _ or - (max 64)
actionYes
paramsNosave: JSON describing the expected ARGS
sourceNosave: Luau program (same environment as run)
dry_runNorun: edit DM only: run, then roll back
sessionNoStudio session GUID or unique prefix (default: the active hub; required for writes when several Studios are connected)
wait_msNoWait this long for completion before returning a {job_id,status:"running"} handle (default 25000)
timeout_msNoExecutor deadline in ms (default 30000)
undo_labelNorun: ChangeHistory waypoint name (default "skill: <name>")
descriptionNosave: one-line description
source_fileNosave: instead of source: absolute path of a file holding the Luau (read by the bridge; UTF-8, BOM ok, ≤ 4 MB). No shell/JSON escaping touches it
geometry_policyNowarn (default; env STUDIO_LIVE_GEOMETRY_POLICY): overlapping/nested parts → result.geometry + warnings; reject: such a run is rolled back (error geometry_violation); off: no check. Play DMs check only when set
response_formatNorun: as for the run tool

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses important behavioral traits beyond the annotations: builtins are read-only, saving over a builtin overrides it, deleting the override restores it, and builtins themselves cannot be deleted. It also notes that run geometry is checked exactly as for `run`. The annotations already mark the tool destructive, and the description aligns with and expands on that.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-organized: it opens with a one-line definition and purpose, then compactly lists action signatures, then covers builtin behavior. Minor redundancy exists in repeating the run-result equivalence, but overall every sentence earns its place by conveying action-specific or behavioral information not present in the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (15 parameters, five actions, no output schema), the description covers the main action modes, return shapes, builtin handling, override semantics, and geometry behavior. It leans on the sibling `run` for the full run result shape, which is acceptable given the context signals, but it is not entirely self-contained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 93% schema description coverage, the schema does most parameter documentation work. The description adds action-specific parameter grouping (e.g., save uses source | source_file, run uses dm/undo_label/dry_run/etc.) and clarifies the relationship between source and source_file. This goes beyond the flat schema definitions and helps an agent assemble valid calls per action.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the resource as a Luau program library on disk and enumerates five distinct actions (list, get, save, delete, run) with their signatures and result shapes. It also contrasts with the sibling `run` by noting skills run in the same environment and result shape as `run`, making the tool's purpose and scope unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives practical guidance: 'Save programs you will run again; run them with args' and explicitly states that execution matches the `run` tool. It also explains builtin read-only behavior and override/restore semantics. It does not explicitly say when to prefer this over direct `run`, but the context strongly implies the distinction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 9 tool updatesv0.1.0
    • First observedcloud
    • First observedevents
    • First observedinput
    • First observedjob
    • First observedlook
    • First observedobserve
    • First observedplaytest
    • First observedrun
    • First observedskills

TDQS

A4.4/5.0

Scored across 9 tools

Disambiguation4/5

Each tool has a distinct primary job—observe for structured reads, look for vision QA, run for scripted execution, playtest for session lifecycle, input for synthetic control, events for journal backfill, skills for reusable programs, job for async tracking, cloud for Open Cloud. A couple of near-overlaps remain (observe status/logs versus playtest status and the events journal), but descriptions generally call out when to use each.

Naming Consistency5/5

All tool names are lowercase single words in a consistent shorthand style (observe, look, run, playtest, input, events, skills, job, cloud). No casing or separator conventions are mixed, making the naming predictable.

Tool Count5/5

Nine top-level tools is well-scoped for a Studio automation server; each tool covers a necessary capability, and the broad observe modes are bundled under one read tool rather than fragmented into many MCP tools. Nothing feels redundant at the top level.

Completeness5/5

The surface covers the full loop: inspect (observe), visually verify (look), mutate and script (run), control playtests (playtest), send input (input), consume events (events), persist reusable programs (skills), manage long jobs (job), and access Open Cloud (cloud). The run tool's S API also provides an escape hatch for anything not explicitly modeled, so there are no obvious dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    A production MCP integration that lets AI agents control Roblox Studio to autonomously build, test, and debug Roblox games. Provides 39 tools for explorer control, script management, terrain generation, and autonomous testing.
    1
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables AI assistants to control Roblox Studio by running Luau code, creating and editing instances, reading the scene tree, and managing scripts via an MCP server with a long-polling plugin bridge.
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Enables AI agents to inspect, create, edit, debug, and playtest projects inside the Roblox editor via 29 lean tools, with push-based SSE transport, editor-safe script edits, and batched undoable writes.
    34
    946 npm
    6
    MIT