owmcp
Drives the Opiumware Roblox executor attached to a live Roblox client, letting an agent run Luau code or script files, read console logs, browse and grep decompiled client scripts, capture and inspect remote network traffic (FireServer, InvokeServer, OnClientEvent), and take screenshots of the Roblox window to iteratively test and verify scripts in-game.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@owmcprun print(game.Players.LocalPlayer.Name) and show me the output"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
owmcp
owmcp is an MCP server that lets AI agents drive the Opiumware Roblox executor on macOS. It enables an agent to write a script, run it, check the logs or a screenshot, and iterate autonomously.
Tools
Tool | Description |
| Check live Roblox instances (game, player, port) and decompiler state. |
| Run Luau code or a file. Returns return values, prints, and errors with tracebacks. Cancels the run on timeout. |
| Read the Roblox console. Filter by level, regex, time, limit, or "since the last execute". |
| List, search, and read decompiled client scripts, mirrored as |
| Start, stop, and read captures of remote traffic (FireServer, InvokeServer, OnClientEvent), including arguments, return values, and calling scripts. Hooks only exist during capture. |
| Capture the Roblox window (even if covered) to visually verify ESP, UI, and menus. |
Large outputs are saved to temporary files ($TMPDIR/owmcp/). Tools return the file path and a preview to keep the agent's context small.
Related MCP server: Roblox Executor MCP Server
Requirements
macOS
Opiumware attached to Roblox
Node.js 18+
Screen Recording permissions for your MCP client's host app (only needed for
screenshot).
Setup
git clone https://github.com/bzzimmy/owmcp && cd owmcp
npm installConfigure your MCP client (e.g., Claude Desktop or ~/.config/mcp/mcp.json):
{
"mcpServers": {
"opiumware": {
"command": "node",
"args": ["/absolute/path/to/owmcp/build/index.js"]
}
}
}Keep the server process alive between calls. Network captures and the decompile cache live in-process.
Allow tool calls to run for up to ~5 minutes (
executetakes timeouts up to 300s).Open Roblox, attach Opiumware, join a game, and call
statusto verify the connection.
How it works
Scripts are sent to Opiumware's local TCP server (ports 8390–8399). They are wrapped in lua/runtime.luau, which captures output and posts the results back to a local HTTP callback.
You never need to paste code into the executor.
Rejoining a game does not break the connection.
Logs are read directly from
~/Library/Logs/Roblox.Decompiling uses Opiumware's bundled decompiler, which starts automatically if needed.
Development
npm run build # Compile TypeScript
npm run lint # Run ESLint, tsc, and luau-lspRebuild after pulling updates. Clients run build/index.js.
Disclaimer
Executors violate the Roblox Terms of Use and can lead to account bans. The execute tool runs arbitrary code in your client, so only connect agents you trust.
License
MIT © Ben Zimmermann
Available Tools
6 toolsexecuteExecute LuauA
Run Luau in the Roblox client (Opiumware, thread identity 8) and get back return values, print/warn output, and errors with traceback.
Just
returnvalues; tables, Instances, Vector3s etc. are serialized automatically. Don't JSONEncode.Each run has its own global scope. Use getgenv() to keep state between runs. _G is the executor's; the game's is getrenv()._G.
Full executor API is available (getgc, hookmetamethod, decompile, firesignal, ...). To run a local file use the file parameter; the executor's own readfile/loadfile only see Opiumware's workspace folder.
Heavy scans (getgc, GetDescendants on large trees) must yield (task.wait() every few thousand items) or they freeze the game.
Long-running code (loops, listeners) should task.spawn and return right away. On timeout the run itself is cancelled, but threads it spawned keep running.
Errors returned here are NOT in the console. Errors/prints after the run returns (spawned/deferred code) only show in logs: use logs sinceExecute=true.
Separate execute calls made in parallel may run in any order.
Large results are saved to a temp file whose path is returned.
| Name | Required | Description | Default |
|---|---|---|---|
| code | No | Luau source to run. Provide either code or file. | |
| file | No | Path to a local .lua/.luau file to run instead of code. | |
| port | No | Opiumware port of the Roblox instance (default: first found, see status). | |
| timeout | No | Seconds to wait for the result. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: thread identity 8, per-run global scope, executor-vs-game _G distinction, requirement to yield on heavy scans, spawned threads surviving a timeout cancellation, and large results being written to a temp file. These are non-obvious runtime traits an agent could not infer elsewhere.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action and return shape are front-loaded in the first sentence, followed by tightly scoped bullets. Every bullet carries operational information (scope, yielding, timeout side effects, ordering) with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description compensates by describing exactly what comes back (return values, serialized tables/Instances/Vector3s, print/warn, traceback, temp-file path for large results). Combined with the timeout and logging caveats, an agent has everything needed to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds genuine meaning: code vs file are framed as alternatives, and the timeout is explained behaviorally ('the run itself is cancelled, but threads it spawned keep running'), which the schema's plain 'Seconds to wait for the result' does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Run Luau in the Roblox client') plus the exact return payload: return values, print/warn output, and errors with traceback. It also implicitly separates itself from the 'logs' sibling by noting that post-return output lands there instead.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete routing conditions: use the file parameter for local files, use logs sinceExecute=true for prints/errors emitted after the run returns, and note that parallel calls may interleave unpredictably. The getgenv()/_G/getrenv() state guidance tells the agent when state must be persisted rather than re-declared.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
logsRead console logsA
Read the Roblox client console (print/warn/error from all scripts, including the game's own), newest last.
Debug loop: after
execute, call with sinceExecute=true to see everything logged since that run started (deferred errors, spawned loops), or new=true for everything since the previous logs call.Errors from code run via execute are tagged "mcp:"; from the Opiumware editor "Opiumware:". Filter with grep.
Repeated consecutive lines are collapsed to (xN). Long messages are truncated; when anything is cut, the full results are saved to a temp file you can grep.
| Name | Required | Description | Default |
|---|---|---|---|
| new | No | Only entries logged since the previous logs call. | |
| grep | No | Case-insensitive regex matched against message and stack. | |
| level | No | Levels to include. "engine" adds Roblox engine channels (assets, network, ...), usually noise. | |
| limit | No | Max entries shown (newest). | |
| seconds | No | Only entries from the last N seconds. | |
| sinceExecute | No | Only entries logged since the most recent execute call started. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full burden and does so well: it discloses ordering (newest last), output-shaping behaviors (repeated lines collapsed to (xN), long messages truncated with full results saved to a greppable temp file), and error provenance tagging ('mcp:<line>' vs 'Opiumware:<line>'). These are non-obvious operational traits an agent could not infer from the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose leads in the first clause, followed by three tight bullet-style lines covering the debug loop, error tagging, and output shaping behavior. Every sentence carries information; there is no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a six-parameter tool with no annotations and no output schema, the description covers ordering, truncation, collapsing, and provenance tagging well enough to call it correctly. It stops short of describing the shape of individual returned entries, which is the only notable remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond it: it explains what sinceExecute actually captures (everything since that run started, including deferred errors and spawned loops) and frames grep as the way to filter the provenance tags. That is genuine added context on top of the schema's own descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read the Roblox client console') and scopes it precisely to print/warn/error from all scripts, including the game's own. This clearly separates it from siblings like network, status, and screenshot without the agent needing to open any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit conditional guidance: call with sinceExecute=true after an execute run to catch deferred errors and spawned loops, or new=true for everything since the previous logs call. That is clear routing between parameter modes, though there is no explicit statement of when NOT to use this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
networkCapture remote trafficA
Capture remote traffic between the client and server: outgoing FireServer/InvokeServer (with InvokeServer return values and the calling script) and incoming OnClientEvent.
start: begin a fresh capture (replaces the previous log). stop: stop recording and remove the hooks (the log is kept). read: view the log (works while running).
Hooks only exist while capturing, so stop as soon as you have what you need.
read without
remotegives a per-remote summary (count, rate, latest args); withremoteit lists individual calls with arguments, collapsing consecutive identical calls.Calls made by your own executed code are marked [by executor].
Long results are truncated and the full log is saved to a temp file.
| Name | Required | Description | Default |
|---|---|---|---|
| port | No | Opiumware port (default: first found). | |
| limit | No | read: max rows shown (remotes in the summary, newest calls with remote=). | |
| action | Yes | ||
| remote | No | read: case-insensitive regex on remote path/method; lists matching calls. | |
| exclude | No | read: case-insensitive regex of remotes to hide (e.g. noisy ones). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses that start replaces the previous log, that hooks exist only during capture, that the log is retained after stop, that executor-originated calls are tagged [by executor], and that long results are truncated to a temp file. Missing peripheral details like auth or port behavior, but the key lifecycle and side-effect traits are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core capability is stated first, then bulleted semantics for each action and each read mode. Every line carries information — no filler, no repetition of the schema, and the bullets make it scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description properly explains the return shape (per-remote summary vs. individual calls, truncation with temp-file fallback). Combined with the action semantics, an agent has enough to invoke and interpret results, though a note on how the temp file path is returned would close the last gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema already describes all five parameters (80% coverage), yet the description adds genuine meaning: the two read modes (summary vs. individual calls collapsed), what collapsing means, and concrete use of `exclude` (hiding noisy remotes). It exceeds the baseline-3 expectation for high-coverage schemas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource ("Capture remote traffic between the client and server") and enumerates exactly what is captured — outgoing FireServer/InvokeServer with return values and calling script, plus incoming OnClientEvent. This level of specificity makes it trivially distinguishable from siblings like logs or scripts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly documents when to use each action (start/stop/read), notes that read works while running, and gives actionable advice ("stop as soon as you have what you need"). It stops short of explicitly routing against alternative sibling tools such as logs or scripts, so it lands at 4 rather than 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
screenshotScreenshotA
Capture the Roblox window as the player sees it (works even if other windows cover it). Use to verify visual scripts (ESP, UI, menus). Each image costs ~1K tokens, so use sparingly.
| Name | Required | Description | Default |
|---|---|---|---|
| port | No | Opiumware port of the Roblox instance (default: first Roblox window). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and delivers real value: it discloses that capture bypasses window occlusion and that each image costs ~1K tokens, a meaningful rate/cost signal. It omits any permission or error behavior and does not state what is actually returned, so it stops short of fully covering the operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero waste: the capture semantics come first, then usage, then the cost caveat. Every clause contributes information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-optional-parameter read tool with no output schema, the description supplies the key operational facts (what is captured, why to call it, token cost). It could note that the result is image content and any failure modes, but nothing essential to invoking it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% ('Opiumware port of the Roblox instance'), so the sole optional parameter is fully documented in structured form. The description adds no port-selection guidance, which is the baseline expectation when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Capture the Roblox window') plus a distinguishing qualifier ('as the player sees it, works even if other windows cover it'). This cleanly separates it from siblings like logs, network, scripts, execute, and status, which retrieve non-visual data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete use case ('Use to verify visual scripts (ESP, UI, menus)') and a frugality constraint ('use sparingly') driven by token cost. It stops short of naming an alternative or an explicit when-not-to-use condition, so it is clear context rather than full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scriptsGame scriptsA
Read the game's client-side code (LocalScripts, ModuleScripts; server Scripts never reach the client). Decompiled sources are cached and saved as .luau files in a temp folder mirroring the game tree, so you can also grep/read them directly.
list: script paths (optionally filtered).
search: regex across all decompiled sources, returns path:line matches. The first search decompiles everything (a few seconds).
read: one script's source with line numbers, 150 lines at a time (use startLine/endLine).
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | read: full script path as shown by list/search. | |
| port | No | Opiumware port (default: first found). | |
| limit | No | search: max matches shown. | |
| query | No | search: case-insensitive regex. | |
| action | Yes | ||
| filter | No | list/search: case-insensitive regex on script paths, e.g. ReplicatedStorage. | |
| endLine | No | read: last line (default startLine + 149). | |
| startLine | No | read: first line (default 1). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden and does well: it states the scope limitation ('server Scripts never reach the client'), the caching side effect (decompiled sources saved as .luau in a temp folder mirroring the game tree), the first-call latency, and the pagination window (150 lines at a time). It does not mention permissions/auth or failure modes, but the read-oriented nature is clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose before the bulleted per-action breakdown, and every sentence carries information (scope, caching, latency, pagination). Slightly dense, but nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter tool with no annotations and no output schema, the description supplies the missing context an agent needs: what the actions do, that results are also readable on disk, the shape of search matches, and the read windowing. Only minor gaps remain (exact auth/port semantics, failure behavior when nothing decompiles).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 88%, so the baseline is 3, and the description adds real meaning on top: read defaults to 150-line windows governed by startLine/endLine, path must be the full path as shown by list/search, and search output is path:line matches. This clarifies parameter interrelationships (query vs limit, filter vs path) that the schema only documents field-by-field.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource (the game's client-side code: LocalScripts, ModuleScripts) and enumerates the three concrete operations (list/search/read) with what each returns. None of the siblings (screenshot, status, execute, logs, network) overlap with script source inspection, so an agent can place this tool unambiguously.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Each action is given a clear purpose ('list: script paths (optionally filtered)', 'search: regex across all decompiled sources', 'read: one script's source'), which serves as when-to-use guidance for the sub-commands. It also flags the cost condition ('The first search decompiles everything (a few seconds)'), but offers no explicit exclusions or comparison to an alternative tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
statusStatusA
Check the connection: which Roblox instances (Opiumware ports) are live, what game/player each is in, and whether the decompiler is running. Call this first or when something isn't working.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses what is inspected (ports, game/player, decompiler state), implying a non-mutating read, and it tells the agent to run it before other operations. It never explicitly states it is read-only/side-effect free, which is the one gap for a no-annotation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero waste. The subject of the check is front-loaded and the usage instruction follows immediately; every clause adds information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description compensates by enumerating the reported fields. It is complete enough to invoke correctly, though it omits any hint of response shape (e.g., per-instance list vs. summary) that an agent might want.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, which is the baseline-4 case. There are no argument semantics to explain, and the description correctly implies a parameterless probe.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Check') and resource ('the connection') and enumerates exactly what is reported: live Roblox instances/Opiumware ports, the game/player per instance, and whether the decompiler is running. Among siblings (screenshot, execute, logs, scripts, network), it is unmistakably the diagnostic/health tool, so an agent can distinguish it without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Call this first or when something isn't working' gives explicit conditions for invocation, which is unusually useful for a diagnostic tool. It does not name a specific alternative tool to use instead, so it falls short of full when/when-not routing, but the trigger conditions are clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.1.0- First observed
execute - First observed
logs - First observed
network - First observed
screenshot - First observed
scripts - First observed
status
TDQS
Scored across 6 tools
Each tool targets a clearly distinct domain: screenshot (visual capture), status (connectivity), execute (code execution), logs (runtime output), scripts (source decompilation), and network (traffic capture). Descriptions explicitly differentiate overlapping areas like execute's immediate output vs logs' deferred console, leaving no ambiguity in tool selection.
All names are lowercase single words without underscores or camelCase, following a consistent format. Minor deviation: 'execute' is a verb while the others are nouns, but the uniform style keeps the set predictable.
Six composite tools with sub-actions (list/search/read for scripts; start/stop/read for network) provide a well-scoped surface for a focused Roblox executor. Each tool earns its place, and the count sits comfortably within the 3–15 ideal range.
The set covers connection status, code execution, log inspection, script decompilation/search/read, network capture, and visual verification—a near-complete client-side lifecycle. Minor gaps include no direct script writing/editing tool or instance management beyond status, but execute can work around these.
Maintenance
Related MCP Connectors
Operate Linux, macOS and Windows from your LLM. Every action runs through an auditable allowlist.
Use your own Mac from ChatGPT, Claude or Codex: files, commands, documents, and a browser.
Drive real Android & iOS devices and web browsers from natural language for mobile + web QA. 290+ tools across device control, app management, automation sessions, browser automation, and flow recording / replay. Bearer-auth — get a token at robotactions.com → Profile → API Tokens.
Discover, inspect and run 63,000+ agent tools from one balance. Pay per call, no subscriptions.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to interact with a running Roblox game client to execute Lua code, inspect scripts, spy on remotes, and more.139 npm216MIT
- AlicenseAqualityBmaintenanceEnables AI agents to execute Lua code, inspect scripts, spy on remotes, and interact with a running Roblox game client through an MCP interface.7222139 npm1MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to interact with a running Roblox game client, including executing Lua code, inspecting scripts, and spying on remotes, with a local dashboard for monitoring and control.139 npmMIT
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to interact with a running Roblox game client, including executing Lua code, inspecting scripts, searching instances, intercepting remote events via Sigma Spy, and controlling the GUI.139 npm1MIT