Opal Emu MCP
An MCP server that lets LLM agents autonomously play retro games by driving a real, unmodified OpalEmu (EmulatorJS) page in a Playwright-controlled browser.
list_roms — list recognized ROM files in the configured local ROM directory (names, likely systems, sizes).
load_rom — boot a game by name from the library, or by Base64 + fileName; returns a first screenshot. Replaces the current game and its unsaved progress.
reset_emulator — hard-reset the loaded game and get a screenshot.
get_current_screen — fetch the last captured frame as PNG without advancing the game.
control_emulator — press or release a controller button (a, b, up, start, l2, etc.), optionally for players 0–3; hold directions across skip_frames via state "down"/"up".
skip_frames — advance at least N core frames, then pause and return a screenshot with the actual frame count.
Optional MCP Apps viewer — clients that support it can render a live interactive control surface backed by the same tools.
The emulator auto-pauses when no tool call is in flight, and all tools except list_roms are registered as MCP Apps tools.
Plays Sega games through the OpalEmu retro emulator, with tools to list and load Sega ROMs, control the controller, capture screenshots, and advance frames in Sega console games.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Opal Emu MCPLoad the first available ROM and show me the screen."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
An MCP server that lets LLM agents (Claude, GPT, etc.) autonomously play retro games, by wrapping OpalEmu, a browser-based, EmulatorJS-powered retro emulator supporting 18 systems, with a Node.js bridge.
This project does not modify OpalEmu. It drives a real, unmodified OpalEmu page in a Playwright-controlled browser and exposes it to agents as 6 MCP tools, plus (if the client supports it) a live, interactive MCP Apps viewer.
Credit
All emulation is OpalEmu (source) running EmulatorJS cores. This project only adds the MCP bridge around it; it contributes no emulation code of its own.
Related MCP server: MCP GameBoy Server
License
AGPL-3.0, the same license as OpalEmu (see LICENSE).
Architecture
LLM Agent (Claude, GPT, ...)
│ MCP protocol (stdio)
▼
Node.js MCP server (this project)
│ │
│ WebSocket │ HTML resource (MCP Apps, optional)
▼ ▼
Playwright-controlled Chromium MCP client's own sandboxed iframe
│ runs the real OpalEmu page (client/mcp-app.ts, a *separate*
│ (served by this project's browser context; talks back to this
│ own Express server) server's tools via the App Bridge,
▼ never touches the emulator directly)
window.EJS_emulator
(EmulatorJS instance, the game
actually runs here)Two independent browser contexts are involved, and it's worth being explicit about why:
The Playwright tab (
client/agent.tsinjected into it) is where the game actually runs. It driveswindow.EJS_emulatordirectly and is the only place OpalEmu's real, stateful emulator instance exists.The MCP Apps viewer (
client/mcp-app.ts), if the connecting MCP client supports it, renders in the client's own sandboxed iframe, a completely different, unrelated browser context. It never touches the emulator directly; every button press or screenshot request goes throughApp.callServerTool(), which the host proxies to this server's real tools, the same tools the LLM calls. It's a control surface, not a second embed of the OpalEmu page (embedding the raw page there would boot a second, disconnected, unloaded emulator instance).
Auto-pause
The emulator is paused whenever no tool call is in flight. A button press resumes it briefly; skip_frames runs until at least the requested number of core frames have advanced and reports the measured count. Frame polling can overshoot the target. The emulator pauses again before returning a screenshot, so it does not keep running while an agent is thinking.
Tools
Tool | Description |
| Lists ROM files available on the server (from |
| Loads a ROM by |
| Hard-resets the current game. |
| Returns the last captured frame as a PNG, without advancing the game. |
| Presses or releases a button ( |
| Advances until at least the requested number of core frames has elapsed, then reports the actual count. |
All tools except list_roms are also registered as MCP Apps app tools, so a supporting client can render the live viewer regardless of which one is called first.
Setup
Prerequisites:
Node.js 20+
A built checkout of OpalEmu. By default this project looks for it as a sibling directory (
../OpalEmu/dist); use--opalemu-dist <path>for any other layout. It only ever reads from there, never writes.
# 1. Build OpalEmu itself (the emulator this wraps)
git clone https://github.com/thevalmarch/opalemu ../OpalEmu
cd ../OpalEmu && npm install && npm run build && cd -
# 2. Build this project
npm ci
npx playwright install chromium
npm run buildIf OpalEmu's build isn't found, the server says so explicitly at startup, including the path it looked in.
ROMs
No ROMs are included, and none are downloaded. You supply your own game files, which you should already legally own. Drop them into roms/ (or point elsewhere with --roms-dir <path>) and list_roms will pick up anything with a recognized extension. Nothing in roms/ is committed to git.
Local ROM files larger than 128 MiB are rejected before loading. Direct romBase64 input is limited to 6 MiB decoded so its Base64 and JSON-RPC message fit below the MCP SDK's default 10 MiB stdio limit. Symlinks in the ROM directory are not listed or loaded. Large disc images may need a future streaming path; this release keeps ROM transfer bounded in memory.
Run standalone
npm start # headed browser by default, so you can watch it play
npm start -- --headless # for CI / headless environmentsCLI flags (all optional): --opalemu-dist <path>, --roms-dir <path>, --playwright-profile-dir <path>, --http-port <port> (default 4173), --ws-port <port> (default 4174), --headless. The browser profile defaults to .playwright-profile/ in this checkout.
The HTTP and WebSocket listeners bind to 127.0.0.1. The Playwright browser receives a fresh bridge token at startup. The token is not served by the HTTP page and is not intended for other clients. Do not expose these ports through a proxy or port forward.
Connect to Claude Desktop
Add to claude_desktop_config.json:
{
"mcpServers": {
"opalemu": {
"command": "node",
"args": ["/absolute/path/to/opalemu-mcp/dist/index.js"]
}
}
}Use absolute paths: Claude Desktop spawns MCP servers without a working directory set.
Compatibility with OpalEmu
Originally tested against OpalEmu commit 776874a, whose package version was 1.1.0. The published v1.1.0 tag is the following commit, 800755a. The v1.1.0 hardening tests also passed with a clean local OpalEmu v1.2.3 checkout at 3df680cd11b40f4590fb64c879e14ea253d5a0c1.
OpalEmu and this project are separate repos with independent versions on purpose: different concerns, different release cadences. In practice, most of what this project depends on isn't OpalEmu-specific at all: button indices, frame counting, screenshot capture, and gameManager.restart() all come from EmulatorJS's stable CDN bundle, a third-party dependency OpalEmu itself just configures. OpalEmu releases mostly don't touch any of that.
The coupling that is real, and worth knowing about before bumping the sibling checkout:
dist/index.html's structure.src/http/static.tsinjectsagent.jswith a literal</body>string-replace. Breaks if OpalEmu's build output changes shape.The
drop-event loading contract.load_romworks by dispatching a synthetic drag-drop event that OpalEmu'suseDragDrop.tslistens for onwindow. Breaks if OpalEmu changes how files get loaded (e.g. drops the drag-drop path in favor of file-input-only).The extension→system list in
src/roms/store.ts. A small local discovery list determines whatlist_romsandload_rom({name})can find. OpalEmu's in-page detection still chooses the actual system. When OpalEmu adds formats, local discovery must be reviewed too.COOP/COEP headers. Mirrored from OpalEmu's own vite/vercel config. Breaks threaded cores (n64, psx, etc.) if OpalEmu's requirements change.
When you update the OpalEmu checkout: rerun npm run test:e2e and npm run test:mcp against it (they boot a real OpalEmu build and exercise the full tool path), then bump the version line above if they pass.
Development
npm run dev # run from source via tsx, no build step
npm run test:e2e # scripted check against a real ROM, no LLM/MCP client needed
npm run test:mcp # spawns the real server and talks real MCP stdio JSON-RPC to it
npm run test:mcp-app # verifies the MCP Apps ui:// resource is registered and well-formed
npm run test:security # network, bridge, and ROM safety regression tests
npm run test:browser # boots the local OpalEmu build and checks the authenticated browser bridgetest:e2e and test:mcp need a ROM in roms/; they use whichever one they find first. test:mcp-app checks the UI resource without a ROM. Build first with npm run build before running test:mcp-app, since it spawns the compiled server. test:security does not need a ROM or browser. test:browser needs a built OpalEmu checkout and Playwright Chromium, but no ROM.
The emulator and MCP smoke scripts accept the server's --http-port and --ws-port flags. The two MCP smoke tests default to 4193/4194 so they do not collide with a running instance; pass the flags explicitly if you need something else:
npm run test:mcp -- --http-port 5000 --ws-port 5001Known limitations
Ambiguous disc formats.
.bin,.iso, or.chdfiles that OpalEmu can't confidently identify (e.g. it can't tell PSX from Sega CD) normally prompt the user with a picker dialog. There's no automated path through that dialog here, soload_romwill time out on such a file. Unambiguous formats (cartridge-based systems, clearly-identified discs) are unaffected.First load per system downloads a core. EmulatorJS cores (5-30MB) download from
cdn.emulatorjs.orgon first use per system;load_romaccounts for this with a generous timeout, and a persistent browser profile (.playwright-profile/) means it only happens once.Stale browser profile data. Cached assets in an older profile may prevent a game from booting after an OpalEmu or EmulatorJS update. Use
--playwright-profile-dir <new-path>to try a fresh profile without deleting the old one.MCP client deadlines. A first
load_rommay need up to 100 seconds, while some MCP clients default to a shorter tool-call deadline. Increase the client's timeout for cold core downloads if it supports that setting.One browser tab, one game at a time. A second
load_romcall replaces the current session; there's no multi-instance support.Frame targets are approximate.
skip_framesmeasures core frames after each display callback and may advance beyond the requested count. Its result states the actual count; it fails if the target is not reached before timeout.EmulatorJS assets are controlled by OpalEmu. The separate OpalEmu checkout currently references the floating
stable/dataCDN path. This project cannot pin its loader and cores without changing OpalEmu or rewriting its built assets. Review that upstream dependency before deployments that require reproducible emulator assets.
Security
See SECURITY.md for private vulnerability reporting guidance. The local HTTP and WebSocket ports are intended only for the Playwright browser started by this process.
Available Tools
6 toolscontrol_emulatorPress or release a buttonADestructive
Presses or releases a controller button. The emulator briefly resumes just long enough for the input to register, then pauses again and returns a PNG screenshot. This advances the game and may affect progress. To hold a direction across skip_frames calls, send state="down" once and state="up" later.
| Name | Required | Description | Default |
|---|---|---|---|
| state | Yes | "down" to press, "up" to release. | |
| button | Yes | Button to press/release. | |
| player | No | Controller port, 0-indexed. Defaults to player 1. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true and idempotentHint=false, but the description explains why: the emulator resumes, input registers, then it pauses and returns a PNG, and the game 'advances and may affect progress.' This adds concrete behavioral context beyond the structured hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the action, then side effects, then the hold pattern. No filler or repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite no output schema, the description states the return (a PNG screenshot), the emulator state transitions, and the progress side effect. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% (enums document button/state, defaults document player), so baseline is 3; the description earns an extra point by clarifying the down/up lifecycle needed to hold a direction across calls, which the schema's terse enum descriptions do not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Presses or releases a controller button') that an agent can immediately distinguish from siblings like get_current_screen or skip_frames. The added detail about resuming the emulator to register input sharpens the purpose further.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names a sibling (skip_frames) and gives the pattern for holding a direction across calls (state=down once, state=up later), which is real routing guidance. It does not, however, cover when to prefer this over other siblings such as get_current_screen, so it stops short of full when/when-not coverage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_current_screenGet current screenARead-onlyIdempotent
Returns the last captured emulator frame as a PNG image without advancing the game. Fails if no game is loaded.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description adds genuinely new behavior: the return type (PNG image), that game state is not advanced, and the failure condition when no game is loaded. It stops short of describing caching or image dimensions, but the key state and error semantics are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero filler; the return value is front-loaded and the constraint and failure mode follow. Every clause carries information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only tool with full annotation coverage, the description supplies everything remaining: return format (PNG), state behavior (no advance), and the error case (no game loaded). No output schema is needed given the return type is stated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate and the baseline of 4 applies. No parameter text is needed or missing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb+resource ('Returns the last captured emulator frame') and immediately disambiguates from the sibling skip_frames by stating it works 'without advancing the game.' An agent can pick this over skip_frames without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'without advancing the game' gives clear context for when this is the right choice versus frame-advancing siblings, but it never names an alternative tool explicitly or states when not to use it. Clear context, no exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_romsList ROMsARead-onlyIdempotent
Lists recognized ROM files in the configured local ROM directory. Returns full filenames, names, likely systems, and sizes. Use a full filename with load_rom; a unique name without its extension also works. Recognition does not guarantee a game will boot.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent, non-destructive, and closed-world behavior. The description adds value beyond them by enumerating the returned fields (filenames, names, likely systems, sizes) and by warning that "Recognition does not guarantee a game will boot," an important expectation-setting caveat not encoded anywhere else.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the core purpose, then returns, then the handoff to load_rom, then the boot caveat. Every sentence carries distinct information and none is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no input parameters and no output schema, the description carries the burden of describing results, and it does so adequately while also flagging the reliability caveat. An agent has everything needed to call it and interpret the response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4 and there is nothing for the description to disambiguate. The description correctly avoids inventing parameters and instead describes the shape of the results.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ("Lists recognized ROM files") and scopes it to "the configured local ROM directory," making clear this is the discovery tool rather than a load/control tool. It is easily separable from siblings like load_rom or get_current_screen.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It routes the agent explicitly to load_rom and explains which identifier form (full filename, or unique name without extension) is accepted there, which is real usage guidance. It stops short of stating when-not to call it, but for a zero-argument listing tool the context is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
load_romLoad ROMADestructive
Loads a ROM into the emulator and starts the game, paused, ready for the next tool call. Provide either "name" (from list_roms) or "romBase64" + "fileName" for a ROM not in the library. Returns the first PNG screenshot once the game has booted. Replaces the current game and its unsaved progress. Local ROM files are limited to 128 MiB; direct Base64 ROMs are limited to 6 MiB decoded. Some disc images need multiple files or manual system selection and cannot load through this tool. First use of a system may download a core and can take up to 100 seconds; MCP clients with shorter call deadlines may time out.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | A ROM name or fileName from list_roms(). | |
| fileName | No | Required with romBase64; use a plain filename with an extension, such as game.nes. | |
| romBase64 | No | Base64-encoded ROM bytes, up to 6 MiB decoded over MCP stdio. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true and non-idempotent, and the description reinforces and extends them: it names exactly what is destroyed ('the current game and its unsaved progress'), gives size ceilings (128 MiB local, 6 MiB decoded Base64), and warns that first use of a system may download a core taking up to 100 seconds and time out MCP clients. This is behavioral context well beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and post-condition, then alternatives, then constraints and caveats. Dense but every sentence carries a distinct, actionable fact; nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still tells the agent what comes back ('the first PNG screenshot once the game has booted'), and covers destructive effect, input modes, size limits, unsupported disc images, and latency/timeout risk. Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the three parameters are already documented; the description nonetheless adds the mutual-exclusion rule (either 'name' or 'romBase64'+'fileName') and the decoded-size limit that the schema's maxLength does not express in human terms. It stops short of clarifying conflict handling if both 'name' and 'romBase64' are supplied.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Loads a ROM into the emulator') plus the resulting state ('starts the game, paused, ready for the next tool call'). It also names the sibling it draws its input from (list_roms), so an agent can distinguish it from list_roms/reset_emulator without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use 'name' versus 'romBase64' + 'fileName', and adds an exclusion: some disc images needing multiple files or manual system selection 'cannot load through this tool.' That is when/when-not guidance rather than mere context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reset_emulatorReset emulatorADestructive
Hard-resets the loaded game and returns a PNG screenshot. Unsaved game progress may be lost. Fails if no game is loaded.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=false; the description adds the concrete consequence ('Unsaved game progress may be lost') and the failure mode when no game is loaded, which is real value beyond the structured flags. It does not cover whether the reset is reversible or what state the emulator ends in.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action, then the risk, then the precondition. No filler; every clause conveys usable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, yet the description states the return type (PNG screenshot), and the mutation risk is covered alongside destructive annotations. Adequate for a zero-parameter tool, with only minor gaps about post-reset state.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; baseline 4 applies. The description correctly introduces no phantom arguments.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (hard-resets) and resource (the loaded game), plus the return artifact (PNG screenshot). It is clearly distinguishable from list_roms, load_rom and get_current_screen, though it never explicitly names a sibling or contrast case.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a precondition ('Fails if no game is loaded') which implies the tool only makes sense with an active ROM, but says nothing about when to prefer a reset over control_emulator operations or how it relates to load_rom.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
skip_framesSkip framesADestructive
Runs until at least the requested number of core frames have advanced, then pauses and returns a PNG screenshot with the actual frame count. Frame polling can overshoot the target. Long waits may use fast-forward and fail on timeout.
| Name | Required | Description | Default |
|---|---|---|---|
| count | Yes | Number of frames to advance. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and destructiveHint=true, so mutation is established. The description adds real behavioral context beyond that: polling can overshoot the requested count, long waits may fast-forward, and the operation can fail on timeout. It does not explain the destructiveness of lost prior state, which is the one gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core behavior and outcome, then the two caveats. Every sentence carries distinct information with no padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly carries the return burden by naming the PNG screenshot and the actual frame count. It also covers overshoot and timeout failure. The only omission is any precondition about the emulator needing a loaded ROM, which is left to sibling context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single parameter is self-documented, so the baseline would be 3. The description earns above baseline by clarifying that count is a minimum rather than an exact advance ('at least the requested number', 'can overshoot'), which is genuine semantics the schema does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource (advance core frames until a target is met, then pause) and names its return artifact (PNG screenshot plus actual frame count). This distinguishes it from the read-only sibling get_current_screen, which only captures the current view without advancing state.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied — advance frames when you need the emulator to progress — but no alternative is named and there is no explicit when-not guidance (e.g., 'use get_current_screen if you do not want to advance state'). The behavioral notes hint at appropriate conditions but do not route the agent between siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v1.1.0- Changed
load_rom7 fields changed- changed
Input schema / properties / fileName / descriptionPrevious value: -"Required with romBase64 — needs a real extension (e.g. \"game.nes\") for system detection."New value: +"Required with romBase64; use a plain filename with an extension, such as game.nes." - added
Input schema / properties / fileName / maxLengthAdded value: +255 - added
Input schema / properties / fileName / minLengthAdded value: +1 - added
Input schema / properties / name / maxLengthAdded value: +255 - added
Input schema / properties / name / minLengthAdded value: +1 - changed
Input schema / properties / romBase64 / descriptionPrevious value: -"Base64-encoded ROM bytes, for a ROM not in the library."New value: +"Base64-encoded ROM bytes, up to 6 MiB decoded over MCP stdio." - added
Input schema / properties / romBase64 / maxLengthAdded value: +8388608
6 tool updates
v0.1.0- First observed
control_emulator - First observed
get_current_screen - First observed
list_roms - First observed
load_rom - First observed
reset_emulator - First observed
skip_frames
TDQS
Scored across 6 tools
Each tool has a clearly distinct purpose: listing ROMs, loading, resetting, capturing the current screen, sending controller input, and advancing frames. The three tools that return screenshots (get_current_screen, control_emulator, skip_frames) are differentiated by whether they advance the game, and the descriptions make these boundaries explicit.
All tool names follow a consistent snake_case verb_noun pattern: list_roms, load_rom, reset_emulator, get_current_screen, control_emulator, skip_frames. There are no mixed conventions or vague standalone verbs.
Six tools is well-scoped for an emulator control server. Each tool earns its place, covering discovery, loading, reset, screen capture, input, and frame advancement without redundancy.
The core emulation loop is covered, but save/load state operations are absent, meaning agents cannot preserve progress before risky actions or reset. An explicit stop/unload tool and system selection for problematic disc images are also missing, creating notable gaps.
Maintenance
Related MCP Connectors
Hosted MCP server connecting claude.ai, ChatGPT and other AI apps to your own computer
MCP server for building and testing AI agents with multi-model experimentation and insights.
Cloud-hosted MCP server for durable AI memory
Related MCP Servers
- AlicenseDqualityDmaintenanceAn MCP server that enables LLMs to 'see' what's happening in browser-based games and applications through vectorized canvas visualization and debug information.16 npm54MIT
- AlicenseAqualityDmaintenanceA Model Context Protocol server that enables LLMs to interact with a GameBoy emulator, providing tools for controlling the GameBoy, loading ROMs, and retrieving screen frames.1336MIT
- AlicenseNot gradedqualityDmaintenanceA Model Context Protocol server that allows LLMs to interact with Game Boy games through PyBoy emulation, providing capabilities to load ROMs, control games, capture screens, save/load states, and maintain game knowledge.1MIT
- AlicenseAqualityCmaintenanceAn MCP server for the Pyxel retro game engine that enables AI models to autonomously run, verify, and iterate on retro game programs. It includes tools for visual verification through screenshots, sprite and layout analysis, and audio rendering.829MIT