JavaScript VM
Provides a tool to run JavaScript snippets inside a fresh isolated V8 isolate, capturing stdout, stderr, return value, and errors, with support for top-level await and a bridged fetch API.
Provides a tool to run TypeScript snippets by compiling them with esbuild to JavaScript and executing them in the same isolated V8 VM, with type erasure and exact stack trace line numbers.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@JavaScript VMrun this JavaScript: console.log([1,2,3].map(x => x*2))"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
JS-VM-MCP-Server
Usage
{
"mcpServers": {
"JavaScript VM": {
"command": "npx",
"args": [
"--allow-git",
"all",
"-y",
"github:Quicksilver0218/JS-VM-MCP-Server"
]
}
}
}Related MCP server: Netmind Code Interpreter
Tools
There are 2 available tools:
run_javascript— runs the snippet as-is.run_typescript— compiles the snippet with esbuild (loader: "ts",target: "esnext"), then runs the emitted JavaScript through the exact same path asrun_javascript.
Both VM tools behave identically; run_typescript just strips the types off first.
Run
Executes the submitted snippet inside a fresh isolated-vm V8
isolate and returns the captured stdout, stderr, the top-level return value and any failure.
The first line of the response is always Code execution finished in #.### seconds..
Per call:
Limit | Value |
V8 heap | 512 MB |
Wall clock | 30 seconds |
Buffered output | 262,144 characters per stream |
The isolate is a separate heap with no access to the host process, so there is no require,
process, timer or file-system API. Each call gets a brand new isolate, so nothing is
shared between runs. The snippet is evaluated as the body of an async function, which means
top-level await and top-level return both work. import/export module syntax is not
supported in either tool.
fetch
The isolate has no network stack of its own, so fetch is bridged to the host process, which
performs the actual request and copies the buffered response back into the isolate:
Limit | Value |
Timeout per request | 15 seconds (rejects with |
Request/response body | 8,388,608 bytes |
Only http: and https: URLs are allowed and the host sends no cookies. fetch, Headers,
Response, AbortController, TextEncoder and TextDecoder are minimal shims:
response.text()/json()/arrayBuffer()/clone() work, but response.body is null
(no streams, no blob()), URL/Request/FormData/Blob do not exist, and the isolate has
no timers, so setTimeout and AbortSignal.timeout are unavailable.
run_typescript caveats:
Type annotations, interfaces, type aliases and generics are erased, and line numbers in runtime stack traces stay exact.
enum,namespaceand constructor parameter properties emit new code, which can push a runtime stack trace past the line you wrote.esbuild reports compile errors on the line numbers you submitted.
Available Tools
2 toolsrun_javascriptRun JavaScriptA
Execute JavaScript inside a hardened V8 isolate (isolated-vm) and return everything it printed.
Each call gets a brand new isolate, so no state, memory or globals are shared between runs. The isolate is a separate heap with no access to the host process: there is no require, process, timer or file-system API available.
Limits, per call:
512 MB of V8 heap.
30 seconds of wall time; the isolate is disposed when the limit is hit.
fetch is available: the isolate has no network stack of its own, so each request is performed by the host process and the response is copied back in. Per request:
10 seconds of wall time; a slower request rejects with a
TimeoutError.Bodies in either direction are capped at 8388608 bytes.
http:andhttps:URLs only; the host sends no cookies and follows redirects unlessredirectsays otherwise.
Language notes:
The first line of the response reports how long the run took.
The snippet runs as the body of an async function, so top-level
awaitand top-levelreturnare both allowed.console.log/console.info/console.debugare captured as stdout;console.warn,console.errorandconsole.traceare captured as stderr.fetch,Headers,Response,AbortController,TextEncoderandTextDecoderare minimal shims:response.text(),response.json(),response.arrayBuffer()andresponse.clone()work, butresponse.bodyisnull(no streams, noblob()),URL/Request/FormData/Blobdo not exist, and the isolate has no timers, sosetTimeoutandAbortSignal.timeoutare unavailable.A top-level
returnvalue is reported back underresult:. - Uncaught exceptions, syntax errors and their stack traces are reported back undererror:, and the tool is flagged as an error.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | JavaScript code to run |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only openWorldHint=true as an annotation, the description carries the full behavioral burden and does so thoroughly: it discloses isolate isolation, no shared state, memory and wall-time limits, fetch behavior and timeouts, URL restrictions, output capture channels, and error reporting. It also warns about missing APIs and shim limitations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but front-loaded and mostly dense with important constraints. It is appropriately sized for a complex sandboxed-execution tool, though there is minor redundancy (timers are mentioned twice, and the host network proxy role is restated).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although there is no output schema, the description explains the response structure: the first line reports elapsed time, stdout and stderr capture, a `result:` channel, and an `error:` channel with error flagging. Combined with the runtime limits and network constraints, it gives an agent everything needed to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% (the single `code` parameter is documented), so the baseline is 3. The description adds meaningful runtime semantics for that parameter: the code runs as the body of an async function, top-level await and return are allowed, console methods are mapped to stdout/stderr, and return/error handling is explained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence gives a specific verb (Execute), resource (JavaScript), and execution environment (hardened V8 isolate), and says it returns printed output. It does not explicitly differentiate itself from the sibling run_typescript, so an agent must infer that this is for JavaScript rather than TypeScript.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the runtime environment and limits but never says when to choose run_javascript over run_typescript or when not to use it. No selection guidance or exclusions are provided despite there being a clear sibling tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_typescriptRun TypeScriptA
Execute TypeScript inside a hardened V8 isolate (isolated-vm) and return everything it printed.
Each call gets a brand new isolate, so no state, memory or globals are shared between runs. The isolate is a separate heap with no access to the host process: there is no require, process, timer or file-system API available.
Limits, per call:
512 MB of V8 heap.
30 seconds of wall time; the isolate is disposed when the limit is hit.
fetch is available: the isolate has no network stack of its own, so each request is performed by the host process and the response is copied back in. Per request:
10 seconds of wall time; a slower request rejects with a
TimeoutError.Bodies in either direction are capped at 8388608 bytes.
http:andhttps:URLs only; the host sends no cookies and follows redirects unlessredirectsays otherwise.
Language notes:
The first line of the response reports how long the run took.
The snippet is compiled with esbuild (types stripped) and then executed exactly like
run_javascript: same isolate, same limits, same output capture.import/exportmodule syntax is not supported: the isolate has no module resolver, so the snippet has to be self-contained.Constructs that emit new code (
enum,namespace, constructor parameter properties) make esbuild produce extra lines, so a runtime stack trace can point past the line you wrote. Plain type annotations keep line numbers exact.The snippet runs as the body of an async function, so top-level
awaitand top-levelreturnare both allowed.console.log/console.info/console.debugare captured as stdout;console.warn,console.errorandconsole.traceare captured as stderr.fetch,Headers,Response,AbortController,TextEncoderandTextDecoderare minimal shims:response.text(),response.json(),response.arrayBuffer()andresponse.clone()work, butresponse.bodyisnull(no streams, noblob()),URL/Request/FormData/Blobdo not exist, and the isolate has no timers, sosetTimeoutandAbortSignal.timeoutare unavailable.A top-level
returnvalue is reported back underresult:. - Uncaught exceptions, syntax errors and their stack traces are reported back undererror:, and the tool is flagged as an error.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | TypeScript code to run |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only openWorldHint in the annotations, the description carries the full burden and does so thoroughly: per-call isolation and no shared state, heap and wall-time limits with disposal behavior, fetch proxying with its own timeout/byte caps/redirect rules, and the exact stdout-vs-stderr console mapping. It also documents how results and failures surface (result:, error:, error flag), which is behavior no annotation or schema conveys.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the one-line purpose, then tightly scoped sections (limits, fetch, language notes). It is long, and some fetch-shim detail could be trimmed, but for a sandboxed-execution tool nearly every bullet prevents a real failure mode, so the length is largely earned.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although there is no output schema, the description explicitly covers the return shape (first line timing, 'result:' for the return value, 'error:' plus stack traces on failure). Combined with the limit and shim disclosures, an agent has everything needed to call this correctly on the first try.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and there is a single 'code' param, so the baseline is 3. The description goes beyond the schema by explaining the semantics of that code string — it runs as an async function body, top-level await and return are legal, and it must be self-contained — which materially changes what the agent writes.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Execute TypeScript inside a hardened V8 isolate') with the exact runtime named, so the agent immediately knows this runs TS rather than plain JS. It also names the sibling mechanism ('executed exactly like run_javascript: same isolate, same limits'), making the distinction between the two tools concrete.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear operating context (type-stripping via esbuild, self-contained snippets) and explicit exclusions ('import'/'export' not supported, no module resolver; no timers, no URL/Request/FormData/Blob). What it does not do is state directly when to prefer this over run_javascript — the routing is implied by the TypeScript-vs-JavaScript distinction rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.0.1- First observed
run_javascript - First observed
run_typescript
TDQS
Scored across 2 tools
The two tools are distinguished primarily by language (JavaScript vs TypeScript), which is a clear and predictable selection criterion. However, TypeScript is close to a superset of JavaScript, so a plain JS snippet often runs fine in either tool, leaving a narrow band of genuine overlap.
Both names follow the same verb_noun snake_case pattern (run_javascript, run_typescript) with a single shared verb and a language noun. The naming is fully predictable and symmetric.
Two tools is on the thin side, but the server's scope is narrowly 'execute code in a sandbox,' and one tool per supported language is a defensible, minimal surface. No tool feels redundant or padded.
Core execution, output capture, error reporting, and fetch are covered, which is the full lifecycle for a stateless code sandbox. Gaps are deliberate constraints rather than omissions: no session/state persistence, no module resolution, no package installation, and no way to pass structured inputs into a run.
Maintenance
Related MCP Connectors
Execute code in 8 languages (Python, JS, TS, Go, Java, C++, C, Bash) in gVisor sandboxes.
Hosted one-shot code execution: send a snippet plus optional files, get stdout, stderr, exit...
Run Python code in a secure sandbox without local setup. Declare inline dependencies and execute s…
- mcp-serverOAuthai.cdbx
Build Apps and run code in 30 languages — sandboxed, with persistent sessions for agent loops.
Related MCP Servers
- FlicenseBqualityDmaintenanceProvides a secure, isolated JavaScript execution environment with configurable time and memory limits for safely running code from Claude.136 npm5-
- AlicenseNot gradedqualityDmaintenanceEnables secure cloud-based execution of code across 14+ programming languages within a sandboxed environment. It supports file management, standard input/output handling, and automatic generation of visual artifacts like plots and charts.MIT
- AlicenseAqualityBmaintenanceEnables AI assistants to securely execute Python and C++ code snippets inside disposable, isolated Docker containers and receive structured results.31MIT
- FlicenseNot gradedqualityCmaintenanceExecutes Python code in isolated subprocesses with resource limits and timeout, designed for safe execution in Docker containers.-