Skip to main content
Glama

JS-VM-MCP-Server

Usage

{
  "mcpServers": {
    "JavaScript VM": {
      "command": "npx",
      "args": [
        "--allow-git",
        "all",
        "-y",
        "github:Quicksilver0218/JS-VM-MCP-Server"
      ]
    }
  }
}

Related MCP server: Netmind Code Interpreter

Tools

There are 2 available tools:

  • run_javascript — runs the snippet as-is.

  • run_typescript — compiles the snippet with esbuild (loader: "ts", target: "esnext"), then runs the emitted JavaScript through the exact same path as run_javascript.

Both VM tools behave identically; run_typescript just strips the types off first.

Run

Executes the submitted snippet inside a fresh isolated-vm V8 isolate and returns the captured stdout, stderr, the top-level return value and any failure. The first line of the response is always Code execution finished in #.### seconds..

Per call:

Limit

Value

V8 heap

512 MB

Wall clock

30 seconds

Buffered output

262,144 characters per stream

The isolate is a separate heap with no access to the host process, so there is no require, process, timer or file-system API. Each call gets a brand new isolate, so nothing is shared between runs. The snippet is evaluated as the body of an async function, which means top-level await and top-level return both work. import/export module syntax is not supported in either tool.

fetch

The isolate has no network stack of its own, so fetch is bridged to the host process, which performs the actual request and copies the buffered response back into the isolate:

Limit

Value

Timeout per request

15 seconds (rejects with TimeoutError)

Request/response body

8,388,608 bytes

Only http: and https: URLs are allowed and the host sends no cookies. fetch, Headers, Response, AbortController, TextEncoder and TextDecoder are minimal shims: response.text()/json()/arrayBuffer()/clone() work, but response.body is null (no streams, no blob()), URL/Request/FormData/Blob do not exist, and the isolate has no timers, so setTimeout and AbortSignal.timeout are unavailable.

run_typescript caveats:

  • Type annotations, interfaces, type aliases and generics are erased, and line numbers in runtime stack traces stay exact.

  • enum, namespace and constructor parameter properties emit new code, which can push a runtime stack trace past the line you wrote.

  • esbuild reports compile errors on the line numbers you submitted.

Available Tools

2 tools
run_javascriptRun JavaScriptA

Execute JavaScript inside a hardened V8 isolate (isolated-vm) and return everything it printed.

Each call gets a brand new isolate, so no state, memory or globals are shared between runs. The isolate is a separate heap with no access to the host process: there is no require, process, timer or file-system API available.

Limits, per call:

  • 512 MB of V8 heap.

  • 30 seconds of wall time; the isolate is disposed when the limit is hit.

fetch is available: the isolate has no network stack of its own, so each request is performed by the host process and the response is copied back in. Per request:

  • 10 seconds of wall time; a slower request rejects with a TimeoutError.

  • Bodies in either direction are capped at 8388608 bytes.

  • http: and https: URLs only; the host sends no cookies and follows redirects unless redirect says otherwise.

Language notes:

  • The first line of the response reports how long the run took.

  • The snippet runs as the body of an async function, so top-level await and top-level return are both allowed.

  • console.log/console.info/console.debug are captured as stdout; console.warn, console.error and console.trace are captured as stderr.

  • fetch, Headers, Response, AbortController, TextEncoder and TextDecoder are minimal shims: response.text(), response.json(), response.arrayBuffer() and response.clone() work, but response.body is null (no streams, no blob()), URL/Request/FormData/Blob do not exist, and the isolate has no timers, so setTimeout and AbortSignal.timeout are unavailable.

  • A top-level return value is reported back under result:. - Uncaught exceptions, syntax errors and their stack traces are reported back under error:, and the tool is flagged as an error.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYesJavaScript code to run

TDQS

A3.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With only openWorldHint=true as an annotation, the description carries the full behavioral burden and does so thoroughly: it discloses isolate isolation, no shared state, memory and wall-time limits, fetch behavior and timeouts, URL restrictions, output capture channels, and error reporting. It also warns about missing APIs and shim limitations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but front-loaded and mostly dense with important constraints. It is appropriately sized for a complex sandboxed-execution tool, though there is minor redundancy (timers are mentioned twice, and the host network proxy role is restated).

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although there is no output schema, the description explains the response structure: the first line reports elapsed time, stdout and stderr capture, a `result:` channel, and an `error:` channel with error flagging. Combined with the runtime limits and network constraints, it gives an agent everything needed to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% (the single `code` parameter is documented), so the baseline is 3. The description adds meaningful runtime semantics for that parameter: the code runs as the body of an async function, top-level await and return are allowed, console methods are mapped to stdout/stderr, and return/error handling is explained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence gives a specific verb (Execute), resource (JavaScript), and execution environment (hardened V8 isolate), and says it returns printed output. It does not explicitly differentiate itself from the sibling run_typescript, so an agent must infer that this is for JavaScript rather than TypeScript.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the runtime environment and limits but never says when to choose run_javascript over run_typescript or when not to use it. No selection guidance or exclusions are provided despite there being a clear sibling tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_typescriptRun TypeScriptA

Execute TypeScript inside a hardened V8 isolate (isolated-vm) and return everything it printed.

Each call gets a brand new isolate, so no state, memory or globals are shared between runs. The isolate is a separate heap with no access to the host process: there is no require, process, timer or file-system API available.

Limits, per call:

  • 512 MB of V8 heap.

  • 30 seconds of wall time; the isolate is disposed when the limit is hit.

fetch is available: the isolate has no network stack of its own, so each request is performed by the host process and the response is copied back in. Per request:

  • 10 seconds of wall time; a slower request rejects with a TimeoutError.

  • Bodies in either direction are capped at 8388608 bytes.

  • http: and https: URLs only; the host sends no cookies and follows redirects unless redirect says otherwise.

Language notes:

  • The first line of the response reports how long the run took.

  • The snippet is compiled with esbuild (types stripped) and then executed exactly like run_javascript: same isolate, same limits, same output capture.

  • import/export module syntax is not supported: the isolate has no module resolver, so the snippet has to be self-contained.

  • Constructs that emit new code (enum, namespace, constructor parameter properties) make esbuild produce extra lines, so a runtime stack trace can point past the line you wrote. Plain type annotations keep line numbers exact.

  • The snippet runs as the body of an async function, so top-level await and top-level return are both allowed.

  • console.log/console.info/console.debug are captured as stdout; console.warn, console.error and console.trace are captured as stderr.

  • fetch, Headers, Response, AbortController, TextEncoder and TextDecoder are minimal shims: response.text(), response.json(), response.arrayBuffer() and response.clone() work, but response.body is null (no streams, no blob()), URL/Request/FormData/Blob do not exist, and the isolate has no timers, so setTimeout and AbortSignal.timeout are unavailable.

  • A top-level return value is reported back under result:. - Uncaught exceptions, syntax errors and their stack traces are reported back under error:, and the tool is flagged as an error.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYesTypeScript code to run

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With only openWorldHint in the annotations, the description carries the full burden and does so thoroughly: per-call isolation and no shared state, heap and wall-time limits with disposal behavior, fetch proxying with its own timeout/byte caps/redirect rules, and the exact stdout-vs-stderr console mapping. It also documents how results and failures surface (result:, error:, error flag), which is behavior no annotation or schema conveys.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the one-line purpose, then tightly scoped sections (limits, fetch, language notes). It is long, and some fetch-shim detail could be trimmed, but for a sandboxed-execution tool nearly every bullet prevents a real failure mode, so the length is largely earned.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although there is no output schema, the description explicitly covers the return shape (first line timing, 'result:' for the return value, 'error:' plus stack traces on failure). Combined with the limit and shim disclosures, an agent has everything needed to call this correctly on the first try.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and there is a single 'code' param, so the baseline is 3. The description goes beyond the schema by explaining the semantics of that code string — it runs as an async function body, top-level await and return are legal, and it must be self-contained — which materially changes what the agent writes.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Execute TypeScript inside a hardened V8 isolate') with the exact runtime named, so the agent immediately knows this runs TS rather than plain JS. It also names the sibling mechanism ('executed exactly like run_javascript: same isolate, same limits'), making the distinction between the two tools concrete.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear operating context (type-stripping via esbuild, self-contained snippets) and explicit exclusions ('import'/'export' not supported, no module resolver; no timers, no URL/Request/FormData/Blob). What it does not do is state directly when to prefer this over run_javascript — the routing is implied by the TypeScript-vs-JavaScript distinction rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.0.1
    • First observedrun_javascript
    • First observedrun_typescript

TDQS

A4.2/5.0

Scored across 2 tools

Disambiguation4/5

The two tools are distinguished primarily by language (JavaScript vs TypeScript), which is a clear and predictable selection criterion. However, TypeScript is close to a superset of JavaScript, so a plain JS snippet often runs fine in either tool, leaving a narrow band of genuine overlap.

Naming Consistency5/5

Both names follow the same verb_noun snake_case pattern (run_javascript, run_typescript) with a single shared verb and a language noun. The naming is fully predictable and symmetric.

Tool Count4/5

Two tools is on the thin side, but the server's scope is narrowly 'execute code in a sandbox,' and one tool per supported language is a defensible, minimal surface. No tool feels redundant or padded.

Completeness4/5

Core execution, output capture, error reporting, and fetch are covered, which is the full lifecycle for a stateless code sandbox. Gaps are deliberate constraints rather than omissions: no session/state persistence, no module resolution, no package installation, and no way to pass structured inputs into a run.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers