Skip to main content
Glama
okenjioxx

Roblox Executor MCP Server

by okenjioxx

Time a Luau snippet over repeated runs (micro-benchmark)

profile-code
Destructive

Benchmark a Luau snippet by running it multiple times and timing each execution with os.clock(), returning min, max, average, and total wall-clock stats for comparing code speed.

Instructions

Micro-benchmark a Luau snippet by running it runs times and timing each run with os.clock(), returning the wall-clock statistics so you can measure how fast (or how variable) a piece of code is. Use this to compare two implementations, to find out how expensive a function call / loop / property access really is, or to confirm a fix actually made something faster. The code is COMPILED ONCE via loadstring (a syntax error is returned cleanly as { error } and nothing is run); then each of the runs invocations is timed individually inside its own pcall so a runtime error in one run is counted but does not abort the benchmark. Timing measures the run only — compile time is excluded. Note os.clock() resolution is coarse, so for very cheap snippets raise runs or wrap a loop inside your code. Requires loadstring and os.clock (both guarded). Returns { runs, totalMs, avgMs, minMs, maxMs, errorCount, firstError? } or { error }. Signature: { code: string, runs: any?, threadContext: number? }. Phase: act; cost=high; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval, validated-source. Produces: structured-result. Verify with: assert-state. Safety: MUTATING; writes executor workspace filesystem. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
codeYesThe Luau snippet to benchmark. Compiled once with loadstring, then executed `runs` times. It may do anything (call a function, run a loop, read properties); any value it returns is ignored — only the elapsed time per run is measured. Wrap an inner loop here if a single execution is too cheap to time accurately.
runsNoHow many times to execute the compiled snippet (default 1, clamped 1..100000). More runs give a more stable average for cheap code but take longer. Each run is timed and pcall-guarded independently.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv2.0.0-spies.2

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial behavior beyond the annotations: loadstring compilation happens once and syntax errors return { error } without running anything, each run is individually pcall-guarded so a runtime error is counted but does not abort, compile time is excluded from timing, and os.clock() resolution is coarse so cheap snippets need more runs or an inner loop. It also discloses the dependency on guarded loadstring/os.clock and the mutating filesystem side effect, consistent with destructiveHint=true.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose and usage scenarios are front-loaded and dense with useful detail. The trailing metadata block (Phase/cost/idempotency/Requires/Produces/Verify with/On failure) is boilerplate padding that does not add tool-selection value, but the body itself earns its length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, and the description fully compensates by naming the return shape ({ runs, totalMs, avgMs, minMs, maxMs, errorCount, firstError? } or { error }) and the error path. Combined with the mutation/safety disclosure, an agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so baseline is 3, but the description adds real value: it explains that `runs` is clamped 1..100000 (the schema only shows a loose min/max), that returned values are ignored, and that an inner loop can be wrapped in `code` when a single run is too cheap to time. threadContext is left unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: micro-benchmark a Luau snippet by running it `runs` times and timing each run with os.clock(). The scoping detail (compile once, time each run, exclude compile time) makes it clearly distinct from execution siblings like run-luau or execute, which run code but do not measure it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete when-to-use scenarios: compare two implementations, find the cost of a function call/loop/property access, confirm a fix made something faster. It does not name an alternative tool or state exclusions (e.g., memory profiling belongs to measure-memory), so the routing guidance stops short of fully explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools