Skip to main content
Glama
okenjioxx

Roblox Executor MCP Server

by okenjioxx

Roblox Executor MCP Server

License: MIT Node.js: 20+ TypeScript

An MCP server that connects an AI client to a live Roblox game. The model calls a tool, the server runs Luau in the game, and hands back structured data. The agent can reverse-engineer scripts, walk the instance tree, spy on remotes, scan memory, hook functions, orchestrate work across many connected clients, and more.

It ships 289 tools across 22 categories, a dashboard with ten tabs, persistent playbooks and session traces, a token-gated bridge, and a Luau scripting surface (mcp.*) that lets one in-game script call any of the server's tools — sequentially, in parallel, batched, or across N clients at once. Schemas are introspectable at runtime via mcp.help(name?) so a script never has to guess what arguments a tool takes.

Start here

This project is the bridge between your AI assistant and a running Roblox executor client. It includes the local Node.js server, browser dashboard, and in-game Luau connector. You supply an MCP-compatible AI client and a compatible executor; executor-specific features depend on the capabilities available in that client. The Cobalt integration is documented for Potassium, with live validation still required.

flowchart LR
    AI[AI client] -->|MCP| Server[Local Node.js server]
    Dashboard[Browser dashboard] --> Server
    Server <-->|WebSocket bridge| Connector[Luau connector]
    Connector <--> Roblox[Running Roblox client]

Use it to explore an instance tree, read properties, inspect scripts and functions, capture remote calls, run Luau, or save a repeatable workflow. The dashboard lets you inspect the same connected clients and run tools directly from your browser.

Related MCP server: Roblox Executor MCP Server

What's in the box

Tools (289 across 22 categories)

The big ones you'll reach for first:

  • Run code. run-luau executes Luau and returns JSON; eval-expression is the one-liner. Everything else is well-tested Luau you didn't have to write.

  • script (persistent VM with mcp.*). Write one Luau program that can ALSO call any other tool inline as mcp.<tool>(args) and use the result. Globals you define persist across calls (REPL-style); vm-reset wipes the VM.

  • script-fanout. Run one script on N connected clients in parallel; per-client {result, output} returned with a summary.

  • Reverse engineering. GC walking, closure constants/upvalues/protos, bytecode disassembly, call graphs, duplicate-function detection, filtergc. 34 tools.

  • Remotes. Inventory, signatures, selectable Cobalt/Ketamine incoming/outgoing capture, persistent filters, pause/resume, reversible block/ignore, active-rule inspection/reset, Ketamine GUI controls, and generated call code. 12 tools.

  • Instrumentation. Hook-and-log, count calls, spoof returns, trace durations.

  • Closures. Complete Volt closure primitives: classify/hash/clone/wrap/invoke/hook/restore, retained function handles, stack visibility, environments, constants, upvalues, and protos.

  • Execution-footprint audit. One bounded read-only Luau report for virtual-input provenance, getfenv/global leaks, closure and hook identity, script/source exposure, executor fingerprints, evidence confidence, and truncation telemetry.

  • Actors and Lua states. Actor discovery/execution, full LuaStateProxy inspection/Execute/Event support, communication channels, parallel-context checks, and bounded actor/channel/state event monitors.

  • Hidden surfaces. Actor scripts, nil-parented instances, hidden GUIs, gethui, detached remotes.

  • Discovery. discover-player-values auto-ranks candidate money/score/XP paths from leaderstats / Player / ReplicatedStorage with a scored heuristic walk.

  • Playbooks. Save/list/run/delete named, parameterized Luau snippets persisted to ~/.executor-mcp/playbooks/.

  • Sessions. Every tool call appends to a JSONL session trace; session-list/show/replay browse and re-execute past traces.

  • Discovery aids. list-tools browses the catalog by category. suggest-tools ranks matches by past success.

  • Definition intelligence. Every tool receives a compiled signature, documented fields, defaults/constraints/examples, prerequisites, capability requirements, side effects, verification paths, recovery guidance, and a measurable quality grade. tool-quality-audit checks the whole catalog without a game client.

  • AI planning. tool-plan turns a natural-language goal into a schema-aware discover→act→verify workflow with ranked alternatives.

  • Agent context. agent-context bootstraps the current clients, selection, game, executor, and next actions in one read-only call.

  • Agent runtime. agent-run executes explicit workflows with dry runs, mutation approval, $steps.* references, retries, and automatic verification; agent-memory stores verified facts and successful workflow episodes.

  • World Brain. observe-world fuses the live character, camera, visible GUI, nearby objects, interactables, and tools into bounded semantic handles; resolve-entity safely revalidates or rediscovers stale handles.

  • Verified adaptive tasks. smart-task adds plan/preview/execute modes, hard budgets, loop detection, typed recovery branches, and real assert-state postconditions. explain-failure classifies errors and ranks safe fallbacks without blindly repeating mutations.

  • Rollback and learning. state-transaction restores explicitly captured reversible state, world-delta streams bounded event changes, and teach-mode turns a user demonstration into a conservative reviewable playbook.

Run list-tools once connected for the full catalog, or GET /api/tools/schema for JSON schemas, or GET /mcp.d.luau for Luau type declarations any editor with a Luau LSP can consume.

See docs/architecture/actors-closures.md for the capability-first Actor, LuaStateProxy, channel, event-monitor, and closure workflows. See docs/architecture/execution-footprint-audit.md for the target-resolution, evidence, scoring, privacy, and performance contracts of the footprint auditor. See docs/architecture/tool-definition-quality.md for the library-wide schema, contract, safety, discovery, recovery, and quality compiler.

Dashboard

Open http://127.0.0.1:16384/ once the server's running. Ten tabs, flat-dark, sub-100ms live updates over WebSocket:

  • Clients — connected games with PlaceId/JobId chips and click-to-explore.

  • Tools — category-grouped browser of all 289 tools with search.

  • Activity — live tool-call stream with text/category/outcome filters.

  • Intelligence — bounded live perceive→resolve→act→verify/recover timeline with targets, confidence, evidence, rollback, and teaching state.

  • Explorer — Studio-style game tree with real Studio class icons (314 mapped), Properties + Connections panels, paged children with hover prefetch, double-click decompile tabs, a bounded proto/function tree, origin/upvalue metadata, exact line jumps, and cross-script reference navigation.

  • Brief — Place/Game/JobId metadata, surface counts (RemoteEvent/Script/Tool), Local Player info, top remotes from the spy buffer, Discover Values button, Fanout-across-all-clients starter.

  • Spy — engine selector for Cobalt/Ketamine, filtered captured remote calls, and capture JSON copying.

  • Playbooks — list rail + edit pane + parameter form + Run on selected client + Auto-params button that infers ${param} from string/number literals.

  • REPL — Luau textarea with mcp.* autocomplete from /api/tools/schema, Ctrl+Enter to run on the selected client, Save-as-playbook button.

  • Output — terminal-style print/warn/error stream with per-script scoping, source filter, and a 1.5K-line ring buffer.

Scripting (mcp.*)

Inside a script body, mcp is bound to the whole tool surface. Tool names map kebab → camelCase:

local p = mcp.getPlayers()
local r = mcp.searchInstances({ className = "RemoteEvent" })
print(#p .. " players, " .. #r.instances .. " remotes")

-- look up any tool's args at runtime, no guessing — returns
-- { signature, args = {{name, type, optional, nullable, description, constraints, example}, ...},
--   exampleInput, guidance, quality, compiledDescription, ... }
local schema = mcp.help("discover-player-values")

-- batch N independent calls into one round-trip:
local b = mcp.parallel({
  players = function() return mcp.getPlayers() end,
  money   = function() return mcp.discoverPlayerValues({ limit = 5 }) end,
})

-- cross-game pub/sub:
mcp.subscribe("scores", function(payload, fromClientId)
  print("got", payload.score, "from", fromClientId)
end)
mcp.publish("scores", { score = 100 })

mcp.help(name?) is the in-script equivalent of the top-level tool-schema tool: with a name it returns the full per-field detail; with no argument it returns every tool's compact signature. Use it before calling an unfamiliar tool instead of guessing arg shapes.

At the start of a task, call local context = mcp.agentContext() to learn the active client, game, executor, and available capabilities. For an ambiguous objective, then use local plan = mcp.toolPlan({ goal = "find the player's money and verify it" }). The planner returns ranked tools, exact signatures, mutation/client flags, and one or more discover→act→verify workflows. For multi-step tasks, use the selected tools inside one script call and branch on each result rather than assuming a step succeeded.

mcp.parallel is a real coroutine scheduler — every mcp.* call inside any of the passed functions yields a marker, the scheduler collects markers across all coroutines per round and batches them into ONE rpc-batch. A 5-step recipe across 5 coroutines runs in ~5 round trips, not 25.

How it works

Hexagonal (ports + adapters). The real logic — which client a session targets, how a tool call gets validated and run, what the errors mean — is plain TypeScript that has no idea WebSockets, the MCP SDK, or pino exist. Tests run with fakes; no socket, no SDK, no game.

domain/          pure types and rules, no dependencies
application/     ports (interfaces) + use-cases + the Tool contract
infrastructure/  adapters: WebSocket bridge, MCP stdio, pino, dashboard, ...
tools/           the tools themselves, each a defineTool() plugin
interface/       main.ts, the only file that knows about concrete adapters

Imports only point inward. tools and infrastructure lean on application, application leans on domain, and nothing in the core reaches back out.

Adding a tool is one file:

export default defineTool({
  name: "get-health",
  category: "Inspection",
  input: z.object({ path: z.string() }),
  async execute({ path }, ctx) {
    const hp = await ctx.runLuau(`return ${path}.Humanoid.Health`);
    return { data: { hp } };
  },
});

The tool never touches the transport and never picks a client. The invoker resolves the active client first and hands you a ctx that's already bound. Longer write-up in docs/architecture/overview.md; decisions are ADRs under docs/adr/.

Cobalt remote spy (Potassium)

You can also select Ketamine with remote-spy input {"operation":"start","engine":"ketamine"}. The AI can capture both directions, block/unblock outgoing remotes and incoming function callbacks, and switch engines one at a time. configure-remote-spy adds persistent capture filters, pause/resume, buffer resizing, and Ketamine GUI visibility/logging; remote-spy can inspect/reset active block and ignore rules. Cobalt shares the capture and rule-management improvements and retains RakNet support. See the dual-spy setup guide for requirements, examples, and limitations.

This build keeps Polaris as the MCP server and uses bundled Cobalt 2.2.5.15 as its remote-spy engine. All capture tools and the dashboard share one Cobalt subscription. Start with ensure-remote-spy using { "mode": "raknet" }, or use the dashboard's Start Cobalt button. remote-spy adds lifecycle/status controls, ranked captures, and call-code generation.

See COBALT_REMOTE_SPY.md for installation, MCP configuration, examples, and validation limits. The bundled Cobalt build includes a callback/handle cleanup fix for Potassium. Live Potassium testing is still required.

Windows dashboard shortcut

For the repair ZIP, extract every file directly into your existing project folder and double-click START-CobaltDashboard.cmd. It installs locked dependencies and opens the dashboard after the compiled server passes its health check. See REPAIR-FIRST.md. MCP clients continue to use dist/interface/launcher.js.

Quick start

Requires Node.js 20+, pnpm 10.27.0, an MCP-compatible AI client, and a Roblox executor that supports the connector's WebSocket and Luau APIs. This source checkout must be built before it can run.

git clone https://github.com/okenjioxx/executor-mcp-roblox.git
cd executor-mcp-roblox
pnpm install --frozen-lockfile
pnpm build
pnpm start        # or pnpm dev for watch mode

For Windows, after running pnpm build, you can double-click START-CobaltDashboard.cmd from the project folder to install locked dependencies and open the dashboard.

Configure your AI client

Add this server to your client's MCP configuration. Replace the example path with the absolute path to your checkout; the exact settings location depends on your client.

{
  "mcpServers": {
    "roblox-executor": {
      "command": "node",
      "args": ["C:/Tools/executor-mcp-roblox/dist/interface/launcher.js"]
    }
  }
}

Restart or reconnect your MCP client after saving the configuration. On macOS or Linux, use a path such as /home/you/executor-mcp-roblox/dist/interface/launcher.js.

The server speaks MCP over stdin/stdout. Point your client (Claude, Cursor, Windsurf, anything that speaks MCP) at node /path/to/executor-mcp-roblox/dist/interface/launcher.js. The launcher starts the owner automatically, reuses an already-running owner, and proxies later MCP stdio sessions to it so multiple host windows do not collide on the bridge port. Every proxied connection gets an isolated logical agent session, so its selected Roblox client cannot overwrite another agent's selection. All agents share the same bounded per-client execution scheduler. It also coordinates simultaneous starts with a user-owned lock, removes stale locks, buffers early MCP messages, and can take over after an owner exits. Logs go to stderr; stdout is the protocol channel.

Launcher tuning is optional. ROBLOX_MCP_LAUNCHER_DEBUG=1 enables startup/proxy diagnostics; ROBLOX_MCP_RUNTIME_DIR changes the lock directory; and the ROBLOX_MCP_LAUNCHER_* timeout/retry variables can tune slow machines without changing the MCP command.

Connecting the game

Paste in your executor (or add to autoexec):

getgenv().BridgeURL = "127.0.0.1:16384"
-- Optional: if the server has ROBLOX_MCP_BRIDGE_TOKEN set, mirror it here.
-- getgenv().BridgeToken = "your-shared-secret"
loadstring(game:HttpGet("http://" .. getgenv().BridgeURL .. "/connector.luau"))()

The connector pulls itself from the server, opens a WebSocket to ws://<BridgeURL>/bridge, sends a hello with its identity + probed capabilities, and from then on runs whatever the server asks and replies with JSON. Wire shapes live in src/domain/protocol/messages.ts.

Verify your connection

  1. Open http://127.0.0.1:16384/ and confirm your game appears in Clients.

  2. Ask your AI client to call list-clients, then select-client for the client you want to inspect.

  3. Call get-game-info to confirm the selected game, then get-instance-tree to begin exploring.

  4. Use list-tools and tool-schema to discover available tools and their arguments.

Troubleshooting

Symptom

What to check

AI client cannot start the server

Run pnpm build, confirm node is available, and use an absolute path to dist/interface/launcher.js.

Dashboard opens but no game appears

Run the connector in your executor after joining Roblox; confirm BridgeURL matches the server's host and port.

Connection is rejected after enabling a token

Set the same value in server-side ROBLOX_MCP_BRIDGE_TOKEN and connector-side getgenv().BridgeToken.

A tool reports an unsupported capability

Run test-capabilities or get-executor-info; support depends on your executor's APIs.

An older build is still running

Stop the old server and reconnect your MCP client; the launcher reuses an existing owner on the same port.

Cobalt captures are missing

Follow the Cobalt setup guide, enable RakNet in Potassium when using RakNet mode, and check spy status.

Automated tests exercise host logic, bridge behavior, and mocked tool execution. They do not establish compatibility with every executor or replace a live Roblox test. See validation notes for the tested scope and known limits.

Configuration

Read once at startup, validated, then passed read-only. No process.env access after that.

Flag

Env var

Default

--port

ROBLOX_MCP_PORT

16384

Bridge + dashboard port.

--host

ROBLOX_MCP_HOST

127.0.0.1

Bind address. Keep on loopback unless you've thought about it.

--session-label

ROBLOX_MCP_SESSION_LABEL

generated

Friendly name for this process.

--no-dashboard

off

Disable the dashboard entirely.

ROBLOX_MCP_BRIDGE_TOKEN

unset

When set, the bridge AND dashboard require this token. Connector reads getgenv().BridgeToken.

ROBLOX_MCP_LOG_LEVEL

info

trace through fatal.

ROBLOX_MCP_LOG_PRETTY

off

1 for human-readable logs.

ROBLOX_MCP_RUNTIME_DIR

~/.executor-mcp

Directory for per-port launcher locks.

ROBLOX_MCP_LAUNCHER_DEBUG

off

1 to log owner discovery, lock, startup, and proxy transitions.

ROBLOX_MCP_LAUNCHER_READY_TIMEOUT_MS

15000

Maximum time to wait for a newly spawned owner to expose health + MCP.

ROBLOX_MCP_LAUNCHER_MAX_START_ATTEMPTS

4

Startup retries after a bind/process race.

ROBLOX_MCP_MAX_CONCURRENT_EVALS

2

Active eval lanes per Roblox client; one lane is reserved for nested mcp.* work.

ROBLOX_MCP_MAX_QUEUED_EVALS

128

Bounded waiting evals per client; overflow returns retryable BRIDGE_OVERLOADED.

ROBLOX_MCP_MAX_QUEUED_SOURCE_BYTES

4194304

Total queued Luau source bytes per client.

ROBLOX_MCP_RPC_BATCH_CONCURRENCY

8

Host workers used inside one in-script RPC batch.

ROBLOX_MCP_MAX_RPC_BATCH_CALLS

128

Calls accepted from one RPC batch before later entries receive a bounded error.

ROBLOX_MCP_MAX_CONCURRENT_RPC_FRAMES

2

Inbound script RPC frames processed per client.

ROBLOX_MCP_MAX_QUEUED_RPC_FRAMES

32

Waiting inbound script RPC frames per client.

ROBLOX_MCP_SCRIPT_DIRS

—

Extra folders execute-file may read.

ROBLOX_MCP_EMBEDDINGS_URL

local

Embeddings endpoint for semantic search (Ollama / OpenAI-compatible).

ROBLOX_MCP_EMBEDDINGS_MODEL

embeddinggemma

Model name passed to the embeddings endpoint.

Other defaults: 30s default per-call timeout, thread identity 8, connector heartbeat every 2s. The default per-script RPC budget is 500 mcp.* calls; scripts can opt in to more via the script tool's rpcBudget input.

The connector independently enforces a second safety layer. Optional executor globals are MCPMaxConcurrentEvals (2), MCPMaxQueuedEvals (96), MCPMaxQueuedSourceBytes (2 MiB), MCPMaxRpcBatchCalls (64), MCPMaxParallelCoroutines (64), MCPOutputBufferLimit (256), MCPOutputBatchLimit (50), MCPOutputMessageLimit (4096), and MCPStreamOutput=false to disable game-log streaming. Overrides are clamped to safe ranges. bridge-status and /api/health expose active, queued, saturated, and rejected load without touching the game.

Persistent storage

The server writes to a few places under ~/.executor-mcp/:

  • playbooks/<name>.json — saved Luau snippets via playbook-save or the dashboard.

  • sessions/<sessionId>.jsonl — append-only trace of every tool call (one line each); read via session-show, replayed via session-replay.

  • embeddings.json — sha256-keyed cache for semantic search; cold-start re-embeds drop from minutes to seconds.

Safety

Read this once.

  • The server runs arbitrary code on your game client. That's the whole point. Only connect AI clients you trust.

  • The bridge binds 127.0.0.1 by default. If you switch --host to 0.0.0.0, keep it behind a LAN, VPN, or SSH tunnel — never the open internet.

  • For shared multi-user machines, set ROBLOX_MCP_BRIDGE_TOKEN to a random string. The bridge then rejects WebSocket handshakes without a matching getgenv().BridgeToken, and the dashboard requires the same token via cookie or X-Executor-MCP-Token header.

  • Tools that mutate game state carry mutatesState: true and say so in their description; the risky surface is easy to spot.

  • session-replay skips originally-failed steps and refuses to replay flagged-mutating tools unless includeMutating:true is set explicitly; it also refuses to recursively call itself.

  • The per-script RPC budget (default 500) caps how much damage a runaway loop in a script can do before the bridge cuts it off.

Layout

src/
  domain/          pure types and rules
  application/     ports, use-cases, the Tool contract
  infrastructure/  adapters: bridge, MCP stdio, dashboard, semantic, playbooks, sessions, config
  tools/           one folder per category, each tool its own defineTool() file
  interface/       main.ts, the composition root
connector/         the in-game Luau connector
assets/            Studio class-icons sprite sheet
docs/              architecture notes + ADRs
test/              unit + integration + helpers

Scripts

Command

pnpm verify

Typecheck, lint, and the full test suite.

pnpm test

Vitest (test:coverage / test:watch variants exist).

pnpm build

Compile to dist/.

pnpm dev

Run the server under tsx watch.

pnpm lint / pnpm format

ESLint and Prettier.

Tests

The core is genuinely easy to test, which was the point of laying it out this way. resolveSelection, the error mapping, ToolInvoker, SessionManager, ScriptBridge, FsSavedScriptsStore, FsSessionLogger, CachedEmbeddingsProvider, and the preflight all run against fake ports — no socket, no SDK, no game. The bridge has full integration tests under test/integration/ that drive a real ws client through the protocol end-to-end, covering rpc-call, rpc-batch, pub/sub, and auth.

Contributing

CONTRIBUTING.md has the setup, the layer rules, how to add a tool, and the PR checklist.

License

MIT. See LICENSE.

Available Tools

291 tools
actor-capabilitiesProbe Actors, LuaStateProxy, channels, and actor eventsB
Read-onlyIdempotent

Read-only matrix for Volt's complete Actors and LuaStateProxy surface, including event objects and aliases. Signature: { threadContext: number? }. Phase: verify; cost=medium; idempotency=read-only. Requires: active-client. Capabilities: getactors. Produces: structured-result. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/idempotentHint/destructiveHint/openWorldHint, so 'idempotency=read-only' and 'Safety: read-only' add no value. The description does add cost=medium and the 'active-client' precondition, which are genuine context beyond the structured fields, but no return-shape or failure behavior detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose leads, which is good, but the dash-separated metadata block is partly redundant ('idempotency=read-only', 'Safety: read-only' duplicate annotations) and 'idempotency=read-only' conflates two distinct properties. Some statements do not earn their place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden but only says 'Produces: structured-result', which is too vague for a probe/matrix tool. 'On failure' guidance pointing to tool-schema helps, but what the matrix actually contains is underspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and there is a single optional threadContext whose schema description already explains it. 'Signature: { threadContext: number? }' merely restates the schema without adding format, default, or scoping meaning, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence names a specific resource: a read-only matrix covering Actors, LuaStateProxy, event objects, and aliases. An agent can distinguish this enumeration/probe surface from siblings like list-actors or get-actor-details. It lacks an explicit contrast to those siblings, hence not a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Phase: verify' and 'Requires: active-client' hint at a workflow position and a precondition, but the description never says when to choose this over list-actors, get-actor-details, or get-lua-state-actors. No when-not or alternative routing is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

actor-event-monitorMonitor actor-ready and actor-state-created eventsA
Destructive

WRITES LIVE GAME STATE when starting/stopping. Connect to on_actor_added and/or on_actor_state_created, retain a bounded 200-event buffer, and poll it without repeated game-tree scans. Start/stop require confirm=true. Signature: { action: "start" | "poll" | "stop", key: any?, limit: any?, clear: any?, confirm: any?, threadContext: number?, events: any? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval. Capabilities: getactors. Produces: bounded-event-snapshot. Verify with: assert-state. Safety: MUTATING; changes persistent executor-side observer or hook state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyNoOptional validated input for key.default
clearNoOptional validated input for clear.
limitNoOptional hard result/work budget used to keep output and runtime bounded.
actionYesoperation selector; use one of the schema's allowed values.
eventsNoOptional validated input for events.both
confirmNoExplicit safety acknowledgement; must be true before this state-changing operation runs.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true and idempotentHint=false, and the description adds real context beyond them: it changes persistent executor-side observer/hook state, requires active-client and explicit mutation approval, and needs confirm=true for start/stop. It stops short of describing exactly what observer state is left behind or how to recover it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The key behavioral warning is front-loaded in the first sentence, and the phase/cost/safety lines are terse. The verbatim Signature block partially duplicates the input schema and adds density, but overall it remains compact and scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the 'Produces: bounded-event-snapshot' line and 'Verify with: assert-state' pointer cover the return contract. Requirements, safety, and a failure fallback (inspect tool-schema) are all present, leaving little an agent must infer before calling it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds genuine meaning: confirm must be true before start/stop runs, and the buffer is bounded to 200 (matching the limit cap). The signature block itself mostly repeats schema fields, but the confirm and buffer details go beyond it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a concrete mechanism: connect to on_actor_added / on_actor_state_created, retain a bounded 200-event buffer, and poll it. This tells an agent what the tool does and how it differs from a live-scan approach. It does not explicitly name the nearest sibling (lua-state-event-monitor), so sibling differentiation is only implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives usage context (poll without repeated game-tree scans; start/stop require confirm=true) but never states when to prefer this over alternatives like lua-state-event-monitor or watch-value. The start/poll/stop lifecycle implies usage, but no explicit when-not guidance is present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

agent-contextBuild a live context brief for the AI agentA
Read-onlyIdempotent

READ-ONLY. Bootstrap an AI agent with the current MCP/session/game context in one call. Returns connected clients, this session's active selection resolution, game identity, executor identity/capabilities, and actionable next steps. Use this at the start of a task or after a reconnect instead of separately guessing which client, place, executor, or capability set is active. If multiple clients are connected, the brief tells you to select one; it never silently chooses between distinct accounts. Optional capability probing is safe but slower. The tool never mutates game state. Signature: { includeGameInfo: any?, includeExecutorInfo: any?, includeCapabilities: any?, includeHistory: any?, historyLimit: any?, includeMemory: any?, memoryLimit: any? }. Phase: observe; cost=medium; idempotency=read-only. Requires: none. Produces: agent-guidance. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
memoryLimitNoOptional hard result/work budget used to keep output and runtime bounded.
historyLimitNoOptional hard result/work budget used to keep output and runtime bounded.
includeMemoryNoInclude recent persistent agent-memory entries for continuity across tasks.
includeHistoryNoInclude a compact tail of this session's successful/failed tool history.
includeGameInfoNoInclude PlaceId/JobId/game metadata when this session resolves to a client.
includeCapabilitiesNoRun the larger side-effect-free capability matrix probe; slower but useful for planning advanced tools.
includeExecutorInfoNoInclude executor identity and headline capability flags when a client is active.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only/idempotent, but the description adds non-obvious behavior: it never silently chooses between distinct accounts, capability probing is safe but slower, and the tool never mutates game state. It also gives a failure path (inspect tool-schema), which annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads 'READ-ONLY' and the core purpose efficiently, then structures the rest. However, the trailing 'Phase: observe; cost=medium; idempotency=read-only' and the full signature line partially duplicate the annotations and schema, adding mild redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must describe return contents, and it does: connected clients, selection resolution, game identity, executor identity/capabilities, and next steps. Combined with usage triggers and failure guidance, an agent has everything needed to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all seven optional parameters and their defaults are already documented in the schema. The signature line in the description largely restates the schema (listing them as 'any?') rather than adding format or interaction guidance, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Bootstrap an AI agent with the current MCP/session/game context in one call') and enumerates exactly what the brief contains. It is clearly distinguishable from siblings like get-active-client, get-game-info, and get-executor-info because it is the aggregating one-call bootstrap rather than a single-field getter.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use it ('at the start of a task or after a reconnect') and what it replaces ('instead of separately guessing which client, place, executor, or capability set is active'). It even states the disambiguation behavior when multiple clients are connected, which is a real routing rule an agent needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

agent-memoryRemember and recall successful agent facts/workflowsA
Read-onlyIdempotent

PERSISTENT AGENT MEMORY. Store, recall, or forget compact facts and successful workflow notes across tasks. Memory is scoped optionally by game/place/executor and stored in the existing local playbook store, never in the Roblox game. Use remember after a verified discovery or successful workflow; use recall before planning a similar task. The learn-session operation summarizes the current session's successful tool sequence into a reusable episodic memory. Do not store secrets or tokens. Signature: { operation: "remember" | "recall" | "forget" | "learn-session", key: string?, text: string?, facts: {[string]: any}?, scope: string?, query: string?, limit: any?, fromSeq: number? }. Phase: observe; cost=high; idempotency=read-only. Requires: none. Produces: structured-result. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyNoStable memory key, such as 'shop-button-path'.
textNoHuman-readable fact or workflow note to remember.
factsNoStructured facts to store alongside the note.
limitNoMaximum memories to return.
queryNoRecall filter matched against key, text, facts, and scope.
scopeNoOptional game/place/executor scope label.
fromSeqNoFor learn-session, only use trace records from this sequence onward.
operationYesMemory operation.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnly/idempotent/destructive/openWorld, but the description adds real behavioral context: memory is scoped by game/place/executor, persisted only to the local playbook store (not the game), 'do not store secrets or tokens', and failure recovery via tool-schema. It does not describe the recall return shape or any capacity/retention behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and usage are front-loaded, which is good, but the long 'Signature: { operation: ... }' block is pure duplication of the 100%-covered schema, and the Phase/cost/idempotency/Safety tail repeats annotation data. Roughly a third of the text does not earn its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-param tool with nested facts and no output schema, the description covers operations, scoping, storage location, and safety caveats, and points to tool-schema on failure. The main gap is that recall's structured-result format is never described, which an agent needs to consume the output.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds a little meaning beyond the schema (scope = game/place/executor label; learn-session summarizes the session's successful tool sequence), but the inline signature largely restates what the schema already documents field-by-field.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Front-loads 'PERSISTENT AGENT MEMORY' and names the concrete verbs (store/recall/forget) plus the resource (compact facts and successful workflow notes). It also distinguishes itself from the playbook siblings by stating memory lives in 'the existing local playbook store, never in the Roblox game', so an agent can tell what this tool owns.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit triggers: 'use remember after a verified discovery or successful workflow; use recall before planning a similar task', and explains what learn-session is for. However, it never says when NOT to use it (e.g. vs playbook-save or session-* siblings) and leaves 'forget' without a usage condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

agent-runExecute a verified multi-tool AI workflowA
Destructive

ORCHESTRATES LIVE TOOLS. Execute an explicit discover→act→verify workflow with step IDs, result references, dry-run planning, mutation approval, retries, and automatic contract-based verification. Use agent-context and tool-plan first when the goal is vague, then pass concrete steps. Each later input can reference earlier data with $steps.stepId.data.field or $result.stepId.data.field. By default mutations are blocked and only read-only steps run; set allowMutations=true only when the user authorized state changes. Mutating steps are never retried unless retryMutations=true. If steps is omitted, this returns a schema-aware plan instead of guessing required arguments. This is the main closed-loop execution surface for AI agents. Signature: { goal: string, steps: {{ id: string, tool: string, input: any?, verifyWith: string?, verifyInput: {[string]: any}?, retries: any? }}?, dryRun: any?, allowMutations: any?, retryMutations: any?, verify: any?, stopOnError: any?, maxSteps: any? }. Phase: orchestrate; cost=medium; idempotency=contextual-write. Requires: explicit-mutation-approval. Produces: operation-receipt, step results, verification results, next actions. Verify with: assert-state. Safety: MUTATING; executes caller-selected behavior in the live client, may execute multiple tools; mutations require allowMutations=true. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
goalYesHuman-readable goal used for planning, summaries, and audit output.
stepsNoExplicit ordered steps. Omit to receive a plan only.
dryRunNoValidate and return the workflow without executing anything.
verifyNoRun each step's contract verifier when available.
maxStepsNoHard execution cap for this run.
stopOnErrorNoStop after the first failed or blocked step.
allowMutationsNoPermit tools marked mutatesState=true to execute.
retryMutationsNoPermit retries for mutating steps; disabled by default to avoid duplicate side effects.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructive/mutating, but the description adds critical behavior: mutations are blocked by default, mutating steps are never retried unless retryMutations=true, maxSteps caps the run, stopOnError defaults true, and it requires explicit-mutation-approval. This is well beyond what the structured fields provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The lead sentence is strongly front-loaded, but the body is dense and includes a Signature block that duplicates the 100%-covered input schema, plus a metadata tail (Phase, cost, idempotency, Produces, Verify with) that is more inventory than guidance. Several sentences are load-bearing, but the piece is longer than needed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a high-complexity orchestration tool with no output schema, the description covers execution model, gating, retries, references, and failure recovery ('inspect tool-schema for exact fields'), and names produced artifacts (operation-receipt, step results, verification results). Nested step verification semantics are only lightly addressed but the picture is largely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, and the description adds the `$steps.stepId.data.field` / `$result.stepId.data.field` reference syntax and reiterates default gating behavior (dryRun, allowMutations, retryMutations) in prose. The embedded Signature block largely restates the schema rather than adding new semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('ORCHESTRATES LIVE TOOLS', 'Execute an explicit discover→act→verify workflow') and frames itself as 'the main closed-loop execution surface for AI agents.' It distinguishes itself from nearby siblings like execute, batch-execute, tool-plan, and agent-context by positioning them as prerequisites or components.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use guidance ('Use `agent-context` and `tool-plan` first when the goal is vague, then pass concrete `steps`'), plus exclusions ('set allowMutations=true only when the user authorized state changes') and a fallback behavior when `steps` is omitted. The alternative routing is named, not implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

append-fileAppend content to a file in the executor workspace (UNC appendfile)A
Destructive

Append content to the end of a file in the executor's workspace folder, creating the file first if it does not exist. Existing contents are preserved (unlike write-file, which overwrites). NOTE: this is executor-side file I/O — the connector runs INSIDE the executor, so the path is relative to the executor's workspace directory on the host machine, NOT the Roblox game. Requires the UNC function appendfile(path, content). The call is type-guarded and pcall-wrapped: if appendfile is missing you get { error = 'appendfile is not available in this executor.' }, and any failure returns { error = }. Returns { path, ok = true } or { error }. Signature: { path: string, content: string, threadContext: number?, timeoutMs: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Capabilities: executor filesystem. Produces: operation-receipt. Verify with: file-exists. Safety: MUTATING; writes executor workspace filesystem. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesPath to the file within the executor workspace folder, e.g. 'logs/session.txt'. Created if missing.
contentYesThe text content to append to the end of the file.
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true and non-idempotent, and the description adds substantial context beyond them: type-guarded/pcall-wrapped behavior, exact error shapes ({ error = 'appendfile is not available...' }, { error = <message> }), the required UNC function, prerequisites (active-client, resolved-target, explicit-mutation-approval), and verification via file-exists. This is unusually rich behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and the key write-file contrast in the first two sentences, which is where it matters most. It is long and carries some metadata (phase/cost/idempotency labels) that borders on noise, but nearly every clause contributes decision-relevant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully compensates by documenting return values ({ path, ok = true } or { error }), error conditions, prerequisites, and a verification step. Nothing an agent needs to invoke this mutation correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters. The description restates the signature ({ path, content, threadContext?, timeoutMs? }) but adds no format or constraint detail beyond what the schema provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (append content to the end of a file in the executor's workspace folder) and explicitly differentiates from the sibling write-file by noting write-file overwrites while this preserves contents. An agent can distinguish it without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains exactly when to use it (append to end, create if missing, preserve existing content) and names the alternative (write-file) with the condition that selects it. Also clarifies the executor-side vs Roblox-game scope, removing the most likely misuse.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

assert-stateAssert live game stateA
Read-onlyIdempotent

VERIFY LIVE OUTCOMES in one bounded, read-only execution. Evaluate path existence, property and attribute values, effective GUI state, bounded descendant searches, custom-character distance, camera facing angle, and collection counts. Every result includes expected/actual evidence and errors. Missing, unreadable, or incomplete state always fails instead of being mistaken for success. Use this after actions and consume aggregate.passed, passRatio, and confidence as proof of the real outcome. Signature: { assertions: {{ id: string, kind: any, path: string } | { id: string, kind: any, path: string } | { id: string, kind: any, path: string, property: string, expected: string | number | boolean } | { id: string, kind: any, path: string, property: string, expected: string | number | boolean } | { id: string, kind: any, path: string, property: string, expected: string, caseSensitive: any? } | { id: string, kind: any, path: string, property: string, expected: number } | { id: string, kind: any, path: string, property: string, expected: number } | { id: string, kind: any, path: string, attribute: string, expected: string | number | boolean } | { id: string, kind: any, path: string, expected: any?, effective: any? } | { id: string, kind: any, path: string, expected: any? } | { id: string, kind: any, path: string, selector: { by: any, value: string } | { by: any, value: string, match: any?, caseSensitive: any? } | { by: any, value: string, match: any?, caseSensitive: any? }, expected: any? } | { id: string, kind: any, targetPath: string, operator: "at-most" | "at-least", distance: number, playerName: string?, characterPath: string?, rootPath: string? } | { id: string, kind: any, targetPath: string, maxAngleDegrees: any?, cameraPath: string? } | { id: string, kind: any, path: string, scope: any?, operator: "equals" | "not-equals" | "greater" | "less" | "at-least" | "at-most", count: number, selector: { by: any, value: string } | { by: any, value: string, match: any?, caseSensitive: any? } | { by: any, value: string, match: any?, caseSensitive: any? }? }}, scanLimit: any?, readBudget: any?, timeoutMs: any?, threadContext: number? }. Phase: verify; cost=medium; idempotency=read-only. Requires: active-client. Produces: grounded-evidence, per-assertion evidence, aggregate verification result. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
scanLimitNoMaximum descendants examined by any bounded search.
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
assertionsYesPredicates evaluated together against one live-game snapshot.
readBudgetNoMaximum guarded live reads shared by the whole assertion batch.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the read-only/idempotent annotations: it discloses the fail-closed semantic ("Missing, unreadable, or incomplete state always fails instead of being mistaken for success"), the bounded execution model with scanLimit/readBudget/timeout, the requires-active-client precondition, and cost=medium. These are exactly the behavioral traits annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening prose is tight and front-loaded, but the enormous inline Signature block restates the input schema verbatim and dominates the description, diluting the high-value guidance. Significant redundancy that a reader must wade through.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully names the aggregate fields to consume (passed, passRatio, confidence) and the per-assertion evidence/error semantics, plus cost and preconditions. It stops short of full coverage of batching limits and defaults, but is close to complete for a verify-phase tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents every parameter. The inline "Signature" block merely re-renders the schema shape rather than adding syntax, format, or default meaning, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ("VERIFY LIVE OUTCOMES ... live-game snapshot") and enumerates the assertion families (path existence, property/attribute, GUI state, descendant search, character distance, camera facing, collection counts). An agent can distinguish it from the narrower verify-path-exists sibling without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear timing context ("Use this after actions") and tells the agent to consume aggregate.passed/passRatio/confidence as the outcome proof. It does not, however, name alternatives such as verify-path-exists or get-game-state, so the when-not and alternative-routing guidance is missing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

batch-executeRun many Luau snippets in order in one callA
Destructive

Run several independent Luau snippets in sequence in a single round trip, collecting every result without paying a separate tool call per snippet. Each snippet is COMPILED with loadstring and run under its own pcall, in input order; a compile or runtime error in one snippet is recorded for that snippet only and never aborts the others. For each snippet you get { index, ok, value? , error? } — value is the FIRST return value, encoded so Instances/Vector3/etc. survive serialization; error carries the compile or runtime message when ok is false. Use this to read several instances, probe multiple remotes, or run a short script of steps in one shot. Requires loadstring (guarded). Returns { results = [{ index, ok, value?, error? }] } in input order. Signature: { snippets: {string}, client: string?, agent: string?, threadContext: number?, timeoutMs: number? }. Phase: orchestrate; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; executes caller-selected behavior in the live client. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNoOptional. A stable label for WHICH agent is calling when several share this MCP session (e.g. 'researcher'). Gives that agent its own fair scheduling lane, its own persistent VM on each game, and its own queue budget, so co-tenant agents don't starve or clobber each other.
clientNoOptional. Run on a specific connected client — its clientId OR username — for THIS call only, overriding your session's select-client binding without changing it. Lets multiple agents drive different games at the same time; omit to use your session's selected client.
snippetsYesThe Luau snippets to run, in order. Each is compiled and pcall-guarded independently; to capture a value, `return` it (only the first return value is collected).
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Far exceeds the annotations: it discloses per-snippet loadstring compilation, independent pcall isolation, that one error never aborts the batch, input-order guarantees, serialization-safe value encoding, the guarded loadstring requirement, and the mutation/approval/safety profile. This is exactly the behavioral context annotations alone (destructive/readOnly hints) cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core behavior and ordered logically, but the result shape is effectively stated twice ({ index, ok, value?, error? } and Returns { results = [...] }), which is mild redundancy rather than waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates fully by documenting the return shape, ordering, error isolation, and prerequisites. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real semantics: only the FIRST return value is collected, snippets execute in input order, and the value is encoded so Instances/Vector3 survive. That meaning goes beyond the raw schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Run several independent Luau snippets in sequence in a single round trip') and immediately frames the distinction from single-snippet siblings: one call instead of a separate tool call per snippet. An agent can tell it apart from execute/run-luau without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete usage contexts ('read several instances, probe multiple remotes, or run a short script of steps in one shot') and states the requirement of loadstring. It does not explicitly name alternatives like execute or run-luau and when NOT to batch, so it stops short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

block-functionNeutralize a function so it does nothing (MUTATES STATE via hookfunction)A
Destructive

WRITES LIVE GAME STATE — INSTALLS A PERSISTENT GLOBAL HOOK. Replaces a target function with a no-op that ignores its arguments, returns nothing, and NEVER calls the original. Use this to disable a behavior outright: silence an anticheat heartbeat/tick, stop a function that kicks the player, neutralize a telemetry reporter, or freeze a damage routine. Distinct from spoof-function-return (which forces a specific return value) — block-function simply makes the call a complete no-op (returns nil/nothing). WORKFLOW (stateful — survives across tool calls via getgenv().__mcp_blockedFns, keyed by functionPath): 1. action='start' with functionPath — resolves the target, captures the original, installs the no-op stub. Returns { started, key }. 2. action='stop' with the same functionPath — restores the original function. Returns { stopped }. CAVEATS: the hook is GLOBAL and PERSISTS until you stop it (or the client restarts). Because the original never runs, blocking a function the game depends on can break gameplay, and a live function hook CAN TRIP ANTICHEAT. Always stop when done. Requires hookfunction, newcclosure, and getgenv; restoration uses hookfunction(target, original) with a restorefunction fallback. Returns { error } if a capability is missing, the target cannot be resolved, or there is already an active block for fetch/stop. Signature: { action: "start" | "stop", functionPath: string?, threadContext: number? }. Phase: act; cost=medium; idempotency=idempotent-write. Requires: active-client, resolved-target, explicit-mutation-approval. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYes'start' installs the no-op stub on functionPath; 'stop' restores the original function. Use the SAME functionPath for both so they address the same registry entry.
functionPathNoLuau expression resolving to the function to neutralize, e.g. 'getsenv(game.Players.LocalPlayer.PlayerScripts.AntiCheat).tick' or 'getconnections(game.Players.LocalPlayer.Idled)[1].Function'. Evaluated as `return <functionPath>` and must resolve to a function. REQUIRED for 'start'. For 'stop' it is the registry key identifying which block to restore, so it must match the string used at start.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, but the description adds substantial context beyond that: the hook is GLOBAL and persists until stopped or client restart, state lives in getgenv().__mcp_blockedFns keyed by functionPath, a live hook CAN TRIP ANTICHEAT, and required capabilities (hookfunction, newcclosure, getgenv) plus failure returns are disclosed. This is rich behavioral detail that the annotations alone do not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the critical MUTATES STATE warning and the no-op contract before the workflow and caveats, which is well-structured. It is on the long side, with some meta-boilerplate (Phase/cost/idempotency line, 'Signature:' restatement, and a generic tool-schema pointer) that could be trimmed without loss.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a stateful, destructive mutation tool with no output schema, the description covers the required capabilities, the start/stop lifecycle, persistence and registry semantics, anticheat risk, and failure modes. An agent has everything needed to invoke it and to clean up correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by clarifying the stateful relationship between the two actions — that 'stop' must address the same registry key created at 'start' — which is the key semantic hook for correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Replaces a target function with a no-op'), the resource (a function), and the exact semantics (ignores arguments, returns nothing, never calls the original). It explicitly distinguishes itself from the sibling spoof-function-return, so an agent can route correctly without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use scenarios (silence an anticheat heartbeat, stop a kick function, neutralize telemetry, freeze damage), names the alternative and the condition that selects it (spoof-function-return forces a value; block-function is a no-op), and lays out the full start/stop workflow with the requirement to use the same functionPath.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

block-packetsBlock outgoing RakNet packets by criteriaA
Destructive

WRITES LIVE GAME STATE — installs a RakNet send hook that DROPS matching outgoing packets. A packet is blocked when its Size >= minSize (if set) and/or its payload contains containsHex (if set); with neither criterion every outgoing packet is blocked (dangerous). action='start' installs the hook, 'stop' removes it. Requires the raknet library. WARNING: blocking outgoing traffic can break the game, freeze replication, or disconnect the client — use narrow criteria and stop promptly. Signature: { action: "start" | "stop", minSize: number?, containsHex: string?, threadContext: number? }. Phase: act; cost=medium; idempotency=idempotent-write. Requires: active-client, explicit-mutation-approval. Capabilities: RakNet packet APIs. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; performs external network or socket I/O. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYes'start' installs the blocking hook; 'stop' removes it.
minSizeNoBlock packets whose Size is at least this many bytes. Omit to not filter by size.
containsHexNoBlock packets whose payload contains this hex byte sequence (e.g. '1a2b'). Omit to not filter by content.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, openWorldHint=true and readOnlyHint=false, but the description adds real value beyond them: what blocking actually does to the game ('break the game, freeze replication, or disconnect the client'), the required library and approval prerequisites, mutation classification, and a verification path (assert-state). The only divergence is the 'idempotency=idempotent-write' taxonomy tag against idempotentHint=false, which reads as metadata shorthand rather than a behavioral contradiction of the tool's mutating/destructive profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The destructive warning and core mechanism are front-loaded, then criteria, action semantics, prerequisite, signature, and a metadata tail (phase/cost/idempotency/verify/failure). Slightly dense and run-on, but every clause carries actionable information and nothing important is buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating, hook-installing tool with no output schema, the definition covers prerequisites, safety consequences, matching semantics, action lifecycle (start/stop), and post-call verification ('Produces: operation-receipt', 'Verify with: assert-state'), plus a failure fallback to tool-schema. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the per-parameter docs already exist and baseline is 3. The description adds the parameter-interaction semantics the schema omits: with neither minSize nor containsHex set, every outgoing packet is blocked, which is the single most important parameter-level fact for a correct call, plus an explicit signature line.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('installs a RakNet send hook that DROPS matching outgoing packets') and names the exact mechanism and scope, so it is immediately distinguishable from siblings like block-remote, block-function, send-packet, or packet-spy. The title and first sentence align without redundancy.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives strong operating context: the matching rule (Size >= minSize and/or payload contains containsHex), the dangerous no-criteria default, and the advice to 'use narrow criteria and stop promptly'. It also lists prerequisites (active-client, explicit-mutation-approval, raknet library) and phase/cost. It does not explicitly name an alternative tool (e.g., packet-spy for observation), so it falls just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

block-remoteSet a remote's block stateA
Destructive

WRITES LIVE GAME STATE. Reversibly sets the selected spy's block state for an observed remoteId or Luau remotePath, in the selected direction. blocked=false undoes it. Ketamine supports outgoing remotes and incoming RemoteFunction callbacks; incoming RemoteEvent blocking returns an explicit unsupported error. Both-direction requests are validated before changing any state. Signature: { engine: "cobalt" | "ketamine"?, remotePath: string?, remoteId: string?, direction: any?, blocked: any?, threadContext: number? }. Phase: act; cost=medium; idempotency=idempotent-write. Requires: active-client, resolved-target, explicit-mutation-approval. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
engineNoRemote-spy backend. Only one engine runs per client; starting another stops the current engine. Defaults to cobalt.cobalt
blockedNoOptional validated input for blocked.
remoteIdNoOptional text value for remote id.
directionNoCapture/control direction; incoming includes client callbacks.Outgoing
remotePathNoOptional dotted Roblox instance/value path resolved in the active client.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.8/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description asserts "idempotency=idempotent-write," which directly contradicts the annotation idempotentHint=false. An agent reasoning about safe retries would be actively misled by this conflict, which is the rubric's contradiction case despite the otherwise rich behavioral detail (mutation warning, engine support matrix, both-direction pre-validation).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the critical mutation warning and the core action, then metadata. It is dense but mostly earns its place; the only waste is mild redundancy between the opening "WRITES LIVE GAME STATE" and the later "Safety: MUTATING; writes live game/client state."

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter mutating tool with no output schema, the description covers the action, reversibility, prerequisites, success artifact (operation-receipt), verification step (assert-state), and failure handling. An agent has everything needed to call and validate it, aside from the idempotency inconsistency noted above.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description exceeds that by adding engine-dependent semantics not in the schema: Ketamine's support for outgoing remotes and incoming RemoteFunction callbacks, the unsupported incoming-RemoteEvent case, and the guarantee that Both-direction requests are validated before any state change.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first words name a specific verb and resource ("Reversibly sets the selected spy's block state") and scope it to an observed remoteId or remotePath in a chosen direction. This distinguishes it cleanly from siblings like monitor-remote, ignore-remote, and fire-remote, which are the other remote-spy manipulation tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states prerequisites (active-client, resolved-target, explicit-mutation-approval), the undo path (blocked=false), and a when-NOT condition (incoming RemoteEvent blocking returns an explicit unsupported error). It does not name an alternative sibling for filtering rather than blocking, so it falls short of an explicit alternative-routing statement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bridge-statusBridge & multi-session statusA
Read-onlyIdempotent

Report the bridge/session state as seen from this MCP session, WITHOUT touching any game. Shows: this session's { id, label, selection }; which client this session currently resolves to after applying its selection (the active client, or why none resolves); and the full roster of connected Roblox executor clients (clientId, username, userId, placeId, executor). Use this to debug multi-session routing — e.g. to confirm that your session and another session are pointed at different games, or to see whether any client is connected at all. Includes live per-client queue/concurrency pressure and rejected-overload counts when the transport exposes them. Returns { session, active, bridgeLoad, clients } and never runs Luau. Signature: {}. Phase: observe; cost=low; idempotency=read-only. Requires: none. Produces: diagnostic-report. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/no-destructive, but the description adds real value beyond them: it never runs Luau, never touches the game, degrades gracefully ('when the transport exposes them' for queue pressure), and points to tool-schema on failure. That is rich behavioral context rather than restating the hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and scope, and most sentences earn their place. The tail ('Phase: observe; cost=low; idempotency=read-only... Safety: read-only') largely duplicates the readOnlyHint/idempotentHint annotations, adding some redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by naming the return shape '{ session, active, bridgeLoad, clients }' and what each region contains. For a zero-param read diagnostic, an agent has everything needed to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4 and the empty schema is consistent with the description's 'Signature: {}'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: report bridge/session state as seen from this MCP session, without touching any game. It enumerates exactly what is reported (this session's id/label/selection, resolved active client, full client roster), which separates it cleanly from siblings like get-active-client, list-clients, and select-client.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use context: debug multi-session routing, confirm two sessions point at different games, or check whether any client is connected at all. It does not name the sibling tools that are alternatives (e.g. get-active-client for a narrower view), so it stops short of full when/when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

build-call-graphBuild call graph / function tree (IDA call graph)A
Read-onlyIdempotent

Build an IDA-style call graph (function tree) rooted at a target Luau function. Each Luau function carries its nested protos — the functions it can construct and call — so this recurses through getprotos breadth-first to map the callee tree. Returns a FLATTENED node list (each node knows its parentIndex and depth) so you can reconstruct the tree, plus per-node proto/upvalue counts and source/line from debug.info. Use it to understand how a closure fans out into helpers, spot deeply nested logic, and pick disassembly targets. Resolve the root via a Luau expression (e.g. "getrenv().game.PlayerScripts.Main.someFunc" or a global). Requires getprotos; caps depth and node count. Signature: { functionPath: string, maxDepth: any?, maxNodes: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Produces: operation-receipt. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
maxDepthNoHow many proto levels deep to recurse from the root (default 3, max 6). Root is depth 0.
maxNodesNoHard cap on total nodes (including the root) to keep the graph bounded (default 120).
functionPathYesLuau expression that resolves to the ROOT function (e.g. "getgenv().myFunc" or "require(game.ReplicatedStorage.Mod).start"). Evaluated as `return <expr>`; must yield a function.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint, idempotentHint, non-destructive), yet the description adds substantial behavior beyond them: the BFS recursion over getprotos, the flattened return shape with parentIndex/depth, per-node proto/upvalue counts and source/line, and the fact that depth and node count are capped. It also names prerequisites (active-client, resolved-target) and that it produces an operation-receipt.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core behavior and use cases in the opening sentences, then the operational metadata. There is minor redundancy where 'idempotency=read-only' and 'Safety: read-only' repeat the annotation hints, keeping it from a 5, but overall tight and well ordered.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description carries the return-value burden and does so: flattened node list with parentIndex/depth for reconstruction, plus per-node proto/upvalue counts and source/line. Behavior, prerequisites, safety, and return shape are all present for a medium-cost read-only analysis tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3; the description lifts above it by giving a concrete resolution example for functionPath ('getrenv().game.PlayerScripts.Main.someFunc' or a global) and restating that depth/node-count are capped. It does not add anything for maxDepth defaults or threadContext beyond what the schema already documents.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Build an IDA-style call graph (function tree) rooted at a target Luau function') and immediately scopes it with the mechanism (recursing through getprotos breadth-first). It is clearly distinguishable from siblings like get-function-protos or get-closure-protos, which are point lookups rather than a rooted tree walk.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly gives use cases: 'understand how a closure fans out into helpers, spot deeply nested logic, and pick disassembly targets.' That is clear when-to-use context. It stops short of naming an alternative tool to prefer in other situations (e.g. get-function-protos for a single level), so no explicit exclusion routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cache-invalidatecache.invalidate — drop an instance from the executor's instance cacheA
Destructive

Invalidate the executor's cached reference to a single Instance via cache.invalidate(inst). After invalidation the next time the game indexes that instance it receives a FRESH reference rather than the cached one — the classic technique for de-syncing a server-trusted object so your subsequent edits to the cached copy go unnoticed, or for forcing the executor to re-wrap a part you have been tampering with. The target is resolved from a Luau path/expression via loadstring('return ' .. expr). Requires the cache library (type(cache)=='table') with cache.invalidate — both are type-guarded and the call is pcall-wrapped, returning { error } when missing or on failure. Mutates live executor state. Returns { ok } or { error }. Signature: { instancePath: string, threadContext: number?, timeoutMs: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Produces: structured-result. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
instancePathYesLuau path/expression resolving to the Instance to invalidate, e.g. 'game.Workspace.Boss' or 'game.Players.LocalPlayer.Character.HumanoidRootPart'. Evaluated as `return <expression>`.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations by disclosing the loadstring-based target resolution, the type-guarded cache-library requirement, the pcall wrapping, and the { ok } / { error } return shape on failure. It also restates the mutating/live-state nature in prose, consistent with destructiveHint=true and readOnlyHint=false.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose is front-loaded and the technique/rationale is genuinely useful, but the trailing metadata block (Phase/cost/idempotency/Requires/Produces/Verify with/Safety/On failure) is dense boilerplate that partly restates the annotations and schema. Slightly over-stuffed but well ordered.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating, non-idempotent tool with no output schema, the description covers prerequisites (cache library), failure semantics, mutation warning, and the return contract { ok } / { error }, so an agent has everything needed to invoke and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (instancePath, threadContext, timeoutMs) are already documented in the schema, including the 'return <expression>' evaluation. The description's signature recap and loadstring note add no meaning beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource: invalidating the executor's cached reference to a single Instance via cache.invalidate(inst), and clarifies the effect (next index returns a FRESH reference). This is functionally distinct from the other cache-* siblings (cache-is-cached, cache-replace) and an agent can tell what it does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear use cases ('de-syncing a server-trusted object so your subsequent edits go unnoticed', 'forcing the executor to re-wrap a part you have been tampering with') plus prerequisites and verification guidance. It does not, however, explicitly name or exclude alternatives such as cache-replace or cache-is-cached, so the when-vs-which guidance is incomplete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cache-is-cachedcache.iscached — check whether an instance is in the executor cacheA
Read-onlyIdempotent

Report whether the executor currently holds a cached reference for a single Instance via cache.iscached(inst). Pairs with cache-invalidate / cache-replace: use it to confirm an instance is cached before invalidating it, or to verify that an invalidate actually dropped the cached reference. The target is resolved from a Luau path/expression via loadstring('return ' .. expr). Read-only. Requires the cache library (type(cache)=='table') with cache.iscached — both are type-guarded and the call is pcall-wrapped, returning { error } when missing or on failure. Returns { cached } or { error }. Signature: { instancePath: string, threadContext: number?, timeoutMs: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
instancePathYesLuau path/expression resolving to the Instance to check, e.g. 'game.Workspace.Boss'. Evaluated as `return <expression>`.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnly/idempotent/non-destructive, but the description goes further by disclosing that the cache library and cache.iscached are type-guarded, the call is pcall-wrapped, and the tool returns { error } rather than throwing when the library is missing or the call fails. That failure semantics is genuinely new information beyond the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose, then usage, then mechanics, then return/failure shape — a logical order. It is somewhat padded by repeated metadata ('Read-only' appears twice, plus 'Phase: observe; cost=medium; idempotency=read-only'), which costs it a point but does not bury the core information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description supplies return shapes ({ cached } or { error }), prerequisites (active-client, resolved-target, cache library present), failure handling, and even the phase/cost profile. An agent has everything needed to call and interpret this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description recaps the signature and the loadstring('return ' .. expr) resolution mechanism, but that largely restates what the instancePath schema description already says ('Evaluated as `return <expression>`'), and it adds no meaning for threadContext or timeoutMs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource ('Report whether the executor currently holds a cached reference for a single Instance') and names the exact API call it wraps (cache.iscached(inst)), which an agent can distinguish from the many other cache and inspection siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use it: 'confirm an instance is cached before invalidating it, or to verify that an invalidate actually dropped the cached reference', and names the paired siblings cache-invalidate / cache-replace. Both the when and the alternatives are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cache-replacecache.replace — swap one instance for another in the executor cacheA
Destructive

Replace the executor's cached reference for one Instance with another via cache.replace(a, b). After the swap, every script that indexes instance A through the cache transparently receives instance B instead — a powerful redirection primitive for impersonating one object with another (e.g. pointing a checkpoint, hitbox, or remote wrapper at a substitute you control). Both targets are resolved from Luau path/expressions via loadstring('return ' .. expr). Requires the cache library (type(cache)=='table') with cache.replace — both are type-guarded and the call is pcall-wrapped, returning { error } when missing or on failure. Mutates live executor state. Returns { ok } or { error }. Signature: { instancePath: string, replacementPath: string, threadContext: number?, timeoutMs: number? }. Phase: act; cost=medium; idempotency=idempotent-write. Requires: active-client, resolved-target, explicit-mutation-approval. Produces: structured-result. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
instancePathYesLuau path/expression resolving to the Instance whose cached reference is replaced (A). Evaluated as `return <expression>`.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
replacementPathYesLuau path/expression resolving to the Instance to substitute in (B). Evaluated as `return <expression>`.

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds strong behavioral detail beyond the annotations: the call is type-guarded and pcall-wrapped, returns { error } when the library is missing or on failure, and mutates live executor state. However, it claims 'idempotency=idempotent-write' while the annotations declare idempotentHint=false — a direct conflict on retry safety for a destructive tool, which misleads an agent into believing repeat calls are harmless. This contradiction caps the score despite the otherwise valuable disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core action is front-loaded in the first sentence, which is good. But the text is long and densely packed with generated metadata boilerplate (Phase/cost/Requires/Produces/Verify/Safety/On failure), and it redundantly restates the return shape and the 'mutates state' claim, so several sentences do not fully earn their place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex, path-resolving, state-mutating tool with no output schema, the description covers prerequisites, failure behavior, the { ok }/{ error } return, and a verification step (assert-state). That is nearly everything an agent needs to invoke it correctly, missing only the routing against sibling cache tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (including optional timeoutMs/threadContext) are already documented in the schema. The description adds only the resolution mechanism (loadstring('return ' .. expr)) and a compact signature line, which the schema largely supplies already, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource+mechanism: it swaps the executor's cached reference for one Instance with another via cache.replace(a, b). An agent can distinguish it from cache-invalidate and cache-is-cached without opening either schema, and the redirection semantics are spelled out concretely (checkpoint/hitbox/remote-wrapper impersonation).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context and preconditions: requires the cache library (type(cache)=='table') with cache.replace, an active client, resolved targets, and explicit mutation approval. It explains the redirection use case. It stops short of naming the sibling alternative (cache-invalidate) or stating when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

call-closureCall any function directly with arguments and capture its resultA
Destructive

ACTS ON LIVE GAME STATE — EXECUTES THE FUNCTION. Resolve a Luau expression to a function and CALL it with an ordered list of typed arguments, returning every value it produces. This lets you invoke internal, hidden, or anonymous functions on demand: a remote's OnClientEvent handler (getconnections(remote.OnClientEvent)[1].Function), a metamethod (getrawmetatable(game).__namecall), a member pulled from a script env (getsenv(script).someFunc), a constant/upvalue you extracted, or any closure found in the GC. Each argument is a typed value; use kind='raw' for non-primitive arguments (Vector3, Color3, Enum, CFrame, tables, Instance references, ...). The call is fully pcall-guarded, so a function that errors reports its message instead of aborting. WARNING: this genuinely runs the target function with the arguments you supply — it MAY cause side effects, mutate game state, fire remotes to the server, or trip anti-cheat. Only call functions you understand. Returns { Target, ok, returns, returnCount, truncated, argCount } on success, or { Target, ok=false, error, argCount } when the function raised. Signature: { functionPath: string, args: {{ kind: "string" | "number" | "boolean" | "nil" | "raw", value: string | number | boolean? }}?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Capabilities: debug closure primitives. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; executes caller-selected behavior in the live client. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
argsNoOrdered list of positional arguments to pass to the function. Omit or pass [] to call it with no arguments. Each entry is a typed value; use kind='raw' for anything that isn't a plain string/number/boolean/nil.
functionPathYesLuau expression resolving to the function to call, e.g. 'getsenv(game.Players.LocalPlayer.PlayerScripts.Main).update', 'getrawmetatable(game).__namecall', 'getconnections(game.Workspace.Part.Touched)[1].Function', or 'getgenv().myGlobalFn'. Evaluated as `return <functionPath>`; the result must be a function.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructive/non-idempotent, but the description adds substantial context beyond them: pcall-guarding and error reporting, the exact return shape on success and failure, side-effect/anti-cheat warnings, and the 'kind=raw' evaluation semantics at call time. This meaningfully de-risks a dangerous mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Long, but front-loaded with the critical 'EXECUTES THE FUNCTION' warning and each detail (guarding, return shape, raw semantics, safety) earns its place for a mutating tool. Some enumerations of example expressions are redundant with the schema, preventing a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description fully specifies the return object on both success and failure, plus failure guidance ('inspect tool-schema'). For a live-state mutating tool with 100% param coverage, nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description goes slightly further by surfacing the signature and stressing the kind='raw' path for non-primitive arguments — a subtlety that governs correct invocation. It largely mirrors the schema, so it earns only modest credit above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (call/invoke) and resource (a resolved Luau function/closure) with explicit scope. It names the concrete patterns it handles (remote handlers, metamethods, env members, GC closures) and is clearly distinct from siblings like eval-expression or run-luau, which evaluate code rather than invoke a resolved function object.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives rich context on when to use it — 'invoke internal, hidden, or anonymous functions on demand' with several concrete examples of what those are. It does not explicitly contrast against close siblings such as invoke-closure or execute, so the agent must infer the boundary, but the usage context is otherwise clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

camera-controlRead or move the current cameraA
Destructive

WRITES LIVE GAME STATE. Read the active camera or set its CFrame, position/look-at target, FieldOfView, and optional CameraType. Use this to reproduce camera movement and aim states in the local client. The tool only changes the local CurrentCamera; it does not move the character or replicate a camera state to the server. For a one-shot read, use action=get. For movement, action=setCFrame requires position and either lookAt or rotation in degrees. Existing camera properties are returned so the change is auditable. Signature: { action: "get" | "setCFrame" | "setFov", position: { x: number, y: number, z: number }?, lookAt: { x: number, y: number, z: number }?, rotation: { pitch: number, yaw: number, roll: number }?, fov: number?, cameraType: "Fixed" | "Attach" | "Watch" | "Track" | "Follow" | "Custom" | "Scriptable" | "Orbital"?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval. Produces: structured-result. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
fovNoFieldOfView for setFov or alongside setCFrame.
actionYesCamera operation.
lookAtNoWorld target for setCFrame; creates CFrame.lookAt(position, lookAt).
positionNoWorld position for setCFrame.
rotationNoEuler rotation in degrees for setCFrame when lookAt is not supplied.
cameraTypeNoOptional Enum.CameraType to apply after changing the CFrame.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare a mutating, non-idempotent, destructive write profile, and the description corroborates and enriches this: it emphasizes WRITES LIVE GAME STATE, clarifies the write is confined to the local camera, notes idempotency=contextual-write, and states that existing properties are returned so the change is auditable. It also points at assert-state for verification. This is well beyond what the structured fields alone convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is front-loaded with the key mutating warning, which is good, but it is padded with a full Signature block that duplicates the input schema and a boilerplate tail (Phase/cost/Produces/Verify/Safety/On failure) that repeats the opening 'writes live game state'. Several sentences do not add information beyond the schema and annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, and the description compensates by noting that existing camera properties are returned for auditability, and by specifying prerequisites, idempotency, and a verification path. It is complete enough for a 7-param, nested-object mutation, though return-shape detail is only summarized.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds cross-parameter semantics the schema does not: setCFrame requires position plus either lookAt or rotation in degrees. That constraint reasoning is genuinely useful beyond the per-field descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource (the active/CurrentCamera) and concrete operations (read, set CFrame/position/look-at/FOV/CameraType). It even scopes it against the environment by stating it only changes the local CurrentCamera and does not move the character. An agent can tell exactly what this tool does and how it differs from generic instance/character tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit routing is given: action=get for a one-shot read, action=setCFrame for movement, with preconditions (requires active-client and explicit-mutation-approval). It also states what the tool is NOT for (no character movement, no server replication). It stops short of naming a specific alternative sibling tool for those excluded cases, so it falls just under fully explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

can-signal-replicateCheck whether a signal can be replicated to the serverA
Read-onlyIdempotent

Query whether an RBXScriptSignal is one the Roblox engine permits to be replicated to the server, using the executor's cansignalreplicate. When CanReplicate is true, the replicate-signal tool can fire this signal so the server receives it as if the game client raised it natively — the basis for driving server-side logic that is normally gated behind engine-internal signals. When false, replicate-signal will be rejected and you must use a different vector. Use this to pre-flight a signal before attempting replication. Pass the instance that owns the signal plus the signal member name (or leave signalName empty if instancePath already resolves to the signal). Requires the cansignalreplicate executor function; degrades with a clear { error } if unavailable. Returns { Signal, Instance?, CanReplicate, Note }. Signature: { instancePath: string, signalName: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: getconnections. Produces: structured-result. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
signalNameNoName of the RBXScriptSignal member on the instance (e.g. 'OnClientEvent', 'Touched', 'Changed'). Leave empty/omit if instancePath already evaluates to the signal.
instancePathYesLua expression resolving to the Instance that owns the signal (e.g. 'game.ReplicatedStorage.MyRemote', 'game.Workspace.Door'). Evaluated as `return <instancePath>`. Leave signalName empty if this expression already resolves to the RBXScriptSignal itself.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only/idempotent, but the description adds substantial context beyond them: the executor-function dependency, graceful degradation with a structured { error }, and the required preconditions (active-client, resolved-target). This is exactly the kind of runtime behavior the schema cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded and logically ordered: purpose, behavior on true/false, then preconditions and failure mode. The trailing metadata block (Phase/cost/idempotency/capabilities/safety) partly duplicates the annotations, which costs it the top score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The return shape ({ Signal, Instance?, CanReplicate, Note }) is described even without an output schema, the failure mode is covered, and executor dependency is called out. Nothing an agent needs to invoke and interpret this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both parameters are already fully documented in the schema. The description restates the ownership/mutual-exclusion semantics ('pass the instance plus the signal member name, or leave signalName empty') but adds no syntax or format detail beyond what the schema provides; baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Query whether an RBXScriptSignal ... permits to be replicated') and explicitly names the sibling it feeds into (replicate-signal). An agent can distinguish this pre-flight check from the actual replication tool without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('pre-flight a signal before attempting replication') and both branches of the outcome: when true, replicate-signal fires it; when false, replicate-signal will be rejected and a different vector is required. Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

capture-log-outputCapture all print/warn/error log output over a window (MUTATES STATE via LogService)A
Read-onlyIdempotent

WRITES LIVE GAME STATE — INSTALLS A PERSISTENT LOG CONNECTION. Connects LogService.MessageOut and records every line the game emits (print, warn, error, and engine messages) into a ring buffer, so you can later read everything that was logged during a window of play. This is the best way to watch a game's own console output over time — see what a script prints when you trigger an action, catch errors/stack traces as they happen, or correlate warnings with behavior. It uses a signal CONNECTION (LogService.MessageOut), NOT a function hook, so it is low-risk compared to the hook-based instrument tools. WORKFLOW (stateful — survives across tool calls via getgenv().__mcp_logCapture): 1. action='start' — connects MessageOut to a handler that pushes { message (first 500 chars), messageType, t } into a 1000-entry ring buffer. Returns { started }. 2. action='fetch' — returns the captured messages (newest-bounded by limit) WITHOUT clearing them; poll while you play. Returns { count, returned, entries }. 3. action='stop' — Disconnects the connection and clears state. Returns { stopped, captured }. CAVEATS: the connection PERSISTS until you stop it (or the client restarts) and fires on every logged line, so a noisy game fills the 1000-entry buffer (oldest dropped). Each message is truncated to 500 chars. Always stop when done. Requires game:GetService('LogService') and getgenv. Returns { error } if a capability is missing, or { notRunning } for fetch/stop when nothing is active. Signature: { action: "start" | "fetch" | "stop", limit: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-result. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNofetch only: maximum number of messages to return, newest first (default 200, capped at the 1000-entry buffer).
actionYes'start' connects LogService.MessageOut and begins capturing log lines (persistent until stopped); 'fetch' returns captured messages WITHOUT clearing them (poll while you play); 'stop' disconnects and clears all captured state.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnlyHint, destructiveHint=false, idempotentHint), but the description adds far more: the connection persists across calls until stopped, it survives across tool calls via getgenv().__mcp_logCapture, it fires on every logged line, the 1000-entry ring buffer drops oldest entries, messages are truncated to 500 chars, and failure shapes are { error } and { notRunning }. This is exactly the extra context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded and well-sectioned (workflow, caveats, error returns), and every sentence carries operational information. It is dense and lengthy, however, with some redundancy between the workflow block and the trailing Signature/Phase/Cost summary, which keeps it short of a 5 on conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a stateful, multi-action tool with no output schema, the description supplies return shapes for all three actions ({ started }, { count, returned, entries }, { stopped, captured }), the error variants, prerequisites (game:GetService('LogService'), getgenv), and the persistence caveat. Nothing an agent needs to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already documents action, limit, and threadContext, so baseline is 3. The description adds marginal value beyond the schema — it clarifies limit is newest-bounded against the 1000-entry buffer and that fetch does not clear captured messages — which nudges it above baseline but does not introduce syntax or defaults the schema lacks.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — connect LogService.MessageOut and record every emitted line into a ring buffer — and immediately scopes it against alternatives ('best way to watch a game's own console output over time', 'low-risk compared to the hook-based instrument tools'). An agent can distinguish it from siblings like get-console-output without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit three-step lifecycle (start/fetch/stop) with what each returns, names the concrete scenarios that call for it (see what a script prints when you trigger an action, catch stack traces as they happen), states when to stop, and contrasts it with hook-based tools. Conditions for use and non-use are both present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check-callerCheck whether the current call originates from the executorA
Read-onlyIdempotent

Call Volt checkcaller in the active executor thread. This normally reports true for a direct MCP call and is most useful as a capability/behavior probe before installing a hook. Signature: { threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-observation, operation-receipt. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, non-destructive behavior, so the description adds useful context beyond that: active-client requirement, medium cost, produced artifacts, and failure guidance pointing to tool-schema. It does repeat the read-only annotation, which slightly limits its added value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded and the remaining operational details are compactly presented. Some fields (signature, idempotency=read-only, Safety: read-only) are redundant with structured data, but the text is not bloated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter probe with rich annotations and no output schema, the description covers requirements, outputs, cost, and failure handling. However, the produced observation ('structured-observation, operation-receipt') is vague and the return value is only implied by 'normally reports true.'

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already documents threadContext as an optional Roblox thread identity with default behavior. The description only restates the signature without adding any semantic meaning beyond what the schema provides, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: calling Volt checkcaller to determine whether the current call originates from the executor. It clearly distinguishes the probe's behavior ('normally reports true for a direct MCP call'), though it does not name a sibling alternative to contrast with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear usage context: 'most useful as a capability/behavior probe before installing a hook.' This tells the agent when the tool is appropriate, but provides no explicit exclusions or alternative tools for different conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clear-queue-on-teleportClear teleport-queued code (sUNC clearqueueonteleport)A
Destructive

WRITES EXECUTOR STATE — clears any code previously queued with queueonteleport, via clearqueueonteleport(). After this, nothing will auto-run on the next teleport. Requires clearqueueonteleport. The call is type-guarded and pcall-wrapped: if clearqueueonteleport is missing you get { error = 'clearqueueonteleport is not available in this executor.' }. Returns { cleared } or { error }. Signature: { threadContext: number?, timeoutMs: number? }. Phase: act; cost=medium; idempotency=idempotent-write. Requires: active-client, explicit-mutation-approval. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses a great deal beyond the annotations: the error envelope when the function is missing, the pcall/type-guard wrapping, the return shape { cleared } or { error }, and the MUTATING safety profile — all of which align with destructiveHint=true. However, it asserts 'idempotency=idempotent-write' while the annotations declare idempotentHint=false, a direct conflict that an agent could rely on when deciding whether to retry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The first sentence is well front-loaded with the key action and its effect, but the definition then piles on a metadata block (phase/cost/idempotency/requires/produces) plus a generic 'On failure: inspect tool-schema…' instruction that is boilerplate rather than tool-specific. It is dense and mostly earns its place, but not tight enough for a 4-5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by specifying the return contract ({ cleared } or { error }), the failure envelope, the preconditions, and a verification path — everything an agent needs to invoke and check this mutation correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters are already documented as optional deadline/thread-identity fields, so the schema does the heavy lifting. The description's restatement ('Signature: { threadContext: number?, timeoutMs: number? }') adds no semantics beyond the schema, matching the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('clears any code previously queued with queueonteleport'), names the underlying executor function, and contrasts directly with the sibling queue-on-teleport ('After this, nothing will auto-run on the next teleport'). An agent can distinguish it from siblings without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides the precondition ('Requires: active-client, explicit-mutation-approval'), the phase ('act'), and a verification step ('Verify with: assert-state'), which tells the agent when this call is appropriate and what must happen around it. It does not, however, name an explicit alternative tool or an when-not-to-use condition, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clear-remote-spy-logsClear remote-spy historyB
Destructive

WRITES LIVE GAME STATE. Clears the selected engine's MCP buffer and GUI history while retaining block/ignore settings and continuing capture. Call IDs remain monotonic. Signature: { engine: "cobalt" | "ketamine"?, threadContext: number? }. Phase: act; cost=medium; idempotency=idempotent-write. Requires: active-client, explicit-mutation-approval. Produces: bounded-event-snapshot, operation-receipt. Verify with: assert-state. Safety: MUTATING; changes persistent executor-side observer or hook state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
engineNoRemote-spy backend. Only one engine runs per client; starting another stops the current engine. Defaults to cobalt.cobalt
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

B3.4/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description claims 'idempotency=idempotent-write', but the annotations declare idempotentHint=false. These directly conflict on a stated behavioral attribute, so this is an Annotation Contradiction. Otherwise the description does add value (state is MUTATING, persistent executor-side observer state changes, block/ignore retained, monotonic call IDs), but the explicit contradiction triggers the score-1 rule.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is well front-loaded ('WRITES LIVE GAME STATE. Clears...'), and the labeled metadata block is easy to scan. However, boilerplate restatement of the signature and the generic tail ('On failure: inspect tool-schema for exact fields, defaults...') do not earn their place for this specific tool, diluting conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the description covers the essentials: safety profile, prerequisites, produced artifacts (bounded-event-snapshot, operation-receipt), and a verification route (assert-state). The only gap is the unreliable idempotency claim, which is a correctness issue rather than a missing-information one.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the enum values, default, and thread-identity semantics are already documented in the schema. The description merely restates the signature ('{ engine: "cobalt" | "ketamine"?, threadContext: number? }') without adding format or constraint detail beyond what the schema provides, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Clears the selected engine's MCP buffer and GUI history.' It also scopes the operation by noting what is retained (block/ignore settings) and that capture continues, which distinguishes it from siblings like block-remote, ignore-remote, and get-remote-spy-logs. An agent can identify the action without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides real invocation context: 'Phase: act', 'Requires: active-client, explicit-mutation-approval', and a follow-up action 'Verify with: assert-state.' This tells the agent when the tool is eligible and what prerequisites must hold. It does not, however, name any alternative tool or an explicit when-not-to-use condition, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clear-selectionClear this session's client selectionA
Read-onlyIdempotent

Drop THIS session's client binding so it no longer pins a specific client or account. Afterwards tools auto-resolve: with exactly one client connected they use it; with several different accounts you must select-client again. Use this to reset routing before re-selecting, or to hand control back to auto-resolution. Signature: {}. Phase: observe; cost=low; idempotency=read-only. Requires: none. Produces: operation-receipt. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only, idempotent, and non-destructive, so the bar is lower. The description earns credit for disclosing post-condition routing behavior and the produced operation-receipt, though the repeated "idempotency=read-only / Safety: read-only" lines duplicate annotation data rather than adding to it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The first two sentences are front-loaded and high value. The trailing metadata block (Phase/cost/idempotency/Requires/Produces/Safety) is somewhat boilerplate and partially restates annotations, but it stays compact and the failure pointer is useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, annotation-covered operation with no output schema, the description supplies everything an agent needs: scope, post-state behavior, producing an operation-receipt, and what to do on failure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the rubric this is the baseline 4. The "Signature: {}" note correctly signals an empty argument set, matching the empty input schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Drop THIS session's client binding") with explicit scope, and clarifies it loosens a pin rather than disconnecting a client. An agent can distinguish it from select-client without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use it ("reset routing before re-selecting, or to hand control back to auto-resolution") and describes the post-state routing rules (auto-resolve with one client; must select-client again with several accounts). It also names the relevant sibling, select-client.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clear-semantic-indexClear the semantic script indexA
Read-onlyIdempotent

Drop the cached semantic index for THIS session's active client. This clears a SERVER-SIDE embedding cache only — it does not touch the game or any live script. Resolves the active client from this session's selection; if one is resolved its index is cleared, otherwise nothing happens. Use it to force the next semantic-search-scripts call to re-harvest and re-embed from scratch (e.g. after the game's scripts changed). Signature: {}. Phase: observe; cost=low; idempotency=read-only. Requires: none. Produces: operation-receipt. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive), but the description adds real value: it clarifies this only clears a SERVER-SIDE embedding cache and does not touch the game or live scripts, and explains the active-client resolution with a no-op fallback. This goes meaningfully beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and scope, then supporting detail. The trailing boilerplate ('Phase: observe; cost=low; idempotency=read-only... Safety: read-only... On failure: inspect tool-schema') partly repeats annotation data and is generic, but the overall structure is tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-param, no-output-schema tool whose annotations already carry the safety profile, the description supplies the missing pieces: cache-only scope, session-based client resolution, no-op behavior, and the intended follow-up effect. Sufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters, so the baseline is 4. The description confirms the empty signature ('Signature: {}'), aligning with the schema and leaving no ambiguity about invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Drop the cached semantic index') and scopes it precisely to 'THIS session's active client'. This clearly distinguishes it from siblings like semantic-search-scripts, get-semantic-index-stats, and the generic cache-invalidate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the trigger ('force the next semantic-search-scripts call to re-harvest and re-embed from scratch, e.g. after the game's scripts changed') and the related sibling it affects. Lacks explicit when-not guidance, but the context for use is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

click-buttonClick a GuiButtonA
Destructive

WRITES LIVE GAME STATE. Click a Roblox TextButton or ImageButton by firing its GUI signals via firesignal, exactly as if the local player clicked it — so any handler connected to the button runs. Use when direct UI activation is needed inside the active client. Resolves the path to a GuiButton, then either fires the single named action signal or, when action is omitted, fires every standard click signal (Activated, MouseButton1Down, MouseButton2Down, MouseButton1Click, MouseButton2Click). Requires the executor's firesignal; degrades with a clear { error } if unavailable. Returns { Path, Fired, ok } or { error }. Signature: { path: string, action: string?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Capabilities: firesignal. Produces: structured-result. Verify with: get-gui-text. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesThe instance path to the Button
actionNoThe specific signal to fire (e.g., 'Activated', 'MouseButton1Click'). If omitted, fires all standard click signals.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, but the description adds what actually happens: it writes live game state, runs any connected handler, requires the executor's firesignal, and degrades with a clear { error } when unavailable. It also names the failure path and a verification tool (get-gui-text), which is well beyond the annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the critical warning 'WRITES LIVE GAME STATE' and flows logically through mechanism, usage, behavior, and return shape. The trailing Phase/cost/capabilities metadata (e.g. idempotency=contextual-write, Capabilities: firesignal) is somewhat redundant and could be trimmed, keeping it just under the top mark.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by describing the return shape ({ Path, Fired, ok } or { error }), the prerequisites, the degradation path, and a verification step. An agent has everything needed to invoke it correctly and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning: it spells out the signature and lists the exact signals fired when action is omitted (Activated, MouseButton1Down, etc.), which the schema only summarizes. It does not add much for threadContext beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (click) and resource (Roblox TextButton or ImageButton) and explains the exact mechanism: firing GUI signals via firesignal so connected handlers run. This distinguishes it from siblings like fire-signal and virtual-input, which operate at a lower level.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear condition: 'Use when direct UI activation is needed inside the active client,' plus prerequisites (active-client, resolved-target, explicit-mutation-approval). It does not explicitly name the alternative tool (e.g. virtual-input or fire-signal) an agent should pick instead, so it stops short of full when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clone-functionClone a function and retain a reusable closure referenceA
Read-onlyIdempotent

Clone a function with clonefunction/clonefunc. The function result is retained in getgenv().__mcp_closure_refs and returned as a reusable Reference expression because functions cannot cross JSON. Signature: { functionPath: string, threadContext: number?, key: any? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Produces: created-handle. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyNoOptional stable key for getgenv().__mcp_closure_refs. Generated when omitted.
functionPathYesLuau expression resolving to the target function.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), so the bar is lower, and the description earns credit by disclosing the non-obvious retention mechanism: results land in getgenv().__mcp_closure_refs because functions can't cross JSON. It also names a produced handle type, though it doesn't explain lifecycle/release of the handle.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose in the first sentence, then packages operational metadata into compact semicolon-delimited clauses. Information-dense with little waste, though the trailing 'On failure: inspect tool-schema...' clause is generic boilerplate.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description explains the return value (a reusable Reference expression stored in a named global), covers prerequisites, and flags a failure fallback. Complete enough for a 3-param tool, with only minor gaps around handle lifecycle.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and each parameter is well documented in the schema, so baseline 3 applies. The description's signature line (functionPath, threadContext, key) merely restates what the schema already provides and adds no new semantic detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (clone) and resource (function) and explains the mechanism and outcome: the result is retained in getgenv().__mcp_closure_refs and returned as a reusable Reference. This is clearly distinct from sibling closure tools (inspect-closure, get-closure-protos, invoke-closure), which don't produce reusable handles.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides operating context via 'Phase: observe', 'Requires: active-client, resolved-target', and 'Produce: created-handle', which imply when the tool is usable. However, it never states when to prefer ou clone-function over alternatives like inspect-closure or get-closure-protos, nor any exclusions, so it only implies usage rather than routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clone-instanceClone a live InstanceA
Destructive

WRITES LIVE GAME STATE. Resolve a Luau expression to a source Instance, deep-copy it via source:Clone() (which duplicates the instance and all of its descendants), and optionally parent the clone into the game tree. Useful while debugging for duplicating a template Part/Model/GUI, spawning extra copies of an item, or capturing a snapshot of a subtree before mutating the original. The clone starts parentless; it is parented LAST and only if parentPath is provided. NOTE: :Clone() only succeeds when the source's Archivable property is true — a clone of a non-Archivable instance returns nil and yields a clean error. WARNING: a parented clone immediately affects the running game and may replicate. Returns { Source, Clone, Parented, ok } or { error }. Signature: { instancePath: string, parentPath: string?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Produces: created-handle. Verify with: get-instance-properties. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
parentPathNoOptional Luau expression resolving to the Instance that should become the clone's Parent, e.g. 'game.Workspace' or 'game.Players.LocalPlayer.PlayerGui'. Evaluated as `return <parentPath>`. Set after the clone is made. Omit to leave the clone parentless (it still exists in memory).
instancePathYesLuau expression resolving to the source Instance to clone, e.g. 'game.ReplicatedStorage.Templates.Coin', 'game.Workspace.Model', or 'game.Players.LocalPlayer.PlayerGui.Main'. Evaluated as `return <instancePath>`.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond annotations: the clone starts parentless and is parented LAST only if parentPath is given, :Clone() returns nil on non-Archivable sources, and a parented clone immediately affects the running game and may replicate. These are rich behavioral disclosures the destructiveHint/readOnlyHint flags alone would not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the critical warning (WRITES LIVE GAME STATE) and proceeds through mechanism, use cases, caveats, return shape, and metadata. Dense and largely waste-free, though the signature/phase/requires metadata lines make it longer than strictly necessary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, but the description supplies the return shape ({ Source, Clone, Parented, ok } or { error }), failure behavior, prerequisites, and a verification tool (get-instance-properties). Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters in detail. The description adds minor sequencing context (parenting happens after cloning, threadContext defaults to server), but mostly restates what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (deep-copy an Instance), explains the exact mechanism (resolve a Luau expression, source:Clone() duplicates the instance and all descendants, optionally parent the clone), and is clearly distinguishable from siblings like create-instance and destroy-instance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete when-to-use scenarios (duplicating a template Part/Model/GUI, spawning extra copies, snapshotting a subtree before mutating) and states prerequisites (active-client, resolved-target, explicit-mutation-approval). It does not explicitly name a sibling alternative, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

closure-capabilitiesProbe the complete closure/reflection primitive surfaceB
Read-onlyIdempotent

Read-only capability matrix for the official Volt closure library plus the MCP's useful debug closure operations. Reports the selected alias for every primitive without calling it. Signature: { threadContext: number? }. Phase: verify; cost=medium; idempotency=read-only. Requires: active-client. Capabilities: debug closure primitives. Produces: structured-result. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds two useful behavioral facts beyond the annotations: primitives are NOT invoked ('without calling it') and an active-client is required. However, the remaining 'Phase/cost/idempotency/Safety' lines largely restate the annotation hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The first sentence is a strong, front-loaded purpose statement, but it is followed by a templated block (Phase, cost, idempotency, Capabilities, Produces, Safety, On failure) that repeats annotation data and is only loosely integrated. The 'On failure' pointer to tool-schema is useful, but overall the entry is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-required-param, read-only probe with a fully documented schema, this is adequate. Yet there is no output schema and 'Produces: structured-result' gives no sense of what the capability matrix contains, leaving the agent unable to anticipate the return shape.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single parameter is optional, so the baseline is 3. The description only echoes the schema with 'Signature: { threadContext: number? }' and adds no syntax, default, or semantic detail beyond it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('reports the selected alias for every primitive without calling it') and identifies the scope as the official Volt closure library plus debug closure operations. It is distinguishable from introspection siblings like get-closure-protos, though it does not name a sibling to contrast against.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Phase: verify' and 'Requires: active-client' imply when the tool is relevant and what precondition is needed, but there is no explicit statement of when to choose this over tool-schema, list-tools, or get-closure-protos. Usage is implied rather than directed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

comm-channel-monitorMonitor messages on an Actor communication channelA
Destructive

WRITES LIVE GAME STATE when starting/stopping. Connect to Channel.Event, retain a bounded 200-message buffer, and poll by monitor key. Start/stop require confirm=true. Signature: { action: "start" | "poll" | "stop", key: any?, limit: any?, clear: any?, confirm: any?, threadContext: number?, id: any? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval. Capabilities: create_comm_channel. Produces: bounded-event-snapshot. Verify with: assert-state. Safety: MUTATING; changes persistent executor-side observer or hook state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoOptional validated input for id.
keyNoOptional validated input for key.default
clearNoOptional validated input for clear.
limitNoOptional hard result/work budget used to keep output and runtime bounded.
actionYesoperation selector; use one of the schema's allowed values.
confirmNoExplicit safety acknowledgement; must be true before this state-changing operation runs.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, but the description goes further: it flags that start/stop write live game state, that executor-side observer state persists, that explicit mutation approval is required, that the buffer is bounded to 200 messages, and where to look on failure. The remaining vagueness is what exactly 'writes live game state' mutates and whether poll itself is side-effect free.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The lead sentence is front-loaded and the operational facts (confirm requirement, buffer size, key) come early. However, the inline Signature block duplicates the input schema field-for-field, and tag soup like 'Phase: act; cost=medium; idempotency=contextual-write' plus 'Capabilities: create_comm_channel' adds volume without much decision value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating, 7-parameter tool with no output schema, the description covers mutation consequences, approval prerequisites, buffer bounding, and a failure-recovery pointer (tool-schema, assert-state), which is enough to call it safely. It stops short of describing what a poll returns beyond 'bounded-event-snapshot'.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds real semantics: which actions (start/stop) require confirm=true, that polling is keyed by monitor key, and that limit is a hard budget tied to the 200-message buffer. The signature block itself mostly restates schema fields, so it does not go beyond a 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: connect to Channel.Event, retain a bounded 200-message buffer, and poll by monitor key. It is clearly distinguishable from generic execute/eval siblings, but it never names or contrasts with the closest siblings (actor-event-monitor, lua-state-event-monitor, get-comm-channel), leaving the agent to infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: start/stop require confirm=true and polling is keyed, which tells the agent how to call it but not when to prefer it over sibling monitors or the one-shot get-comm-channel. No explicit when-not or alternative routing is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compare-gc-snapshotsCapture / compare GC snapshotsA
Read-onlyIdempotent

Take a census of the garbage collector (getgc) and diff it over time to find what was allocated or freed between two points - a leak-hunting / behaviour-attribution tool. Workflow: call with action='capture' (a baseline), perform some action in-game (open a menu, fire a remote, etc.), then call with action='compare' (same snapshotName) to see what changed. Snapshots are stored CLIENT-SIDE in getgenv().__mcp_gc_snapshots[name], so they live in the target game and are naturally isolated per game / per session. capture returns: { action='capture', name, ts, counts={ function, table, thread, total }, fnSampled, truncated }. compare returns: { action='compare', name, baselineTs, nowTs, countDeltas={ function, table, thread, total }, newFunctions=[..pointers..], newFunctionCount, freedApprox (functions no longer present), sampledOnly, truncated }. Caps: at most 200000 GC objects are walked per pass and at most 4000 function pointers are remembered; results note when truncation occurred, so treat large freed/new counts near the cap as approximate. Signature: { action: "capture" | "compare", snapshotName: any?, threadContext: number? }. Phase: verify; cost=medium; idempotency=read-only. Requires: active-client. Capabilities: getgc. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYes'capture' stores a baseline census of the GC under snapshotName. 'compare' diffs the current GC against the previously stored snapshot of that name and reports new/freed functions and per-type count deltas.
snapshotNameNoName/slot for the snapshot (default 'default'). Use distinct names to keep several independent baselines around at once.default
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the readOnly/destructive/idempotent annotations by disclosing where state lives (getgenv().__mcp_gc_snapshots[name], per-game/per-session isolation) and hard caps (200000 objects walked, 4000 function pointers remembered) with truncation flags to treat large counts as approximate. Requires/preconditions (active-client, getgc capability) are also stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and workflow, then return shapes and caps, so an agent can stop reading early. It is dense rather than bloated, though the trailing 'Phase/cost/idempotency/Safety: read-only' block partly restates annotation data already provided structurally.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the full burden of describing return values and does so precisely for both action='capture' and action='compare' (counts, fnSampled, countDeltas, newFunctions, freedApprox, truncated). Nothing an agent needs to call or interpret this tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3; the description adds sequencing semantics the schema lacks — capture must precede compare, and compare must reuse the same snapshotName, with distinct names keeping independent baselines. threadContext is not elaborated further than the schema already does.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('take a census of the garbage collector (getgc) and diff it over time') and names the analyst goal ('leak-hunting / behaviour-attribution tool'). It is clearly separable from siblings like list-gc-functions or list-gc-tables, which enumerate rather than diff snapshots.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit three-step workflow (capture baseline, perform an in-game action, compare with the same snapshotName), which tells the agent precisely when each action applies. It does not, however, name any alternative sibling (e.g. measure-memory, list-gc-functions) or state when this tool is the wrong choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compare-instancesTest whether two references point at the same underlying instanceA
Read-onlyIdempotent

Resolve two Luau expressions and report whether they reference the SAME underlying Roblox instance via compareinstances. This matters because a game (or an executor proxy/clone) can hand you a wrapped userdata whose identity differs from the real Instance even though == or :GetFullName() looks identical — and some anticheat hands out decoy/newproxy objects. compareinstances unwraps those and compares true instance identity, so you can confirm 'is this captured remote actually game.ReplicatedStorage.RemoteEvent, or a lookalike?'. Non-mutating. Requires compareinstances; returns { Same, TypeA, TypeB } or { error } when the capability is missing or either expression fails to evaluate. Signature: { pathA: string, pathB: string, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathAYesFirst Luau expression to compare, e.g. 'game.ReplicatedStorage.RemoteEvent' or a captured handle like 'getgenv().__capturedRemote'. Evaluated as `return <pathA>`.
pathBYesSecond Luau expression to compare against pathA, e.g. 'getreg()[123]' or 'getrawmetatable(game).__index'. Evaluated as `return <pathB>`.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, and the description reinforces rather than contradicts them. Beyond that it discloses the capability dependency (requires compareinstances), the failure path (error when capability missing or an expression fails to evaluate), and a partial return shape ({ Same, TypeA, TypeB }) — useful behavioral context a bare annotation set would not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is well front-loaded, but the trailing metadata block is redundant: 'Non-mutating', 'idempotency=read-only', and 'Safety: read-only' restate the same fact three times, and the signature repeats the schema. It could be trimmed without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, yet the description explains the return shape and error behavior, and it names the capability requirement. For a 3-parameter read-only tool this is essentially complete; only minor gaps like timeout/error field detail remain, and it points to tool-schema for those.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents pathA, pathB and threadContext with examples. The description's signature line merely restates the schema and adds no syntax or format meaning beyond it, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (resolve/compare) and resource (two Luau expressions referencing an underlying Roblox instance). It further distinguishes its identity-comparison semantics from `==` or :GetFullName(), which a naive agent might otherwise use, so the tool's niche is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete motivating scenario ('confirm "is this captured remote actually game.ReplicatedStorage.RemoteEvent, or a lookalike?"') and states the prerequisite capability and non-mutating nature. It lacks an explicit when-not or a named alternative sibling, so it's clear context without full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

configure-remote-spyConfigure spy capture and GUIA
Destructive

WRITES LIVE GAME STATE. Configures an already-running spy without restarting or expiring IDs. Use engine=ketamine for Ketamine. Patch persistent MCP capture filters, pause/resume with capture.enabled, resize bounded history, or control Ketamine GUI visibility/logging. Omitted settings are preserved; resetFilters clears capture filters. Returns effective settings and capabilities. Pausing/filtering capture does not unblock remotes or change network delivery. Reads/configuration never load a spy; start it with remote-spy first. Cobalt accepts capture/buffer settings but rejects Ketamine GUI settings. Signature: { engine: "cobalt" | "ketamine"?, max: number?, capture: { enabled: boolean?, direction: "Incoming" | "Outgoing" | "Both"?, nameFilter: string?, method: "FireServer" | "InvokeServer" | "OnClientEvent" | "OnClientInvoke"?, classFilter: "RemoteEvent" | "UnreliableRemoteEvent" | "RemoteFunction"?, blockedOnly: boolean? }?, resetFilters: boolean?, guiVisible: boolean?, guiLogging: boolean?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval. Produces: bounded-event-snapshot. Verify with: assert-state. Safety: MUTATING; changes persistent executor-side observer or hook state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
maxNoRetained MCP calls; resizing preserves newest history and IDs.
engineNoRemote-spy backend. Only one engine runs per client; starting another stops the current engine. Defaults to cobalt.cobalt
captureNoPersistent capture patch; omitted fields keep their current values. Filters apply only to future MCP captures.
guiLoggingNoKetamine only: enable/disable new GUI log entries independently of MCP capture; disabling also clears queued GUI entries.
guiVisibleNoKetamine only: show/hide its GUI without unloading the spy.
resetFiltersNoClear persistent capture filters before applying this patch; preserve pause state, history, and network controls.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true and non-idempotent, but the description adds substantial non-structured context: omitted settings are preserved, resetFilters clears filters while preserving pause/history, and pausing/filtering does not unblock remotes or alter network delivery. It also discloses required preconditions (active-client, explicit-mutation-approval) and that permanent observer/hook state changes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the critical 'WRITES LIVE GAME STATE' warning, then scope, then behavior. Slightly bloated because the inline Signature block restates the input schema field-for-field, which the agent can already read, costing it a fifth point but not much readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-param, nested-object mutation tool with no output schema, the description covers everything an agent needs: preconditions, engine branching, preservation/reset semantics, side-effect boundaries ('does not unblock remotes'), return content, and a failure recovery pointer to tool-schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3 and the schema carries the field-level meaning. The description still adds cross-field semantics the schema can't express — omitted fields are preserved, resetFilters ordering, the Cobalt/Ketamine acceptance split, and per-engine applicability of GUI fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource and scope: 'Configures an already-running spy without restarting or expiring IDs.' It explicitly distinguishes itself from the lifecycle starter by stating 'Reads/configuration never load a spy; start it with remote-spy first.' An agent can tell it apart from remote-spy/ensure-remote-spy without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use (patch filters, pause/resume, resize history, GUI control), names the prerequisite alternative ('start it with remote-spy first'), and states an engine-specific exclusion ('Cobalt accepts capture/buffer settings but rejects Ketamine GUI settings'). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

count-function-callsCount how many times a function is called (MUTATES STATE via hookfunction)A
Destructive

WRITES LIVE GAME STATE — INSTALLS A PERSISTENT GLOBAL HOOK. Lightweight call-frequency counter for any function: hook a target so that every invocation bumps an integer counter, then read the counter, then restore the original. Unlike hook-and-log-function (which records full args/returns), this captures ONLY a count, so it is the cheapest way to answer 'is this function actually being called, and how often?' — ideal for confirming an anticheat tick fires, measuring how hot a code path is, or verifying a remote handler runs. WORKFLOW (stateful — survives across tool calls via getgenv().__mcp_callCounts, keyed by functionPath): 1. action='start' with functionPath — resolves the target, captures the original, installs a counting hook that transparently calls the original and increments a counter. Returns { started, key }. 2. action='fetch' with the same functionPath — returns { calls } captured so far WITHOUT stopping. Poll to watch live. 3. action='stop' with the same functionPath — restores the original and clears the entry. Returns { stopped, calls }. CAVEATS: the hook is GLOBAL and PERSISTS until you stop it (or the client restarts), adds (small) overhead on every call, and a live function hook CAN TRIP ANTICHEAT — always stop when done. The counting work is pcall-isolated and the hook always calls through to the original, so behavior is unchanged. Requires hookfunction, newcclosure, and getgenv; restoration uses hookfunction(target, original) with a restorefunction fallback. Returns { error } if a capability is missing, the target cannot be resolved, or there is no active counter for fetch/stop. Signature: { action: "start" | "fetch" | "stop", functionPath: string?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYes'start' installs the counting hook on functionPath; 'fetch' returns the call count so far (hook stays live); 'stop' restores the original function and clears the counter. Use the SAME functionPath for all three so they address the same registry entry.
functionPathNoLuau expression resolving to the function to count, e.g. 'getsenv(game.Players.LocalPlayer.PlayerScripts.Main).heartbeat', 'getrawmetatable(game).__namecall', or 'getconnections(game.Workspace.Part.Touched)[1].Function'. Evaluated as `return <functionPath>` and must resolve to a function. REQUIRED for 'start'. For 'fetch'/'stop' it is the registry key identifying which running counter to act on, so it must match the string used at start.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations by disclosing that the hook is global and persists until stopped or client restart, adds per-call overhead, and can trip anticheat, plus the pcall-isolation and call-through guarantee and the hookfunction/restorefunction fallback restoration path. This is exactly the mutation-safety context an agent needs before invoking a destructiveHint tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the critical warning ('WRITES LIVE GAME STATE — INSTALLS A PERSISTENT GLOBAL HOOK') and the numbered workflow is easy to follow. It is long and carries some boilerplate meta footer (Phase/cost/idempotency/Safety) that slightly dilutes density, but nearly every sentence carries operational value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex stateful mutation tool with no output schema, it documents each action's return shape ({started,key}, {calls}, {stopped,calls}), the error conditions, the persistence model, and the required capabilities — nothing an agent needs to invoke and manage it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so the baseline is 3, but the description adds real value by explaining that action='fetch' polls without stopping, that functionPath doubles as the registry key across all three actions, and the underlying getgenv().__mcp_callCounts keying. threadContext is left unexplained but the schema already documents it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Lightweight call-frequency counter for any function') and explicitly distinguishes itself from the sibling hook-and-log-function by contrasting what each captures. An agent can tell exactly what this does and how it differs without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the alternative (hook-and-log-function) with the tradeoff that selects it, and gives concrete when-to-use scenarios ('confirming an anticheat tick fires, measuring how hot a code path is, or verifying a remote handler runs'). It also spells out the three-action workflow and that the same functionPath must be reused.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

count-signal-connectionsCount connections on a signal (breakdown)A
Read-onlyIdempotent

Resolve an RBXScriptSignal and return a fast numeric breakdown of everything connected to it: Total connections, how many are Lua vs Foreign (engine/C-side), how many are currently Enabled vs Disabled, and how many carry a Lua Function. Use this as a lightweight first pass before list-signal-connections when you only need counts (e.g. "how many handlers are on this RemoteEvent?", "are any of this signal's connections disabled?"). Does NOT describe each function, so it is much cheaper than the full listing. Requires the executor's getconnections; degrades to a clear { error } if unavailable. Returns { Signal, Instance?, Total, Lua, Foreign, Enabled, Disabled, WithFunction }. Signature: { instancePath: string, signalName: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: getconnections. Produces: created-handle. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
signalNameNoName of the RBXScriptSignal member on the instance (e.g. 'Touched', 'Changed', 'OnClientEvent'). Leave empty/omit if instancePath already evaluates to the signal.
instancePathYesLua expression resolving to the Instance that owns the signal (e.g. 'game.Workspace.Door', 'game.ReplicatedStorage.RemoteEvent'). Evaluated as `return <instancePath>`. Leave signalName empty if this expression already resolves to the RBXScriptSignal itself.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the read-only/idempotent safety profile, but the description adds operational context annotations cannot express: it depends on the executor's getconnections, degrades to a clear { error } when that is unavailable, and requires an active-client/resolved-target context at cost=medium. It also names the exact return shape, which matters since there is no output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose, breakdown fields, and routing guidance are all front-loaded in the first three sentences, which is exactly the information an agent needs. The trailing metadata block (phase/cost/idempotency/Safety: read-only) partially duplicates the annotations and the schema, adding length without new signal, which keeps it just short of a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a count-only read tool with no output schema, the description supplies the return field list, the dependency (getconnections), the failure mode ({ error }), the execution context requirements, and a pointer to tool-schema for exact fields on failure. An agent has everything needed to call it correctly and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters are fully documented in the schema itself, including the instancePath/signalName interaction. The description's 'Signature' line restates the parameters without adding format or fallback semantics beyond what the schema provides, so baseline 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Resolve an RBXScriptSignal and return a fast numeric breakdown') and enumerates exactly what is counted (Total, Lua vs Foreign, Enabled vs Disabled, WithFunction). It distinguishes itself from its closest sibling by name, so an agent can tell it apart from list-signal-connections without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly frames itself as a 'lightweight first pass before list-signal-connections when you only need counts,' and supplies concrete trigger questions ('how many handlers are on this RemoteEvent?'). It also states the inverse condition ('Does NOT describe each function, so it is much cheaper than the full listing'), which routes the agent to the alternative when descriptions are needed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create-comm-channelCreate and retain an Actor communication channelA
Destructive

WRITES LIVE GAME STATE. Call create_comm_channel(name?), retain the Channel in getgenv().__mcp_comm_channels, and return only its serializable identifier/metadata. Requires confirm=true. Signature: { name: any?, confirm: boolean?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval. Capabilities: create_comm_channel. Produces: created-handle. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoOptional validated input for name.
confirmNoExplicit safety acknowledgement; must be true before this state-changing operation runs.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag readOnlyHint=false and destructiveHint=true, and the description adds substantially more: it writes live game/client state, requires confirm=true and explicit mutation approval, discloses the retention side-effect (stores the handle in getgenv().__mcp_comm_channels), specifies the reduced return surface, and gives failure guidance. This is well beyond what the annotations convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The critical warning is front-loaded and each clause carries real information (signature, phase/cost, requirements, verification, safety). It is dense rather than padded, though the 'On failure: inspect tool-schema...' boilerplate and the phase/cost tokens are template filler that dilute the signal slightly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating tool with no output schema, the description covers the safety contract, preconditions, side effects, verification, and error path. It is close to complete; only explicit alternative-tool routing is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents name, confirm, and threadContext in detail. The description restates the signature and marks confirm as required, but adds no syntax, format, or constraint detail beyond the schema — the baseline 3 for full schema coverage is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description leads with a specific verb+resource ('WRITES LIVE GAME STATE', create_comm_channel) and states the exact effect: create a channel, retain it in getgenv().__mcp_comm_channels, and return only its serializable identifier/metadata. That distinguishes it cleanly from siblings like get-comm-channel and fire-comm-channel.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives prerequisites (active-client, explicit-mutation-approval, confirm=true) and a verification step (assert-state), which implies when the tool is appropriate. But it never names the sibling alternatives (get-comm-channel, fire-comm-channel, run-on-actor) or states when-not to use this tool, so routing guidance is only partial.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create-instanceCreate a new Instance in the live gameA
Destructive

WRITES LIVE GAME STATE. Construct a brand-new Instance via Instance.new(className), optionally set its Name and a list of initial properties, and optionally parent it into the game tree. Useful while debugging for spawning a Part, a Highlight/BillboardGui ESP marker, a Folder, a Value object (IntValue/StringValue/BoolValue), or any other class. PROCESS: Instance.new is pcall-guarded (returns a clean error if className is invalid); Name is set if provided; each property is set independently and pcall-guarded so one bad property does not abort the others (failures are collected into propErrors); Parent is set LAST (only if parentPath is given) so all properties are applied before the instance becomes live in the tree. If you do NOT pass parentPath the instance is created but left parentless (nil) — it still exists in memory and can be parented later via set-instance-property on its Parent. WARNING: a parented instance immediately affects the running game and may replicate. Returns { Created, ClassName, Parented, propErrors, ok } or { error }. Signature: { className: string, name: string?, parentPath: string?, properties: {{ name: string, value: { kind: "string" | "number" | "boolean" | "nil" | "raw", value: string | number | boolean? } }}?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Produces: created-handle. Verify with: get-instance-properties. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoOptional Name to assign to the new instance. If omitted the engine's default name for the class is kept.
classNameYesThe Roblox class name to instantiate, e.g. 'Part', 'Folder', 'Highlight', 'BillboardGui', 'IntValue', 'StringValue', 'ScreenGui'. Passed to Instance.new(className); an unknown/non-creatable class returns a clean error.
parentPathNoOptional Luau expression resolving to the Instance that should become the new instance's Parent, e.g. 'game.Workspace', 'game.Players.LocalPlayer.PlayerGui', or 'game:GetService("ReplicatedStorage")'. Evaluated as `return <parentPath>`. Set LAST, after all properties. Omit to leave the instance parentless.
propertiesNoOptional list of initial properties to assign before parenting, e.g. [{ name: 'Anchored', value: { kind: 'boolean', value: true } }, { name: 'Size', value: { kind: 'raw', value: 'Vector3.new(4,1,4)' } }]. Omit or pass [] for none.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations already declaring destructive=true and non-idempotent, the description goes well beyond: pcall-guarded className validation, per-property pcall isolation collecting propErrors, Parent applied LAST, parentless-by-default behavior, and the replication warning. This is rich behavioral context an agent cannot derive from structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The live-state warning is correctly front-loaded and the PROCESS paragraph is dense with useful detail. The embedded 'Signature:' line and phase/cost/idempotency meta largely restate schema and annotations, which is minor waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, but the description documents the return shape ({ Created, ClassName, Parented, propErrors, ok } or { error }), failure handling, and a verification step. An agent has everything needed to invoke and interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, giving a baseline of 3, but the description adds ordering semantics beyond per-parameter docs — properties are applied before Parent, and parentPath is evaluated as a Luau expression. The duplicated signature line is redundant but the process detail adds real meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Construct a brand-new Instance via Instance.new(className)' with optional Name, properties, and parenting. It is distinguishable from siblings like clone-instance, set-instance-property, and destroy-instance by its creation semantics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains when to use it with concrete debugging examples (Part, ESP marker, Folder, Value objects) and lists prerequisites (active-client, resolved-target, explicit-mutation-approval) plus a verification tool. It does not explicitly contrast with clone-instance for existing instances, leaving some alternative selection to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

crypt-base64-decodeBase64-decode a string via the executor crypt libraryA
Read-onlyIdempotent

Decode a Base64 string back to its raw bytes using the executor's crypt library. This is a pure, side-effect-free compute that runs entirely in-game. The decoder probes BOTH the flat form (crypt.base64decode) and the namespaced form (crypt.base64.decode), using whichever the executor provides. Requires a crypt table; on an executor without it (no crypt, or no base64 decoder under either name) it returns { error } instead of throwing. The call is pcall-guarded so malformed Base64 degrades to a clean error. Returns { decoded } or { error }. Signature: { data: string, threadContext: number?, timeoutMs: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-result. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
dataYesThe Base64-encoded string to decode.
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations by disclosing that the decoder probes both flat and namespaced forms, requires a crypt table, returns { error } rather than throwing when crypt is absent, and is pcall-guarded so malformed Base64 degrades to a clean error. This is exactly the runtime behavior an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core operation and error-handling semantics well. It is slightly padded with trailing metadata (Phase/cost/idempotency/safety restated alongside annotations) and a generic 'On failure: inspect tool-schema' boilerplate sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description specifies the return shapes ({ decoded } or { error }), the failure behavior, and the executor dependency, leaving nothing essential missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents data, timeoutMs, and threadContext with meaning. The description repeats the signature but adds no syntax or format detail beyond what the schema provides, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (decode) and resource (Base64 string to raw bytes) plus the exact library used (executor's crypt library). It is immediately distinguishable from the sibling crypt-base64-encode.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description establishes context (pure, side-effect-free, runs in-game) but never states when to prefer this tool over alternatives or when it is inapplicable. Usage is implied by the operation rather than explicitly spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

crypt-base64-encodeBase64-encode a string via the executor crypt libraryA
Read-onlyIdempotent

Encode an arbitrary string to Base64 using the executor's crypt library. This is a pure, side-effect-free compute that runs entirely in-game. The encoder probes BOTH the flat form (crypt.base64encode) and the namespaced form (crypt.base64.encode), using whichever the executor provides. Requires a crypt table; on an executor without it (no crypt, or no base64 encoder under either name) it returns { error } instead of throwing. The call is pcall-guarded so a malformed input degrades to a clean error. Returns { encoded } or { error }. Signature: { data: string, threadContext: number?, timeoutMs: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-result. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
dataYesThe raw string to Base64-encode.
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations (readOnlyHint/idempotentHint/destructiveHint), disclosing that it probes both flat and namespaced crypt forms, that a missing crypt table yields { error } rather than throwing, and that the call is pcall-guarded so malformed input degrades cleanly. This is exactly the kind of failure-mode context annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in a strong opening sentence, but the tail is padded with schema-duplicating signature text and annotation-duplicating metadata (idempotency=read-only, Safety: read-only) plus generic filler ('inspect tool-schema for exact fields...'). Roughly half the text is redundant boilerplate.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by naming the return shapes ({ encoded } or { error }) and the failure path, and it states the active-client requirement. It is essentially complete for a simple read-only compute tool, with only the boilerplate tail adding noise rather than substance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents data, timeoutMs, and threadContext fully. The description repeats the signature but adds no syntax, format, or constraint meaning beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (encode) and resource (string to Base64) and names the underlying mechanism (executor `crypt` library), which cleanly distinguishes it from the sibling crypt-base64-decode. An agent can identify the operation without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage through its purpose and metadata (Phase: observe, Requires: active-client), but never states when to reach for this tool versus alternatives like crypt-hash or execute. There is no explicit when-to-use/when-not guidance, so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

crypt-decryptSymmetrically decrypt data with crypt.decryptA
Read-onlyIdempotent

Decrypt a ciphertext with the executor's crypt.decrypt(data, key, iv, mode?). Supply the same Base64 key and iv that were used to encrypt, plus the optional cipher mode if a non-default mode was used. The function returns the recovered plaintext. This is a pure, side-effect-free compute that runs entirely in-game. Requires crypt.decrypt; on an executor without it, it returns { error } instead of throwing. The call is pcall-guarded so a wrong key/iv/mode degrades to a clean error. Returns { plaintext } or { error }. Signature: { data: string, key: string, iv: string, mode: string?, threadContext: number?, timeoutMs: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-result. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
ivYesThe Base64-encoded initialization vector that was used to encrypt.
keyYesThe Base64-encoded symmetric key that was used to encrypt.
dataYesThe ciphertext to decrypt.
modeNoOptional cipher mode; must match the mode used at encrypt time if it was non-default.
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/destructive, but the description adds real value beyond them: it discloses the pure side-effect-free nature, the executor requirement for crypt.decrypt, and the pcall-guarded failure contract where wrong key/iv/mode yields { error } rather than a throw.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core is front-loaded well, but the trailing metadata string ('Phase: observe; cost=medium; idempotency=read-only... Safety: read-only') repeats what the annotations and earlier sentences already convey, adding bulk without new meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly explains return values ({ plaintext } or { error }) and failure handling. It covers the key facts an agent needs, though it does not detail output format nuances or timeout/threadContext usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds a genuine cross-parameter constraint absent from the schema: key and iv must be the SAME Base64 values used at encrypt time, and mode must match if non-default.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (decrypt), resource (ciphertext), and even the underlying executor function crypt.decrypt(data, key, iv, mode?). It is clearly distinguishable from crypt-encrypt and other crypt-* siblings by stating it recovers plaintext.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives prerequisite guidance ('supply the same Base64 key and iv that were used to encrypt', plus the mode if non-default), but never names alternatives or when-not-to-use conditions. Usage is implied rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

crypt-encryptSymmetrically encrypt data with crypt.encryptA
Read-onlyIdempotent

Encrypt a string with the executor's crypt.encrypt(data, key, iv?, mode?). The key is a Base64 string (see crypt-generate-key). An optional Base64 initialization vector (iv) and an optional cipher mode (e.g. 'CBC', 'CTR') may be supplied; when omitted the executor picks/derives them. The function returns the ciphertext plus the iv actually used, both of which are reported. This is a pure, side-effect-free compute that runs entirely in-game. Requires crypt.encrypt; on an executor without it, it returns { error } instead of throwing. The call is pcall-guarded so a bad key/iv degrades to a clean error. Returns { ciphertext, iv } or { error }. Signature: { data: string, key: string, iv: string?, mode: string?, threadContext: number?, timeoutMs: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-result. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
ivNoOptional Base64-encoded initialization vector; omit to let the executor derive one.
keyYesThe Base64-encoded symmetric key (see crypt-generate-key).
dataYesThe plaintext string to encrypt.
modeNoOptional cipher mode (e.g. 'CBC', 'CTR', 'CFB', 'OFB', 'ECB'); omit for the executor default.
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive), yet the description adds substantial context: pure/side-effect-free in-game compute, pcall-guarded so bad key/iv degrades to a clean error, and returns { ciphertext, iv } or { error }. This is rich, non-redundant behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded and clear, but the entry is long: it restates the full signature (already in schema) and appends metadata lines (Phase, cost, idempotency, Safety, Produces, and an 'On failure' pointer) that partly duplicate annotations and structured data.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description explains the return shape ({ ciphertext, iv } or { error }), error/degradation behavior, prerequisites, and the executor's fallback for omitted params. An agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description notes Base64 key/iv, mode examples (CBC/CTR), and that the executor derives iv/mode when omitted, but most of this is already documented in the schema properties, so it adds only marginal value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Encrypt a string with the executor's crypt.encrypt') and even gives the call signature. It is clearly distinguishable from siblings like crypt-decrypt, crypt-hash, and crypt-base64-encode.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear context: a pure, side-effect-free compute that runs in-game, requires an active client, and depends on `crypt.encrypt` existing (returning { error } otherwise). It also routes to crypt-generate-key for obtaining the key. It doesn't contrast when to prefer this over crypt-decrypt or other crypto siblings, so no explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

crypt-generate-bytesGenerate random bytes with crypt.generatebytesA
Read-onlyIdempotent

Generate size cryptographically random bytes using the executor's crypt.generatebytes(size). The bytes are returned as a Base64 string (decode with crypt-base64-decode if you need the raw bytes). This is a pure, side-effect-free compute that runs entirely in-game. The default size is 16; the request is clamped to a sane 1..1024 range. Requires crypt.generatebytes; on an executor without it, it returns { error } instead of throwing. The call is pcall-guarded. Returns { bytes } or { error }. Signature: { size: any?, threadContext: number?, timeoutMs: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-result. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
sizeNoNumber of random bytes to generate (default 16, clamped to 1..1024).
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so the bar is lower, yet the description adds real behavior: default size 16, clamp to 1..1024, pcall-guarded, returns { error } rather than throwing on unsupported executors, and pure side-effect-free in-game compute. That is meaningful context beyond the structured safety hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening sentences are well front-loaded, but the trailing meta block (Phase/cost/idempotency/Requires/Produces/Safety plus a repeated signature) largely restates annotations and the schema, padding the definition without adding selection value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly explains the return contract ({ bytes } or { error } as Base64) and the failure mode, which is what an agent needs to invoke and interpret it. Slightly long-winded but functionally complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already documents size with its default and 1..1024 clamp plus timeoutMs/threadContext. The description repeats the clamp/default without adding syntax beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (generate) and resource (cryptographically random bytes) and names the underlying primitive crypt.generatebytes. It also points to the complementary crypt-base64-decode sibling for raw bytes, so an agent can place it in the crypto family without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context: bytes come back Base64, decode with crypt-base64-decode if raw bytes are needed, and it requires an executor exposing crypt.generatebytes. It does not explicitly contrast with crypt-generate-key or crypt-hash, but the usage conditions are clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

crypt-generate-keyGenerate a random symmetric key with crypt.generatekeyA
Read-onlyIdempotent

Generate a cryptographically random symmetric key using the executor's crypt.generatekey(). The key is returned as a Base64 string suitable for use with crypt-encrypt / crypt-decrypt. This is a pure, side-effect-free compute that runs entirely in-game. Requires crypt.generatekey; on an executor without it, it returns { error } instead of throwing. The call is pcall-guarded. Returns { key } or { error }. Signature: { threadContext: number?, timeoutMs: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-result. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: discloses it is a pure, side-effect-free in-game compute, names the required `crypt.generatekey` capability, and states the failure contract (returns { error } instead of throwing, pcall-guarded). This is exactly the behavioral context annotations can't convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and return shape, which is good, but the trailing metadata (`idempotency=read-only`, `Safety: read-only`) duplicates the readOnlyHint/idempotentHint annotations and adds boilerplate weight without new information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return contract itself ("Returns { key } or { error }") and names the fallback tool (tool-schema) for deeper details. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both params are already documented with their own descriptions. The description restates the signature and default-omission behavior but adds no new syntax or constraint meaning, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Generate a cryptographically random symmetric key") and names the downstream siblings it feeds (crypt-encrypt / crypt-decrypt). An agent can distinguish it from crypt-generate-bytes or crypt-hash without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for use: the key is Base64-suitable for crypt-encrypt / crypt-decrypt, and it notes the active-client prerequisite. It does not explicitly contrast against the nearest alternative (crypt-generate-bytes), so it stops short of full when/when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

crypt-hashHash a string with crypt.hash (sha1/sha256/sha384/sha512/md5)A
Read-onlyIdempotent

Compute a cryptographic digest of a string using the executor's crypt.hash(data, algorithm). This is a pure, side-effect-free compute that runs entirely in-game. Supported algorithms are sha1, sha256 (default), sha384, sha512, and md5; the result is the hex digest string returned by the executor. Requires crypt.hash; on an executor without it (no crypt table or no hash function) it returns { error } instead of throwing. The call is pcall-guarded so an unsupported algorithm degrades to a clean error. Returns { hash, algorithm } or { error }. Signature: { data: string, algorithm: any?, threadContext: number?, timeoutMs: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-result. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
dataYesThe string to hash.
algorithmNoThe hash algorithm to use (default 'sha256').sha256
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, yet the description adds genuinely new behavioral facts: it runs entirely in-game via the executor's crypt.hash, it is pcall-guarded, and on an executor lacking crypt.hash it returns { error } instead of throwing. Those failure semantics are not derivable from the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core behavior and error semantics are front-loaded in the first two sentences. The trailing 'Phase/cost/idempotency/Safety/On failure' block is somewhat redundant with the annotations and footer-ish, but the overall structure remains scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description steps in to define the return shape ({ hash, algorithm } or { error }) and the failure path, and it points to tool-schema for exact fields. An agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents data, algorithm (with enum + default), timeoutMs, and threadContext. The description restates the signature and adds only the note that the result is a hex digest string, which is marginal beyond the schema baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Compute a cryptographic digest of a string') and enumerates the exact algorithm surface (sha1/sha256/sha384/sha512/md5) plus the default. This clearly separates it from adjacent siblings like crypt-encrypt, crypt-base64-encode, and generic execution tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives useful context for invocation: phase=observe, cost=medium, requires active-client, and that the tool degrades gracefully rather than throwing. It does not explicitly contrast against alternatives such as run-luau or execute-lua-state, which is the only real gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete-fileDelete a file from the executor workspace (UNC delfile)A
Destructive

Delete a single file from the executor's workspace folder. NOTE: this is executor-side file I/O — the connector runs INSIDE the executor, so the path is relative to the executor's workspace directory on the host machine, NOT the Roblox game. This is destructive and cannot be undone. Requires the UNC function delfile(path). The call is type-guarded and pcall-wrapped: if delfile is missing you get { error = 'delfile is not available in this executor.' }, and any failure (missing file, permission) returns { error = }. Returns { path, ok = true } or { error }. Signature: { path: string, threadContext: number?, timeoutMs: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Capabilities: executor filesystem. Produces: operation-receipt. Verify with: file-exists. Safety: MUTATING; writes executor workspace filesystem. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesPath to the file to delete within the executor workspace, e.g. 'logs/old.txt'.
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=false, but the description adds substantial context beyond them: the operation cannot be undone, it is type-guarded and pcall-wrapped, delfile-missing yields a specific error object, and failures (missing file, permission) return { error = <message> }. It also states the return shape and the approval requirement, which annotations do not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose and scope are front-loaded in the first two sentences, which is good, but the description then trails into a long block of metadata-style fields (Phase, cost, idempotency, Capabilities, Produces, On failure) and a generic closing instruction to inspect tool-schema. Several of these clauses are filler relative to selection and invocation, so it is more verbose than it needs to be.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return contract itself ({ path, ok = true } or { error }) plus error-case behavior, prerequisites, and a verification sibling. Nothing an agent needs to invoke and interpret this destructive tool correctly appears to be missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents path, timeoutMs and threadContext. The description restates the signature and reinforces that path is executor-relative rather than game-relative, which is useful disambiguation, but it adds no syntax, format, or default detail beyond the schema. Baseline 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Delete a single file from the executor's workspace folder') and immediately scopes it against siblings by clarifying it is executor-side I/O, not the Roblox game. An agent can distinguish it from delete-folder, read-file, write-file and append-file without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear operational context: the path is relative to the executor workspace on the host, not the game, and it lists prerequisites (active-client, resolved-target, explicit-mutation-approval) plus a verification step (file-exists). It does not explicitly contrast with the closest sibling, delete-folder, so it stops short of full when/when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete-folderDelete a folder from the executor workspace (UNC delfolder)A
Destructive

Delete a folder (and typically its contents) from the executor's workspace. NOTE: this is executor-side file I/O — the connector runs INSIDE the executor, so the path is relative to the executor's workspace directory on the host machine, NOT the Roblox game. This is destructive and recursive on most executors; it cannot be undone. Requires the UNC function delfolder(path). The call is type-guarded and pcall-wrapped: if delfolder is missing you get { error = 'delfolder is not available in this executor.' }, and any failure (missing folder, permission) returns { error = }. Returns { path, ok = true } or { error }. Signature: { path: string, threadContext: number?, timeoutMs: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Capabilities: executor filesystem. Produces: operation-receipt. Verify with: file-exists. Safety: MUTATING; writes executor workspace filesystem. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesPath to the folder to delete within the executor workspace, e.g. 'data/snapshots'.
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations by disclosing that deletion is recursive on most executors, irreversible, type-guarded and pcall-wrapped, with concrete error payloads and return shapes. Annotations cover the safety flags but the description adds substantial operational detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and the critical NOT-the-game warning are front-loaded, and the metadata (phase, idempotency, requires, produces) is organized. It is dense and includes some boilerplate, but every part is scannable and earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully specifies return values, error cases, prerequisites, and a verification tool. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents path, timeoutMs, and threadContext. The description restates the signature and the executor-relative path semantics, adding little beyond what the schema provides — baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (delete) and resource (folder) with explicit scope: the executor's workspace, not the Roblox game. This cleanly distinguishes it from siblings like delete-file and destroy-instance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context (executor-side file I/O, phase=act, requires explicit-mutation-approval) and a verification path (file-exists). It implies when to use it, but does not explicitly contrast against delete-file for the folder-vs-file decision.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

destroy-instanceDestroy a live Instance (IRREVERSIBLE)A
Destructive

WRITES LIVE GAME STATE. Resolve a Luau expression to an Instance and permanently remove it from the game via inst:Destroy(). Destroy() unparents the instance and ALL of its descendants, disconnects their events, and locks their Parent so they can never be re-parented — the subtree is gone. The instance's full path and ClassName are captured BEFORE destruction so the response records exactly what was removed. WARNING: THIS IS IRREVERSIBLE — there is no undo; you cannot get the instance back (use clone-instance first if you might need a copy). Destroying a player's Character, a critical service child, or a script can break the running game and may replicate to the server. Only destroy instances you are certain about. Returns { Destroyed, ClassName, ok } or { error }. Signature: { instancePath: string, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Produces: operation-receipt. Verify with: verify-path-exists. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
instancePathYesLuau expression resolving to the Instance to destroy, e.g. 'game.Workspace.UnwantedPart', 'game.Players.LocalPlayer.PlayerGui.Ad', or 'game.Workspace:FindFirstChild("Trap")'. Evaluated as `return <instancePath>`. Resolve the EXACT instance — Destroy also removes every descendant.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, but the description adds substantial context beyond them: exactly what Destroy does (unparents the subtree, disconnects descendant events, locks Parent so re-parenting is impossible), the irreversibility, and the server-replication risk of destroying Characters or script instances. It also discloses the return shape and that path/ClassName are captured before destruction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The warning and mutation status are correctly front-loaded, and the destroy semantics are explained efficiently. However, the trailing meta-block (Phase/cost/idempotency/Requires/Produces/On failure: inspect tool-schema...) is boilerplate, and the irreversibility warning is stated twice, so a few sentences do not fully earn their place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description supplies the return contract ({ Destroyed, ClassName, ok } or { error }), the mutation and replication semantics, required preconditions, and a follow-up verification tool. Nothing an agent needs in order to invoke this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters thoroughly (including the Luau-expression evaluation and the descendant caveat). The description adds a signature recap and the reminder to resolve the EXACT instance, which is a small amount of reinforcement rather than new semantic content.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (destroy/remove) and resource (a live Instance), explains the mechanism (inst:Destroy()), and distinguishes itself from the sibling clone-instance by naming it as the alternative. An agent can separate this from set-instance-property or create-instance without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use conditions are given: 'Only destroy instances you are certain about', 'use clone-instance first if you might need a copy', plus named prerequisites (active-client, resolved-target, explicit-mutation-approval) and a verification step (verify-path-exists). The condition selecting the alternative tool is stated outright.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

diff-instance-snapshotSnapshot an instance subtree and diff what an action changedA
Read-onlyIdempotent

Answer 'what did clicking this button / firing this remote / opening this menu actually change in the game tree?' by snapshotting an instance subtree, performing the action in-game, then diffing the before/after. Workflow: (1) call action='snapshot' on a root (default game.Workspace) to capture a baseline; (2) do the thing in-game (click, fire a remote, walk somewhere, wait for a wave to spawn); (3) call action='compare' with the SAME name to see exactly which instances were ADDED, REMOVED, or CHANGED. How it works: it walks root:GetDescendants() (capped at maxInstances) and builds a per-instance signature from a handful of cheap properties (ClassName, Name, Parent, and whichever of Position/Transparency/Value/Text/Visible/Anchored/Health/Enabled exist). 'changed' entries report the before and after signatures so you can see precisely which property moved (a part teleported, a value bumped, a label retexted, a GUI shown). State is stored CLIENT-SIDE in getgenv().__mcp_snapshots[name], so baselines persist across tool calls and are naturally isolated per game / per session. Use distinct names to keep several independent baselines. snapshot returns: { action='snapshot', name, root, captured, truncated }. compare returns: { action='compare', name, root, counts={ added, removed, changed }, added=[paths], removed=[paths], changed=[{ path, before, after }], truncated=bool }. Non-mutating and fully pcall-guarded (locked/destroyed instances are skipped, never aborting the scan). Each detail list is capped at 200 entries; the exact counts are always reported. Signature: { action: "snapshot" | "compare", root: any?, name: any?, maxInstances: any?, threadContext: number? }. Phase: verify; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-result. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoSlot/name for the snapshot (default 'default'). compare diffs against the baseline stored under this same name. Use distinct names to keep several independent before/after pairs around at once.default
rootNoLuau expression for the root instance of the subtree to snapshot/diff (default 'game.Workspace'). Examples: 'game.Workspace', 'game.Players.LocalPlayer.PlayerGui', 'game.Workspace:FindFirstChild("Map")'. Evaluated as `return <root>`. Pick the smallest subtree that contains what you expect to change — a tighter root means a faster, less noisy diff.game.Workspace
actionYes'snapshot' captures a baseline signature map of the subtree under `name`. 'compare' re-walks the subtree now and diffs it against the stored baseline of that `name`, reporting added/removed/changed instances.
maxInstancesNoMax descendants to walk per pass (default 4000). The walk stops at this many and flags `truncated`; a truncated baseline plus a truncated compare can produce spurious added/removed near the cap, so prefer a tighter root over a huge cap.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, and the description adds substantial non-obvious behavior: state lives client-side in getgenv().__mcp_snapshots[name] and persists across calls, per-game/per-session isolation, the maxInstances cap and its truncation risk, pcall-guarding that skips locked/destroyed instances, and the 200-entry detail cap with exact counts always reported.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the driving question, then the numbered workflow, then mechanics, then return shapes; nearly every sentence carries information. Minor redundancy appears in the tail ('Non-mutating', 'idempotency=read-only', 'Safety: read-only' restate the same trait) and the signature/phase line duplicates schema content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates fully by spelling out both return payloads (snapshot's {action,name,root,captured,truncated} and compare's counts/added/removed/changed shapes). For a 5-param verify-phase tool, an agent has everything needed to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description contributes extra meaning: it clarifies that 'name' is the pairing key shared by snapshot and compare, that 'root' is an evaluated Luau expression with examples, and that a tighter root yields a faster/less noisy diff and mitigates truncation near maxInstances.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('snapshotting an instance subtree, performing the action in-game, then diffing') framed around a concrete question the agent would ask. It is clearly separable from siblings like compare-instances, watch-value, and get-instance-tree because it pairs a pre/post snapshot around an in-game action rather than a live subscription or a static tree dump.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit three-step workflow (snapshot → do the thing in-game → compare with the SAME name) and advises picking the smallest root and distinct names for independent baselines. It stops short of naming alternative sibling tools or stating when this approach is the wrong choice, so it is clear context but not full when/when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

disassemble-functionDisassemble function (IDA-style function view)A
Read-onlyIdempotent

Produce a full IDA-style dump of a single Luau function in one call: debug.info (name/source/line/params/ups), whether it is a Lua or C closure (islclosure/iscclosure), constant/upvalue/proto counts, the executor function hash (getfunctionhash) for duplicate-detection, and — optionally — the full constant and upvalue tables with their typeof and a string value (Instances become GetFullName, functions become tostring). Complements inspect-closure by bundling the hash and the structural counts into one reverse-engineering-focused view. Resolve the target via a Luau expression that yields a function; each list is capped at 200 entries. Signature: { functionPath: string, includeConstants: any?, includeUpvalues: any?, threadContext: number? }. Phase: observe; cost=high; idempotency=read-only. Requires: active-client, resolved-target. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
functionPathYesLuau expression that resolves to the function to disassemble (e.g. "getgenv().myFunc" or "require(game.ReplicatedStorage.Mod).start"). Evaluated as `return <expr>`; must yield a function.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
includeUpvaluesNoInclude the full upvalues table (capped at 200) with Index/Type/Value. Default true.
includeConstantsNoInclude the full constants table (capped at 200) with Index/Type/Value. Default true.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only/idempotent/non-destructive, so the bar is lower, yet the description adds real context beyond them: each list is capped at 200 entries, it requires an active client and a resolved target, cost is high, and it returns a structured-observation. It does not describe pagination or truncation behavior when caps are hit, which is the main remaining gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded and dense with information, but the back half is a metadata block (Phase/cost/idempotency/Requires/Produces/Safety) that partly duplicates annotations ("idempotency=read-only", "Safety: read-only") and the schema's own signature. The generic "On failure: inspect tool-schema" tail is boilerplate that doesn't earn its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-shape burden and does so reasonably: it enumerates the debug.info fields, closure kind, counts, hash, and optional tables, plus prerequisites and caps. It does not sketch the actual response envelope or how errors surface, which a high-complexity tool could still benefit from.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds meaning the schema lacks: values in the constant/upvalue tables are coerced for display (Instances become GetFullName, functions become tostring), and it restates the call signature. It still doesn't clarify threadContext semantics beyond the schema's own text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("full IDA-style dump of a single Luau function in one call") and enumerates exactly what is bundled: debug.info, closure kind, counts, function hash, and optional constant/upvalue tables. It explicitly differentiates from siblings by naming inspect-closure and positioning itself as the bundling alternative to get-function-hash/get-closure-constants.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for use (reverse-engineering-focused view, resolve target via a Luau expression that yields a function) and names inspect-closure as the complementary tool. However, "complements inspect-closure" is not an explicit when-to-prefer-this-vs-that rule, and no exclusions or preconditions for choosing between the two are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

discover-characterDiscover a standard or custom character modelA
Read-onlyIdempotent

Find the active player's character even when Players.LocalPlayer.Character, Humanoid, or HumanoidRootPart is missing or custom-named. Checks the normal hierarchy first, then performs a bounded Workspace model search for Humanoids and likely root parts. Use this after get-local-player-info reports a missing character; consume the returned resolved paths instead of assuming game.Players.LocalPlayer.Character.HumanoidRootPart. Signature: { playerName: string?, scanWorkspace: any?, limit: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum Workspace instances examined during fallback search.
playerNameNoPlayer name to inspect; defaults to Players.LocalPlayer.
scanWorkspaceNoSearch Workspace for custom models when the standard character is missing.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description goes beyond that by disclosing the fallback strategy ('checks the normal hierarchy first, then performs a bounded Workspace model search') and the bounded nature of the search, which is genuinely useful operationally.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core guidance is front-loaded in the first two sentences, which is where the value sits. The trailing metadata block (phase, cost, idempotency, safety) partly duplicates the annotations, adding some redundancy without derailing readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by noting that resolved paths are returned and that failure routing points to tool-schema for exact fields. It covers the essentials for a read-only discovery call, though the exact shape of the returned candidates is not described.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so limit, playerName, scanWorkspace, and threadContext are already documented. The description's signature line restates parameter names/types without adding semantic detail such as bounds or defaults that the schema lacks, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Find the active player's character') and precisely scopes the edge cases it handles: missing Humanoid/HumanoidRootPart or custom-named models. An agent can immediately distinguish it from get-local-player-info and generic instance lookups.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Use this after get-local-player-info reports a missing character' and tells the agent to consume the returned resolved paths rather than assuming the conventional path. The trigger condition and the sibling it follows are spelled out, so nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

discover-player-valuesDiscover Candidate Money / Score / Level Value PathsA
Read-onlyIdempotent

Auto-discovery for 'where's the money/score/XP path?'. Walks LocalPlayer (esp. leaderstats), PlayerGui, ReplicatedStorage and ReplicatedFirst for IntValue/NumberValue/StringValue/BoolValue/Folder-of-values, scores each by name keywords (money/coin/cash/gold/score/xp/level/exp/kills/wins/...), container weight (leaderstats >> everything else), kind (numeric > string > bool), and magnitude. Returns a single ranked list of { path, class, name, value, score, reasons[] } so the AI doesn't have to grind through the tree. Pure read; no remotes fired, no state mutated. Use this as the first probe on any unfamiliar game. Signature: { limit: number?, minScore: number?, extraRoots: {string}? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum candidates to return, ranked by score (default 50).
minScoreNoDrop candidates below this score; default 0 returns everything.
extraRootsNoAdditional Luau path expressions to scan, e.g. ['game.Workspace.GameValues']. Defaults already cover LocalPlayer/leaderstats/PlayerGui/ReplicatedStorage/ReplicatedFirst.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/idempotentHint/destructiveHint=false, but the description adds real value: it discloses the scan scope (LocalPlayer/leaderstats/PlayerGui/ReplicatedStorage/ReplicatedFirst), that no remotes are fired, the ranking heuristics, and the precondition "Requires: active-client". Minor redundancy between "Pure read", "Safety: read-only", and the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then scoring, then usage and constraints — every major sentence earns its place. Slightly bloated by the trailing meta-block (Phase/cost/idempotency/Produces/Safety) which partially duplicates the annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description supplies the return shape ({ path, class, name, value, score, reasons[] }), the failure path (inspect tool-schema), and a bounded-result contract, leaving nothing an agent needs to invoke it correctly unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description explains the scoring dimensions (name keywords, container weight, kind, magnitude), which gives concrete meaning to what minScore filters on and why leaderstats candidates rank highest. extraRoots is only listed in the signature block, not further explained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: auto-discovering money/score/XP value paths, with exact scan roots and a keyword-based scoring method. An agent can distinguish it from siblings like find-string-in-tables, search-gc-value, scan-number-range, and read-path-value without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Use this as the first probe on any unfamiliar game" gives a clear when-to-use trigger and implies manual tree traversal as the alternative. However, no sibling tools are named as alternatives, so the routing is contextual rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

draw-clearRemove ALL Drawing objects (wipe the overlay)A
Destructive

Tears down the entire MCP-managed Drawing overlay: iterates every handle stored in getgenv().__mcp_drawings, calls :Remove() on each, empties the registry, and additionally calls Drawing.clear() when the executor exposes it (to sweep any stray objects this tool did not create). Use it as a one-shot reset between ESP sessions. Requires the Drawing table (type-guarded); each :Remove() and the optional Drawing.clear() are pcall-guarded so a single bad handle never aborts the wipe. On an executor without it, it returns { error = "Drawing is not available in this executor." }. Returns { cleared, failed, drawingClearCalled } or { error }, where 'cleared' is the count of registry handles successfully removed. Signature: { threadContext: number?, timeoutMs: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval. Capabilities: Drawing. Produces: operation-receipt. Verify with: list-drawings. Safety: MUTATING; changes executor drawing overlay state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations (destructiveHint/readOnlyHint/idempotentHint) by disclosing the pcall-guarding so a bad handle won't abort the wipe, the exact error payload when Drawing is unavailable, the return shape ({cleared, failed, drawingClearCalled}), and the active-client plus explicit-mutation-approval prerequisites. This is rich operational context an agent cannot get from structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core behavior is front-loaded and dense but mostly earns its place, covering mechanism, guards, and return values. The trailing metadata block (Phase/cost/Requires/Produces/Verify/Safety) and the closing 'inspect tool-schema for exact fields' line are somewhat boilerplate, keeping it just short of a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description fully specifies return values in both success and failure cases, the executor prerequisite, and the mutation approval requirement. For a destructive overlay-wipe tool whose safety profile is already in annotations, nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both optional parameters (threadContext, timeoutMs) are fully documented in the schema. The description's 'Signature: { threadContext, timeoutMs }' adds nothing beyond what the schema already provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb and resource ('Tears down the entire MCP-managed Drawing overlay') and details the exact mechanism (iterating __mcp_drawings handles, calling :Remove(), emptying the registry, optional Drawing.clear()). This clearly distinguishes it from the single-object sibling draw-remove and the read-only list-drawings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit context: 'Use it as a one-shot reset between ESP sessions,' and names list-drawings as the verification step. It does not explicitly name draw-remove as the alternative for removing individual objects, so the when-not guidance is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

draw-createCreate a Drawing object (ESP/debug overlay)A
Destructive

Creates a new on-screen overlay object via the executor Drawing library (used for ESP boxes, tracers, name tags, and debug HUDs) and registers it so it survives across tool calls. Calls Drawing.new(type) for type in {Line, Text, Circle, Square, Quad, Triangle, Image}, applies the given properties, stores the handle in getgenv().__mcp_drawings under a new integer id, and returns { id, type } — pass that id to draw-update / draw-remove. PROPERTIES: each property value is a Luau EXPRESSION STRING evaluated via loadstring, so rich types work: Color = 'Color3.new(1,0,0)', From = 'Vector2.new(10,10)', To = 'Vector2.new(200,200)', Position = 'Vector2.new(100,100)', Visible = 'true', Thickness = '2', Text = '"hello"', Size = '18'. Common props: Visible(bool), Color(Color3), Transparency(number 0..1), ZIndex(number). Per type — Line: From, To, Thickness; Text: Text, Size, Position, Center, Outline; Circle: Center, Radius, NumSides, Thickness, Filled; Square: Size, Position, Thickness, Filled. Requires the Drawing table (type-guarded); on an executor without it, it returns { error = "Drawing is not available in this executor." }. Every call is pcall-guarded. Returns { id, type, applied[] } or { error }. Signature: { type: "Line" | "Text" | "Circle" | "Square" | "Quad" | "Triangle" | "Image", properties: any?, threadContext: number?, timeoutMs: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval. Capabilities: Drawing. Produces: created-handle. Verify with: list-drawings. Safety: MUTATING; changes executor drawing overlay state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
typeYesThe Drawing object type to create via Drawing.new(type): one of Line, Text, Circle, Square, Quad, Triangle, Image.
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
propertiesNoMap of property name -> Luau expression STRING to assign after creation. Values are evaluated via loadstring, so use Luau syntax: Color = 'Color3.new(1,0,0)', From = 'Vector2.new(10,10)', Visible = 'true', Thickness = '2', Text = '"label"'. Unknown/failed properties are reported in 'applied' with ok=false but do not abort the create.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the mutation profile (destructiveHint=true, idempotentHint=false), and the description goes well beyond them: type-guard on the `Drawing` table, the exact error payload when unavailable, pcall-guarding on every call, the fact that handles persist across tool calls, and that failed properties appear in `applied` with ok=false without aborting the create.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening sentence is well front-loaded and dense with useful detail, but the body is a very long single block that re-lists the full signature already present in the schema and appends machine metadata (Phase, cost, Capabilities, Produces, Safety) as one run-on string. Several clauses restate structured data rather than adding new information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return contract itself ({ id, type, applied[] } vs { error }) and covers failure modes, persistence, and executor availability. For a mutating 4-parameter tool this is complete enough that an agent could invoke it correctly without opening the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds genuine meaning the schema lacks: the loadstring/Luau-expression-string convention for `properties` and per-type supported properties (Line: From/To/Thickness; Text: Text/Size/Position/Center/Outline; Circle: Center/Radius/NumSides/Thickness/Filled). It adds value beyond the schema without fully enumerating Quad/Triangle/Image props.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('creates a new on-screen overlay object via the executor `Drawing` library') and enumerates exactly what it produces (handle stored in getgenv().__mcp_drawings, returns { id, type }). It is unmistakably distinct from the sibling draw-update / draw-remove / draw-clear / list-drawings tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent downstream ('pass that id to draw-update / draw-remove') and names the verification step ('Verify with: list-drawings'), plus states prerequisites (active-client, explicit-mutation-approval). It does not describe when not to use it or when a non-Drawing overlay approach would be preferable, so it stops short of full 5-level guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

draw-removeRemove a single Drawing object by idA
Destructive

Destroys one Drawing overlay created by draw-create: looks the handle up by its integer id in getgenv().__mcp_drawings, calls handle:Remove(), and clears the registry slot so list-drawings no longer reports it. Use it to clean up a single ESP element while leaving the rest of the overlay intact. Requires the Drawing table (type-guarded); the :Remove() call is pcall-guarded. If the id is unknown it returns { removed = false, error }; on an executor without it, it returns { error = "Drawing is not available in this executor." }. Returns { id, removed } or { error }. Signature: { id: number, threadContext: number?, timeoutMs: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval. Capabilities: Drawing. Produces: structured-result. Verify with: list-drawings. Safety: MUTATING; changes executor drawing overlay state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesInteger id of the Drawing object to remove (as returned by draw-create / list-drawings).
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond annotations by disclosing the registry lookup, pcall-guarded :Remove() call, type guard on the Drawing table, the clearing of the registry slot, and distinct return shapes for unknown id and missing-executor cases. Annotations only cover the safety profile, so this adds substantial value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and use case, then supporting detail. It carries some metadata boilerplate (phase, cost, capabilities, produces, verify-with) that partly duplicates structured fields, keeping it from being maximally tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully covers return values ({ id, removed } / { removed = false, error } / { error }) and prerequisites (Drawing availability, approval), leaving no gap for a mutating tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all three parameters are already documented. The description restates the signature (id required, threadContext/timeoutMs optional) but adds no new semantics beyond the schema; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Destroys/Remove) and resource (one Drawing overlay created by draw-create), and clarifies the exact mechanism. It is clearly distinguishable from draw-clear (all) and list-drawings (read).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to use it to clean up a single ESP element while leaving the rest of the overlay intact, which implies the contrast with draw-clear. It gives clear context, though it does not name the sibling alternative outright.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

draw-updateUpdate properties of an existing Drawing objectA
Destructive

Mutates a Drawing overlay previously created by draw-create, looking the handle up by its integer id in getgenv().__mcp_drawings and assigning each given property. Use it to move a tracer (To = 'Vector2.new(400,300)'), recolor an ESP box (Color = 'Color3.new(0,1,0)'), toggle visibility (Visible = 'false'), or change a label (Text = '"BOSS"'). PROPERTIES: each value is a Luau EXPRESSION STRING evaluated via loadstring, identical to draw-create. Unknown/failed properties are reported per-property in 'updated' (ok=false) without aborting the rest. Requires the Drawing table (type-guarded). If the id is unknown it returns { error }; on an executor without it, it returns { error = "Drawing is not available in this executor." }. Returns { id, updated[] } or { error }. Signature: { id: number, properties: {[string]: any}, threadContext: number?, timeoutMs: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval. Capabilities: Drawing. Produces: structured-result. Verify with: list-drawings. Safety: MUTATING; changes executor drawing overlay state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesInteger id of the Drawing object (as returned by draw-create / list-drawings).
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
propertiesYesMap of property name -> Luau expression STRING to assign. Values are evaluated via loadstring, so use Luau syntax: Color = 'Color3.new(0,1,0)', To = 'Vector2.new(400,300)', Visible = 'false', Text = '"label"'. Failed properties are reported in 'updated' with ok=false.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructive/readOnly/idempotent hints, and the description adds substantial context on top: per-property failure isolation (ok=false without aborting), the exact { error } shapes for unknown id and missing Drawing executor, the loadstring evaluation model, and the phase/cost/idempotency classification. This is well beyond what structured fields provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose and lookup mechanism are front-loaded and the examples are efficient. However the tail is loaded with meta-boilerplate ('Phase: act; cost=medium; idempotency=contextual-write...') and a generic 'inspect tool-schema' instruction that do not earn their space in a description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by specifying the return shape ({ id, updated[] } or { error }) and the failure modes. For a nested-object mutation tool with a type-guarded dependency, the description covers everything an agent needs to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3; the description nonetheless adds real meaning by stressing that property values are Luau EXPRESSION STRINGS evaluated via loadstring and by listing worked examples (Color = 'Color3.new(0,1,0)'). It also restates the signature, which is minor redundancy but not harmful.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Mutates a Drawing overlay previously created by draw-create') and names the lookup mechanism (getgenv().__mcp_drawings by integer id). The sibling contrast with draw-create is explicit, so an agent can distinguish this from create/list/remove drawing tools without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete usage contexts — moving a tracer, recoloring an ESP box, toggling visibility, changing a label — and names prerequisites (must be created by draw-create) and a verification sibling (list-drawings). It never states when NOT to use it or which alternative to pick for bulk/instance property changes, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dump-function-envDump function envB
Read-onlyIdempotent

Find function by query and inspect environment table keys/values. Signature: { query: string, limit: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoOptional hard result/work budget used to keep output and runtime bounded.
queryYessearch text used to filter and rank bounded results.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so safety is covered; the description adds genuinely new context: 'Requires: active-client', 'cost=medium', and 'Produces: structured-observation'. It does repeat the read-only/idempotent profile ('idempotency=read-only', 'Safety: read-only'), which is redundant but not contradictory.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded in the first sentence, which is good, but the remaining metadata block wastes space on duplicated facts ('idempotency=read-only' and 'Safety: read-only' say the same thing, both echoing the annotations). Field/value pairs are terse but the redundancy costs it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and 'Produces: structured-observation' is only a vague gesture at return shape. Still, for a read-only, annotation-covered tool the description supplies the key operational facts (precondition, cost, phase), so an agent has enough to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters with defaults and constraints. The description's restated signature ('limit: any?') is actually less precise than the schema's 'number' typing and adds no new semantics. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb pair and resource: 'Find function by query and inspect environment table keys/values.' An agent can tell this reads a function's environment table. However, it never distinguishes itself from the close sibling 'get-function-env', so the boundary is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use / when-not-to-use guidance or named alternative. 'Phase: observe' hints at workflow position but does not tell the agent when this beats get-function-env, get-script-env, or dump-table. The fallback note is about failure handling, not selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dump-tableRecursively dump a Luau tableA
Read-onlyIdempotent

Resolve a Luau expression to a TABLE and recursively encode its contents to a chosen depth. Scalar, Instance and function values are encoded via the shared encoder; nested tables are recursed into until maxDepth, after which they collapse to 'table: (truncated)'. Each level is capped at maxKeys (the cap is noted via a per-level Truncated flag), and cycles are detected so self-referential tables won't loop forever. Ideal for reading config tables, getgenv()/getrenv() subtables, a ModuleScript's return value, or any captured upvalue table you already have a handle on. WARNING: this ACTS ON THE LIVE GAME — it evaluates your expression in the running client, which may trigger __index metamethods or other side effects while iterating. Returns { Target, Depth, Table:, Truncated } or { error }. Signature: { tablePath: string, maxDepth: any?, maxKeys: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
maxKeysNoMaximum number of keys to encode PER table level (default 100, max 500). When a level has more keys than this, the extras are dropped and that level is flagged Truncated = true.
maxDepthNoHow many levels of nested tables to recurse into (default 2, max 5). Beyond this depth, nested tables are rendered as 'table: <addr> (truncated)' rather than expanded.
tablePathYesLuau expression resolving to a table, e.g. 'getgenv()', 'getrenv()._G', 'require(game.ReplicatedStorage.Config)', 'getrawmetatable(game)', or 'debug.getupvalues(someFn)[1]'. Evaluated as `return <tablePath>` and must yield a table.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, but the description adds material context: cycle detection, per-level truncation flagging, maxDepth collapse behavior, and a WARNING that live evaluation may trigger __index metamethods or side effects. It also states the return shape and error form. This is well beyond what the annotations convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded and the body is informative, but the tail (Phase/cost/idempotency/Requires/Produces/Safety) partly duplicates the annotations (idempotency=read-only vs idempotentHint=true) and the closing 'inspect tool-schema...' line is filler. Mixed redundancy keeps it out of the top band.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a moderately complex read tool with no output schema, the description covers the evaluation model, recursion/truncation/cycle rules, side-effect risk, and the exact return shape ({ Target, Depth, Table, Truncated } or { error }). Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so maxKeys, maxDepth, tablePath and threadContext are already documented. The description's Signature line and truncation notes re-state the semantics rather than adding new meaning beyond the schema. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Resolve a Luau expression to a TABLE and recursively encode its contents to a chosen depth.' The recursion, depth-collapse, key capping and cycle detection clearly separate it from table-scanning siblings like find-tables-by-key, list-gc-tables, and find-table-references. An agent can identify exactly what this tool does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete use cases: 'reading config tables, getgenv()/getrenv() subtables, a ModuleScript's return value, or any captured upvalue table.' This tells the agent when the tool is a fit. However it names no alternative tool or exclusion condition, so the when-not side is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ensure-remote-spyStart the selected remote spyA
Destructive

WRITES LIVE GAME STATE. Starts the selected engine (default cobalt), stopping the other spy on this client first. Cobalt is bundled; Ketamine is downloaded from a pinned upstream commit and hash-checked before loading. Ketamine requires hookfunction, hookmetamethod, getnamecallmethod, getcallbackvalue, setfenv, and crypt.hash. Idempotently attaches one capture observer. RakNet is Cobalt-only. max bounds the MCP buffer, separate from GUI history. Use remote-spy operation=restart for mode changes. Signature: { engine: "cobalt" | "ketamine"?, mode: any?, max: number?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval. Produces: bounded-event-snapshot. Verify with: assert-state. Safety: MUTATING; changes persistent executor-side observer or hook state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
maxNoMCP buffer capacity, clamped to 10..5000. Omit to retain the current capacity (500 on first start).
modeNoCobalt auto prefers supported RakNet; luau selects standard hooks. Ketamine accepts auto/luau only. Changing an active mode requires restart.auto
engineNoRemote-spy backend. Only one engine runs per client; starting another stops the current engine. Defaults to cobalt.cobalt
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations (destructiveHint=true, readOnlyHint=false, idempotentHint=false) by disclosing engine-specific requirements (Ketamine needs hookfunction, crypt.hash, etc.), that RakNet is Cobalt-only, that Ketamine is hash-checked on download, the one-observer idempotent attach, and that state is persistent executor-side. 'Safety: MUTATING' plus 'changes persistent observer/hook state' is exactly the mutation context an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the mutation warning and the core action, and most sentences carry real content. The trailing metadata block (Phase, cost, idempotency, Requires, Produces, Verify, Safety) is dense but somewhat boilerplate-heavy relative to the rest.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-param, no-output-schema mutation tool, the description supplies return summary ('Produces: bounded-event-snapshot'), verification path (assert-state), failure guidance (inspect tool-schema), and all engine constraints. Nothing needed to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already documents max, mode, engine, and threadContext thoroughly, so the description's inline signature is largely redundant. The only marginal addition is 'max bounds the MCP buffer, separate from GUI history', which nuances the schema's capacity note. Baseline 3 fits.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Leads with a specific verb+resource ('Starts the selected engine'), names the default (cobalt), and states the side effect of stopping the other spy. This distinguishes it cleanly from remote-spy, configure-remote-spy, and get-remote-spy-logs without needing their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes mode changes to 'remote-spy operation=restart' and states prerequisites (active-client, explicit-mutation-approval) plus a verification step (assert-state). It does not spell out when NOT to start a spy at all, but the alternative for the main confusion case is named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

eval-expressionEvaluate a single Luau expressionA
Read-onlyIdempotent

Evaluate ONE Luau expression on the active client and return its typeof plus an encoded value. Convenience wrapper over run-luau for quick reads like 'workspace.Gravity', '#game.Players:GetPlayers()', or 'game.PlaceId'. The expression is pcall-guarded, so a runtime error is reported as { ok = false, error } instead of failing the call. Pass an expression, not statements (no 'return', no ';'). Note: this is PURE Luau — to compose with other server tools (e.g. mcp.getPlayers(), mcp.searchInstances({...})) use the script tool instead. Signature: { expression: string }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-result. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
expressionYesA single Luau expression to evaluate (e.g. 'workspace.Gravity', 'game.PlaceId').

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover safety (readOnly/destructive/idempotent/openWorld), but the description adds substantial context beyond them: pcall guarding, the exact failure representation ({ ok = false, error }), purity (PURE Luau), the active-client requirement, and a stated cost/phase. This is materially more than the annotations convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and a concrete example before the constraints. It trail-loads some metadata (Phase/cost/idempotency/Safety) that partly duplicates the annotations, but it is well organized and most sentences earn their place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden and does so: typeof plus encoded value, and the error envelope on failure. It also names the active-client prerequisite and the fallback tool, so nothing needed to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents the single `expression` param and the baseline would be 3. The description adds real semantic value the schema lacks: it must be an expression, not statements, and no 'return' or ';' are permitted. That constraint meaningfully raises it above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Evaluate) and resource (ONE Luau expression) with the return shape ('typeof plus an encoded value'). It explicitly positions itself against siblings, calling out run-luau and the `script` tool by name, so an agent can distinguish it without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('quick reads like workspace.Gravity'), when-not with alternatives ('to compose with other server tools ... use the `script` tool instead'), and a hard input constraint ('pass an expression, not statements'). This is the full when/when-not/alternatives triad.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

executeExecute Code in the Roblox Game ClientA
Destructive

Execute Luau in the active Roblox client WITHOUT waiting for it to finish. The code is COMPILED FIRST via loadstring (a syntax error is returned cleanly as { error } and nothing runs), then handed to task.spawn so it runs on its own thread; this tool returns { scheduled = true } the moment the thread is started — it does NOT wait for completion and does NOT return the code's output, return value, or runtime errors. Use this for fire-and-forget side effects. When you need the value(s) your code produces, use run-luau or execute-and-wait instead. Requires loadstring and the task library (both guarded). Returns { scheduled = true } or { error }. Signature: { code: string, client: string?, agent: string?, threadContext: number?, timeoutMs: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval, validated-source. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; executes caller-selected behavior in the live client. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYesThe Luau code to execute in the Roblox client. Compiled with loadstring, then spawned on its own thread. This tool does NOT return output — use run-luau or execute-and-wait if you need data back.
agentNoOptional. A stable label for WHICH agent is calling when several share this MCP session (e.g. 'researcher'). Gives that agent its own fair scheduling lane, its own persistent VM on each game, and its own queue budget, so co-tenant agents don't starve or clobber each other.
clientNoOptional. Run on a specific connected client — its clientId OR username — for THIS call only, overriding your session's select-client binding without changing it. Lets multiple agents drive different games at the same time; omit to use your session's selected client.
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare a mutating, non-idempotent, destructive write, and the description still adds real context beyond them: compilation via loadstring before any run, syntax errors returned as { error } with nothing executed, execution on a task.spawn thread, the exact return shape { scheduled = true }, and the loadstring/task library requirement. It also flags the mutating nature and required approvals.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core behavior distinction and every sentence carries information, but the trailing metadata block (Phase/cost/idempotency/Requires/Produces/Verify/Safety) is dense boilerplate that partly duplicates the annotations and LOADSTRING info already stated earlier.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description fully specifies the return contract ({ scheduled = true } or { error }), the failure path, and where to look for exact fields. For a single-required-param mutating tool, nothing needed to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already carries parameter meaning (agent/client/timeoutMs/threadContext are all documented in-schema). The description's inline signature restates the field names but adds no semantics beyond what the schema provides, so the baseline 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource+scope ('Execute Luau in the active Roblox client WITHOUT waiting') and explicitly differentiates itself from siblings run-luau and execute-and-wait by the fire-and-forget behavior. An agent can pick it apart from the alternatives without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use ('Use this for fire-and-forget side effects') and when-not ('When you need the value(s) your code produces, use run-luau or execute-and-wait instead'), naming both alternatives and the selecting condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

execute-and-waitExecute code and wait for the resultA
Destructive

Run Luau in the active Roblox client and WAIT for what happened, returning a structured result: { ok, returnValue, output, error? }. Unlike the fire-and-forget 'execute' tool, this reports success/failure, any error message, the FIRST value your code returns (encoded so Instances/Vector3/etc. survive), and the print()/warn() output it emitted. Output capture connects game:GetService("LogService").MessageOut to a buffer for the duration of the run, then disconnects — so it sees logs even when an executor routes print() to the Roblox console rather than swapping the global. The code is COMPILED FIRST via loadstring (a syntax error is reported as { ok = false, error } and nothing runs), then executed under pcall; a runtime error is reported in error, never as a tool failure. Use this for quick experiments, debugging, or calling a function and inspecting what it gives back in one round trip. Signature: { code: string, client: string?, agent: string?, threadContext: number?, timeoutMs: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval, validated-source. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; executes caller-selected behavior in the live client. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYesThe Luau code to run. It may 'return' a value (the FIRST return comes back in `returnValue`, encoded) and may print()/warn() (captured into `output`). Do not JSON-encode anything yourself.
agentNoOptional. A stable label for WHICH agent is calling when several share this MCP session (e.g. 'researcher'). Gives that agent its own fair scheduling lane, its own persistent VM on each game, and its own queue budget, so co-tenant agents don't starve or clobber each other.
clientNoOptional. Run on a specific connected client — its clientId OR username — for THIS call only, overriding your session's select-client binding without changing it. Lets multiple agents drive different games at the same time; omit to use your session's selected client.
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations (destructive/idempotent hints): it discloses the output-capture mechanism (LogService.MessageOut buffered then disconnected), compile-first via loadstring with syntax errors returned as { ok=false, error }, execution under pcall, and the crucial contract that runtime errors are reported in `error` and never as a tool failure. That is exactly the behavioral context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core behavior and error contract efficiently, with each sentence doing work (capture mechanism, compile-then-pcall, sibling distinction). It is dense and slightly long, and the trailing 'Phase/cost/Requires/Produces/Verify' metadata block plus the pointer to tool-schema is boilerplate padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating execution tool with no output schema, the description fully specifies the return object ({ ok, returnValue, output, error? }), the error semantics, the safety posture, and required preconditions (active-client, explicit-mutation-approval). Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the per-parameter descriptions (especially agent and client) are already rich, so the baseline is 3. The description's 'Signature: { code, client, agent, threadContext, timeoutMs }' restates parameter names without adding meaning beyond the schema, so it does not exceed the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Run Luau in the active Roblox client') plus the key scope distinction ('WAIT for what happened') and the exact return shape. It explicitly names the sibling it is not ('fire-and-forget execute'), so an agent can differentiate without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the alternative ('execute') and the condition that selects this one (fire-and-forget vs. reporting result). Adds concrete scenarios: 'quick experiments, debugging, or calling a function and inspecting what it gives back in one round trip.' Nothing left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

execute-fileExecute a Luau file on the active clientA
Destructive

Read a Luau script from the SERVER host filesystem and execute its contents on the active Roblox client, returning the first value the script returns (decoded automatically — return what you want back). The path is read through an allow-list sandbox: only files inside the configured roots (the server's working directory, ~/Documents, and any extra script directories set in the config) can be read, and symlinks that escape those roots are rejected. A path outside the allow-list, or a missing file, returns an error without running anything. Use run-luau when you already have the source inline. Signature: { path: string, threadContext: number?, timeoutMs: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; writes executor workspace filesystem. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesAbsolute or relative path to a .lua/.luau file inside an allow-listed root.
timeoutMsNoPer-call deadline in milliseconds. Server default if omitted.
threadContextNoRoblox thread identity to run under (e.g. 2 = game scripts, 8 = elevated). Server default if omitted.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: describes the allow-list sandbox (configured roots, symlink escape rejection), that out-of-allow-list or missing files error without running anything, that the return value is auto-decoded, and the MUTATING safety profile and operation-receipt output. This enriches the destructiveHint=true annotation rather than merely restating it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core read-execute-return behavior, then constraint and routing details. Some material is boilerplate envelope (Phase/cost/idempotency/Signature) that partly duplicates structured fields, but nothing is truly wasted and the ordering serves the agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description explains the return value (first decoded return value) and the error path. Combined with requires/produces/verify and safety notes, an agent has enough to call and verify this mutating executor correctly, so nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already defines path, timeoutMs, and threadContext semantics; the description's signature listing is redundant. It does add meaningful constraint for `path` via the allow-list root/symlink rules and notes 'return what you want back' for the return value, so it slightly exceeds the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States specific verbs and resources: read a Luau file from the SERVER host filesystem and execute its contents on the active Roblox client, returning the first decoded value. It explicitly contrasts with the sibling run-luau by describing the file-vs-inline distinction, so an agent can distinguish it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes between alternatives: 'Use run-luau when you already have the source inline,' and this tool when the source is a file. It also names prerequisites (active-client, resolved-target, explicit-mutation-approval) and the verify step (assert-state), giving clear when-to-use context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

execute-lua-stateSchedule source on a LuaStateProxyA
Destructive

WRITES LIVE GAME STATE. Resolve current/game/expression state and call LuaStateProxy:Execute(source, ...args). Execution is asynchronous and confirmation-gated. Signature: { state: any?, stateExpression: any?, source: string, arguments: any?, confirm: boolean?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval, validated-source. Capabilities: getluastate. Produces: operation-receipt. Verify with: get-lua-state. Safety: MUTATING; executes caller-selected behavior in the live client. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
stateNoOptional state selector or reusable state reference returned by a discovery tool.current
sourceYesLuau source text processed by this tool; it is never inferred or guessed.
confirmNoExplicit safety acknowledgement; must be true before this state-changing operation runs.
argumentsNoOptional ordered typed arguments forwarded to the selected operation.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
stateExpressionNoFor state='expression', a Luau expression resolving to a LuaStateProxy, Actor, or BaseScript.

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well past the annotations by disclosing that execution is asynchronous and confirmation-gated, that it requires explicit mutation approval and an active client, that it produces an 'operation-receipt', and how to verify the result. Annotations only give safety hints; the description adds the approval gating and post-run verification path.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the mutating warning, then signature, preconditions, and verification in a compact tag format. Slightly redundant ('WRITES LIVE GAME STATE' vs 'Safety: MUTATING') and the 'On failure' line is generic boilerplate, but overall efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, but the description compensates by naming the produced artifact ('operation-receipt') and the verification tool. Preconditions, safety, and failure guidance are covered; only deeper result-format detail is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all six parameters, including defaults and enums. The description restates the signature but adds no semantics beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: resolves state and calls LuaStateProxy:Execute(source, ...args), with 'WRITES LIVE GAME STATE' front-loaded. However, it never distinguishes itself from the many sibling executors (execute, execute-and-wait, run-luau, eval-expression), so an agent can't route between them from the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear preconditions ('Requires: active-client, explicit-mutation-approval, validated-source') and names the verification tool ('Verify with: get-lua-state'). No explicit when-not or alternatives to sibling execution tools, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

execution-footprint-auditAudit script execution footprint and local detection exposureA
Read-onlyIdempotent

READ-ONLY, one-shot Luau auditor for local execution exposure. It resolves an optional script and/or closure without invoking it; inventories virtual-input globals and VirtualInputManager/VirtualUser references; distinguishes direct Instance references from cloneref-like alternate references when compareinstances exists; compares input functions with retained MCP closure handles; audits bounded getsenv/getfenv key-name leaks; classifies closure origin/hash/hook state; scans bounded source/constants for input, environment, debug, and hook indicators; and returns findings, unknown checks, confidence, risk score, privacy-safe evidence, and truncation telemetry. It never sends input, calls the target closure, walks getgc/descendants, installs hooks, or writes to game objects or script values. A clean result does NOT prove that server-side or external detection did not occur, and cloneref/clone matches are provenance only—not an undetectability guarantee. Signature: { scriptPath: string?, functionPath: string?, includeSourceScan: any?, includeStackEnvironments: any?, maxStackFrames: any?, maxEvidence: any?, maxEnvironmentKeys: any?, maxSourceChars: any?, maxConstants: any?, maxUpvalues: any?, threadContext: number? }. Phase: verify; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Produces: diagnostic-report. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
scriptPathNoOptional Luau expression resolving to the LocalScript or ModuleScript to audit.
maxEvidenceNoHard cap for findings and source/constant evidence records.
maxUpvaluesNoMaximum target-closure upvalues classified without returning captured values.
functionPathNoOptional Luau expression resolving to a function/closure to audit without invoking it.
maxConstantsNoMaximum target-closure constants inspected internally.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
maxSourceCharsNoMaximum script.Source characters scanned; the report says when this truncates.
maxStackFramesNoMaximum current auditor stack levels inspected with getfenv (0 disables stack probing).
includeSourceScanNoRead and scan a bounded script.Source prefix when the property is accessible.
maxEnvironmentKeysNoMaximum key names retained from each unique environment.
includeStackEnvironmentsNoInspect key names from a bounded number of the auditor's current getfenv stack frames.

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only/idempotent/non-destructive/non-openWorld, yet the description adds substantial non-structured behavior: what it will never do (never sends input, never invokes the closure, never walks getgc/descendants, never installs hooks, never writes), plus the critical caveat that a clean result does not prove no server-side/external detection and that cloneref/clone matches are provenance only, not undetectability guarantees.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The key identity ('READ-ONLY... auditor') is correctly front-loaded, but the core is one very long run-on sentence enumerating many operations, followed by caveats and metadata. It is information-dense rather than padded, yet it is hard to scan and could be split into a purpose line, a returns line, and a caveats line.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully enumerates the return payload (findings, unknown checks, confidence, risk score, privacy-safe evidence, truncation telemetry) and covers prerequisites, safety boundary, and failure handling. What remains thin is the distinction between the many limits and how to choose values, but overall it is complete enough for an agent to call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so each of the 11 parameters is already documented in the schema; baseline is 3. The description restates the signature and mentions bounding generally ('bounded getsenv/getfenv key-name leaks') but adds no semantics beyond what the schema descriptions already provide.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence names a specific verb+resource ('READ-ONLY, one-shot Luau auditor for local execution exposure') and then enumerates exactly what it does: resolves script/closure, inventories virtual-input globals, compares signatures, audits getsenv leaks, classifies closure origin, scans source. It also distinguishes itself from neighbors by ruling out hook installation and getgc/descendant walking, so an agent can separate it from scan-hook-surfaces or list-gc-functions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear invocation context ('Phase: verify; cost=medium; idempotency=read-only', 'Requires: active-client, resolved-target') and a failure path ('inspect tool-schema for exact fields'). It does not name an alternative sibling or state when-not to use it, so it stops short of explicit routing but the context is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

explain-failureExplain and recover from a failed tool callA
Read-onlyIdempotent

READ-ONLY AND NO CLIENT REQUIRED. Deterministically classify a failed MCP/Roblox tool call from its error, handled result, attempted input, and optional context. Returns a standard recovery envelope with cause, exact evidence, confidence, retry safety, a schema-validated correctedInput only when safely derivable, live-registry fallback tools ranked using AI contracts and discovery, an optional useful recovery script, and concrete next actions. Use this after any failed or blocked call. It never invokes a fallback and never recommends repeating an identical failed mutation. Signature: { toolName: string, error: any?, result: any?, attemptedInput: any?, context: {[string]: any}? }. Phase: observe; cost=medium; idempotency=read-only. Requires: none. Produces: agent-guidance, grounded-evidence, failure classification, safe retry policy, ranked recovery plan. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
errorNoThrown error, transport error, domain error object, or error message.
resultNoHandled tool result or partial result associated with the failure.
contextNoOptional structured facts such as selection state, resolvedPath, correctedInput, capability results, or recent observations.
toolNameYesExact name of the tool that failed or returned a handled error.
attemptedInputNoThe exact input passed to the failed tool; used for mutation safety and corrections.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations by enumerating what the recovery envelope contains (cause, evidence, confidence, retry safety, correctedInput, ranked fallback tools, script, next actions) and stating strong behavioral guarantees ('never invokes a fallback and never recommends repeating an identical failed mutation'). This adds real value on top of readOnlyHint/idempotentHint, though it restates read-only twice.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the key READ-ONLY constraint, but the body is dense and partly redundant: read-only is asserted three times (header, idempotency, Safety) and the Signature block duplicates the input schema. The metadata line (Phase/cost/idempotency/Requires/Produces) is boilerplate that could be trimmed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of describing the return value, and it does so thoroughly by listing the envelope's contents. Combined with the safety guarantees and fallback behavior, it is largely complete for a complex five-parameter tool, with only minor overlap against the annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all five parameters. The description's inline Signature restates those same fields and their optionality without adding syntax, format, or constraint detail beyond the schema, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Deterministically classify a failed MCP/Roblox tool call' from its error/result/input/context. This is unmistakably distinct from every sibling (tool-schema, tool-plan, agent-context, etc.), so an agent can route to it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Use this after any failed or blocked call,' giving a clear triggering condition. It also points to a concrete alternative ('inspect tool-schema for exact fields, defaults, constraints') for the schema-lookup case, though it does not enumerate when-not-to-use scenarios beyond that.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

file-existsCheck whether a path is a file or folder in the executor workspace (UNC isfile/isfolder)A
Read-onlyIdempotent

Probe a path in the executor's workspace and report whether it is an existing file and/or an existing folder. NOTE: this is executor-side file I/O — the connector runs INSIDE the executor, so the path is relative to the executor's workspace directory on the host machine, NOT the Roblox game. Requires the UNC functions isfile(path) -> bool and isfolder(path) -> bool. Each is type-guarded and pcall-wrapped INDEPENDENTLY: if a probe's function is missing or errors, its result is reported as false. If NEITHER isfile nor isfolder is available you get { error = 'isfile/isfolder are not available in this executor.' }. Returns { path, isFile, isFolder } or { error }. Signature: { path: string, threadContext: number?, timeoutMs: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: readfile. Produces: structured-result. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesPath within the executor workspace to probe, e.g. 'config.json' or 'data'. Reported as file and/or folder.
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the read-only/idempotent annotations: it names the required UNC functions, explains that each probe is type-guarded and pcall-wrapped independently so a missing/erroring function reports false, and specifies the exact error object when neither is available. This is exactly the runtime behavior an agent cannot infer from annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core behavior and the critical executor-vs-game warning, then structured metadata. It is somewhat verbose, repeating read-only/safety twice and echoing the signature already in the schema, but every section carries usable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully documents the return shapes ({ path, isFile, isFolder } or { error }) and the error case. Prerequisites, capabilities, and failure guidance are all present, so an agent has everything needed to invoke and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents path, timeoutMs, and threadContext, making the baseline 3. The description reinforces that path is workspace-relative and repeats the signature, but adds no format or constraint detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('probe a path ... report whether it is an existing file and/or folder') and scopes it precisely to the executor's workspace, explicitly distinguishing it from the Roblox game. This differentiates it cleanly from siblings like verify-path-exists and read-file without needing the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for when this applies ('executor-side file I/O ... path is relative to the executor's workspace directory on the host machine, NOT the Roblox game') and prerequisites ('active-client, resolved-target'). It does not name the sibling tools (e.g. verify-path-exists, list-files) that a user might confuse it with, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

filter-gcfiltergc — query the GC heap for functions or tables by structural criteriaA
Read-onlyIdempotent

The headline reflection tool: run the executor's UNC filtergc(filterType, options) against the entire live garbage collector to find Lua closures or tables that match a structural fingerprint, without writing any Luau by hand. Far more targeted than a raw getgc() sweep — you describe WHAT you are looking for and the executor returns only the objects that match. For filterType='function' the options are { Name?, Hash?, IgnoreExecutor? (default true), Constants? (array of constants the closure must reference), Upvalues? (array of upvalue values the closure must hold) } — e.g. find the closure that owns the string 'FireServer' and the upvalue 1337. For filterType='table' the options are { Keys? (array of keys that must be present), Values? (array of values that must be present), KeyValuePairs? (record of exact key=value pairs), Metatable? } — e.g. find the player-data table that has a 'Coins' key. Each match is encoded to a compact summary: functions report { source, line, name } (via debug.info) and tables report { address, keyCount }. Output is capped by 'limit'. Requires filtergc (type-guarded; returns { error } where it is unavailable) and every call is pcall-wrapped so a locked object can never abort the query. Returns { filterType, matchCount, truncated, matches } or { error }. Signature: { filterType: "function" | "table", options: { Name: string?, Hash: string?, IgnoreExecutor: boolean?, Constants: {string | number | boolean}?, Upvalues: {string | number | boolean}?, Keys: {string | number | boolean}?, Values: {string | number | boolean}?, KeyValuePairs: {[string]: string | number | boolean}?, Metatable: string | number | boolean? }, limit: number?, threadContext: number?, timeoutMs: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Capabilities: getgc. Produces: structured-result. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of matches to encode and return (default 100). Hitting this sets truncated=true.
optionsYesThe UNC filtergc criteria. Only the fields relevant to filterType are emitted into the Luau.
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
filterTypeYesWhat kind of GC object to search for: 'function' (Lua closures) or 'table'. This selects which option set below is meaningful.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only/idempotent/non-destructive, but the description adds substantial behavior: every call is pcall-wrapped so a locked object can't abort the query, filtergc is type-guarded and returns {error} where unavailable, output is capped by 'limit', and matches are encoded to compact summaries. This is rich disclosure well beyond the structured hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded and well organized, but the description is bloated: it reprints the entire parameter signature that the input schema already defines verbatim, plus phase/cost/idempotency/capability metadata. Several sentences duplicate structured data rather than earning their place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the full return-value burden and does so: it documents the {filterType, matchCount, truncated, matches} shape, the per-type match encoding, the limit/truncation behavior, and the {error} path. An agent has everything needed to call and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so each option (Name, Hash, Constants, Upvalues, Keys, Values, KeyValuePairs, Metatable, IgnoreExecutor) is already documented. The description groups options by filterType and gives examples, but that largely restates the schema's per-field 'function only'/'table only' annotations, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (run filtergc) and resource (the live GC heap), and precisely scopes it to finding Lua closures or tables matching a structural fingerprint. It also distinguishes itself from the raw getgc() sweep and from sibling lookup tools, so an agent can identify it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete when-to-use framing ('you describe WHAT you are looking for') and contrasts with a raw getgc() sweep, plus worked examples ('find the closure that owns the string FireServer'). However, it never names the closest siblings (find-functions-by-constant, find-tables-by-key, scan-closures-by-name) to route the agent, so the alternative-selection guidance is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find-bytecode-size-outliersFind bytecode size outliersA
Read-onlyIdempotent

Scan scripts and rank by getscriptbytecode size to find complex/high-value targets quickly. Signature: { limit: any?, includeModules: any?, includeLocalScripts: any?, threadContext: number? }. Phase: observe; cost=high; idempotency=read-only. Requires: active-client. Capabilities: getscriptbytecode. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoOptional hard result/work budget used to keep output and runtime bounded.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
includeModulesNoWhether to include modules in the bounded result.
includeLocalScriptsNoWhether to include local scripts in the bounded result.

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only/idempotent/non-destructive, so the bar is lower, and the description still adds real context: 'cost=high', 'Requires: active-client', and 'Produces: bounded-candidates'. These are behavioral traits not present in the annotations. It does not disclose pagination or output shape, keeping it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening sentence is well front-loaded, but the trailing metadata block repeats information already carried by annotations ('idempotency=read-only' and 'Safety: read-only' duplicate readOnlyHint/destructiveHint). The 'any?' signature is low-value filler. Some sentences do not fully earn their place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only scanning tool with no output schema, the definition supplies the key operating context: phase, cost, required client, and that output is bounded candidates. What is missing is any description of the returned ranking/format, but the core invocation context is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all four parameters with defaults and constraints. The description's inline 'Signature' repeats the parameter names but with vague 'any?' types, adding no meaning beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb+resource: 'Scan scripts and rank by getscriptbytecode size.' It adds a goal ('find complex/high-value targets quickly') that helps distinguish it from generic bytecode tools. It does not explicitly name a sibling alternative (e.g., search-bytecode, find-functions-by-complexity), so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Phase: observe; cost=high' tags imply this is an expensive reconnaissance step, and the stated goal hints at when it is useful. However, there is no explicit when-to-use-vs-alternatives guidance, and no exclusions or prerequisites beyond 'Requires: active-client'. Usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find-constants-xrefFind constants xrefA
Read-onlyIdempotent

Find getgc functions containing a target constant and return compact xrefs. Signature: { constantQuery: string, limit: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Capabilities: debug closure primitives. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoOptional hard result/work budget used to keep output and runtime bounded.
constantQueryYessearch text used to filter and rank bounded results.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, non-destructive behavior, so the bar is lower. The description still adds real context beyond them: bounded-candidates output, cost=medium, the active-client prerequisite, and a failure-path pointer to tool-schema. That is genuine added value over the structured hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose sentence, followed by a compact key=value spec block. Dense but every line carries signal (phase, cost, prerequisite, failure path). The inline signature duplicates the schema slightly, which is the only minor waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does explain what is produced (bounded-candidates, compact xrefs) and covers prerequisites and failure handling. For a read-only search tool this is close to complete, though it does not describe result shape or how ranking/bounding works in practice.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for all three parameters, so the schema already documents limit and threadContext semantics. The description only restates the signature with optionality markers, adding no format or constraint detail beyond the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence gives a specific verb+resource+scope: find getgc functions containing a target constant, returning compact xrefs. It is distinguishable from narrower siblings like get-closure-constants or find-functions-by-constant via the 'getgc' scope, though it never names an alternative to route between them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Phase: observe; cost=medium; Requires: active-client' implies when this tool fits operationally, but there is no explicit when-to-use versus when-not, and no mention of the closest alternatives (find-functions-by-constant, find-upvalue-xref, get-closure-constants). Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find-detached-instancesFind detached instances (in registry, off the DataModel)A
Read-onlyIdempotent

Walk every instance the executor can see (getinstances) and report the ones that are DETACHED from the running game — i.e. they exist in the executor's instance registry but pcall(inst.IsDescendantOf, inst, game) returns false, so they are not reachable from the DataModel. This catches objects that were Destroy()'d or reparented to nil yet kept alive by a lingering reference, plus anything an anti-detection script has stashed off the hierarchy. An optional className filter narrows the walk to a single ClassName (matched via IsA so subclasses are included). Unlike find-hidden-instances (which uses an ancestor-walk to test reachability), this tool uses IsDescendantOf against game directly, supports a class filter, and always reports a full byClass breakdown across ALL detached instances. Returns { totalDetached, byClass, truncated, samples: [{ class, name }] }. Requires getinstances; degrades with a clear error otherwise. Signature: { className: string?, limit: any?, maxScan: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax number of detailed sample instances to return in the samples list (default 200, clamped to 2000). The totalDetached count and byClass tally always cover every detached instance found regardless of this limit.
maxScanNoMax number of instances to examine from getinstances() before stopping (default 60000, clamped to 500000). Protects against huge games; if hit, `truncated` is set true.
classNameNoOptional ClassName to filter to, e.g. 'RemoteEvent', 'LocalScript', 'ScreenGui', 'Part'. Matched with inst:IsA(className) so subclasses are included. Omit to report detached instances of every class.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly/idempotent/non-destructive), and the description adds substantial context on top: it names the underlying mechanism (getinstances + IsDescendantOf), explains what kinds of objects it surfaces (Destroy()'d, reparented to nil, anti-detection stashes), discloses the truncation flag behavior, and the dependency failure mode. This is well beyond what the annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core definition and the sibling contrast before the return shape and metadata tail. The trailing boilerplate (Phase/cost/idempotency/Requires/Produces/Safety/On failure) partially duplicates the annotations and is somewhat verbose, but the substantive content is dense and earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description supplies the return shape ({ totalDetached, byClass, truncated, samples }), the prerequisite (getinstances), the failure behavior, and the capability boundary versus the sibling tool. Nothing an agent needs to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters are already documented in the schema, including limit/maxScan clamping and className IsA matching. The description restates className semantics and the 'always reports full byClass breakdown' behavior but adds no syntax, format, or constraint detail not already in the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (walk/report) and resource (instances detached from the DataModel), and defines 'detached' precisely via pcall(inst.IsDescendantOf, inst, game) returning false. It explicitly contrasts itself with the sibling find-hidden-instances, so an agent can differentiate without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the alternative (find-hidden-instances) and the mechanism that distinguishes them (IsDescendantOf vs ancestor-walk), plus the class-filter and full-byClass capabilities unique to this tool. It also states the prerequisite ('Requires getinstances; degrades with a clear error otherwise'), giving clear when-to-use and when-not conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find-duplicate-functionsFind identical / duplicated functions (IDA find-identical)A
Read-onlyIdempotent

Walk every Luau function in the GC, hash each one with getfunctionhash, and group functions that share the exact same hash — the runtime equivalent of IDA's 'find identical functions'. Reveals clones and copy-pasted code (duplicated module logic, repeated handlers, library functions instantiated many times). Each group reports the shared hash, how many functions carry it, and a few sample functions (ptr/name/source/line). Requires getgc + getfunctionhash; caps the scan and flags truncation. Signature: { minGroup: any?, limit: any?, maxScan: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax duplicate groups to return, largest groups first (default 50).
maxScanNoMax GC functions to scan (default 9000).
minGroupNoOnly report hash groups shared by at least this many functions (default 2 = any duplicate).
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), yet the description adds genuinely new behavior: it caps the scan and flags truncation, declares the producer of bounded candidates, and states the runtime dependencies. The remaining gap is that it doesn't quantify the cap behavior beyond 'caps the scan'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core mechanic in the first clause and keeps the rest dense with useful detail. Minor redundancy exists in the trailing metadata (idempotency and safety both assert read-only), which slightly dilutes an otherwise tight, well-ordered description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only scanner with no output schema, the description explains the return shape well (per-group shared hash, function count, sample ptr/name/source/line), the prerequisite APIs, truncation behavior, and failure guidance. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (minGroup, limit, maxScan, threadContext) are already documented with defaults and bounds. The description only restates the signature and gestures at 'caps the scan', adding no semantics beyond the schema — baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb+resource and mechanism: walk every Luau function in the GC, hash each with getfunctionhash, group by identical hash. It explicitly anchors the concept ('the runtime equivalent of IDA's find identical functions') and names what it surfaces (clones, copy-pasted module logic, repeated handlers), which cleanly separates it from siblings like find-functions-by-complexity or find-upvalue-sharing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for when the tool is useful ('reveals clones and copy-pasted code') plus operational preconditions (requires getgc + getfunctionhash, active-client) and cost/phase signals. It does not, however, explicitly compare against the many near-neighbour xref/complexity scanners, so the routing decision is inferential rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find-event-connectionsFind event connectionsA
Read-onlyIdempotent

Inspect RBXScriptSignal connections by instance path + signal name. Signature: { instancePath: string, signalName: string, limit: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: getconnections. Produces: bounded-candidates, created-handle. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoOptional hard result/work budget used to keep output and runtime bounded.
signalNameYestext value for signal name.
instancePathYesdotted Roblox instance/value path resolved in the active client.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, non-destructive, non-open-world behavior, so the bar is lower; the description still adds cost=medium, the capability (getconnections), and the produced artifacts (bounded-candidates, created-handle), which are not in the annotations. The only slight oddity is "created-handle" against readOnlyHint=true, but handles are internal bookkeeping rather than a mutation, so this is not a contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then a compact signature, phase, preconditions, capabilities, outputs, and failure guidance — no filler sentences. The stacked telegraphic fragments ("Phase: observe; cost=medium; idempotency=read-only") are dense but readable and each line earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description supplies the missing pieces an agent needs: required preconditions, a rough output shape (bounded-candidates, created-handle), the underlying capability, and a pointer to tool-schema for defaults and an invocation example. Only the return-value detail and the boundary against sibling connection tools remain thin.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters and the baseline is 3. The description only restates the signature, and its "limit: any?" / "threadContext: number?" typings are actually vaguer than the schema's number/integer definitions, so it adds no meaning beyond the structured data.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Inspect) and resource (RBXScriptSignal connections) plus the scoping keys (instance path + signal name), so the agent knows exactly what is queried. It does not, however, differentiate itself from close siblings such as list-signal-connections, get-connection-info, or find-instances-with-connections, leaving the agent to infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Phase: observe" and "Requires: active-client, resolved-target" give useful preconditions and hint at when the tool is applicable. But there is no explicit when-to-use vs. alternative statement, no exclusion, and no routing to a sibling for the overlapping connection-listing cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find-functions-by-complexityFind the most complex functions (RE starting points)A
Read-onlyIdempotent

Rank every Luau function in the GC by complexity so you know where to start reverse-engineering — the biggest, most interesting functions usually carry the core logic. For each function it counts the number of constants, upvalues, and nested protos, then returns the top limit by the chosen metric (sortBy), heaviest first. Each entry reports the owning script source, line, function name and pointer, plus the three counts. Pivot from here with get-closure-constants / find-string-xrefs / call-graph tools. Requires getgc + getconstants/getupvalues/getprotos; caps the scan and flags truncation. Signature: { sortBy: any?, limit: any?, maxScan: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax functions to return, heaviest first (default 50).
sortByNoWhich complexity metric to rank by: 'constants' (literals/strings/numbers used, default), 'upvalues' (captured variables), or 'protos' (nested child functions).constants
maxScanNoMax GC functions to scan (default 9000).
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only/idempotent/closed-world, so the bar is lower, and the description adds genuinely new context: it requires getgc + getconstants/getupvalues/getprotos, caps the scan, and flags truncation. It also declares phase/cost and points to tool-schema on failure, though the 'Safety: read-only' line merely repeats the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening sentences are well front-loaded and high-value, but the trailing metadata block restates annotation-derived facts ('idempotency=read-only', 'Safety: read-only') and re-lists the signature that the schema already provides, so not every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by enumerating return fields (script source, line, function name, pointer, three counts) and by disclosing prerequisites, scan caps, and truncation behavior. An agent has everything needed to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: 'heaviest first' ordering for limit, the metric semantics of sortBy, and that maxScan caps the scan and triggers a truncation flag. It still does not fully restate enum/default behavior, so not a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (rank) and resource (every Luau function in the GC) plus the ranking criterion (complexity by constants/upvalues/protos). An agent can distinguish this from siblings like list-gc-functions or find-functions-by-constant without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context ('know where to start reverse-engineering') and explicit follow-up routing to get-closure-constants / find-string-xrefs / call-graph tools. It does not state when *not* to use it versus other GC-scanning siblings, so it falls short of a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find-functions-by-constantFind functions by constantA
Read-onlyIdempotent

Scan getgc() and find closures whose constants contain a target string/number. Useful for locating handlers and hidden logic by magic constants. Signature: { constantQuery: string, limit: any?, includeCClosures: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Capabilities: debug closure primitives. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of matched functions to return (default: 30).
constantQueryYesCase-insensitive constant substring to search for.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
includeCClosuresNoInclude C closures (default: false).

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/destructive/openWorld, and the description partly repeats those (idempotency=read-only, safety=read-only). It does add non-schema context: phase=observe, cost=medium, requires an active client, produces bounded-candidates, and a recovery hint (inspect tool-schema on failure). That prerequisite and cost disclosure earns credit beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first two sentences, but the remaining content is a boilerplate template block (Phase/cost/idempotency/Requires/Capabilities/Produces/Safety) that partly duplicates the annotations and schema, and the Signature line repeats the input schema. Functional but with avoidable redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With full annotation coverage and 100% schema coverage, and no output schema required to be explained, the description is largely complete; it adds phase, cost, prerequisite, and a failure-recovery pointer. The one soft spot is the vague 'Produces: bounded-candidates' return hint, which does not describe actual output shape.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all four parameters (defaults, types, constraints). The 'Signature' line merely restates the parameter names/types and adds no new meaning beyond the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb+resource: scanning getgc() to find closures whose constants contain a target string/number, plus a use case (locating handlers/hidden logic by magic constants). It distinguishes itself from constant-scanning siblings like find-constants-xref by scoping to closures, but never names those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a motivating use case ('locating handlers and hidden logic by magic constants') which implies when it is useful, but provides no when-not guidance and does not name alternatives such as get-closure-constants or find-constants-xref. Usage is inferable but not directed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find-function-xrefsFind xrefs TO a function (IDA 'xrefs to')A
Read-onlyIdempotent

Resolve a target Luau FUNCTION from an expression, then walk every function in the GC and report which ones REFERENCE it — the runtime equivalent of IDA's 'cross references to' on a sub. A referrer references the target if the target appears in its upvalues (closures that captured it) OR among its nested protos (functions that embed it as an inner function). Each xref reports __fnInfo (ptr/name/source/line/nparams/nups) plus 'via' ("upvalue" or "proto"). The target itself is skipped. Requires getgc + getupvalues/getprotos; caps the scan and flags truncation. Pivot from list-gc-functions or lookup-function to get an expression that resolves here. Signature: { functionPath: string, limit: any?, maxScan: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax referrer functions (xrefs) to return (default 100).
maxScanNoMax GC functions to scan (default 9000).
functionPathYesLuau expression that resolves to the TARGET function, e.g. 'getrenv().game.ReplicatedStorage.Modules.Combat.attack' or 'require(path).onHit'. Evaluated as `(loadstring('return '..expr))()`; must yield type=='function' or the tool errors.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only/idempotent/non-destructive, so the bar is lower; the description adds real substance by defining what counts as a reference (upvalue capture or nested proto), noting the target is skipped, listing runtime dependencies (getgc, getupvalues/getprotos), and disclosing that the scan is capped and truncation is flagged.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core behavior and the reference semantics before the signature and metadata tail. The trailing 'Phase/cost/idempotency/Safety' line partly restates the annotations and the 'On failure' pointer is boilerplate, so it is dense but not waste-free.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by describing the return shape (per-xref __fnInfo fields plus the 'via' discriminator) and truncation behavior. An agent has everything needed to call and interpret this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters and their defaults. The description adds only marginal value beyond it (the expression must resolve to a function, scan capping), which meets the baseline-3 expectation when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (walk the GC and report referrers) and resource (a target Luau FUNCTION), and anchors it with the IDA 'cross references to' analogy. It is clearly separable from sibling xref tools such as find-string-xrefs, find-upvalue-xref, and find-global-xrefs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the upstream pivots ('Pivot from list-gc-functions or lookup-function to get an expression that resolves here') and states prerequisites (active-client, resolved-target). It does not state when NOT to use this versus the other xref tools, but the context is clear enough to select correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find-global-xrefsFind functions that reference a global or method by nameA
Read-onlyIdempotent

Find functions that likely call or reference this global or method by name. Luau bytecode stores global lookups and method names (e.g. FireServer, require, loadstring, HttpGet, GetService) as string constants, so this walks every function in the GC and reports each one whose constants contain the exact name string. This is the IDA xref-to-import equivalent: pivot from a sensitive API to all the code that uses it. Each hit reports the owning script source, line, function name and pointer. Requires getgc + getconstants; caps the scan and flags truncation. Signature: { name: string, limit: any?, maxScan: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesThe exact global or method name to look for (e.g. FireServer, require, loadstring, HttpGet, GetService).
limitNoMax matching functions to return (default 100).
maxScanNoMax GC functions to scan (default 9000).
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive), but the description adds genuine context: it requires getgc/getconstants, caps the scan, flags truncation, and reports the exact hit fields (script source, line, function name, pointer). It also notes 'likely' matches, warning the agent results are heuristic rather than exact call edges.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and mechanism, followed by requirements, output, and metadata in a structured order. It is dense but each clause carries signal; the phase/cost/safety metadata block is slightly boilerplate-heavy but not redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by describing the hit shape (owning script source, line, function name, pointer) and the truncation behavior. It also covers requirements, active-client precondition, and failure guidance, leaving nothing an agent needs to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents name/limit/maxScan/threadContext. The description's restated signature and examples add no syntax or defaulting detail beyond what the schema already provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Find functions that ... call or reference this global or method by name') and explains the underlying mechanism (string constants in Luau bytecode). The IDA 'xref-to-import equivalent' framing tells the agent exactly what class of tool this is and distinguishes it from the many other xref siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear intended context ('pivot from a sensitive API to all the code that uses it') and prerequisites ('Requires getgc + getconstants'). It does not explicitly contrast against the closest siblings (find-string-xrefs, find-constant-xref, find-upvalue-xref) or state when-not to use it, so routing between xref variants remains partly inferential.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find-hidden-guisFind hidden GUI containersA
Read-onlyIdempotent

Surface GUI overlays/menus that are hidden away from the normal PlayerGui hierarchy — the classic home of cheat menus, ESP overlays, and anti-detection UIs. Scans every nil-parented instance (getnilinstances) and the children of CoreGui, keeping anything that is a GUI container (LayerCollector / ScreenGui / BillboardGui / SurfaceGui / any GuiBase2d). Each hit reports its name, class, full path, and where it is hidden (location: nil-parented, CoreGui, detached, inside an Actor, etc.). Results are deduped. Requires getnilinstances for the nil sweep; the CoreGui sweep still runs even if getnilinstances is missing. Signature: { threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, but the description adds substantial extra context: graceful degradation when getnilinstances is missing (CoreGui sweep still runs), dedup of results, the location taxonomy of each hit, and the exact instance classes matched. This is well beyond what the annotations convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and scope, then structured metadata (signature, phase, requires, safety, failure). Some trailing tokens (idempotency=read-only, Safety: read-only) restate the annotations and could be trimmed, but overall it is dense and mostly earns its length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description specifies the return fields (name, class, full path, hidden location) and dedup behavior, plus the failure path (inspect tool-schema). Nothing an agent needs to invoke or interpret this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one optional parameter (threadContext) exists and schema coverage is 100%, so the schema fully documents it. The description adds no meaning to threadContext beyond the signature hint, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (surface hidden GUI overlays/menus) with precise scope: nil-parented instances and CoreGui children, filtered to GUI containers. An agent can distinguish it from siblings like list-gui-elements or find-hidden-instances purely from the description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly frames when to use it (anti-cheat/ESP/overlay detection, scanning outside the normal PlayerGui hierarchy), and the getnilinstances dependency is explained. However, it never names an alternative sibling (e.g. get-hidden-ui, list-gui-elements, find-detached-instances) or states when NOT to use it, so routing is inferred rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find-hidden-instancesFind hidden / detached instancesA
Read-onlyIdempotent

Enumerate EVERY instance the executor can see (getinstances — this includes objects that are not reachable from the game DataModel) and report the ones that are hidden: nil-parented, detached from the game tree, or buried inside an Actor / CoreGui. An instance is considered hidden when it is not reachable from game (__inTree is false). Returns a per-class tally (byClass) plus a capped list of samples, each describing the instance's class, name, full path, and exactly how it is hidden (location). This is the broadest sweep for anything an exploit/anti-detection script has stashed off the normal hierarchy. Requires getinstances; degrades with a clear error otherwise. Signature: { limit: any?, maxScan: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax number of detailed sample instances to return (default 500, clamped to 3000). The byClass tally always counts every hidden instance found regardless of this limit.
maxScanNoMax number of instances to examine from getinstances before stopping (default 60000). Protects against huge games; if hit, `truncated` is set true and scannedApprox reflects how many were inspected.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only/idempotent/destructive, and the description adds real behavioral context beyond them: dependency on getinstances with graceful degradation, that results are bounded/capped samples with a full byClass tally, a maxScan cap that sets a truncated flag, and cost/phase metadata. It also describes the exact sample shape (class, name, path, location) and the meaning of 'hidden' (__inTree false), which is not in the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded and the sentences are information-dense, but the trailing metadata block ('Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only.') is boilerplate that largely duplicates the annotations and adds length without new meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, yet the description fully specifies the return shape (per-class byClass tally plus capped samples with class/name/path/location), the definition of a hidden instance, the getinstances dependency, and failure behavior. For a 3-param read-only tool this leaves nothing an agent needs to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already fully documents limit, maxScan, and threadContext with defaults and clamping behavior. The description only restates the signature and mentions the capped-sample and truncation concepts, adding no syntax or semantics the schema does not already provide. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb and resource: 'enumerate EVERY instance the executor can see... and report the ones that are hidden', then defines hidden operationally (nil-parented, detached, buried in Actor/CoreGui; __inTree is false). It explicitly positions itself as 'the broadest sweep', which distinguishes it from narrower siblings like find-detached-instances or get-nil-instances without needing to open their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for use ('broadest sweep for anything an exploit/anti-detection script has stashed off the normal hierarchy') and a prerequisite ('Requires getinstances; degrades with a clear error otherwise'). It stops short of naming when-not-to-use or naming the narrower alternative tools, so the agent must infer the boundary between this and find-detached-instances/get-nil-instances.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find-hidden-remotesFind hidden Remote / Bindable channelsA
Read-onlyIdempotent

Find RemoteEvent / RemoteFunction / UnreliableRemoteEvent / BindableEvent / BindableFunction instances that are NOT in the normal game tree — i.e. nil-parented or detached from the DataModel. A remote/bindable kept alive by a reference but hidden off the hierarchy is a classic backdoor / data-exfiltration channel: an exploit or malicious script fires it to phone home or to receive commands without the object ever appearing in the Explorer. This tool pulls getnilinstances() (the primary source of nil-parented objects) and, when getinstances() is available, also includes any remote/bindable that is not a descendant of game, deduping across both sources. It complements get-nil-instances / find-hidden-instances by narrowing to just the communication-channel classes and giving a ready-to-inspect list. Returns { count, byClass, truncated, samples: [{ class, name, location }] }, capped. Requires getnilinstances; getinstances is used additionally when present. Degrades with a clear error if neither is available. Signature: { limit: any?, maxScan: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax number of detailed sample remotes/bindables to return in the samples list (default 200, clamped to 2000). The count and byClass tally always cover every hidden remote/bindable found regardless of this limit.
maxScanNoMax number of instances to examine from getinstances() before stopping the getinstances pass (default 60000, clamped to 500000). Protects against huge games; if hit, `truncated` is set true. The getnilinstances pass is always fully scanned.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, and the description goes well beyond them: it discloses the two data sources (getnilinstances plus optional getinstances), dedup across both, capping, the exact return shape, a hard dependency requirement, and graceful degradation with a clear error. That is rich behavioral disclosure beyond the annotation set.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose and threat rationale are front-loaded and every early sentence earns its place. The trailing metadata block (phase/cost/idempotency/requires/produces/safety/on-failure) is somewhat repetitive with the annotations and schema, slightly diluting conciseness, but the core is tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, but the description supplies the return shape ({ count, byClass, truncated, samples }), the dependency and degradation behavior, and the scoping rules. Combined with 100% parameter coverage, an agent has everything needed to call and interpret this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents limit, maxScan, and threadContext thoroughly (defaults, clamps, truncation semantics). The description adds only the generic note that output is 'capped' and does not add per-parameter meaning beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (find) plus the exact resource classes (RemoteEvent/RemoteFunction/UnreliableRemoteEvent/BindableEvent/BindableFunction) and the defining scope condition (NOT in the normal game tree, nil-parented/detached). It also names the sibling tools it complements (get-nil-instances / find-hidden-instances) and how it narrows relative to them, so an agent can distinguish it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the investigative context ('classic backdoor / data-exfiltration channel') and explicitly positions itself as a narrower complement to get-nil-instances / find-hidden-instances, which implies when to pick it. It does not state an explicit when-not condition or a full alternative-selection rule, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find-hidden-scriptsFind scripts that are hidingA
Read-onlyIdempotent

Scan ALL scripts (getscripts) and currently-running scripts (getrunningscripts) and report the ones that are trying to hide — i.e. not sitting normally in the game tree: nil-parented, detached from the DataModel, running inside an Actor (parallel-Luau VM), living in CoreGui, or destroyed-but-still-executing. Each result gives the script's name, class, where it actually lives, and whether it is currently running. This is the go-to tool for finding malicious/obfuscated/anti-detection scripts that hide outside the normal hierarchy. Requires getscripts/getrunningscripts; degrades with a clear error otherwise. Signature: { maxScan: any?, threadContext: number? }. Phase: orchestrate; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
maxScanNoMax scripts to examine from getscripts (default 6000). Running scripts are always checked first.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, non-destructive, so safety is covered. The description adds genuinely new behavioral context: the getscripts/getrunningscripts dependency, graceful degradation with a clear error, cost=medium, and the shape of each result. It does not, however, describe pagination or result-size bounds for maxScan=6000.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The hiding criteria and use case are front-loaded and easy to scan, but there is some redundancy at the tail: 'Safety: read-only' duplicates readOnlyHint and 'idempotency=read-only' repeats the same idea twice. Otherwise well organized and information-dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only orchestration tool with no output schema, the description covers the dependency, the failure mode, and what each result contains (name, class, actual location, running status). Enough to call it correctly; only the maxScan tradeoff/truncation behavior is left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents maxScan and threadContext fully. The description only restates the signature and defers to tool-schema for defaults and constraints, adding no meaning beyond the structured fields. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb+resource (scan scripts and report the ones hiding) and enumerates the exact hiding criteria: nil-parented, detached from DataModel, running inside an Actor, living in CoreGui, destroyed-but-still-executing. This distinguishes it cleanly from siblings like find-running-scripts, find-detached-instances, and find-hidden-instances.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the go-to use case (finding malicious/obfuscated/anti-detection scripts) and a prerequisite (requires getscripts/getrunningscripts, degrades with a clear error otherwise). It does not explicitly name sibling alternatives like find-running-scripts or find-hidden-instances, but the context for when to reach for it is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find-instances-with-connectionsFind instances that have connections on a named signalA
Read-onlyIdempotent

Walk every descendant of a root instance and report which ones have an active handler on a specific signal. For each descendant it reads inst[signalName]; if that member is an RBXScriptSignal with at least one connection (per getconnections), it records { Path, ClassName, ConnectionCount }. Results are sorted by ConnectionCount descending. This answers questions like "which Parts in Workspace have a Touched handler?", "which GuiButtons are wired to Activated?", or "who is listening to Changed under this model?". Requires the executor's getconnections; degrades to a clear { error } if unavailable. Scanning is capped at limit instances to stay responsive on huge games. Returns { Signal, Root, ScannedInstances, Truncated, MatchCount, Instances: [...] }. Signature: { root: any?, signalName: string, limit: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Capabilities: getconnections. Produces: bounded-candidates, created-handle. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
rootNoLua expression for the root instance whose descendants are scanned (e.g. 'game', 'game.Workspace', 'game.Workspace.Map'). Evaluated as `return <root>`. Defaults to 'game' (scans the whole DataModel up to the limit). Narrow this for faster, more focused scans.game
limitNoMaximum number of descendant instances to SCAN (not match) before stopping. Defaults to 1000. Increase to cover larger hierarchies (slower); decrease for a quick sample. If the scan stops early because this cap was hit, Truncated is true.
signalNameYesREQUIRED name of the signal member to probe on each descendant (e.g. 'Touched', 'Changed', 'ChildAdded', 'Activated', 'OnClientEvent'). Only instances exposing this member as an RBXScriptSignal with >0 connections are reported.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the readOnly/idempotent annotations by disclosing the dependency on the executor's getconnections, the graceful { error } degradation, the limit-based scan cap that sets Truncated, and the descending sort order. These are the operational traits an agent needs before invoking.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The first sentences are front-loaded and dense with useful detail, and the return shape is spelled out. However, the trailing metadata block (Phase, cost, Requires, Capabilities, Produces, Safety, On failure) is largely boilerplate that dilutes an otherwise tight description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by naming every returned field ({ Signal, Root, ScannedInstances, Truncated, MatchCount, Instances }) and explaining failure behavior. For a medium-cost read-only scan tool, nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the schema already documents root evaluation, the limit-as-scan-cap semantics, and signalName requirements. The description's inline signature and limit note largely restate that, adding little new meaning, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (walk/report) and resource (descendant instances with active signal connections) and names the exact mechanic (reads inst[signalName], checks getconnections). It is clearly distinguishable from siblings like list-signal-connections or scan-connections-by-source, which operate on a different axis.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete motivating questions ('which Parts in Workspace have a Touched handler?', 'which GuiButtons are wired to Activated?') that tell the agent when this tool fits. It does not explicitly name when to prefer an alternative such as find-event-connections or list-instance-signals, so no exclusions are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find-instance-xrefsFind functions that reference an Instance (IDA data xrefs)A
Read-onlyIdempotent

Resolve a target Instance from a Luau expression, then walk every function in the GC and report which ones hold a reference to it — the runtime equivalent of IDA's data cross-references on a global object. A referrer references the instance if it appears in the function's upvalues OR constants. Answers 'which functions read, manipulate, or watch this object?'. Each xref reports __fnInfo (ptr/name/source/line/nparams/nups) plus 'via' ("upvalue" or "constant"). Requires getgc + getupvalues/getconstants; caps the scan and flags truncation. Signature: { instancePath: string, limit: any?, maxScan: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax referrer functions (xrefs) to return (default 100).
maxScanNoMax GC functions to scan (default 9000).
instancePathYesLuau expression that resolves to the target Instance, e.g. 'game.Workspace.Boss' or 'game:GetService("Players").LocalPlayer.Character'. Evaluated as `(loadstring('return '..expr))()`; must yield typeof=='Instance' or the tool errors.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover safety (readOnlyHint, idempotentHint, destructiveHint=false), so the bar is lower. The description adds real behavioral context beyond that: the exact referrer-detection rule (upvalue OR constant), the per-xref return shape (__fnInfo fields plus 'via'), and the scan cap with truncation flagging. It stops short of stating expected runtime cost magnitude or what a truncated result looks like structurally.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The behavioral core is front-loaded and every early sentence earns its place; the signature restatement and boilerplate metadata (Phase/cost/Produces, 'On failure: inspect tool-schema') are somewhat redundant. Still appropriately sized for a medium-cost scan tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by naming the __fnInfo fields and the 'via' discriminator, so an agent knows what comes back. Requirements and constraints are covered; only edge behavior (truncation signal shape, error cases beyond type mismatch) is left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and all four parameters are documented in the schema, so the baseline is 3. The description adds the important semantic that instancePath is a Luau expression yielding an Instance, but re-lists the signature redundantly rather than clarifying unknown limits beyond the schema (e.g. what threadContext changes).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb chain (resolve an Instance, walk every GC function, report referrers) and identifies the resource precisely. The IDA data-xref analogy and the 'upvalues OR constants' criterion distinguish it from sibling xref tools like find-upvalue-xref, find-constants-xref, and find-function-xrefs without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear usage frame ('which functions read, manipulate, or watch this object?') and prerequisites (active-client, resolved-target, requires getgc + getupvalues/getconstants), which tells the agent when it applies. It does not, however, explicitly route the agent away from the narrower per-mechanism siblings (find-upvalue-xref, find-constants-xref) when a single mechanism is what's wanted.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find-module-scriptsFind module scriptsA
Read-onlyIdempotent

Search ModuleScripts by substring against name/full path. Signature: { query: string, limit: any?, threadContext: number? }. Phase: orchestrate; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoOptional hard result/work budget used to keep output and runtime bounded.
queryYessearch text used to filter and rank bounded results.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, non-destructive, non-open-world, so the safety line is redundant; however the description adds real context beyond them: it requires an active client, costs medium, and produces bounded candidates rather than an exhaustive listing. It also points to tool-schema on failure, which is useful operational guidance.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose sentence is well front-loaded, but the remaining text is dense keyword-soup (Phase/cost/idempotency/Requires/Produces/Safety) and the inline signature duplicates the input schema without adding syntax or defaults. Readable but not every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With rich annotations, 100% schema coverage, and no output schema to explain, the description covers the important extras: prerequisite (active client), output nature (bounded candidates), and a failure recovery path. Only the missing guidance on choosing it over sibling search tools leaves a gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters are already documented with defaults and bounds. The description's signature block merely restates them, adding only that the query matches name/full path, which is a modest clarification. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Search ModuleScripts') plus the matching semantics ('by substring against name/full path'), so an agent knows exactly what is queried. It does not name or differentiate itself from near siblings like semantic-search-scripts, script-grep, or get-module-source, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives operating context ('Phase: orchestrate', 'Requires: active-client', cost=medium) that implies when the tool is appropriate, but never says when to prefer it over alternatives such as semantic-search-scripts for fuzzy lookups or script-grep for source-text search. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find-path-referencesFind path referencesA
Read-onlyIdempotent

Find constants containing instance path fragments like ReplicatedStorage/Workspace. Signature: { query: string, limit: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoOptional hard result/work budget used to keep output and runtime bounded.
queryYessearch text used to filter and rank bounded results.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only/idempotent, but the description adds real context: cost=medium, the active-client requirement, and the bounded-candidates output shape. It also points to tool-schema for exact fields on failure, which is useful operational guidance beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence, followed by compact labeled fields (Phase/cost/Requires/Produces/Safety). The inline signature duplicates the schema slightly, but overall it is tight and free of filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, schema-fully-documented tool with no output schema, the description covers purpose, precondition, cost, output shape, and failure handling. Only the lack of sibling routing guidance keeps it from being fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents query, limit, and threadContext fully. The description restates the signature but adds no meaning (e.g., ranking behavior or how limit interacts with results) beyond what structured data already carries. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (find) and a well-scoped resource (constants containing instance path fragments like ReplicatedStorage/Workspace), which is a genuinely distinct scope from sibling xref tools. It does not explicitly name an alternative sibling, so it falls just short of 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides preconditions (Requires: active-client) and a phase marker (observe), which imply context of use. However, it never says when to choose this over siblings like find-constants-xref, find-string-xrefs, or resolve-entity, leaving the routing decision to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find-remote-xrefsFind functions that fire/handle a remote (IDA xrefs)A
Read-onlyIdempotent

Resolve a RemoteEvent / RemoteFunction / BindableEvent / BindableFunction from a Luau expression, then walk every function in the GC and report which ones reference it — the runtime equivalent of cross-referencing a network endpoint. A referrer references the remote if it appears in the function's upvalues OR constants, which surfaces the closures that FireServer/InvokeServer/OnClientEvent-connect this remote. Each xref reports __fnInfo (ptr/name/source/line/nparams/nups) plus 'via' ("upvalue" or "constant"); the response includes the remote's ClassName. Requires getgc + getupvalues/getconstants; caps the scan and flags truncation. Signature: { remotePath: string, limit: any?, maxScan: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax referrer functions (xrefs) to return (default 100).
maxScanNoMax GC functions to scan (default 9000).
remotePathYesLuau expression resolving to a RemoteEvent/RemoteFunction/Bindable, e.g. 'game.ReplicatedStorage.Remotes.AttackEvent' or 'getrenv().game.ReplicatedStorage.Net.RF'. Evaluated as `(loadstring('return '..expr))()`; must yield typeof=='Instance' or the tool errors.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive), yet the description adds real behavioral context: the exact matching rule (appears in upvalues OR constants), the per-xref payload (__fnInfo fields, 'via'), that ClassName is returned, the dependency on getgc/getupvalues/getconstants, and that the scan is capped with truncation flagged.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core explanation is front-loaded and dense, but the trailing metadata block is partly redundant: 'idempotency=read-only' duplicates idempotentHint and 'Safety: read-only' duplicates readOnlyHint/destructiveHint, which is wasted space in an otherwise tight definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-shape burden and does so (__fnInfo contents, 'via', ClassName, truncation flag), plus it covers dependencies and a failure-path pointer to tool-schema. It stops short of explaining how limit and maxScan interact when truncation occurs, which is the one remaining ambiguity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters including defaults and the loadstring evaluation rule. The inline Signature block restates the same fields with less detail than the schema, adding essentially nothing beyond the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource chain: resolve a Remote/Bindable from a Luau expression, then walk GC functions and report referrers. It explicitly names the mechanism (upvalues OR constants) and the concrete payoff (closures that FireServer/InvokeServer/connect the remote), which distinguishes it from siblings like find-function-xrefs or scan-remote-listeners.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The purpose implies when to reach for it (cross-referencing a remote endpoint) and it lists prerequisites ('Requires getgc + getupvalues/getconstants', 'active-client, resolved-target'), but it never names an alternative tool or states when NOT to use it versus scan-remote-listeners, find-event-connections, or find-function-xrefs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find-running-scriptsList currently-running scriptsA
Read-onlyIdempotent

List every script that is currently executing (getrunningscripts), including ones that are not visible in the normal hierarchy — nil-parented, detached from the DataModel, running inside an Actor, or living in CoreGui. For each running script this returns its name, class, full path, where it actually lives (location), and whether it is still reachable from the game tree (inTree). This surfaces active anti-detection / obfuscated logic that is running but hidden from the Explorer. Requires getrunningscripts; degrades with a clear error otherwise. Output is capped at 300 scripts with a truncated flag. Signature: { threadContext: number? }. Phase: orchestrate; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates, operation-receipt. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only/idempotent/non-destructive, and the description adds genuinely new behavior: requires the getrunningscripts capability, degrades with a clear error otherwise, and output is capped at 300 scripts with a truncated flag. It also lists the return fields, which the agent would otherwise not know. Loses a point only because it doesn't explain what happens at the cap boundary or ordering.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The first two sentences are well front-loaded and earn their place. But the trailing meta block ('Phase: orchestrate; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates, operation-receipt. Safety: read-only.') largely duplicates the annotations, and the closing 'On failure: inspect tool-schema...' is generic boilerplate.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden and does so (name, class, path, location, inTree), plus the truncation cap and the capability requirement. An agent has enough to call it correctly; only failure semantics beyond 'clear error' are thin.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single threadContext parameter, so the schema already documents it fully. The description's 'Signature: { threadContext: number? }' merely restates that, adding no semantics beyond the schema — baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('List every script that is currently executing') and enumerates the scope that distinguishes it — nil-parented, detached, Actor-hosted, CoreGui — plus the fields returned. An agent can tell it apart from find-hidden-scripts or list-script-actors without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies when to reach for it ('surfaces active anti-detection / obfuscated logic that is running but hidden from the Explorer'), which is useful context. However, it never names an alternative sibling (find-hidden-scripts, find-detached-instances, get-nil-instances, list-script-actors) or states when this tool is the wrong choice versus those overlapping ones.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find-string-in-tablesFind string VALUES stored in GC tables (runtime string data scan)A
Read-onlyIdempotent

Walk every live Luau GC table and report each string VALUE that matches query. This is the runtime-DATA complement to find-string-xrefs (which scans closures' bytecode constants): it finds strings that exist as actual table data at this moment — server responses, built-up chat/UI text, dynamically constructed remote names, cached config strings, decoded tokens — which may never appear as a source literal. By default contains=true performs a plain (non-pattern) substring search; set contains=false for exact equality. Each hit is recorded as { table, key, value } where 'table' is the container address (tostring), 'key' is the field (tostring), and 'value' is the matched string truncated to its first 200 characters. Pivot from a hit with read-path-value / write-path-value (or find-table-references to see who owns the container). Each table's pairs() iteration is pcall-guarded so locked/proxy tables never abort the scan; GC objects examined are capped by maxScan and results by limit, with a 'truncated' flag. Requires getgc (falls back from getgc(true) to getgc()). Returns { query, contains, matchCount, scannedObjects, truncated, matches } or { error }. Signature: { query: string, contains: any?, limit: any?, maxScan: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of matching strings to return (default 150). Hitting this sets truncated=true.
queryYesThe string to search for among table VALUES, e.g. 'GodMode', 'http', 'BuyItem', a token prefix, or any substring of runtime text. With contains=true (default) any string value containing this matches; with contains=false only string values exactly equal to this match. Case-sensitive, plain text (not a Lua pattern).
maxScanNoMaximum number of GC objects to examine before stopping (default 40000). Hitting this sets truncated=true. Raise for a deeper sweep at the cost of time; lower if scans are slow.
containsNoWhen true (default), match any string value that CONTAINS `query` as a plain substring (string.find with plain=true). When false, require exact string equality. Use exact to pin one specific value.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only/idempotent/no-destruct, but the description adds substantial behavior beyond them: pairs() iteration is pcall-guarded so locked/proxy tables never abort the scan, GC objects are capped by maxScan and results by limit with a 'truncated' flag, getgc(true) falls back to getgc(), and it declares phase/cost. This is rich, non-redundant operational context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and sibling contrast are front-loaded, and the sentences carry dense, useful detail. It is long and includes some redundancy — read-only is asserted in the description body, the 'idempotency=read-only' line, and 'Safety: read-only' — which is partly duplicated by the annotations, but nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description carries the return contract and does so fully ('Returns { query, contains, matchCount, scannedObjects, truncated, matches }'), plus the per-hit shape { table, key, value } with truncation to 200 chars and the failure shape { error }. For a medium-cost, 5-parameter scan tool this leaves no material gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and every parameter already documents its own default and semantics, so the baseline is 3. The description corroborates the contains behavior (plain substring vs exact, case-sensitive, not a Lua pattern) but adds little beyond what the schema already states for limit/maxScan/threadContext.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Walk every live Luau GC table and report each string VALUE that matches query') and explicitly differentiates itself from the sibling find-string-xrefs by scope (runtime data vs bytecode constants). An agent can route between the two without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the alternative (find-string-xrefs) and the exact condition that selects this tool ('runtime-DATA complement... strings that exist as actual table data at this moment'), and gives concrete examples of when this is the right choice. It also explains the contains=true/false branch as a usage decision and offers pivot tools (read-path-value / write-path-value) for follow-up.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find-string-xrefsFind xrefs to a string (IDA xref-to-string)A
Read-onlyIdempotent

Cross-reference a string the way IDA jumps from a string literal to every place it is used. Walks every Luau function in the GC and reports each function that has the query as a string constant in its bytecode — i.e. each function that likely produces or compares against that text. Use exact=true for an exact match, otherwise it does a plain substring search (case-sensitive). Each hit reports the owning script source, line, function name and pointer, plus the first matching constant. Great for pivoting from list-strings to the code that references a remote name, URL, error message, or anti-cheat tag. Requires getgc + getconstants; caps the scan and flags truncation. Signature: { query: string, exact: any?, limit: any?, maxScan: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
exactNoIf true, only match constants that equal the query exactly. If false (default), match any constant that contains the query as a substring (case-sensitive).
limitNoMax matching functions to return (default 100).
queryYesThe string to cross-reference (e.g. a remote name, URL, or error message).
maxScanNoMax GC functions to scan (default 9000).
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description still adds real behavioral context beyond them: the GC-walking scan, the getgc + getconstants dependencies, and the fact that the scan is capped and truncation is flagged. It never contradicts the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose and search semantics are front-loaded well, but the trailing block ('Signature: {...}. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client... Safety: read-only. On failure: inspect tool-schema...') largely repeats the schema and annotations and does not earn its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With annotations, a 100%-covered schema, and no output schema, the description is nearly complete: it describes the return payload (owning script source, line, function name, pointer, first matching constant) and the truncation flag, plus the active-client prerequisite. Minor boilerplate aside, an agent has what it needs to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents exact, limit, maxScan, and threadContext. The description restates the exact/substring behavior and the scan caps without adding syntax or constraints the schema lacks, which lands at the baseline for full-coverage schemas.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The definition states a precise verb+resource ('Cross-reference a string... reports each function that has the query as a string constant in its bytecode') and gives a concrete mechanism (walks every Luau function in the GC). It implicitly separates itself from list-strings by framing the tool as the pivot step after listing, and from find-constants-xref/find-function-xrefs by scoping to string constants. An agent can tell what this does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear usage context ('Great for pivoting from list-strings to the code that references a remote name, URL, error message, or anti-cheat tag') and the exact-vs-substring condition for the exact parameter. It does not, however, name a sibling to prefer or state an explicit when-not-to-use condition, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find-table-referencesFind all GC references to a specific table (reverse ownership scan)A
Read-onlyIdempotent

Answer 'who holds this table alive / who can reach it?' Resolve a Luau expression to a target TABLE, then walk the entire GC and record every OTHER object that references it: (1) for each GC TABLE, pcall-iterate pairs() and if any VALUE is the target, record { container, via='value', key } (the surrounding key text); (2) for each Lua CLOSURE, scan its upvalues (getupvalues, guarded) and if any UPVALUE is the target, record { container, via='upvalue', key } where container is the function's source:line (via debug.info / getinfo). The target table itself is skipped so it never lists itself. This is the inverse of the look-down tools (read-path-value, dump-table): use it to find the owner of a shared state/config table, to discover which closures captured it (so you can hook or inspect them), or to understand why a table is not being collected. Every access is pcall-guarded so locked objects never abort the scan; GC objects examined are capped by maxScan and results by limit, with a 'truncated' flag. Requires getgc; upvalue scanning additionally needs getupvalues (type-guarded — skipped if absent). Returns { target, referenceCount, scannedObjects, truncated, references } or { error }. Signature: { tableExpr: string, limit: any?, maxScan: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of references to return (default 100). Hitting this sets truncated=true.
maxScanNoMaximum number of GC objects to examine before stopping (default 40000). Hitting this sets truncated=true. Raise for a deeper sweep at the cost of time; lower if scans are slow.
tableExprYesLuau expression resolving to the target TABLE whose references you want to find, e.g. 'getgenv().PlayerData', 'require(game.ReplicatedStorage.Config)', '_G.Settings', or any table reference from a prior scan. Evaluated as `return <tableExpr>` and must resolve to a table.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only/idempotent/non-destructive, but the description adds substantial context beyond them: every access is pcall-guarded, the target is skipped to avoid self-listing, upvalue scanning is type-guarded and skipped if getupvalues is absent, requires getgc, and scans/results are capped with a truncated flag. This is exactly the behavioral disclosure the rubric rewards.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core question is front-loaded and the algorithmic detail is valuable, but the trailing block (Signature/Phase/cost/idempotency/Requires/Produces/Safety/On failure) partially duplicates annotations and schema and lengthens the text. Efficient overall, with minor redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex GC-walking tool with no output schema, the description supplies return shapes ({ target, referenceCount, scannedObjects, truncated, references } or { error }), prerequisites (getgc, optional getupvalues), and the truncation contract. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents limit, maxScan, tableExpr, and threadContext thoroughly. The description restates the signature and the limit/maxScan/truncated interplay but adds no syntax or format detail beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource+scope: resolve a Luau expression to a target table, then walk the GC recording every OTHER object that references it. It explicitly positions itself against siblings ('the inverse of the look-down tools (read-path-value, dump-table)'), so an agent can distinguish it without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete when-to-use scenarios (find the owner of a shared state/config table, discover capturing closures, diagnose why a table is not collected) and names the contrasting alternatives. The dual-value vs upvalue modes are laid out, so the agent knows what this returns and when it is the right choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find-tables-by-keyFind GC tables that contain a named string key (structure scan)A
Read-onlyIdempotent

Cheat-Engine-style STRUCTURE scan: walk every live Luau GC table and report each one that contains a string key matching key. This is the complement to search-gc-value (which matches by stored VALUE) — here you match by the NAME of a field, which is ideal when you know a stat/config table exposes a field like 'Coins', 'WalkSpeed', 'GodMode' or 'Config' but you don't yet know the value or the container. For every matching key the tool records { table, matchedKey, sampleValue } where 'table' is the table's address (tostring), 'matchedKey' is the exact key string found, and 'sampleValue' is the encoded current value at that key — so you can immediately see what the field holds. Set contains=true for a substring match on key names (string.find with plain=true). Iteration of each table is pcall-guarded so locked/proxy tables never abort the scan; the number of GC objects examined is capped by maxScan and the result list by limit, with a 'truncated' flag when either cap is hit. Pivot from a hit with read-path-value / write-path-value (using a Luau expression that reaches the container) to read or flip the field. Requires getgc (falls back from getgc(true) to getgc()). Returns { matchCount, scannedObjects, truncated, matches } or { error }. Signature: { key: string, contains: any?, limit: any?, maxScan: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYesThe field name to look for as a KEY in GC tables, e.g. 'Coins', 'Health', 'WalkSpeed', 'GodMode', 'Config'. With contains=false (default) only string keys exactly equal to this are matched; with contains=true any string key whose text includes this substring matches (case-sensitive, plain text).
limitNoMaximum number of matched keys to return (default 100). Hitting this sets truncated=true.
maxScanNoMaximum number of GC objects to examine before stopping (default 40000). Hitting this sets truncated=true. Raise for a deeper sweep at the cost of time; lower if scans are slow.
containsNoWhen true, match any string key that CONTAINS `key` as a plain substring (string.find with plain=true) instead of requiring exact equality. Default false (exact match). Use for fuzzy field discovery.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only/idempotent/non-destructive, yet the description still discloses non-obvious behavior beyond them: iteration is pcall-guarded so locked/proxy tables never abort the scan, getgc falls back from getgc(true) to getgc(), maxScan caps examined objects and limit caps results, and a 'truncated' flag fires when either cap is hit. This is exactly the added context the annotations can't express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose and the sibling contrast before the mechanics, and every sentence carries substance. It is on the long side and the trailing metadata line (Phase/cost/idempotency/Safety) partially repeats the annotations, which slightly dilutes conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description carries the full burden and does so: it names the return shape { matchCount, scannedObjects, truncated, matches }, the per-hit record fields, the error case, and the getgc prerequisite. An agent has everything needed to invoke and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the prose adds meaning beyond it: it explains the cross-parameter interplay of the two caps ('number of GC objects examined is capped by maxScan and the result list by limit'), the exact key-vs-substring matching behavior, and the contents of the returned { table, matchedKey, sampleValue } record.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+mechanism: 'STRUCTURE scan: walk every live Luau GC table and report each one that contains a string key matching `key`.' It explicitly distinguishes itself from the sibling search-gc-value (matches stored VALUE) versus this one (matches field NAME), so an agent can route correctly without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when-to-use scenario ('when you know a stat/config table exposes a field like Coins/WalkSpeed/GodMode but you don't yet know the value or the container'), names the alternative (search-gc-value) and the condition that selects it, and even supplies the follow-up pivot (read-path-value / write-path-value).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find-upvalue-sharingFind functions sharing the same upvalue table (closure clusters)A
Read-onlyIdempotent

Walk every Luau function in the GC and, for each function, inspect its upvalues. Whenever an upvalue is a TABLE, record the owning function under that table's identity. After scanning, report tables that are shared as an upvalue by several distinct functions — these are typically the shared state / private module table of a single closure family, so the functions that share it belong to the same module or factory. This is how you cluster functions that were defined together (a module's methods all close over the same private table). Each entry reports the table identity, how many functions share it, and a few sample functions. Requires getgc + getupvalues; caps the scan and flags truncation. Signature: { minGroup: any?, limit: any?, maxScan: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Capabilities: debug closure primitives. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax shared tables to return, most-shared first (default 40).
maxScanNoMax GC functions to scan (default 9000).
minGroupNoOnly report tables shared as an upvalue by at least this many distinct functions (default 3).
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only/idempotent/non-destructive. The description adds real behavioral context beyond them: it depends on getgc + getupvalues, caps the scan, and flags truncation, plus reports group size and sample functions. This is useful non-obvious operational detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core prose is front-loaded and clear, but the trailing metadata block (Phase, cost, idempotency, Safety, Produces) largely repeats the annotations and adds boilerplate. Some sentences do not earn their place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully explains the return shape (table identity, share count, sample functions) and notes scan caps/truncation and defaults. It is essentially complete for correct invocation, though it could state prerequisites like requiring an active client more explicitly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the four parameters (minGroup, limit, maxScan, threadContext) are already fully documented. The description only restates the signature without adding syntax or behavioral nuance, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: walking Luau functions in the GC and inspecting their upvalues to find tables shared by several functions. It clearly frames the output as closure clusters (module/factory grouping), which distinguishes it from neighbors like find-upvalue-xref and get-function-upvalues.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a concrete when-to-use case — clustering functions defined together via a shared private table. However it does not name an explicit alternative (e.g. find-upvalue-xref or get-function-upvalues) or state when-not to use this tool, leaving that to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find-upvalue-xrefFind upvalue xrefA
Read-onlyIdempotent

Find functions whose upvalue names or values match a query. Signature: { query: string, limit: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Capabilities: debug closure primitives. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoOptional hard result/work budget used to keep output and runtime bounded.
queryYessearch text used to filter and rank bounded results.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint, idempotentHint, destructiveHint=false), and the description adds genuinely new context beyond them: the 'active-client' precondition, cost=medium, and bounded-candidates output shape. The 'idempotency=read-only'/'Safety: read-only' lines partly restate annotations, keeping it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded, but the trailing meta-tags are boilerplate-heavy and redundant ('idempotency=read-only' followed by 'Safety: read-only'). The signature line duplicates the schema, adding length without new information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description should carry more of the return burden, yet 'Produces: bounded-candidates' is vague about result shape and ordering. It covers preconditions and cost adequately for a read-only tool whose annotations already signal safety.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters with defaults and constraints. The description's 'Signature' line merely restates the schema types without adding format, ranking, or default semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Find functions whose upvalue names or values match a query'), which is clear enough to distinguish from adjacent tools like get-closure-upvalues or find-upvalue-sharing. It does not explicitly name a sibling alternative, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives operational context ('Phase: observe', 'Requires: active-client') that implies when the tool is usable, but never states when to prefer it over siblings such as find-upvalue-sharing or get-function-upvalues, nor any exclusion conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fire-click-detectorTrigger a ClickDetectorA
Destructive

WRITES LIVE GAME STATE. Resolve a Luau expression to a ClickDetector and trigger it via the executor's fireclickdetector, exactly as if the local player clicked the part it is attached to. This bypasses MaxActivationDistance and line-of-sight and fires the MouseClick signal, so any server/client logic bound to the detector runs (buy buttons, levers, clickable doors, NPC dialogs). Use it to drive click-gated game flow during automation/testing. Guards that fireclickdetector exists in this executor and that the target is actually a ClickDetector before firing; the fire call is pcall-guarded. WARNING: this mutates the running game and the effect may replicate to the server. Returns { Path, ok } or { error }. Signature: { path: string, distance: any?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Capabilities: fireclickdetector. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesLuau expression resolving to the ClickDetector to fire, e.g. 'game.Workspace.Lever.ClickDetector' or 'game.Workspace.Shop.BuyButton.ClickDetector'. Evaluated as `return <path>` and must resolve to an Instance whose ClassName is 'ClickDetector'.
distanceNoThe distance (studs) to report to the detector as the click origin, passed as the second argument to fireclickdetector(cd, distance). Default 0 (treated as point-blank). Some games read this value to gate behavior; leave at 0 unless you need to emulate a specific click distance.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructive/non-idempotent, but the description adds substantial context beyond them: it bypasses MaxActivationDistance and line-of-sight, fires the MouseClick signal, guards existence/type, wraps the fire in pcall, warns that effects may replicate to the server, and states the return shape ({ Path, ok } or { error }). This is unusually rich behavioral disclosure for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the critical warning ('WRITES LIVE GAME STATE') and keeps the causal explanation tight. It is somewhat long and partially redundant (signature line repeats schema fields; the 'On failure: inspect tool-schema' boilerplate is generic), but every core sentence carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description supplies the return contract, the mutation risk, the replication caveat, and the preconditions (active-client, resolved-target, explicit-mutation-approval). For a destructive, side-effecting tool this covers everything an agent needs to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents path, distance (default 0, studs, gating semantics), and threadContext. The description restates the signature and reinforces that the call is a mutation but adds no syntax or format detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource: 'Resolve a Luau expression to a ClickDetector and trigger it via the executor's fireclickdetector.' It clearly distinguishes itself from neighbors like fire-signal, click-button, and fire-proximity-prompt by naming the exact mechanism (MouseClick signal) and effect (server/client logic bound to the detector runs).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context: 'Use it to drive click-gated game flow during automation/testing,' and describes conditions it bypasses (MaxActivationDistance, line-of-sight). It stops short of explicitly naming alternatives (e.g. click-button for GUI, fire-proximity-prompt) or stating when NOT to use it, so it is clear but not fully routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fire-comm-channelSend typed values through an Actor communication channelA
Destructive

WRITES LIVE GAME STATE. Resolve get_comm_channel(id) and call Channel:Fire(...typed arguments). Requires confirm=true because channel listeners may execute arbitrary live behavior. Signature: { id: string, arguments: any?, confirm: boolean?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval. Capabilities: create_comm_channel. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYestext value for id.
confirmNoExplicit safety acknowledgement; must be true before this state-changing operation runs.
argumentsNoOptional ordered typed arguments forwarded to the selected operation.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint=false, destructiveHint=true), the description discloses that confirm=true is mandatory, that channel listeners may execute arbitrary live behavior, that idempotency is 'contextual-write', and that it produces an operation-receipt verifiable with assert-state. This is exactly the extra behavioral context the annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The lead sentence is well front-loaded, but the text is a dense run of metadata tags and repeats itself: the signature duplicates the schema, and 'Safety: MUTATING; writes live game/client state' restates both the opening line and the destructiveHint annotation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the description covers prerequisites, the required confirmation, the resulting operation-receipt, and a failure-recovery path (inspect tool-schema). Only a note on reversibility/rollback of the written live state is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters including the confirm safety acknowledgement and the typed argument shape. The description restates the signature and optionality but adds no format, default, or constraint detail beyond what the schema provides, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action (fire typed values) on a specific resource (an Actor communication channel) and situates it against siblings by naming the prerequisite get_comm_channel(id) and the related create_comm_channel capability. An agent can distinguish it from the many other fire-* and invoke-* siblings without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a concrete usage recipe ('Resolve get_comm_channel(id) and call Channel:Fire(...)'), states the confirm=true requirement, and lists prerequisites (active-client, explicit-mutation-approval, phase=act). It does not, however, compare against alternatives such as fire-connection, fire-signal, or fire-remote, so selection guidance is implicit rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fire-connectionFire one specific connection on a signalA
Destructive

MUTATES CLIENT STATE. Invokes a single connection's :Fire(...) (or :Defer(...) when defer=true) so that ONLY the one targeted handler runs with the supplied arguments — unlike fire-signal, which triggers every connection. Identify the connection by its index within getconnections(signal) (use list-signal-connections to find indices). This is client-side only and does not reach the server. Useful for isolating and testing a single suspected handler. The connection's Fire/Defer methods may be absent on foreign/C connections, so the call is pcall-guarded. Requires getconnections; degrades with a clear { error } if unavailable. Returns { Signal, Instance?, ConnectionIndex, Fired, Deferred } or { error }. Signature: { instancePath: string, signalName: any?, connectionIndex: any?, defer: any?, args: {{ kind: "string" | "number" | "boolean" | "nil" | "instance" | "raw", value: string | number | boolean? }}?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Capabilities: getconnections. Produces: created-handle, operation-receipt. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
argsNoOrdered list of arguments to pass to the targeted connection's handler. Omit or pass [] to fire with no arguments.
deferNoWhen true, call connection:Defer(...) to schedule the handler on the next resumption cycle instead of running it synchronously. When false (default), call connection:Fire(...) to run it immediately.
signalNameNoName of the RBXScriptSignal member on the instance (e.g. 'Touched', 'Changed'). Leave empty/omit if instancePath already evaluates to the signal.
instancePathYesLua expression resolving to the Instance that owns the signal (e.g. 'game.Workspace.Door'). Evaluated as `return <instancePath>`. Leave signalName empty if this expression already resolves to the signal itself.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
connectionIndexNoZero-based index of the target connection within getconnections(signal), matching the Index reported by list-signal-connections. Defaults to 0 (the first connection).

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructive=true, readOnly=false, idempotent=false, but the description adds far more: client-side only (does not reach the server), pcall-guarded because Fire/Defer may be absent on foreign/C connections, dependency on getconnections with a defined degradation path ({ error }), and the exact return shape. This is rich behavioral context beyond the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the critical 'MUTATES CLIENT STATE' warning and the core semantics in the first two sentences. It is longer than needed, with a duplicated inline signature plus phase/capability/verification boilerplate that partly repeats the schema, but the ordering keeps the important content first.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter mutating tool with no output schema, the description covers prerequisites, side effects, failure behavior, and the return object shape. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents instancePath, signalName, connectionIndex, defer, args, and threadContext in detail. The description's inline signature restates those fields and adds the pointer to list-signal-connections, which the schema's connectionIndex description already contains. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource and scope ('Invokes a single connection's :Fire(...)'), and explicitly distinguishes itself from the sibling fire-signal ('unlike fire-signal, which triggers every connection'). An agent can differentiate it without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('isolating and testing a single suspected handler'), names the alternative (fire-signal), and routes the agent to the prerequisites (list-signal-connections to find indices). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fire-lua-state-eventFire a LuaStateProxy's cross-state EventA
Destructive

WRITES LIVE GAME STATE. Resolve current/game/expression state and call state.Event:Fire(...typed arguments). Requires confirm=true because listeners may mutate live behavior. Signature: { state: any?, stateExpression: any?, arguments: any?, confirm: boolean?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval. Capabilities: getluastate. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
stateNoOptional state selector or reusable state reference returned by a discovery tool.current
confirmNoExplicit safety acknowledgement; must be true before this state-changing operation runs.
argumentsNoOptional ordered typed arguments forwarded to the selected operation.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
stateExpressionNoFor state='expression', a Luau expression resolving to a LuaStateProxy, Actor, or BaseScript.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructive=true/readOnly=false/idempotent=false, but the description adds substantial context beyond them: the rationale that listeners may mutate live behavior, an idempotency qualifier of 'contextual-write', a 'Produces: operation-receipt' result, a 'Verify with: assert-state' follow-up, and a failure-handling pointer. This is rich disclosure a caller needs before mutating live state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the critical fact ('WRITES LIVE GAME STATE') and uses compact labeled segments that are easy to scan. However, the mutation warning is repeated ('WRITES LIVE GAME STATE' and 'Safety: MUTATING; writes live game/client state') and the signature line duplicates a fully-covered schema, adding redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the description covers everything an agent needs: preconditions (active-client, confirm, approval), safety profile, the produced operation-receipt, a verification step, and a failure fallback to tool-schema. Nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all five parameters (including the enum and the arguments maxItems/typed-argument structure) are fully documented in the schema. The description's signature line merely restates the parameter names without adding semantic meaning beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource: firing a LuaStateProxy's cross-state Event, and specifies it resolves current/game/expression state before invoking state.Event:Fire with typed arguments. The explicit mention of LuaStateProxy/cross-state lets an agent distinguish it from sibling firing tools like fire-connection, fire-signal, and fire-remote.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete preconditions: confirm=true is required, plus 'Requires: active-client, explicit-mutation-approval' and a Phase=act label. It clearly states the context in which this runs, though it never names an alternative tool or a when-not-to-use condition, leaving routing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fire-proximity-promptTrigger a ProximityPromptA
Destructive

WRITES LIVE GAME STATE. Resolve a Luau expression to a ProximityPrompt and trigger it via the executor's fireproximityprompt, exactly as if the local player walked up and held the prompt's key to completion. This bypasses distance, hold-duration, line-of-sight and Enabled checks and fires the prompt's Triggered signal, so any server logic bound to it runs (open a door, buy an item, pick up an object). Use it to drive interaction-gated game flow during automation/testing. Guards that fireproximityprompt exists in this executor and that the target is actually a ProximityPrompt before firing; the fire call is pcall-guarded. WARNING: this mutates the running game and the effect may replicate to the server. Returns { Path, ok } or { error }. Signature: { path: string, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Capabilities: fireproximityprompt. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesLuau expression resolving to the ProximityPrompt to fire, e.g. 'game.Workspace.Door.Attachment.ProximityPrompt' or 'game.Workspace.Shop.BuyPart.ProximityPrompt'. Evaluated as `return <path>` and must resolve to an Instance whose ClassName is 'ProximityPrompt'.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructive=true and readOnly=false, but the description adds substantial context beyond them: which checks are bypassed, that the Triggered signal fires and server logic may run, that guards and a pcall wrap the call, and that the effect may replicate to the server. Return shape is also disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the critical 'WRITES LIVE GAME STATE' warning, and the core mechanics follow efficiently. However, the trailing metadata block (Phase, cost, idempotency, Capabilities, Produces, Requires, Verify, On failure) is boilerplate that partially duplicates the annotations and warning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return format ({ Path, ok } or { error }), failure guidance, guard behavior, and the mutation warning. An agent has everything needed to invoke and interpret the call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented in the schema. The description restates the signature { path, threadContext } but adds no syntax or constraints beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (trigger/fire) and resource (ProximityPrompt) with concrete semantics: 'exactly as if the local player walked up and held the prompt's key to completion.' This distinguishes it from siblings like fire-signal, fire-connection, fire-remote, and fire-click-detector, which fire different object types.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use it: 'Use it to drive interaction-gated game flow during automation/testing,' and clarifies the shortcut (bypasses distance/hold/los/Enabled checks). It lacks explicit naming of an alternative sibling tool for the same task, but the context of use is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fire-remoteFire a RemoteEvent / invoke a RemoteFunction / BindableA
Destructive

ACTS ON LIVE STATE — CAN REACH THE SERVER. Resolve a Luau expression to a remote object (RemoteEvent / UnreliableRemoteEvent / RemoteFunction / BindableEvent / BindableFunction) and call it with the chosen mode and arguments. Modes: FireServer / InvokeServer (RemoteEvent / RemoteFunction → travel to the SERVER), FireAllClients / FireClient (server→client, only valid from the server side), Fire (BindableEvent → local listeners), and Invoke (BindableFunction → local handler, returns a value). InvokeServer and Invoke return the call's return values. The call is pcall-guarded. WARNING: FireServer, InvokeServer and FireClient cross the network boundary and trigger REAL server-side / other-client logic (granting items, taking damage, purchases, etc.) — only use this to test a game you own/control, never to affect other players. Returns { Remote, Mode, ok, ReturnValues? } or { error }. Signature: { remotePath: string, mode: "FireServer" | "InvokeServer" | "FireAllClients" | "FireClient" | "Fire", args: {{ kind: "string" | "number" | "boolean" | "nil" | "raw", value: string | number | boolean? }}?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
argsNoOrdered list of arguments to send. For FireClient the FIRST argument must be the target Player (kind='raw', e.g. 'game.Players.SomePlayer'). Omit or pass [] for no arguments.
modeYesWhich method to call on the remote — must match the object's class. 'FireServer' (RemoteEvent / UnreliableRemoteEvent → server, no return). 'InvokeServer' (RemoteFunction → server, RETURNS the server's return values). 'FireAllClients' / 'FireClient' (RemoteEvent server→client; only valid when this code runs on the server). 'Fire' (BindableEvent → local listeners, same Lua VM). Note: a BindableFunction's Invoke method is not exposed here — this tool targets remotes plus BindableEvent:Fire.
remotePathYesLuau expression resolving to the remote, e.g. 'game:GetService("ReplicatedStorage").Remotes.BuyItem', 'game.ReplicatedStorage.Events.Damage', or 'getRemote()'. Evaluated as `return <remotePath>`. Must resolve to a RemoteEvent, UnreliableRemoteEvent, RemoteFunction, BindableEvent or BindableFunction.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations (destructiveHint=true, idempotentHint=false) by disclosing that the call is pcall-guarded, that specific modes cross the network boundary and fire REAL server logic (grants, damage, purchases), that InvokeServer/Invoke return values while FireServer does not, and by naming the return shape and the required-mutation-approval precondition.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The critical warning and purpose are front-loaded, and the mode list is compact and scannable. It is long and repeats mode semantics in both prose and the signature line, but every block (warning, modes, signature, safety, verification) carries actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex, network-crossing mutation tool with no output schema, the description supplies the return shape ({ Remote, Mode, ok, ReturnValues? } or { error }), the phase/cost/idempotency metadata, required preconditions, and a verification path (assert-state). Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, but the description reinforces meaning: the full signature, the per-mode enum behavior (which returns a value, which is server-only), and the FireClient first-arg-is-target-Player constraint. It adds framing value beyond the schema, though much of the mode/param detail is duplicated there.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (call/invoke) and resource (remote object: RemoteEvent / UnreliableRemoteEvent / RemoteFunction / BindableEvent / BindableFunction) resolved from a Luau expression. It enumerates the exact modes and their semantics, letting an agent distinguish it from neighbors like fire-connection, fire-signal, and fire-comm-channel without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear when-to-use context, including per-mode validity ('only valid from the server side') and an explicit when-not: use only to test a game you own/control, never to affect other players, since these calls trigger real server-side logic. It does not explicitly route to alternative sibling tools, but the conditions that select this tool are clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fire-signalFire all local connections of a signalA
Destructive

MUTATES CLIENT STATE. Invokes firesignal(signal, ...) to synchronously fire EVERY local Lua connection currently attached to an RBXScriptSignal, passing the supplied arguments. This is client-side only: it runs the handlers in your own client and does NOT send anything to the server or other players. Useful for exercising a game's own event handlers (e.g. simulating a 'Touched' or a custom BindableEvent) while debugging UI/gameplay logic, without performing the real physical action. NOTE: firesignal invokes the connection functions directly and does NOT respect a connection's Disabled state — disabled connections still run. (A real event, or set-connection-state, does respect Enabled/Disabled.) Requires the executor's firesignal; degrades with a clear { error } if unavailable. Returns { Signal, Instance?, Fired, ArgCount } or { error }. Signature: { instancePath: string, signalName: any?, args: {{ kind: "string" | "number" | "boolean" | "nil" | "instance" | "raw", value: string | number | boolean? }}?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Capabilities: getconnections, firesignal. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
argsNoOrdered list of arguments to pass to the connection handlers, mirroring the real arguments the event would normally fire with. Omit or pass [] to fire with no arguments.
signalNameNoName of the RBXScriptSignal member on the instance (e.g. 'Touched', 'Changed', 'OnClientEvent'). Leave empty/omit if instancePath already evaluates to the signal.
instancePathYesLua expression resolving to the Instance that owns the signal (e.g. 'game.Workspace.Door'). Evaluated as `return <instancePath>`. Leave signalName empty if this expression already resolves to the signal itself.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: states it runs only in the caller's client and does NOT reach the server/other players, discloses the important caveat that disabled connections still fire (contrasted with real events), notes the executor-capability dependency with a graceful { error } fallback, and names the return shape. Annotations only supply the safety flags, so this adds substantial non-redundant context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the most important constraint ('MUTATES CLIENT STATE') and organizes caveat/dependency/return info into distinct sentences. The trailing signature block duplicates the 100%-covered schema and is the one piece that does not earn its space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating tool with no output schema, the description covers safety, client-side scope, disabled-connection behavior, required capability, failure mode, and return shape. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the nested arg `kind` enum is already fully documented in the schema. The description's inline signature restates field names without adding format or edge-case detail beyond it, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('synchronously fire EVERY local Lua connection currently attached to an RBXScriptSignal'), and its scope ('client-side only') distinguishes it from fire-connection (single connection) and replicate-signal. An agent can identify it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete use context ('exercising a game's own event handlers ... while debugging UI/gameplay logic, without performing the real physical action') and names an alternative behavior ('A real event, or set-connection-state, does respect Enabled/Disabled'). It stops short of explicitly naming sibling tools like fire-connection, so the routing guidance is clear but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-active-clientShow which Roblox client this session targetsA
Read-onlyIdempotent

Report which live Roblox client THIS session currently resolves to (read-only), including its current selection and whether resolution is ambiguous or offline. Use it to confirm routing before running code, especially when multiple games are connected or after a rejoin. Signature: {}. Phase: observe; cost=low; idempotency=read-only. Requires: none. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered; the description restates it but also adds useful behavioral detail beyond the annotations, namely that the result reports current selection and can indicate ambiguous or offline resolution. The 'On failure: inspect tool-schema' pointer is also helpful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is correctly front-loaded, but the trailing metadata block is partly redundant: 'idempotency=read-only' and 'Safety: read-only' restate the readOnlyHint/idempotentHint annotations, and 'Requires: none' is derivable from the empty schema. 'Phase: observe; cost=low' does add value, so the padding is moderate rather than severe.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-argument read-only observation tool with no output schema, the description covers what it returns (client selection, ambiguity, offline state), when to reach for it, and where to look for field details, which is nearly everything an agent needs. Only the exact return shape is left to the tool-schema pointer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters (baseline 4) and schema coverage is 100%, so there are no parameter semantics to explain. The empty 'Signature: {}' correctly signals a no-argument call.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: it reports which live Roblox client THIS session resolves to, including selection state and ambiguity/offline status. The phrase 'THIS session currently resolves to' implicitly separates it from list-clients and select-client, but no sibling is named explicitly, so it stops short of the top score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear usage condition: 'Use it to confirm routing before running code, especially when multiple games are connected or after a rejoin.' That is concrete when-to-use context, but there is no explicit when-not guidance and no alternative tool (e.g. list-clients) is named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-actor-detailsInspect Actor(s) and their scriptsA
Read-onlyIdempotent

Drill into Actor instances (parallel-Luau VMs) and report their contents. Actors run code in isolated VMs outside the normal serial scheduler, so they are a common place to hide logic. Give an actorPath to inspect ONE Actor in detail, or omit it to summarize EVERY Actor (getactors). For each Actor this returns its name, where it lives (in tree / nil-parented / detached / inside CoreGui), how many descendants it has, and a list (capped at 30 per Actor) of the LuaSourceContainer scripts inside it — each with name, class, full path, and whether the script is currently executing (running, determined by membership in getrunningscripts). This complements list-actors by adding the running flag, the descendant count, and single-target resolution from a Luau expression. Requires getactors; getrunningscripts is used additionally when present (running falls back to false if it is unavailable). Returns a single Actor object when actorPath is given, otherwise { actorCount, truncated, actors: [...] }. Degrades with a clear error if getactors is missing or actorPath does not resolve to an Actor. Signature: { actorPath: string?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: getactors. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
actorPathNoOptional Luau expression resolving to a single Actor to inspect, e.g. 'workspace.MyModel.Actor' or 'getactors()[1]'. Evaluated as `return <actorPath>` and validated to be an Instance that IsA('Actor'). If omitted, ALL actors from getactors() are summarized.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the read-only/idempotent safety profile, but the description goes well beyond them: it discloses dependency on getactors, the additional use of getrunningscripts, the fallback to running=false when it is unavailable, the 30-script-per-Actor cap, and graceful degradation 'with a clear error' when getactors is missing or actorPath does not resolve.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and behavior are front-loaded and information-dense, but the trailing boilerplate block (Signature, Phase, cost, idempotency, Requires, Capabilities, Produces, Safety, On failure) is verbose metadata that partly duplicates the prose. Efficient, but not tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully carries the return contract, specifying a single Actor object when actorPath is given versus { actorCount, truncated, actors: [...] }, plus the per-Actor fields returned. Dependencies, failure behavior, and phase are all covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already fully documents actorPath (Luau expression, IsA('Actor') validation) and threadContext. The description restates the actorPath omit-vs-provide semantics but adds no syntax or format detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Drill into Actor instances (parallel-Luau VMs) and report their contents') and explicitly distinguishes itself from the sibling list-actors by noting it 'complements list-actors by adding the running flag, the descendant count, and single-target resolution'. An agent can tell what it does and how it differs from neighbors without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes between the two modes: 'Give an actorPath to inspect ONE Actor in detail, or omit it to summarize EVERY Actor', and names the alternative tool it complements (list-actors). The condition selecting each behavior is stated outright.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-anticheat-surfacesLightweight anti-cheat surface summaryA
Read-onlyIdempotent

Fast, lightweight defensive recon that summarizes the most common anti-cheat surfaces WITHOUT a heavy GC walk (complements, and does not duplicate, the deeper scan-hook-surfaces tool). It checks: (1) how many connections are attached to RunService's Heartbeat, Stepped, and RenderStepped — the signals anti-cheats most often use for per-frame validation loops — via getconnections; (2) whether game's raw metatable is locked (isreadonly on getrawmetatable(game)), which gates __index/__namecall hooking; (3) the count of nil-parented instances (getnilinstances), where detached watchdog scripts/objects often hide; and (4) any getgenv() global names that look anti-cheat-related (matching detect/ban/kick/anticheat/flag/cheat, case-insensitive), which can reveal an exploit's own loader or a leaked server-side guard name. Use this only as lightweight ambient context after execution-footprint-audit; it reports observable surfaces and does not prove detection or provide concealment. Each probe degrades gracefully and is pcall-guarded; missing executor functions are reported as unavailable rather than failing the call. Returns { runServiceConnections, gameMetatableReadonly, nilInstanceCount, suspiciousGlobals, notes } or { error }. Signature: { threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-observation, diagnostic-report. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only/idempotent/non-destructive, but the description adds real behavioral context: each probe is pcall-guarded and degrades gracefully, missing executor functions are reported as unavailable rather than failing, and it states what the tool does not do (no concealment, no proof of detection).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and the four probes are front-loaded and informative, but the trailing meta line ('Phase: observe; cost=medium; idempotency=read-only... Safety: read-only. On failure: inspect tool-schema...') is boilerplate that duplicates the annotations and template filler rather than earning its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description supplies the return shape ({ runServiceConnections, gameMetatableReadonly, nilInstanceCount, suspiciousGlobals, notes } or { error }), prerequisites (active-client), and failure behavior. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single threadContext parameter is fully documented in the schema, so the description adds no meaning beyond it — the 'Signature: { threadContext: number? }' line merely restates the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('summarizes the most common anti-cheat surfaces') and explicitly distinguishes itself from the sibling scan-hook-surfaces tool. The four enumerated probes make the scope unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('only as lightweight ambient context after execution-footprint-audit') and when-not ('does not prove detection or provide concealment'), and names the alternative it complements rather than duplicates.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-call-stackWalk the current Luau call stack (debug.info frames)A
Read-onlyIdempotent

Unwind the current Luau call stack from the executing thread and return one frame per level: { level, name, source, line }. Walks debug.info(level, 'nsl') (name / source / currentline) from level 1 outward, stopping at the first level that yields nil or after maxLevels frames. Use it to see who is calling the code you just ran — the chain of functions leading into the current execution — which is invaluable when reasoning about a hook callback or a deferred task's origin. Requires debug.info (or debug.getinfo) — type-guarded and pcall-wrapped, returning { error } on an executor that lacks it. Returns { frames } or { error }. Signature: { maxLevels: number?, threadContext: number?, timeoutMs: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-observation, operation-receipt. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
maxLevelsNoMaximum number of stack frames to walk before stopping (default 20).
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/idempotent/destructive=false, and the description adds real behavioral detail beyond them: the walk stops at the first nil level or after maxLevels frames, requires debug.info (type-guarded and pcall-wrapped), and returns { error } on executors lacking it. Failure and dependency behavior is disclosed, though no rate-limit or execution-cost specifics beyond cost=medium.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening sentences are tight and front-loaded, but the trailing block (Phase/cost/idempotency/Safety/On failure) largely duplicates the annotations (read-only, idempotent) and adds boilerplate. Useful content is present but not every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description specifies the return contract ({ frames } or { error }, per-frame shape), the failure mode, and the dependency requirement. For a low-complexity, zero-required-param read tool this is nearly complete; only cross-tool routing is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and each parameter is described in the schema, so the schema already does the heavy lifting. The description lists the signature and notes optionality but adds no syntax, format, or default nuance beyond what the schema provides; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (unwind), resource (the current Luau call stack from the executing thread), and return shape ({ level, name, source, line }), with the mechanism named (debug.info(level, 'nsl')). An agent can distinguish this from sibling stack/call-graph tools like get-thread-stack or build-call-graph.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear context for use — seeing who is calling the code you just ran, useful for hook callbacks or deferred task origins. However it never names the alternative tools (e.g. get-thread-stack, get-stack) or states when not to use it, so routing vs siblings is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-closure-constantsRead a closure's constant poolA
Read-onlyIdempotent

Resolve a Luau expression to a function and dump its constant pool via getconstants. Constants are the literal values the bytecode references — strings, numbers, table keys, global names, and embedded function/method names — so they are the fastest way to fingerprint what a closure does (e.g. spotting a 'FireServer' literal, a remote name, a damage value, or a URL) without reading its source. This is the by-reference companion to the GC-wide scanners (find-functions-by-constant / find-constants-xref): use it once you already hold the function (a remote handler, a metamethod, getsenv(script).fn, etc.). Each entry reports its 1-based Index, Luau Type, and an encoded Value. Requires the executor's getconstants (debug.getconstants); if unavailable a clean { error } is returned. Returns { Target, Info, Constants:[{Index,Type,Value}], Count } or { error }. Signature: { functionPath: string, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: debug closure primitives. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
functionPathYesLuau expression resolving to a function whose constants you want, e.g. 'getsenv(game.Players.LocalPlayer.PlayerScripts.Main).update', 'getrawmetatable(game).__namecall', or 'getconnections(game.ReplicatedStorage.Remote.OnClientEvent)[1].Function'. Evaluated as `return <functionPath>`.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description goes beyond them by disclosing the hard dependency on the executor's getconstants (debug.getconstants), the clean { error } fallback when unavailable, the prerequisite state (active-client, resolved-target), and a cost rating of medium.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with purpose, then usage, output shape, and metadata tags in a logical order. It is dense but somewhat long, and the Signature line partly duplicates the schema; nearly every sentence earns its place, but there is minor redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully specifies the return shape ({ Target, Info, Constants:[{Index,Type,Value}], Count } or { error }), the invocation dependency, and failure behavior. An agent has everything needed to call it and interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents functionPath with examples and threadContext fully. The description only restates the signature '{ functionPath: string, threadContext: number? }', adding little beyond what structured fields provide. Baseline 3 is correct when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Resolve a Luau expression to a function and dump its constant pool via getconstants.' It further defines what a constant is (strings, numbers, table keys, global names, embedded function labels), which distinguishes this tool from siblings like get-closure-upvalues or get-closure-protos.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly positions the tool as 'the by-reference companion to the GC-wide scanners (find-functions-by-constant / find-constants-xref): use it once you already hold the function.' This names the alternatives and the precise condition that selects this tool over them, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-closure-protosList a closure's nested function prototypesA
Read-onlyIdempotent

Resolve a Luau expression to a function and enumerate its nested prototypes (protos) via getprotos. Protos are the inner functions defined inside a closure — closures it creates, callbacks it registers, helper functions in its body. Walking protos lets you drill into a top-level script function to reach the exact inner handler you care about (e.g. an anonymous OnClientEvent callback) without scanning the whole GC, and to map a script's internal call structure. For each proto this returns the full __fnInfo (IsLua/IsC, Name, Source, ShortSource, LineDefined, NumParams, IsVararg, NumUpvalues, Pointer) so you can immediately feed a Pointer/path back into the other closure tools. Requires the executor's getprotos (debug.getprotos); if unavailable a clean { error } is returned. Returns { Target, Info, Protos:[__fnInfo...], Count } or { error }. Signature: { functionPath: string, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: debug closure primitives. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
functionPathYesLuau expression resolving to a function whose nested prototypes you want, e.g. 'getscriptclosure(game.Players.LocalPlayer.PlayerScripts.Main)', 'getsenv(script).init', or 'getrawmetatable(game).__namecall'. Evaluated as `return <functionPath>`.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint, idempotentHint, destructiveHint false), so the bar is lower, and the description adds real context: it depends on the executor's debug.getprotos and returns a clean { error } when unavailable. It also discloses the return shape and the __fnInfo fields, which goes beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Core purpose and behavior are front-loaded in the first sentences. The trailing metadata block (Phase/cost/idempotency/Requires/Capabilities/Produces/Safety/On failure) is somewhat formulaic and 'Safety: read-only' repeats the annotation, but it remains scannable and does not bury the main point.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully carries the return contract: { Target, Info, Protos:[__fnInfo...], Count } or { error }, plus the enumerated __fnInfo fields and how to feed a Pointer back into other closure tools. Nothing an agent needs to invoke and interpret it is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents functionPath and threadContext thoroughly, including example expressions. The description only restates the signature and the notion of resolving a Luau expression, adding little beyond the structured schema; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Resolve a Luau expression to a function and enumerate its nested prototypes') and defines the concept ('inner functions defined inside a closure'). It implicitly separates itself from GC-wide scanning and from the other closure tools, but never names the very close sibling get-function-protos, so an agent must infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear when-to-use framing: drill into a top-level script function to reach a specific inner handler (e.g. an anonymous OnClientEvent callback) 'without scanning the whole GC', and to map a script's internal call structure. It implies the alternative (GC-wide scanning) but does not explicitly state when-not to use it or name a competing sibling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-closure-upvaluesRead a closure's upvalues (captured variables)A
Read-onlyIdempotent

Resolve a Luau expression to a function and dump its upvalues via getupvalues. Upvalues are the variables a closure captured from its enclosing scope — the live state a function carries with it (config tables, cached remotes, counters, references to other functions, flags). Inspecting them reveals hidden state that source code alone does not show, which is invaluable when reverse-engineering a handler or locating a kill-switch/flag to flip. This is the by-reference companion to find-upvalue-xref. Each entry reports its 1-based Index, Luau Type, and an encoded Value; the same Index is what set-closure-upvalue mutates. Requires the executor's getupvalues (debug.getupvalues); if unavailable a clean { error } is returned. Returns { Target, Info, Upvalues:[{Index,Type,Value}], Count } or { error }. Signature: { functionPath: string, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: debug closure primitives. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
functionPathYesLuau expression resolving to a function whose upvalues you want, e.g. 'getsenv(game.Players.LocalPlayer.PlayerScripts.Main).step', 'getrawmetatable(game).__index', or 'getconnections(workspace.Part.Touched)[1].Function'. Evaluated as `return <functionPath>`. Captured variables are editable by Index via set-closure-upvalue.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive), yet the description adds real behavioral context beyond them: it depends on the executor's getupvalues (debug.getupvalues) and returns a clean { error } when unavailable, plus phase/cost/prerequisite metadata. No contradiction with annotations; a solid addition without being exhaustive on limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and the key 'why', and most sentences earn their place. It runs long with a trailing metadata block (Phase/cost/Capabilities/'On failure: inspect tool-schema...') that is somewhat boilerplate, slightly diluting conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description supplies the full return contract ({ Target, Info, Upvalues:[{Index,Type,Value}], Count } or { error }), the field meanings, dependency requirements, and error behavior — everything an agent needs to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both parameters (functionPath, threadContext) already carry detailed schema descriptions. The description only restates the signature and adds no syntax or meaning beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ('Resolve a Luau expression to a function and dump its upvalues via getupvalues') and distinguishes this tool from find-upvalue-xref ('by-reference companion') and set-closure-upvalue (consumes the same Index). It does not, however, differentiate from the very close sibling get-function-upvalues, leaving a real selection ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear motivating context: use it to reveal hidden closure state when reverse-engineering a handler or locating a kill-switch/flag to flip. It names a related tool (find-upvalue-xref, set-closure-upvalue) but offers no explicit when-not conditions or a full alternative matrix.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-comm-channelResolve an existing Actor communication channelB
Read-onlyIdempotent

Call get_comm_channel(id), falling back to the MCP channel registry, and return serializable channel/Event metadata. Signature: { id: string, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Capabilities: create_comm_channel. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYestext value for id.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description still adds real value with 'Requires: active-client' (a prerequisite) and the registry-fallback behavior. However, the 'Capabilities: create_comm_channel' line is noise that sits oddly on a read operation and could confuse routing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The lead sentence is front-loaded and useful, but the body degenerates into template metadata ('cost=medium; idempotency=read-only ... Produces: structured-observation. Safety: read-only') with visible redundancy between idempotency=read-only, Safety=read-only, and the readOnlyHint annotation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 100% schema coverage and full annotations, the prerequisites and fallback are adequately covered, and no output schema means return format needn't be spelled out. Still, the promised 'serializable channel/Event metadata' is left vague and no discriminating context against sibling channel tools is provided.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description's 'Signature: { id: string, threadContext: number? }' merely restates what the schema already documents and adds no format, constraint, or semantic detail beyond it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: resolve an Actor comm channel and return channel/Event metadata, with an explicit fallback to the MCP channel registry. The scope is clear enough to separate it from create-comm-channel and fire-comm-channel, though those siblings are never named.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a 'Phase: observe' label and a failure-handling pointer to tool-schema, but never states when to use this versus create-comm-channel, fire-comm-channel, or comm-channel-monitor. The agent must infer routing from names alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-connection-constantsDump the constants of a connection's handler functionA
Read-onlyIdempotent

Disassemble the Lua function bound to ONE connection on an RBXScriptSignal (selected by zero-based index) and return its constant pool via debug.getconstants. Constants include the literal strings, numbers and referenced globals/methods baked into the closure — invaluable for reverse-engineering what a hidden event handler does (e.g. which RemoteEvent names or HTTP endpoints it touches). Each constant is passed through a value encoder so tables/functions/instances render as readable descriptors instead of raw pointers. Requires getconnections and debug.getconstants; returns a clear { error } if either is missing or the connection has no Lua Function. Returns { Signal, Instance?, ConnectionIndex, Function, Constants: [...], Count }. Signature: { instancePath: string, signalName: any?, connectionIndex: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: debug closure primitives, getconnections. Produces: structured-observation, created-handle. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
signalNameNoName of the RBXScriptSignal member on the instance (e.g. 'Touched', 'Changed', 'OnClientEvent'). Leave empty/omit if instancePath already evaluates to the signal.
instancePathYesLua expression resolving to the Instance that owns the signal (e.g. 'game.Workspace.Door'). Evaluated as `return <instancePath>`. Leave signalName empty if this expression already resolves to the RBXScriptSignal itself.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
connectionIndexNoZero-based index of the connection whose handler function to disassemble, as reported by list-signal-connections. 0 is the first connection. Must be < ConnectionCount or a clear { error } is returned.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only/idempotent/non-destructive, and the description reinforces this while adding real behavioral detail: prerequisite functions, a clear { error } on missing deps or no Lua Function, the value-encoder output behavior, and the exact return shape. The one oddity is 'Produces: created-handle' for an otherwise read-only operation, but that is not a direct contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and rationale in the first two sentences, then layers prerequisites, return shape, and metadata. The trailing block (Phase, cost, Requires, Capabilities, Produces, Safety, On failure) is dense and somewhat templated but each line carries usable information, so it earns most of its length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description supplies the full return shape ({ Signal, ConnectionIndex, Function, Constants, Count }), failure behavior, prerequisites, and the caveat that non-Lua or missing connections produce { error }. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters including the zero-based connectionIndex and the signalName/instancePath relationship. The description largely restates this (signature line, zero-based index note) rather than adding new semantics, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: disassemble the Lua function bound to ONE connection on an RBXScriptSignal and return its constant pool via debug.getconstants. The scope (one connection, selected by index) and the underlying primitive distinguish it from siblings like get-connection-upvalues, get-closure-constants, and get-connection-protos.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for when the tool matters (reverse-engineering hidden event handlers, e.g. which RemoteEvent names or HTTP endpoints are touched) and names prerequisites (getconnections, debug.getconstants) plus the upstream tool (list-signal-connections) for the index. It doesn't explicitly contrast itself against the closest alternatives (e.g. get-connection-upvalues or trace-connection-function), so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-connection-infoInspect a single signal connection in detailA
Read-onlyIdempotent

Return the full Connection metadata for ONE connection on an RBXScriptSignal, selected by zero-based index. Reports Index, Enabled, LuaConnection, ForeignState, whether it has a Function/Thread, the connected thread's status (when present), and — for Lua connections — the handler function's Source script, Name, LineDefined, NumParams, IsVararg and What. Use this to zoom in on a specific listener after list-signal-connections shows you the index of interest. Requires getconnections; degrades with a clear { error } if unavailable or if the index is out of range. Returns { Signal, Instance?, ConnectionCount, Connection: {...} }. Signature: { instancePath: string, signalName: any?, connectionIndex: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: getconnections. Produces: structured-observation, created-handle. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
signalNameNoName of the RBXScriptSignal member on the instance (e.g. 'Touched', 'Changed', 'OnClientEvent'). Leave empty/omit if instancePath already evaluates to the signal.
instancePathYesLua expression resolving to the Instance that owns the signal (e.g. 'game.Workspace.Door'). Evaluated as `return <instancePath>`. Leave signalName empty if this expression already resolves to the RBXScriptSignal itself.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
connectionIndexNoZero-based index of the connection to inspect, as reported by list-signal-connections. 0 is the first connection. Must be < ConnectionCount or a clear { error } is returned.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/no-open-world, and the description still adds real behavioral context: it requires getconnections, degrades with a clear { error } when unavailable or the index is out of range, and enumerates the returned fields and their conditions (e.g. thread status only when present).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then usage, failure behavior, and return shape before the structured metadata tail. The signature/phase/cost block is somewhat boilerplate-heavy but each line is informative; nothing is truly wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully documents the return shape ({ Signal, Instance?, ConnectionCount, Connection: {...} }) and enumerates the nested fields. It also covers failure modes and prerequisites, so an agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already carries parameter meaning; the description's 'Signature' line merely restates the four parameters and adds no syntax or format detail beyond the schema. Baseline 3 is correct when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource+scope: 'Return the full Connection metadata for ONE connection on an RBXScriptSignal, selected by zero-based index.' It explicitly contrasts with the sibling list-signal-connections (bulk listing vs. single-connection zoom), so an agent can distinguish them without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States exactly when to use it ('after list-signal-connections shows you the index of interest') and names the alternative tool. Prerequisites (getconnections) and the out-of-range/unavailable conditions are also documented, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-connection-protosList the nested child functions of a connection's handlerA
Read-onlyIdempotent

Enumerate the inner (proto) functions defined inside the Lua function bound to ONE connection on an RBXScriptSignal (selected by zero-based index) via debug.getprotos. Protos are the nested closures a handler creates — callbacks, deferred tasks, helper lambdas — so this maps out the sub-functions you may want to hook or inspect next. Each proto is reported via the standard function descriptor (Source, Name, LineDefined, NumParams, IsVararg, What, Pointer). Requires getconnections and debug.getprotos; returns a clear { error } if either is missing or the connection has no Lua Function. Returns { Signal, Instance?, ConnectionIndex, Function, Protos: [...], Count }. Signature: { instancePath: string, signalName: any?, connectionIndex: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: getconnections. Produces: structured-observation, created-handle. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
signalNameNoName of the RBXScriptSignal member on the instance (e.g. 'Touched', 'Changed', 'OnClientEvent'). Leave empty/omit if instancePath already evaluates to the signal.
instancePathYesLua expression resolving to the Instance that owns the signal (e.g. 'game.Workspace.Door'). Evaluated as `return <instancePath>`. Leave signalName empty if this expression already resolves to the RBXScriptSignal itself.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
connectionIndexNoZero-based index of the connection whose handler protos to enumerate, as reported by list-signal-connections. 0 is the first connection. Must be < ConnectionCount or a clear { error } is returned.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so the description is not obligated to restate safety. It still adds value by disclosing preconditions (requires getconnections and debug.getprotos, active-client, resolved-target) and failure behavior (clear { error } if dependencies are missing, no Lua Function, or index out of range). Minor redundancy in restating 'idempotency=read-only' and 'Safety: read-only'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in a strong opening sentence, and the return shape and error behavior follow. The trailing metadata block (Phase, cost, Requires, Capabilities, Produces, Safety, On failure) is partly boilerplate that duplicates annotations, keeping it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully carries the return contract ({ Signal, Instance?, ConnectionIndex, Function, Protos, Count }), the per-proto descriptor fields, the error case, and all preconditions. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so each parameter (signalName, instancePath, threadContext, connectionIndex) is already fully documented in the schema, including defaults and bounds. The description only echoes the zero-based index selection and repeats the signature, adding no new semantic detail. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise verb+resource (enumerate nested proto functions inside the Lua handler bound to one connection) and distinguishes itself from siblings like get-connection-info, get-function-protos and scan-proto-functions by scoping to a single connection selected by index. An agent can identify the tool without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains the purpose ('maps out the sub-functions you may want to hook or inspect next'), states the connectionIndex must come from list-signal-connections, and lists prerequisites (getconnections, debug.getprotos). It stops short of naming when NOT to use it or pointing at the closer alternatives such as get-function-protos or get-closure-protos.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-connection-upvaluesDump the upvalues of a connection's handler functionA
Read-onlyIdempotent

Read the captured upvalues of the Lua function bound to ONE connection on an RBXScriptSignal (selected by zero-based index) via debug.getupvalues. Upvalues are the variables a closure captured from its enclosing scope — typically shared state, config tables, cached services or other functions — so this reveals the live context a hidden event handler operates on. Each upvalue is reported with its typeof and a readable encoded value (instances become full names, tables/functions become descriptors). Requires getconnections and debug.getupvalues; returns a clear { error } if either is missing or the connection has no Lua Function. Returns { Signal, Instance?, ConnectionIndex, Function, Upvalues: [...], Count }. Signature: { instancePath: string, signalName: any?, connectionIndex: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: debug closure primitives, getconnections. Produces: structured-observation, created-handle. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
signalNameNoName of the RBXScriptSignal member on the instance (e.g. 'Touched', 'Changed', 'OnClientEvent'). Leave empty/omit if instancePath already evaluates to the signal.
instancePathYesLua expression resolving to the Instance that owns the signal (e.g. 'game.Workspace.Door'). Evaluated as `return <instancePath>`. Leave signalName empty if this expression already resolves to the RBXScriptSignal itself.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
connectionIndexNoZero-based index of the connection whose handler upvalues to read, as reported by list-signal-connections. 0 is the first connection. Must be < ConnectionCount or a clear { error } is returned.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive), so the description's value-add is its disclosure of encoding behavior (instances become full names, tables/functions become descriptors), tool dependencies (getconnections, debug.getupvalues), and explicit error behavior when those are missing. This goes meaningfully beyond the annotations, though the trailing 'Safety: read-only' line merely repeats them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose and the meaning of upvalues before moving to encoding, dependencies, return shape, and the signature. It is dense but well-ordered; the trailing metadata block (Phase/cost/Capabilities/Produces/Safety) is somewhat redundant against the annotations and slightly inflates length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema available, the description fully specifies the return shape ({ Signal, Instance?, ConnectionIndex, Function, Upvalues, Count }), the required capability set, and the failure modes. An agent has everything needed to invoke it correctly and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters including the zero-based connectionIndex and the instancePath/signalName relationship. The description adds the signature and reinforces the index semantics but contributes little beyond what the schema already states, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb and resource: 'Read the captured upvalues of the Lua function bound to ONE connection on an RBXScriptSignal', plus the mechanism (debug.getupvalues). The 'ONE connection' scoping and zero-based index distinguish it from sibling upvalue tools like get-function-upvalues and get-closure-upvalues without needing to open either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides useful context for why one would call it ('reveals the live context a hidden event handler operates on') and names preconditions ('Requires: active-client, resolved-target') and failure conditions. However, it never explicitly names an alternative tool or a when-not-to-use condition, so the usage guidance remains implied rather than directive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-connector-diagnosticsConnector and runtime self-reportA
Read-onlyIdempotent

Probe the live connector/runtime from inside the Roblox client and return a self-report describing the execution environment. Every read is pcall-guarded, so an unavailable function is reported as null rather than failing the call. Returns { threadIdentity, executor={ name, version }, hasWebSocket, gcInfoKB, genvKeyCount, hasReg, hasRenv, capabilities={ getgc, hookfunction, getnilinstances, getactors, getluastate, run_on_actor, clonefunction, newcclosure, cloneref, compareinstances, getcallingscript, getscriptclosure, getsenv, getfenv, mouse1click, keypress, getcallbackvalue, firesignal, getconnections } }. Use this to confirm what the connector can do before relying on reflection/hooking tools, or to diagnose why a client is behaving unexpectedly. Signature: { threadContext: number? }. Phase: verify; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-observation, created-handle, diagnostic-report. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, and closed-world, so the safety profile is covered. The description adds real value by disclosing the pcall-guard semantics (unavailable functions report as null rather than failing) and the Requires: active-client precondition, but it repeats the safety/idempotency facts already in annotations via the 'Phase/cost/idempotency/Safety' metadata block.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded and the long return-shape enumeration is genuinely informative. However the trailing metadata ('Phase: verify; cost=medium; ... Produces: ... On failure: inspect tool-schema...') is somewhat boilerplate and dilutes the signal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of describing the return value and does so exhaustively, plus it names the precondition (active-client) and the failure mode. A minor oddity is listing 'created-handle' as a product for a read-only tool, but overall an agent has enough to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single optional threadContext parameter is fully documented in the schema. The description only restates the signature ('{ threadContext: number? }') without adding syntax or behavioral meaning beyond it, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (probe) and resource (live connector/runtime) and describes the self-report returned. It distinguishes itself from a category of tools ('before relying on reflection/hooking tools'), but never names a concrete sibling like get-executor-info or test-capabilities that an agent would weigh against it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives two clear usage contexts: confirming connector capabilities before reflection/hooking work, and diagnosing unexpected client behavior. It does not state when NOT to use it or name an explicit alternative among the many overlapping diagnostics siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-console-outputGet the roblox developer console output from the Roblox Game ClientB
Read-onlyIdempotent

Read recent Roblox developer console logs from the active client. Use limit and logsOrder to control volume and ordering. Signature: { limit: any?, logsOrder: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of results to return (default: 50, to avoid overwhelming output)
logsOrderNoThe order of the logs to return (default: NewestFirst)NewestFirst
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, and the description's 'Safety: read-only' / 'idempotency=read-only' merely restate them. It does add genuinely new context: 'cost=medium', the 'active-client' prerequisite, and an on-failure fallback to tool-schema. Return format and pagination behavior remain undisclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, which is good, but several lines are redundant boilerplate: the 'Signature' line duplicates the schema verbatim, and the safety/idempotency clauses repeat the annotations. It is short but not every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-optional-param read tool there is no output schema, so the description should convey return shape; 'Produces: structured-observation' is too vague to tell the agent what log entries look like. The prerequisite and fallback guidance partially compensate, making this adequate but with a real gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three params (with defaults and the logsOrder enum) are already documented. The description only summarizes 'limit and logsOrder' and omits threadContext entirely, adding essentially nothing beyond the schema — the expected baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read recent Roblox developer console logs from the active client'), so the agent knows exactly what it returns. It does not, however, distinguish itself from the near-identical sibling capture-log-output, which likely overlaps in intent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Requires: active-client' gives a precondition, and 'Use limit and logsOrder to control volume and ordering' implies usage. But there is no when-to-use guidance relative to alternatives like capture-log-output or get-remote-spy-logs, leaving the agent to infer the choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-custom-assetGet an rbxasset:// URL for a workspace file (UNC getcustomasset)A
Destructive

Turn a file in the executor's workspace into a content URL (rbxasset://...) usable as an asset id for images, sounds, meshes, etc. inside the game. NOTE: this is executor-side file I/O — the connector runs INSIDE the executor, so the path is relative to the executor's workspace directory on the host machine, NOT the Roblox game. This is marked state-mutating because on most executors getcustomasset COPIES the file into the Roblox content cache as a side effect. Requires the UNC function getcustomasset(path) -> string. The call is type-guarded and pcall-wrapped: if getcustomasset is missing you get { error = 'getcustomasset is not available in this executor.' }, and any failure returns { error = }. Returns { path, asset } or { error }. Signature: { path: string, threadContext: number?, timeoutMs: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Produces: structured-observation, operation-receipt. Verify with: assert-state. Safety: MUTATING; writes executor workspace filesystem. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesPath to the file within the executor workspace to expose as an asset, e.g. 'images/logo.png'.
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: explains the state-mutation is a side-effect file copy into the Roblox content cache, requires the UNC getcustomasset function, is type-guarded and pcall-wrapped, and specifies exact error shapes and return values. This is rich behavioral context an agent cannot get from the schema alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose and the critical executor-vs-game caveat in the first two sentences. The trailing metadata block (phase, cost, idempotency, produces/verify) is somewhat templated and partly restates annotation hints, but remains dense and mostly informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates fully by documenting the return shape ({ path, asset } or { error }), failure modes, and prerequisites (active-client, explicit-mutation-approval). An agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so path, threadContext, and timeoutMs are already documented in the schema. The description restates the signature but adds no new syntax, format, or constraint detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: turn a workspace file into an rbxasset:// content URL usable as an asset id. It clearly distinguishes itself from the game-side file tools by stressing the path is relative to the executor's workspace, not the Roblox game.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for when to reach for it (exposing an executor-workspace file as an in-game asset id) and explains the executor-vs-game path semantics. It does not name alternative sibling tools or state explicit exclusions, so it stops short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-executor-infoIdentify the executor and its headline capabilitiesA
Read-onlyIdempotent

In-game probe that reports WHICH executor is hosting the connector and a small capability map. Calls identifyexecutor() (guarded — it may return a name and/or version, or nothing) with getexecutorname()/getexecutorinfo() fallbacks, then flags a handful of marquee functions (getgc, hookfunction, getnilinstances, getactors, …) true/false. Call this first on a new client to confirm you are on a full-featured executor before reaching for reflection/hooking tools. Signature: {}. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-observation, diagnostic-report. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses the internal probe chain: identifyexecutor() is guarded and may return a name, version, or nothing, with getexecutorname()/getexecutorinfo() fallbacks. It also notes required active-client context, read-only safety, diagnostic output, and on-failure guidance to inspect tool-schema, all without contradicting the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded: the purpose comes first, followed by mechanism, usage, and metadata. It is somewhat over-stuffed with fragments such as 'Signature: {}.' and repeats the read-only safety in both the idempotency and safety lines, but the structure remains clear and useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-parameter, read-only executor probe with no output schema, the description covers purpose, internal fallback behavior, first-use guidance, required client state, and failure handling. Nothing an agent needs in order to select and invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and the description states 'Signature: {}', matching the empty input schema. No additional parameter semantics are needed, so the baseline score for a no-parameter tool is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: an in-game probe that reports which executor is hosting the connector and a small capability map. It distinguishes the tool from sibling reflection/hooking tools by positioning it as a first-call capability check. An agent can tell exactly what this tool returns without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says to call this first on a new client before reaching for reflection/hooking tools, which gives clear usage context and relative ordering. However, it does not name a specific alternative sibling tool or state when not to call it, so it stops short of the highest bar.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-fast-flagRead a Roblox FastFlag value (sUNC getfflag)A
Read-onlyIdempotent

Read the current value of a Roblox engine FastFlag by name via the executor's getfflag(name). FastFlags (FFlag/DFFlag/FInt/etc.) are the runtime feature toggles the Roblox client reads at startup; getfflag returns the current value as a string, or nil when the flag is unknown. Requires getfflag. The call is type-guarded and pcall-wrapped: if getfflag is missing you get { error = 'getfflag is not available in this executor.' }. Returns { name, value } or { error }. Signature: { name: string, threadContext: number?, timeoutMs: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesThe FastFlag name, e.g. 'DFIntTaskSchedulerTargetFps' or 'FFlagDebugDisplayFPS'.
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/idempotentHint/destructiveHint=false, so the safety profile is covered; the description goes beyond them by disclosing that the call is type-guarded and pcall-wrapped, the exact error object when getfflag is absent, and that unknown flags return nil. The cost/phase metadata adds further context, though the closing 'inspect tool-schema' line is boilerplate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose and error behavior are front-loaded and useful, but the 'Phase/cost/idempotency/Safety' block partially duplicates the annotations (read-only, idempotent) and the final 'On failure: inspect tool-schema...' sentence is generic filler that does not earn its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly supplies the return shape ({ name, value } or { error }) and the missing-dependency failure mode, plus prerequisites. It is nearly complete for a 3-param read tool; only per-parameter semantics for threadContext/timeoutMs lean entirely on the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3: the schema already documents name, timeoutMs, and threadContext including defaults. The description only restates the signature and gives one example flag name, adding little beyond structured fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Read the current value of a Roblox engine FastFlag by name') and how it is obtained (executor's getfflag(name)). The scoping detail — FastFlags are runtime feature toggles read at startup — makes it clearly separable from the sibling set-fast-flag and from generic value reads like read-path-value.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Preconditions are stated ('Requires getfflag', 'Requires: active-client'), which implies when the call is usable, but there is no explicit guidance on when to prefer this over alternatives such as set-fast-flag or the broader runtime-inspection tools. Usage is inferable rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-fps-capRead the current FPS cap (sUNC getfpscap)A
Read-onlyIdempotent

Read the executor's current render frame-rate cap via getfpscap(). Returns the cap as a number; 0 means uncapped. Requires getfpscap. The call is type-guarded and pcall-wrapped: if getfpscap is missing you get { error = 'getfpscap is not available in this executor.' }. Returns { fpsCap } or { error }. Signature: { threadContext: number?, timeoutMs: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnly/idempotent/non-destructive), yet the description adds substantive behavior beyond them: the call is type-guarded and pcall-wrapped, the exact failure shape ({ error = 'getfpscap is not available...' }), the return shape ({ fpsCap } or { error }), and a cost=medium hint. This is unusually rich disclosure for a read tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded and information-dense, with the return contract stated early. It is somewhat over-stuffed with template-style metadata ('Phase: observe; cost=medium; idempotency=read-only') and a trailing 'inspect tool-schema' pointer that is meta rather than descriptive, but nothing is seriously wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return contract itself (the numeric cap, 0 = uncapped, or an error object) and names the required runtime capability and active-client precondition. An agent has everything needed to invoke and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters (threadContext, timeoutMs) are already documented in the schema. The description only restates the signature ('{ threadContext: number?, timeoutMs: number? }') without adding format, default, or interaction semantics, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (read) and resource (render frame-rate cap via getfpscap()), and clarifies the return semantics ('Returns the cap as a number; 0 means uncapped'). It is immediately distinguishable from its obvious sibling set-fps-cap.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It supplies prerequisites ('Requires getfpscap', 'Requires: active-client') and a phase label ('observe'), which implies when it fits. But it never names alternatives (e.g., set-fps-cap or get-render-stats) or states an explicit when-to-use/when-not condition, leaving routing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-function-envDump a function's environment (getfenv)A
Read-onlyIdempotent

Resolve a Luau expression to a function and dump its function environment table via getfenv. This is the _ENV the function reads its globals from: the global table the closure sees (often the script's environment, or a sandboxed proxy). Use it to learn what globals a handler can reach, to spot a sandbox/proxy environment, or to discover sibling functions you can then inspect or hook by reference. Read-only. Requires getfenv. Returns { Target, KeyCount, Truncated, Keys } (keysOnly) or { Target, KeyCount, Truncated, Entries } or { error }. Signature: { functionPath: string, keysOnly: any?, maxKeys: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
maxKeysNoMaximum number of environment keys/entries to return before truncating (default 200).
keysOnlyNoWhen true (default) return only the environment key names — cheap and safe. When false also serialize each value's type and a scalar/string preview via __encVal.
functionPathYesLuau expression resolving to a function, e.g. 'getsenv(game.Players.LocalPlayer.PlayerScripts.Main).update', 'getrawmetatable(game).__namecall', or a function found via the gc-scan tools. Evaluated as `return <functionPath>`.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint, idempotentHint, destructiveHint=false), so the bar is lower, and the description still adds real context: it names the hard dependency on getfenv being available, states the prerequisites (active-client, resolved-target), and describes the two return shapes. That is useful disclosure beyond the annotations, though the repeated 'Read-only / idempotency=read-only / Safety: read-only' adds little.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action correctly in the first sentence, but the tail is padded with metadata blocks ('Phase: observe; cost=medium; idempotency=read-only', 'Produces: structured-observation', 'Safety: read-only') that restate the annotations and return shape. Those lines cost tokens without adding decision-relevant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly takes on return-value disclosure by enumerating both response shapes plus the error case. It also names prerequisites and a failure fallback (inspect tool-schema). Complete enough to call correctly, with only the sibling ambiguity as a residual gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents functionPath syntax, keysOnly behavior, maxKeys truncation, and threadContext defaults. The description's signature line and its '(keysOnly)' / 'Entries' hint only map return shape to keysOnly, which the schema already states. This is baseline-3 territory: no meaningful semantics added beyond structured fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: resolve a Luau expression to a function and dump its function environment table via getfenv, with a clarifying gloss that this is the _ENV the closure reads globals from. That is far more specific than a tautology. It does not, however, differentiate itself from the very similar siblings 'dump-function-env', 'get-script-env', or 'set-function-env', so an agent cannot tell from the text alone which environment tool to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives three concrete when-to-use cases: learn what globals a handler can reach, spot a sandbox/proxy environment, and discover sibling functions to inspect or hook by reference. That is clear usage context rather than implied usage. It names no explicit alternative tool or exclusion condition, which is what separates this from a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-function-hashHash a Luau closure's bytecodeA
Read-onlyIdempotent

Resolve a function and return getfunctionhash(fn), useful for stable comparison and change detection. Volt only supports bytecode hashing for Luau closures; C closures return a clean executor error. Signature: { functionPath: string, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: debug closure primitives. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
functionPathYesLuau expression resolving to the target function.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, non-destructive, non-open-world, so the safety profile is covered. The description adds real value beyond that: cost=medium, required preconditions, and specifically that C closures return a clean executor error rather than a hash. It does not describe the return format in depth, but the added failure behavior is genuinely useful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded well, but the trailing metadata block (Phase/cost/idempotency/Safety) is largely template and partially duplicates the annotations ('idempotency=read-only' and 'Safety: read-only' restate readOnlyHint/idempotentHint). The 'On failure: inspect tool-schema...' pointer is useful but generic, so several tokens do not fully earn their place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only hashing primitive with no output schema, the description covers purpose, preconditions, failure behavior, and safety adequately; annotations carry the safety profile. It is nearly complete, only lacking a hint of the returned value's shape/format, which is a minor gap for a hash tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both functionPath and threadContext are already documented in the schema. The description restates the signature ({ functionPath, threadContext? }) without adding format, syntax, or defaulting detail beyond it. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: resolve a function and return getfunctionhash(fn) for stable comparison and change detection. The title reinforces it as bytecode hashing of a Luau closure. It is clear what it does, though it doesn't explicitly contrast with near-siblings like get-script-hash, disassemble-function, or get-closure-protos.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete use context ('stable comparison and change detection') and a meaningful boundary condition: Volt only hashes Luau closures, C closures return a clean executor error. It also lists prerequisites (active-client, resolved-target). No explicit sibling alternatives are named, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-function-protosGet function protosC
Read-onlyIdempotent

Find a function by query and dump nested proto debug info. Signature: { query: string, maxProtos: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Capabilities: debug closure primitives. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYessearch text used to filter and rank bounded results.
maxProtosNoOptional hard result/work budget used to keep output and runtime bounded.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered structurally. The description mostly repeats these ('idempotency=read-only', 'Safety: read-only') while adding a little genuine context: 'Requires: active-client' and 'cost=medium'. It adds marginal value beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded, but the body is padded with formulaic metadata tags (Phase, cost, idempotency, Capabilities, Produces, Safety) that largely duplicate the annotations and schema. The final sentence defers to tool-schema, so it consumes space without informing the call.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, fully-documented 3-parameter tool with no output schema, the description is adequate: it names the action, prerequisites, and safety profile. It nonetheless omits what 'nested proto debug info' returns and how it differs from the many proto/closure siblings, leaving an agent with a complete-but-unrouted definition.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents query, maxProtos (default 30, budget), and threadContext. The description merely restates the signature ('{ query: string, maxProtos: any?, threadContext: number? }') without adding format, semantics, or examples beyond the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Find a function by query and dump nested proto debug info.' An agent can tell it retrieves proto debug data. However, it does not differentiate itself from close siblings like get-closure-protos, get-connection-protos, lookup-function, or scan-proto-functions, so selection between them is left to guesswork.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only usage constraint is the prerequisite 'Requires: active-client.' There is no guidance on when to prefer this over sibling proto/closure lookups, no when-not condition, and no alternative named. Phase/cost tags imply observation use but do not route the agent among alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-function-upvaluesGet function upvaluesB
Read-onlyIdempotent

Find matching functions in getgc() and dump their upvalues for live reverse engineering. Signature: { query: string, sourceQuery: string?, maxFunctions: any?, maxUpvalues: any?, includeCClosures: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Capabilities: debug closure primitives. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesCase-insensitive match against function name/source.
maxUpvaluesNoMaximum upvalues per function (default: 30).
sourceQueryNoOptional additional source filter.
maxFunctionsNoMaximum number of matched functions to inspect (default: 5).
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
includeCClosuresNoInclude C closures (default: false).

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the 'Safety: read-only' and 'idempotency=read-only' lines merely repeat them. The genuinely additive facts are 'Requires: active-client', 'cost=medium', and 'Produces: structured-observation', which are useful but thin; 'On failure: inspect tool-schema' deflects rather than explains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence and the metadata tags are compact, but the Signature block duplicates the schema verbatim and the trailing 'On failure' sentence is boilerplate that consumes space without adding actionable detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only tool with a fully documented schema and complete safety annotations, the description covers purpose, precondition, and cost adequately. However, with no output schema it gives only 'structured-observation' and does not characterize the returned upvalue data, leaving a modest gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents all six parameters, and the baseline is 3. The inline Signature block just restates the same parameter names and types, adding no meaning (e.g. no explanation of how query and sourceQuery interact) beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Find matching functions in getgc() and dump their upvalues,' naming the GC enumeration source and the object acted on (upvalues). This distinguishes it reasonably from siblings like get-closure-upvalues and find-upvalue-sharing, though it never names an alternative to route between them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for live reverse engineering' plus 'Phase: observe' and 'Requires: active-client' imply the context of use, but there is no explicit when-to-use/when-not statement and no mention of sibling alternatives such as get-closure-upvalues or find-upvalue-sharing that an agent might otherwise pick.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-game-infoGet current Roblox place and universe metadataA
Read-onlyIdempotent

Read identifying metadata for the game the active client is in: PlaceId, GameId (universe), JobId (server instance), PlaceVersion, the place/universe name where readable, and the current player count. Every read is pcall-guarded, so a restricted field is reported as null rather than failing the call. Use this to confirm which game/server you are attached to before running place-specific code. Signature: {}. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-observation, diagnostic-report. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description still adds genuinely new behavior: every read is pcall-guarded and restricted fields come back as null rather than failing the call, plus the active-client prerequisite and a medium cost signal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The field list and pcall/null behavior are front-loaded and earn their place, but the trailing meta-template is padded: "Safety: read-only" and "idempotency=read-only" merely restate the annotations, and "Phase: observe; cost=medium; Produces: ..." is boilerplate that dilutes an otherwise tight description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden and does so by naming the fields an agent will receive, along with the null-on-restricted-field fallback and the failure pointer to tool-schema. For a zero-parameter read tool this is close to complete, with only the absence of any sibling routing as a gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters and the schema is an empty object at 100% coverage, so there is nothing for the description to disambiguate; baseline 4 applies. The description correctly describes the empty signature rather than inventing parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence gives a specific verb (Read) and resource (identifying metadata for the game the active client is in) and enumerates exactly what is returned: PlaceId, GameId, JobId, PlaceVersion, names, player count. It does not, however, name or contrast any sibling such as get-place-details or get-game-state, so an agent must infer the boundary itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Use this to confirm which game/server you are attached to before running place-specific code" states a concrete usage context, and the Requires: active-client line adds a precondition. There is no explicit when-not or named alternative, which keeps it below a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-game-stateGet the default game LuaStateProxyA
Read-onlyIdempotent

Call getgamestate() and return Id/IsActorState/Event/Actors metadata plus a reusable registry Reference. Signature: { threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds value beyond them: 'cost=medium', the active-client prerequisite, the observable output class (structured-observation), and a failure-recovery pointer to tool-schema. It repeats 'read-only' twice but adds genuine context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose then compact key=value metadata. Slightly redundant in restating read-only ('idempotency=read-only' plus 'Safety: read-only') and echoing the signature, but no wasted sentences overall.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description enumerates the returned metadata (Id/IsActorState/Event/Actors plus registry Reference) and states the active-client prerequisite, so an agent has enough to call and interpret it. The only gap is a lack of sibling routing guidance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single optional threadContext parameter is fully documented in the schema. The description's 'Signature: { threadContext: number? }' restates the schema without adding format or defaulting semantics, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a concrete verb and resource (return game-state metadata: Id/IsActorState/Event/Actors plus a registry Reference), which is specific enough to distinguish it from get-game-info and get-lua-state. It stops short of explicitly naming the sibling it is not, so an agent still has to infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It supplies operational context ('Phase: observe', 'Requires: active-client') that implies when the tool is valid, but never states when to prefer it over get-game-info, get-lua-state, or get-lua-state-actors. No explicit exclusions or alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-gui-textRead the text of a GUI elementA
Read-onlyIdempotent

Resolve a Luau expression to a single GUI Instance and return whichever text properties it actually exposes: .Text (the editable/displayed string on TextLabels, TextButtons and TextBoxes), .ContentText (the rendered text after rich-text/markup processing) and .PlaceholderText (the grey hint shown by an empty TextBox). Use this to inspect exactly what a label says or what a player has typed before acting on it. Each property read is independently pcall-guarded, so missing properties are simply omitted rather than erroring — a Frame with no text fields returns just { Path }. Pair with list-gui-elements to first discover the path. Returns { Path, Text?, ContentText?, PlaceholderText? } or { error }. Signature: { path: string, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesLuau expression resolving to the GUI Instance to read, e.g. 'game.Players.LocalPlayer.PlayerGui.Shop.PriceLabel' or 'game:GetService("Players").LocalPlayer.PlayerGui.Login.UsernameBox'. Evaluated as `return <path>`.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint, idempotentHint, destructiveHint=false), yet the description adds real behavioral detail: each property is independently pcall-guarded, absent properties are omitted rather than erroring, and a textless Frame returns just { Path }. It also documents the exact return shape and failure form.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and the key property semantics are front-loaded in the opening sentences with zero waste. The trailing metadata lines (Phase, cost, Requires, Produces, Safety, On failure) are boilerplate-heavy but structured and skimmable, slightly diluting an otherwise tight description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although no output schema exists, the description fully specifies the return shape ({ Path, Text?, ContentText?, PlaceholderText? } or { error }) and the prerequisites (active-client, resolved-target). Combined with the safety and idempotency metadata, an agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both path (with Luau expression examples and evaluation semantics) and threadContext are already documented in the schema. The description restates the signature but adds little parameter meaning beyond what the schema provides, matching the baseline for full coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (resolve/read) and resource (GUI text properties) and enumerates exactly which properties are returned (.Text, .ContentText, .PlaceholderText). It also names the sibling list-gui-elements as the discovery counterpart, so an agent can distinguish it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context ('inspect exactly what a label says or what a player has typed before acting on it') and a workflow prerequisite ('Pair with list-gui-elements to first discover the path'). It stops short of an explicit when-not or a direct contrast with the write sibling set-gui-text, so it is strong but not fully explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-hidden-uigethui — inspect the executor's hidden protected UI containerA
Read-onlyIdempotent

Return a shallow tree of the executor's hidden, protected GUI container via gethui(). This is the parent that executors hand back for ScreenGuis they want kept away from CoreGui/PlayerGui and shielded from the game's anti-cheat — the usual home of cheat menus, ESP layers, and overlays. The tool walks gethui():GetChildren() and returns a { name, class, children } tree capped at depth 3 with a per-node child cap, reporting how many children were truncated. Requires gethui — type-guarded and pcall-wrapped, returning { error } when missing or on failure. Returns { root: { name, class, childCount, children } } or { error }. Signature: { maxChildren: number?, threadContext: number?, timeoutMs: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
maxChildrenNoMaximum children shown per node (default 50). Extra children are counted in truncatedChildren.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only/idempotent/non-destructive, so the bar is lower, yet the description adds substantial depth: walk behavior (gethui():GetChildren()), depth cap of 3, per-node child cap with truncation counts, and the guard/error contract ({ error } when gethui is missing). That is actionable behavioral detail beyond the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and return shape, and most sentences carry real information. It is dense but a few metadata fragments (phase/cost/idempotency/safety triad) and the "inspect tool-schema for exact fields" line are somewhat redundant with the annotations and schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return contract itself and does so fully: { root: { name, class, childCount, children } } or { error }, plus the depth cap and truncation reporting. Nothing an agent needs to call and interpret this tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three optional parameters and their defaults; baseline is 3. The description restates the signature and notes defaults/constraints but adds no syntax or interaction semantics beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Return a shallow tree of the executor's hidden, protected GUI container via gethui()") and explains exactly what that container is — the parent executors hand back for anti-cheat-shielded ScreenGuis. This clearly distinguishes it from siblings like find-hidden-guis and list-gui-elements, which survey surfaces rather than walk the gethui() root.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives real invocation context: "Phase: observe; cost=medium; idempotency=read-only. Requires: active-client." and the failure path ("Requires gethui — type-guarded and pcall-wrapped"). However, it never names an alternative or a when-not condition, so the agent must infer when to prefer this over find-hidden-guis or list-gui-elements.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-hwidRead the host hardware ID (sUNC gethwid)A
Read-onlyIdempotent

Read the host machine's hardware identifier (HWID) via the executor's gethwid(). This is the stable per-machine fingerprint executors commonly use for key-system/license binding. Requires gethwid. The call is type-guarded and pcall-wrapped: if gethwid is missing you get { error = 'gethwid is not available in this executor.' }. Returns { hwid } or { error }. Signature: { threadContext: number?, timeoutMs: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnly/idempotent/non-destructive, but the description adds real value beyond them: it is type-guarded and pcall-wrapped, cites the exact error string returned when gethwid is unavailable, notes cost=medium, and states the active-client prerequisite. It does not, however, describe timeout or thread-identity behavior in any depth.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and return shape are front-loaded, and the sentences are dense but information-bearing. Minor redundancy appears where 'idempotency=read-only', 'Safety: read-only', and the annotations repeat each other.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-param read tool with no output schema, the description cnumbers the return contract ({ hwid } or { error }), the executor dependency, and the failure mode, which is sufficient. It lacks any cross-reference to related lookup tools, a small gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (threadContext, timeoutMs) are already documented in the schema. The description merely reprints the signature with no added semantics or format guidance. Baseline 3 applies when the schema carries the load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read the host machine's hardware identifier (HWID) via the executor's gethwid()') and immediately explains its use case (stable per-machine fingerprint for key/license binding). No sibling tool reads HWID, so the agent can distinguish it easily.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the use case (license/key binding), but never explicitly says when to prefer this over alternatives such as get-executor-info or get-game-info, nor any when-not conditions. Prerequisites ('Requires gethwid', 'Requires: active-client') are stated, which is more than nothing but not routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-instance-countsGame-size profile: instance counts by ClassNameA
Read-onlyIdempotent

In-game shape/size profile. Walks game:GetDescendants() once (pcall-guarded, capped) and tallies every Instance by its ClassName, returning the heaviest classes first. This is the quickest way to understand how big and how 'shaped' a place is — e.g. tens of thousands of Parts, a forest of UI Frames, a swarm of scripts, or an unusual pile of a single odd class. The descendant walk is capped (maxScan, default 200000) and sets truncated=true if it hits the cap (the counts then reflect only what was scanned). The returned class list is capped to topN entries. Requires nothing beyond a live game; everything is guarded. Returns { totalInstances, scanned, truncated, distinctClasses, topClasses: [{ class, count }] } (sorted by count desc) or { error }. Signature: { topN: any?, maxScan: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-observation, diagnostic-report. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
topNNoHow many of the heaviest ClassName buckets to return, sorted by count descending (default 25). The full distinctClasses count is always reported even when the list is trimmed to topN.
maxScanNoMaximum number of descendants to visit (default 200000). The walk stops at this cap and sets truncated=true; raise it for a huge place if you need exact totals, lower it to bound cost.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, non-destructive behavior, so the bar is lower. The description still adds real value beyond them: the pcall-guarded single descendant walk, the maxScan cap, and the truncated=true flag that means counts reflect only what was scanned. It also states read-only safety and error return, though it doesn't quantify expected cost or latency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded and the examples are useful, but the tail packs redundant schema information (signature, defaults) and metadata boilerplate (phase, cost, produces, requires) that mostly repeats the structured fields already present.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden and does so: it specifies the return object shape ({ totalInstances, scanned, truncated, distinctClasses, topClasses }) and the error case, plus the active-client requirement and truncation semantics. Nothing an agent needs to invoke or interpret it is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description restates maxScan (default 200000) and the topN trim, which the schema already documents in comparable detail, adding little beyond it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: walks game:GetDescendants() once and tallies every Instance by ClassName, returning heaviest classes first. This is clearly distinguishable from siblings like get-instance-tree or search-instances, which serve navigation rather than aggregate profiling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context for when it is useful ('the quickest way to understand how big and how shaped a place is'), with concrete examples (parts, UI frames, scripts). It does not name alternative tools or state when not to use it, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-instance-propertiesRead an instance's common properties and attributesA
Read-onlyIdempotent

Resolve a single instance from a dotted path (e.g. 'game.Workspace.Part') and read a useful core set of its properties plus all of its attributes. Each property is pcall-read independently, so ones that don't exist on the class are simply omitted rather than failing the call. Returns { path, fullName, className, properties, attributes } with each property as { type, value }, or { error } if the path does not resolve. For listing many instances use get-instance-tree instead. Signature: { path: string }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesDotted path to the instance, starting at 'game' (e.g. 'game.Players.LocalPlayer.PlayerGui').

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/idempotentHint/destructiveHint=false, so safety is covered. The description adds real non-obvious behavior: each property is pcall-read independently so non-existent properties are omitted rather than failing the call, and an unresolvable path returns { error }. It also notes cost=medium. Strong, though it does not discuss the 'core set' selection criteria.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and scope, then a concrete example, then the return contract. The trailing boilerplate block ('Phase: observe; cost=medium; idempotency=read-only', 'On failure: inspect tool-schema...') partially restates annotation data but is compact and scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return contract itself: { path, fullName, className, properties, attributes }, per-property { type, value }, and the { error } case. Combined with prerequisites and the sibling pointer, an agent has everything needed to call and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'path' parameter is fully documented in the schema, so the baseline is 3. The description adds only a redundant example ('game.Workspace.Part') that duplicates the schema's own 'game.Players.LocalPlayer.PlayerGui' example, contributing no new semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (resolve + read) and resource (a single instance's core properties and attributes), with a concrete dotted-path example. It explicitly distinguishes itself from the sibling get-instance-tree, which handles the many-instance case, so an agent can select correctly without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the alternative explicitly ('For listing many instances use get-instance-tree instead') and gives prerequisites ('Requires: active-client, resolved-target'). The single-instance vs. many-instance boundary is stated outright, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-instance-treeWalk a depth-limited instance hierarchyA
Read-onlyIdempotent

Return a nested { name, class, children } tree under an instance resolved from a dotted path (default 'game'), capped by maxDepth and maxChildren so large containers don't overwhelm the output. Each node lists how many children were truncated. Use this for broad structure exploration; use get-instance-properties to read one instance's values. The path is pcall-resolved, so a bad path returns { error } rather than failing. Signature: { path: string?, maxDepth: number?, maxChildren: number? }. Phase: observe; cost=high; idempotency=read-only. Requires: active-client, resolved-target. Produces: bounded-candidates, structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoDotted path to the root instance (e.g. 'game.Workspace').game
maxDepthNoMaximum traversal depth (1-20, default 3).
maxChildrenNoMaximum children shown per node (1-500, default 50).

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnly/idempotent/non-destructive, but the description adds real behavioral context beyond them: per-node truncation counts, pcall-resolved paths that return { error } instead of failing, cost=high, and required state (active-client, resolved-target). That is exactly the kind of operational disclosure annotations cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core is front-loaded and clear, but several lines are redundant with structured fields: the Signature line duplicates the input schema, and 'idempotency=read-only' plus 'Safety: read-only' repeat the readOnlyHint/idempotentHint annotations. The Phase/Requires/Produces/On-failure block is useful but verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, yet the description supplies the return shape (nested { name, class, children }) and the per-node truncation reporting, plus error behavior and prerequisites. An agent has everything needed to invoke and interpret the result correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so defaults, ranges, and the dotted-path example are already documented in the schema. The description's 'Signature' line mostly restates those fields; the only added nuance is that the path is pcall-resolved. Baseline 3 is correct when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (return/walk), resource (instance hierarchy under a dotted path), the exact return shape ({ name, class, children }), and the bounding parameters. It also distinguishes itself from get-instance-properties, so an agent can tell the two apart without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit routing: 'Use this for broad structure exploration; use get-instance-properties to read one instance's values.' The when-to-use condition and the named alternative are both present, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-local-player-infoRich snapshot of the local player, character, and humanoid (read-only)A
Read-onlyIdempotent

Capture a single, comprehensive read-only snapshot of THIS client's own player (Players.LocalPlayer), the character model it currently controls, and that character's Humanoid. Use this as the first call when debugging anything about 'me' — health/death state, movement (WalkSpeed/JumpPower), spawn position, rig type, or simply to confirm the LocalPlayer's Name/UserId/Team before targeting other tools. It is a pure read: it mutates nothing and fires no remotes. Every field is independently pcall-guarded and reported as nil when absent (e.g. while dead/respawning the character or humanoid may be missing), so a partial state never errors the whole call. Returns { ok, player, character } where: player = { Name, UserId, DisplayName, AccountAge, Team } (Team is the Team's Name or nil). character = { Name, Health, MaxHealth, WalkSpeed, JumpPower, JumpHeight, MoveMagnitude (magnitude of Humanoid.MoveDirection), RigType (tostring of Humanoid.RigType), Position ({x,y,z} of the HumanoidRootPart) } — any of which may be nil. Requires only the base Roblox API; no special executor capabilities. Returns { error } only if the Players service itself cannot be read. Signature: { threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the readOnly/idempotent/destructive annotations: it discloses that it 'mutates nothing and fires no remotes,' that every field is independently pcall-guarded and reported nil when absent (dead/respawning), that it requires only the base Roblox API with no executor capabilities, and that it returns { error } only if the Players service is unreadable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose, usage rule, and return shape are front-loaded and earn their place, but the trailing boilerplate ('Phase: observe; cost=medium; idempotency=read-only; Requires: active-client; Produces...; On failure: inspect tool-schema...') restates annotations and adds bulk without new information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully documents the return contract ({ ok, player, character } with per-field breakdown including nil-ability), plus error and environment requirements. Nothing an agent needs to call and interpret this tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single optional threadContext parameter with 100% schema description coverage, so the schema already carries its meaning. The description only restates the signature ({ threadContext: number? }) without adding semantics, which matches the baseline-3 case for full schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Capture a single, comprehensive read-only snapshot of THIS client's own player (Players.LocalPlayer), the character model it currently controls, and that character's Humanoid.' The scope ('THIS client's own player') clearly distinguishes it from siblings like get-players and discover-character.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a strong when-to-use rule ('Use this as the first call when debugging anything about me — health/death state, movement, spawn position, rig type...') and a routing hint ('before targeting other tools'). It does not name a specific alternative tool or a when-not condition, so it stops short of the explicit alternatives required for a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-lua-stateGet the current, Actor, or script LuaStateProxyA
Read-onlyIdempotent

Call getluastate() for the current state, or getluastate(target) for an Actor/BaseScript expression. Returns serializable Id/IsActorState/Event/Actors metadata plus a reusable registry Reference. Signature: { targetExpression: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Capabilities: getluastate. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
targetExpressionNoOptional validated input for target expression.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/idempotentHint/destructiveHint=false, so the safety profile is covered, but the description adds genuinely new context: phase=observe, cost=medium, produces=structured-observation, and the active-client prerequisite. The redundancy ('idempotency=read-only' and 'Safety: read-only' restate annotations) slightly dilutes it, and it defers failure behavior to tool-schema rather than explaining it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core behavior is front-loaded and clear, but the middle is a run of template-style key=value fragments (phase/cost/idempotency/capabilities/produces/safety) that partly duplicate the annotations, and the trailing 'inspect tool-schema' sentence is generic boilerplate. It is structured but not tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden and does so ('serializable Id/IsActorState/Event/Actors metadata plus a reusable registry Reference'). For a 2-optional-param read tool this is close to complete; only the failure/invocation specifics are outsourced to tool-schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the schema's own description for targetExpression is vague ('Optional validated input for target expression') while the description clarifies it is an Actor/BaseScript expression. That is real added meaning beyond the schema, though the signature line itself only repeats parameter names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Call getluastate() for the current state, or getluastate(target)...'), and the return shape ('Id/IsActorState/Event/Actors metadata plus a reusable registry Reference') makes the output concrete. It is distinguishable from list-lua-states and new-lua-state-proxy by verb, but never names those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives usage context ('Phase: observe', 'Requires: active-client') and explains the two invocation forms (no-arg vs target expression), which implies when it applies. However, it never states when to prefer this over list-lua-states, get-lua-state-actors, or new-lua-state-proxy, so the agent must infer the alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-lua-state-actorsList Actors associated with one LuaStateProxyA
Read-onlyIdempotent

Resolve current/game/expression state, call LuaStateProxy:GetActors(), and return at most 200 Actor paths. Signature: { state: any?, stateExpression: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Capabilities: getactors, getluastate. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
stateNoOptional state selector or reusable state reference returned by a discovery tool.current
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
stateExpressionNoFor state='expression', a Luau expression resolving to a LuaStateProxy, Actor, or BaseScript.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, non-destructive, closed-world, so the safety profile is covered. The description adds real behavioral context beyond them: a hard result cap ('at most 200 Actor paths'), cost=medium, the active-client prerequisite, and a failure-handling pointer to tool-schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and organized into labeled fields (Signature, Phase, Requires, Capabilities, Produces, Safety, On failure), so an agent can scan it quickly. 'Safety: read-only' duplicates readOnlyHint and the signature repeats the schema, minor waste in an otherwise tight definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description still characterizes the return (a capped list of Actor paths), states the prerequisite, and routes the agent to tool-schema on failure. It is largely self-sufficient for a read-only observation tool, though it omits what an Actor path looks like and any pagination behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the enum, defaults, and threadContext bounds are fully documented in the schema. The description restates the signature and the state values (current/game/expression) but adds no syntax or semantics beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb chain (Resolve state, call LuaStateProxy:GetActors(), return Actor paths) and scopes the resource to Actors associated with one LuaStateProxy, which meaningfully narrows it versus the sibling list-actors/list-script-actors. It does not name those siblings explicitly, so it stops short of full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a prerequisite ('Requires: active-client') and a phase label ('Phase: observe'), which imply when it fits. However it never states when to prefer this over list-actors, list-script-actors, or get-lua-state, leaving alternative selection to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-memory-statsLua memory + GC object census (by type)A
Read-onlyIdempotent

In-game memory health probe. Reports the Lua VM's current heap usage via gcinfo() (in KB, guarded) and takes a census of the garbage collector by walking getgc(true) and counting objects by type — function, table, thread, userdata, and other. Optionally adds engine-level memory from game:GetService('Stats') (GetTotalMemoryUsageMb and a few notable categories) when available. Use this to gauge how heavy the client is, to spot a runaway table/closure leak (an unusually large gcObjectCount or byType.table), or to take a baseline before/after running an exploit. The GC walk is capped (maxScan, default 200000) and sets truncated=true if it hits the cap. Requires getgc for the census (reports gcObjectCount=nil if absent); gcinfo and Stats are optional. Returns { luaMemoryKB, gcObjectCount, truncated, byType, engineMemory } or { error }. Signature: { maxScan: any?, threadContext: number? }. Phase: observe; cost=high; idempotency=read-only. Requires: active-client. Produces: structured-observation, diagnostic-report. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
maxScanNoMaximum number of GC objects to visit during the census (default 200000). The walk stops at this cap and sets truncated=true; raise it on a very large game if you need an exact count, lower it to bound cost.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the read-only/idempotent annotations, it discloses real behavior: requires getgc (else gcObjectCount=nil), gcinfo and Stats are optional/degraded gracefully, the walk is capped at maxScan (default 200000) setting truncated=true, and a runaway leak manifests as large gcObjectCount or byType.table. This is rich, non-obvious operational context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with what it measures and why, then constraints. The trailing metadata tail (Signature, Phase, Requires, Produces, Safety, On failure) is somewhat verbose and partly redundant with the annotations (idempotency=read-only, Safety: read-only), but most sentences earn their place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully documents the return shape ({ luaMemoryKB, gcObjectCount, truncated, byType, engineMemory } or { error }), the guard conditions, and the cap behavior, so an agent has everything needed to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so the baseline is 3, but the description adds meaning: it explains maxScan's default, the cap semantics, and the truncated flag consequence, plus that threadContext is optional. It enriches rather than merely restates the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource: reports the Lua VM heap via gcinfo() and takes a GC census by walking getgc(true) and counting objects by type. It clearly distinguishes itself from neighboring memory/render tools like measure-memory and get-render-stats by specifying exactly what is measured and returned.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete when-to-use contexts (gauge client heaviness, spot runaway table/closure leaks, take baseline before/after exploits). It does not explicitly name alternative sibling tools such as measure-memory or filter-gc, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-metamethodRead a single metamethod off an object's raw metatableA
Read-onlyIdempotent

Resolve a Luau expression to ANY value (table, Instance, userdata, etc.), grab its raw metatable via getrawmetatable (bypassing __metatable locks), and read ONE named metamethod from it (e.g. __index, __namecall, __newindex, __call, __tostring). When the metamethod is a function this deep-dumps it: its info (Lua/C, source, line, params, upvalue count via getinfo) PLUS its constants (getconstants) and upvalues (getupvalues). This is the targeted counterpart to get-metatable: use it to drill into the exact function Roblox's security layer routes through — for example reading __namecall off getrawmetatable(game) to find the C closure that backs FireServer/InvokeServer, then inspecting its constants for method-name strings. For non-function metamethods (locked __metatable strings, __index tables, etc.) it returns the encoded value. Requires getrawmetatable; getconstants/getupvalues are best-effort (omitted if the executor lacks them or the metamethod is a C closure). Returns { Target, Method, Type, Function?, Constants?, Upvalues?, Value? } or { error } when there is no metatable, the method is absent, or getrawmetatable is missing. Signature: { objectPath: string, method: string, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: getrawmetatable. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
methodYesName of the metamethod key to read from the raw metatable, e.g. '__index', '__namecall', '__newindex', '__call', '__tostring', '__metatable'. Read off the metatable as mt[method] (does NOT invoke it).
objectPathYesLuau expression resolving to the object whose metatable holds the metamethod, e.g. 'game', 'game.Players.LocalPlayer', 'getrawmetatable(game)', or 'getgenv().SomeProxy'. Evaluated as `return <objectPath>`.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/no-destructive, but the description adds substantial context beyond them: __metatable lock bypass, the getrawmetatable dependency, best-effort getconstants/getupvalues that may be omitted for C closures, and explicit failure cases (no metatable, absent method, missing getrawmetatable).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then the example, then requirements/return shape. It is dense and every clause carries information, though the trailing Signature/Phase/Cost/Capabilities/Safety/On-failure metadata is somewhat boilerplate and lengthens the block.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description spells out the return shape ({ Target, Method, Type, Function?, Constants?, Upvalues?, Value? } or { error }), the dependency requirements, and the best-effort caveats. An agent has everything needed to call and interpret this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds meaning about the method parameter by explaining the divergent behavior based on its value (function metamethods get deep-dumped; non-function ones return an encoded value). It also restates the signature and confirms the read semantics for objectPath.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource chain (resolve expression → grab raw metatable → read one named metamethod) and names the exact sibling it is the targeted counterpart to (get-metatable). An agent can distinguish it from get-metatable, inspect-closure, and get-closure-constants without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use it versus the alternative: 'targeted counterpart to get-metatable: use it to drill into the exact function Roblox's security layer routes through.' It also supplies a concrete motivating scenario (reading __namecall off getrawmetatable(game) to find the closure behind FireServer/InvokeServer).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-metatableGet an object's raw metatableA
Read-onlyIdempotent

Resolve a Luau expression to ANY value (table, Instance, userdata, etc.) and dump its raw metatable via getrawmetatable (bypassing __metatable locks). For each metamethod it reports the key, value type, and — for function metamethods like __index/__namecall/__newindex — the connected function's source/line/params. Also reports whether the metatable is read-only and whether getmetatable is locked (a string __metatable). Use this to understand how a table/instance is protected or proxied, or to find the __namecall/__index that game security routes through. Unlike inspect-instance-metatable this works on any value, not just Instances. Requires getrawmetatable; returns { Target, TargetType, HasMetatable, ReadOnly, LockedMetatableValue, Metamethods } or { error }. Signature: { objectPath: string, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: getrawmetatable. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
objectPathYesLuau expression resolving to the object whose metatable you want, e.g. 'game', 'game.Players.LocalPlayer', 'getrawmetatable(game)', or 'getgenv().SomeTable'. Evaluated as `return <objectPath>`.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover safety (readOnly/idempotent/non-destructive), and the description adds substantial context beyond them: it bypasses __metatable locks, requires getrawmetatable, reports read-only state and locked-metatable values, and discloses prerequisites (active-client, resolved-target). This is a genuinely informative behavioral profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and clearly structured, but the trailing metadata block (Phase/cost/idempotency/Safety) partly repeats what annotations already state (idempotentHint, readOnlyHint). Slight redundancy, otherwise tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description enumerates the return shape ({ Target, TargetType, HasMetatable, ReadOnly, LockedMetatableValue, Metamethods } or { error }) plus failure handling ('inspect tool-schema'). Nothing an agent needs to call or interpret this tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (objectPath, threadContext) are already documented with examples and evaluation semantics in the schema. The description restates the signature but adds no syntax or constraint detail beyond what the schema provides, matching the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('dump its raw metatable via getrawmetatable') and explicitly contrasts scope with the sibling inspect-instance-metatable ('this works on any value, not just Instances'). An agent can distinguish it from the many other metatable/closure tools without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete when-to-use motivations ('understand how a table/instance is protected or proxied', 'find the __namecall/__index that game security routes through') and names the alternative that does not apply (inspect-instance-metatable, Instances only). Routing is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-module-sourceGet module sourceB
Read-onlyIdempotent

Decompile a specific ModuleScript by path. Signature: { modulePath: string, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: decompile. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
modulePathYesLuau expression or dotted path resolving to the target script/module.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds genuinely new context (cost=medium, Phase: observe, Produces: structured-observation) but also restates what annotations already say ("idempotency=read-only", "Safety: read-only"), and the failure guidance just redirects to tool-schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence, and the remaining metadata is compact semicolon-separated key-value pairs. It is slightly redundant (idempotency=read-only duplicating the annotation) but earns most of its space and is easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description carries the burden of describing the return value, yet "Produces: structured-observation" is vague about what a decompiled-source result actually looks like. Safety and prerequisites are covered, but the agent is left guessing about output shape for a decompile operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both modulePath and threadContext are already fully documented in the schema. The description's "Signature" line repeats the parameter names and types without adding format, default, or constraint meaning beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Decompile a specific ModuleScript by path") with clear scope, distinguishing it from siblings like get-script-content or get-script-bytecode that return raw/bytecode rather than decompiled source. It does not explicitly name those siblings, so differentiation is inferential rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides prerequisites ("Requires: active-client, resolved-target") and a phase/cost hint ("Phase: observe; cost=medium"), which implies when it fits in a workflow. However, it names no alternative tool and gives no when-not guidance, so an agent must infer that this is the decompilation path versus get-script-content or disassemble-function.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-nil-instancesList nil-parented (hidden) instancesA
Read-onlyIdempotent

Enumerate instances whose Parent is nil (getnilinstances). Nil-parenting is a classic hiding spot: a remote, script, GUI, or other object is kept alive by a reference but is unreachable from the game tree, so it never shows up in the Explorer or normal descendant scans. This returns the total count, a byClass breakdown (ClassName -> count) computed across ALL nil instances, and a capped samples list (each with class, name, and full path) for inspection. Requires getnilinstances; degrades with a clear error otherwise. Signature: { limit: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax number of sample instances to return in the samples list (default 500, max 3000). The byClass counts and total always cover every nil instance regardless of this limit.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, and the description goes well beyond them: it discloses the return shape (total count, byClass breakdown over ALL nil instances, capped samples), the getnilinstances dependency, graceful degradation with a clear error, cost phase, and active-client requirement. This is rich behavioral context an agent cannot get from the annotations alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is front-loaded with the purpose and the key scoping/hiding-spot rationale, and the return-shape summary is efficient. It loses a little to redundancy: the signature duplicates the schema and 'Safety: read-only' / 'idempotency=read-only' restate the annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description itself explains the returned structure (total, byClass, samples with class/name/path), the dependency and fallback behavior, and the operational requirements, so an agent has everything needed to invoke and interpret it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already fully documented, including the note that byClass/total ignore the limit. The description's 'Signature: { limit: any?, threadContext: number? }' merely restates the schema and adds no new semantic detail, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The definition gives a precise verb+resource ('Enumerate instances whose Parent is nil') and defines the exact condition that makes an instance nil-parented, so an agent can distinguish this from generic hidden-instance scans without opening the schema. The added explanation of why nil parenting is a hiding spot makes the scope unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It supplies clear context for when this is valuable ('classic hiding spot', unreachable from the game tree, invisible to Explorer/descendant scans), which is strong usage framing. However, it never explicitly contrasts this with alternatives in the sibling set such as find-detached-instances, find-hidden-instances, or find-hidden-guis, and gives no exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-place-detailsSnapshot of the current place, server, and workspace (read-only)A
Read-onlyIdempotent

Read identifying and runtime details about the place and server this client is connected to. Use this to capture the exact place/universe/server you are testing in (so a finding can be reproduced), to grab the PlaceId/GameId/JobId for cross-referencing, to check the player count against MaxPlayers, or to confirm whether StreamingEnabled is on (which affects whether parts may be unloaded). It is a pure read: it mutates nothing and fires no remotes. Every field is independently pcall-guarded and reported as nil when unavailable, so a single unreadable property never fails the whole call. Returns { ok, ... } with: PlaceId, GameId, JobId, PlaceVersion, CreatorId, CreatorType (tostring of game.CreatorType), MaxPlayers (Players.MaxPlayers), PlayerCount (#Players:GetPlayers()), DistributedGameTime (workspace.DistributedGameTime, the server clock in seconds), and StreamingEnabled (workspace.StreamingEnabled). Requires only the base Roblox API; no special executor capabilities. Returns { error } only if the game/DataModel itself cannot be accessed. Signature: { threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only/idempotent/non-destructive, and the description adds real value on top: each field is independently pcall-guarded and reported as nil when unavailable, the call only errors if the DataModel itself is inaccessible, and it needs only the base Roblox API with no special executor capabilities. This is a solid behavioral disclosure; it loses a point only because read-only status is restated three times rather than expanding on other traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose, then usage, behavior, and return fields in a logical order, and the field list is genuinely informative. It is longer than necessary though, with read-only restated in three places and a boilerplate 'On failure: inspect tool-schema...' line that adds little.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully documents the return shape (ok plus every field with its source, e.g. workspace.StreamingEnabled), the error case, the capability requirements, and the intended use. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single threadContext parameter, so the schema already carries its meaning. The description merely repeats it as 'Signature: { threadContext: number? }' without adding semantics, which matches the baseline 3 for high-coverage schemas.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (read) and resource (identifying and runtime details of the connected place and server), then enumerates the exact fields returned (PlaceId, GameId, JobId, MaxPlayers, PlayerCount, StreamingEnabled). This field-level enumeration makes it clearly distinguishable from siblings like get-game-info or get-game-state without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete use cases: capturing the place/universe/server for reproducibility, grabbing IDs for cross-referencing, checking player count against MaxPlayers, and confirming StreamingEnabled. However, it names no alternatives or when-not conditions, so the agent must infer the boundary against get-game-info on its own.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-playersList every player currently in the server (read-only)A
Read-onlyIdempotent

Enumerate the live roster of the running game by calling game:GetService("Players"):GetPlayers() inside the client and reporting one record per Player. Use this to answer 'who is in this server right now?', to find a specific player's UserId/DisplayName for use in other tools, to see team assignments, or to spot the local player among everyone else. This is a pure read — it does NOT mutate game state and fires no remotes. Each player record contains, every field independently pcall-guarded so one bad property never fails the whole scan: Name, UserId, DisplayName, Team (the Team's Name, or nil when the player is on no team), AccountAge (days), and isLocal (true for the one player equal to Players.LocalPlayer). Requires only the base Roblox API (game:GetService) — no special executor capabilities. Returns { ok, count, localPlayer, players } where localPlayer is the LocalPlayer's Name (or nil), or { error } if the Players service cannot be read. Signature: { threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only/idempotent/non-destructive, yet the description adds substantial extra context: per-field pcall-guarding so one bad property never fails the scan, no special executor capabilities required, no remotes fired, and the exact return shape including the error branch. This is well beyond what structured fields provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded well with the core action and use cases, but the tail is long and partly redundant with annotations (restating read-only, idempotency, cost) plus boilerplate like 'On failure: inspect tool-schema...'. Some sentences could be trimmed without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description carries that burden and does so fully — it enumerates every returned field (Name, UserId, DisplayName, Team, AccountAge, isLocal) and the envelope { ok, count, localPlayer, players } plus the { error } path. An agent has everything needed to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% for the single threadContext parameter, so the schema already documents it fully. The description's 'Signature: { threadContext: number? }' merely restates it without adding syntax or default semantics beyond what the schema says.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — enumerate the live player roster of the running game — and even names the underlying API call. An agent can distinguish it from siblings like get-local-player-info or discover-player-values without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete triggering questions and tasks ('who is in this server right now?', resolving a UserId/DisplayName, checking team assignments, spotting the local player). It also clarifies it is a pure read, but it never names an alternative sibling (e.g. get-local-player-info) or states when NOT to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-remote-signatureProbe a single remote/bindable's shape (arg types, listeners, callback)A
Read-onlyIdempotent

Read-only inspection of ONE remote or bindable to learn how it is used WITHOUT firing it or installing any hook. Resolves remotePath to an Instance, reports its ClassName, then probes class-appropriately: - RemoteEvent / UnreliableRemoteEvent: reads getsignalarguments(remote.OnClientEvent) to surface the recent argument types the server sent to the client, and getconnections(remote.OnClientEvent) to count how many local listeners are attached (both guarded). - RemoteFunction: reads getcallbackvalue(remote, 'OnClientInvoke') and, if it is a function, reports its debug.info (source/line/name/params) so you can locate the handler the client registered. - BindableEvent: same OnClientEvent-style getsignalarguments + getconnections probe but on .Event. - BindableFunction: getcallbackvalue(remote, 'OnInvoke') -> function info. Use this after list-remotes to understand a specific remote before firing it (fire-remote) or before spying it (monitor-remote). It tells you the argument shape and whether anything is listening, which is exactly what you need to craft a valid call. Each probe is optional and pcall-guarded: a field is omitted (with a *Error note) when the executor lacks the capability or the read fails — the tool never throws. Requires (optionally) getsignalarguments, getconnections, getcallbackvalue, and debug.info; missing ones simply skip that field. Returns { ok, path, class, signalArguments?, connectionCount?, callback?, notes }, or { error }. Signature: { remotePath: string, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
remotePathYesLuau expression resolving to the RemoteEvent / RemoteFunction / UnreliableRemoteEvent / BindableEvent / BindableFunction to inspect, e.g. 'game:GetService("ReplicatedStorage").Remotes.BuyItem' or 'game.ReplicatedStorage:WaitForChild("DataRemote")'. Evaluated as `return <remotePath>` and must resolve to an Instance of one of those classes.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, non-destructive, and closed-world behavior, but the description goes further: it explains that each probe is optional and pcall-guarded, that the tool never throws, that missing executor capabilities cause fields to be omitted with a *Error note, and it gives the exact return shape. This is rich behavioral context beyond structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose and then structured into class-specific probe explanations. It is longer than typical but earns much of its length through necessary behavioral detail. Some generic boilerplate near the end ('On failure: inspect tool-schema...', phase/cost metadata) is less essential, keeping it from a perfect score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden and does so completely: it lists the exact return object fields, explains optional fields and error notes, covers prerequisites ('Requires: active-client, resolved-target'), and describes failure behavior. An agent has enough context to invoke it correctly without external references.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already fully documents both parameters. The description repeats the signature and clarifies that remotePath must resolve to an Instance of specific classes, but adds no new syntax, defaults, or constraints beyond what the schema provides. A baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb and resource: 'Read-only inspection of ONE remote or bindable to learn how it is used WITHOUT firing it or installing any hook.' It immediately distinguishes the tool from siblings like fire-remote and monitor-remote by stating what it does not do, and it explains the class-specific probing behavior in detail.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use it: 'Use this after list-remotes to understand a specific remote before firing it (fire-remote) or before spying it (monitor-remote).' It names the alternatives and the condition that selects them, leaving little to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-remote-spy-logsRead remote-spy capturesA
Read-onlyIdempotent

Reads bounded captures from the selected engine without loading it. Query filters by direction, remoteId/path, method, class, blocked-only, name, and afterId cursor affect this read only; use configure-remote-spy for persistent capture filters. Typed argument snapshots preserve nil arity and binary previews. Cobalt supplies available RakNet/actor metadata and results; Ketamine captures request arguments only. gap reports expired history. IDs are scoped to the engine and generation. Signature: { engine: "cobalt" | "ketamine"?, limit: any?, afterId: number?, direction: any?, remotePath: string?, remoteId: string?, nameFilter: string?, method: string?, classFilter: "RemoteEvent" | "UnreliableRemoteEvent" | "RemoteFunction"?, blockedOnly: boolean?, raknetOnly: boolean?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Produces: bounded-event-snapshot, structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoOptional hard result/work budget used to keep output and runtime bounded.
engineNoRemote-spy backend. Only one engine runs per client; starting another stops the current engine. Defaults to cobalt.cobalt
methodNoOptional text value for method.
afterIdNoOptional numeric value for after id.
remoteIdNoOptional text value for remote id.
directionNoCapture/control direction; incoming includes client callbacks.Both
nameFilterNoOptional text value for name filter.
raknetOnlyNoWhether to enable raknet only.
remotePathNoOptional dotted Roblox instance/value path resolved in the active client.
blockedOnlyNoWhether to enable blocked only.
classFilterNoOptional validated input for class filter.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the read-only/idempotent annotations: it discloses the non-loading read semantics, that Cobalt returns RakNet/actor metadata while Ketamine only captures request arguments, that 'gap' reports expired history, that IDs are scoped to engine and generation, and the required context (active-client, resolved-target). It even directs the agent to tool-schema on failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded: the read scope and filter/direction routing come first, and the observe/cost/requires metadata is terse. The full signature restatement duplicates the input schema and the 'Safety: read-only' line duplicates the annotations, which is mild redundancy for a definition this dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 12-parameter read tool with no output schema, it covers return shape qualitatively ('bounded-event-snapshot', per-engine result differences) and the gap/expiry behavior. Missing only explicit pagination rules for the leader result and a clear statement of what happens when no engine is running, so it is close to complete but not exhaustive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description mostly restates the signature block that already exists in the schema; its only added semantic value is the afterId cursor scoping ('IDs are scoped to the engine and generation') and the RakNet/engine behavior split, which the schema does not convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ('Reads bounded captures') and immediately scopes it ('from the selected engine without loading it'). It also differentiates from the adjacent configure-remote-spy (persistent filters) and clear-remote-spy-logs, so an agent can place it among siblings without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says query filters affect this read only and routes persistent filtering to configure-remote-spy. It also gives engine-level guidance ('Only one engine runs per client') and per-engine differences, but does not spell out prerequisites such as needing ensure-remote-spy to have run first, so it stops short of full when/when-not coverage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-render-statsRuntime perf snapshot (FPS, frame time, camera, players)A
Read-onlyIdempotent

In-game performance/runtime snapshot taken at the instant of the call. Reports: the physics frame rate from workspace:GetRealPhysicsFPS(); engine timing from game:GetService('Stats') when available (FrameTime in seconds and HeartbeatTimeMs); the current camera's world position (from its CFrame) and FieldOfView; the live player count from #Players:GetPlayers(); and workspace.DistributedGameTime (server-synchronized clock). Use this for a one-shot health read while debugging (is the client dropping frames? where is the camera? how many players are present?), or to capture a baseline before/after an exploit to see its perf impact. Every field is independently pcall-guarded and reported as null when unavailable, so the tool always returns a partial snapshot rather than failing. Requires nothing beyond a live game. Returns { physicsFPS, frameTimeSeconds, heartbeatTimeMs, camera: { position, fieldOfView }, playerCount, distributedGameTime } or { error }. Signature: { threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-observation, diagnostic-report. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only/idempotent/non-destructive, and the description adds substantial extra context: every field is pcall-guarded and reported as null when unavailable, the tool always returns a partial snapshot rather than failing, and it requires nothing beyond a live game. It also states the on-failure path and the return shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded: the opening sentence gives purpose, followed by reported fields, usage, failure semantics, then metadata. Mostly earns its place, but the tail (Phase/cost/idempotency/Safety: read-only) partly duplicates the annotations and the opening, adding mild redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description correctly carries the burden by enumerating the returned object fields and the { error } fallback, plus prerequisites ('Requires: active-client') and a pointer to tool-schema for exact fields. Nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single optional threadContext parameter is fully documented in the schema (100% coverage), and the description's 'Signature: { threadContext: number? }' merely restates it. Baseline 3 applies when the schema does the heavy lifting and the description adds no new parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('in-game performance/runtime snapshot') and then enumerates exactly what it reports (physics FPS, FrameTime, HeartbeatTimeMs, camera CFrame/FOV, player count, DistributedGameTime). This clearly distinguishes it from siblings like get-memory-stats, get-game-info, and get-fps-cap.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete when-to-use contexts ('one-shot health read while debugging… is the client dropping frames? where is the camera?') and a baseline before/after an exploit. However it names no alternative sibling to use instead or any when-not condition, so it stops short of the top tier.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-script-bytecodegetscriptbytecode — dump a script's compiled bytecodeA
Read-onlyIdempotent

Retrieve the compiled Luau bytecode for a script via the executor's getscriptbytecode. The script is resolved from a Luau expression (typically a path to a LuaSourceContainer, e.g. 'game.ReplicatedStorage.Module') via loadstring('return ' .. expr). Returns the total byte count and a hex preview of the first previewBytes bytes (the full bytecode is not shipped across the bridge). Requires getscriptbytecode — type-guarded and pcall-wrapped, returning { error } when missing or on failure. Returns { byteCount, hexPreview } or { error }. Signature: { scriptPath: string, previewBytes: any?, threadContext: number? }. Phase: orchestrate; cost=high; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: getscriptbytecode. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
scriptPathYesLuau expression resolving to a LuaSourceContainer, e.g. 'game.ReplicatedStorage.Module'. Evaluated as `return <expression>`.
previewBytesNoHow many leading bytecode bytes to include as a hex preview (default 64).
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only/idempotent, but the description adds real behavioral depth: the call is type-guarded and pcall-wrapped, returns { error } when the capability is missing or on failure, and crucially warns that the full bytecode is NOT shipped across the bridge (only a hex preview). That is exactly the kind of context annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core behavior and failure modes are front-loaded and useful, but the trailing meta-block (Phase, cost, idempotency, Requires, Capabilities, Produces, Safety, On failure) and the inline 'Signature' line duplicate structured data already in the schema and annotations, adding bulk without new information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully compensates by naming the return shapes ({ byteCount, hexPreview } or { error }), the truncation behavior, and the required capability. Nothing an agent needs to invoke or interpret the call is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds meaning beyond the schema: it explains that scriptPath is evaluated via loadstring('return ' .. expr) resolving to a LuaSourceContainer, and that threadContext defaults to the server identity. It restates previewBytes/threadContext types in the inline signature line, which is redundant rather than additive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (retrieve/dump) and a precise resource (compiled Luau bytecode for a script), which implicitly separates it from source-oriented siblings like get-script-content and get-module-source. However, it never names an alternative sibling or draws an explicit contrast, so differentiation is left to the reader.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides prerequisites (active-client, resolved-target, getscriptbytecode capability) and a phase/cost label, which gives usable context. But it never states when to prefer this over get-script-content, get-script-hash, or disassemble-function, so the alternative-selection guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-script-closureGet a script's main closure (LocalScript/ModuleScript)A
Read-onlyIdempotent

Resolve a Luau expression to a LocalScript or ModuleScript Instance and retrieve its compiled main function via getscriptclosure (falling back to getscriptfunction if the executor exposes that name instead). The returned closure is the script's top-level chunk WITHOUT running it — so you can statically analyse a script's bytecode (its constants, upvalues, and nested protos) even when it is protected, never executed, or you don't want its side effects. This is the entry point for static RE of a single script: feed the resulting function into get-closure-constants / get-closure-upvalues / get-closure-protos to drill in. Requires the executor's getscriptclosure (or getscriptfunction); if neither exists a clean { error } is returned. Returns { Script, Function:__fnInfo, ConstantCount, UpvalueCount } or { error }. Signature: { scriptPath: string, threadContext: number? }. Phase: orchestrate; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: debug closure primitives. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
scriptPathYesLuau expression resolving to a LocalScript or ModuleScript Instance, e.g. 'game.Players.LocalPlayer.PlayerScripts.Main', 'game.ReplicatedStorage.Modules.Settings', or 'game:GetService("ReplicatedFirst").Loader'. Evaluated as `return <scriptPath>`; the value must be an Instance.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only, idempotent, non-destructive, closed-world. The description adds real behavioral context beyond them: the closure is NOT executed (no side effects), the executor's getscriptclosure vs getscriptfunction fallback is disclosed, and a clean { error } is promised when neither primitive exists. Minor redundancy in restating idempotency=read-only.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and semantics, and most sentences earn their place. The trailing metadata block (Phase/cost/idempotency/Requires/Capabilities/Produces/Safety/On-failure) is somewhat boilerplate and duplicates the annotations, slightly diluting density.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description still specifies the return shape ({ Script, Function:__fnInfo, ConstantCount, UpvalueCount } or { error }) plus the failure mode, fallback primitive, and downstream pipeline. An agent has everything needed to call and consume it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both parameters are already fully documented, including the exact example paths and the 'return <scriptPath>' evaluation rule. The description's signature line and threadContext mention merely restate the schema, adding no new parameter semantics. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (resolve/retrieve) and resource (a script's compiled main closure), and names the exact instance types it accepts (LocalScript/ModuleScript). It positions itself as the entry point of a static-RE pipeline and lists the specific sibling tools it feeds, so an agent can distinguish it from get-closure-constants/upvalues/protos.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly states when to reach for it — static analysis of a script that is protected, never executed, or whose side effects you want to avoid — and routes the agent to the three downstream tools. It does not explicitly exclude alternative routes (e.g. get-script-bytecode or get-script-content) for the same goal, so it stops short of full when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-script-contentGet the content of a script in the Roblox Game ClientA
Read-onlyIdempotent

Get decompiled source for a Roblox script by path or getter code. Use startLine/endLine for a focused range when the full script is large. Signature: { scriptGetterSource: string?, scriptPath: string?, startLine: number?, endLine: number?, threadContext: number? }. Phase: orchestrate; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: decompile. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
endLineNoOptional end line number (1-based, inclusive) to return only a range of lines. Defaults to end of script if startLine is set but endLine is omitted.
startLineNoOptional start line number (1-based) to return only a range of lines from the decompiled script. If omitted, returns the full script.
scriptPathNoThe path to the script to get the content of (e.g. 'game.Workspace.MyScript').
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
scriptGetterSourceNoThe code that fetches the script object from the game (should return a script object, and MUST be client-side only, will not work on Scripts with RunContext set to Server)

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description still adds non-annotation context: cost=medium, produces structured-observation, and the active-client/resolved-target prerequisites. It also points at tool-schema for constraints on failure, which is a helpful escape hatch.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded and the metadata block is compact and scannable. It loses a little to redundancy, since 'idempotency=read-only' and 'Safety: read-only' echo the readOnlyHint/idempotentHint annotations already supplied.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description still covers prerequisites, cost, output category, range usage, and a fallback path to tool-schema. For a read-only decompilation tool whose parameters are fully documented in the schema, this is close to complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters are already documented in the schema, including the client-side-only restriction on scriptGetterSource. The description's inline signature restates optionality but adds nothing about format beyond what the schema provides, so it sits at the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence gives a specific verb and resource: 'Get decompiled source for a Roblox script by path or getter code.' It also distinguishes the two lookup modes (path vs getter code), which is useful. However, it never names or contrasts with nearby siblings like get-module-source or get-script-bytecode, so an agent must infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives one concrete usage rule ('Use startLine/endLine for a focused range when the full script is large') and lists prerequisites ('Requires: active-client, resolved-target'). It does not say when to prefer this over get-module-source, get-script-bytecode, or semantic-search-scripts, so alternative selection is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-script-envDump a script's environment (getsenv)A
Read-onlyIdempotent

Resolve an instance path to a running LocalScript/ModuleScript and dump its script environment table via getsenv. This is the live _ENV of that script: its globals, every top-level local it exposed as a global, and the functions/values it defined. Use it to discover what a script holds (config flags, references, handler functions) so you can then inspect or hook them by reference (e.g. feed getsenv(script).someFunc into inspect-closure or hook-function). The script must be currently running for getsenv to succeed. Requires getsenv. Returns { Script, KeyCount, Truncated, Keys } (keysOnly) or { Script, KeyCount, Truncated, Entries } or { error }. Signature: { scriptPath: string, keysOnly: any?, maxKeys: any?, threadContext: number? }. Phase: orchestrate; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
maxKeysNoMaximum number of environment keys/entries to return before truncating (default 200).
keysOnlyNoWhen true (default) return only the environment key names — cheap and safe. When false also serialize each value's type and a scalar/string preview via __encVal.
scriptPathYesLuau expression resolving to the LocalScript/ModuleScript instance, e.g. 'game.Players.LocalPlayer.PlayerScripts.Main' or 'game:GetService("ReplicatedStorage").Modules.Net'. Evaluated as `return <scriptPath>`; must yield a running script instance.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds genuinely useful behavior beyond that: the running-script requirement, keysOnly's cheap-and-safe default mode, maxKeys truncation, and the exact success/error return shapes. It doesn't discuss performance cost beyond the phase metadata, but the added context is solid.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then use case, then constraints, then return shapes — a sensible order. It is somewhat long and the trailing metadata block (Phase/cost/idempotency/Requires/Produces/Safety) partially duplicates the annotations and schema, but nothing is truly wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by enumerating the three return variants ({Keys}, {Entries}, {error}), the required runtime precondition, the prerequisite capabilities (getsenv, active-client, resolved-target), and failure handling. An agent has everything needed to invoke and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds value by tying parameters to outcomes: keysOnly drives which return shape appears ({Keys} vs {Entries}) and notes keysOnly's cheap/safe default, and it contextualizes the scriptPath semantics. This goes beyond merely restating the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (resolve/dump) and resource (a running LocalScript/ModuleScript's environment table), naming the underlying primitive (getsenv) and the exact contents (globals, exposed top-level locals, defined functions). This clearly differentiates it from env-related siblings like dump-function-env, get-function-env, and list-global-env-keys.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete use case and routes the agent forward: 'Use it to discover what a script holds ... so you can then inspect or hook them ... feed getsenv(script).someFunc into inspect-closure or hook-function.' It also states a hard precondition ('The script must be currently running'). It stops short of an explicit when-not-to-use or a direct contrast with the nearest sibling (dump-function-env).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-script-hashgetscripthash — content hash of a scriptA
Read-onlyIdempotent

Compute the content hash of a script via the executor's getscripthash. The script is resolved from a Luau expression (typically a path to a LuaSourceContainer, e.g. 'game.ReplicatedStorage.Module') via loadstring('return ' .. expr). Handy for detecting when a script's bytecode/source changes between two checks. Requires getscripthash — type-guarded and pcall-wrapped, returning { error } when missing or on failure. Returns { hash } or { error }. Signature: { scriptPath: string, threadContext: number? }. Phase: orchestrate; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
scriptPathYesLuau expression resolving to a LuaSourceContainer, e.g. 'game.ReplicatedStorage.Module'. Evaluated as `return <expression>`.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only/idempotent/no-destroy, but the description adds real behavior: it is type-guarded and pcall-wrapped, returns { error } when getscripthash is missing or the call fails, and depends on an active client and resolved target. This goes beyond the structured fields without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded and most sentences carry load (resolution mechanism, failure shape, prerequisites). It loses a little to duplication with the annotations ('idempotency=read-only' and 'Safety: read-only') and the trailing generic 'inspect tool-schema' boilerplate.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description supplies the return shape ({ hash } or { error }), failure semantics, prerequisites, and the phase/cost profile — enough for an agent to invoke and interpret the call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description's note that the path is evaluated as `return <expression>` and that threadContext defaults to the server default duplicates what the schema already documents, adding little.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Compute the content hash of a script') and spells out the resolution mechanism (Luau expression via loadstring('return ' .. expr)), which cleanly distinguishes it from siblings like crypt-hash or get-function-hash.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Handy for detecting when a script's bytecode/source changes between two checks' gives a concrete use case, and the 'Requires: active-client, resolved-target' line states prerequisites. It stops short of naming an alternative or a when-not condition, so it is not a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-semantic-index-statsSemantic index statisticsA
Read-onlyIdempotent

Report the semantic script index for THIS session's active client WITHOUT touching the game: whether anything is indexed, how many documents are cached, and the embedding model + dimensions in use. Resolves the active client from this session's selection; if no client is resolved it returns an empty, not-indexed summary. Use it to confirm an index exists before searching, or to see which embeddings backend is active. Signature: {}. Phase: observe; cost=low; idempotency=read-only. Requires: none. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds meaningful context beyond that: it enumerates what is reported (indexed-or-not, cached document count, embedding model + dimensions) and states the empty-summary fallback when no client resolves. It does not contradict any annotation. It could disclose more about the failure surface, but the 'On failure' pointer to tool-schema partly compensates.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose and return shape are front-loaded and information-dense. The trailing 'Signature / Phase / cost / idempotency / Requires / Produces / Safety / On failure' boilerplate is padded and partly redundant with the annotations, which costs a point on a short description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-param read-only observation tool with no output schema, the description covers purpose, scope, return contents, and the empty-result behavior. It does not enumerate the exact output fields, but it points the agent to tool-schema on failure, which is an adequate completeness level.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters, so the per-rubric baseline is 4. The description notes it resolves the active client from session selection and documents the no-client fallback, which is the only 'input-like' behavior that matters here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Report) and resource (semantic script index statistics) and explicitly names the scope ('THIS session's active client WITHOUT touching the game'). This distinguishes it cleanly from siblings like semantic-search-scripts and clear-semantic-index, which operate on the index rather than reporting on it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use it: 'Use it to confirm an index exists before searching, or to see which embeddings backend is active.' This gives the agent both the prerequisite-check use case and the informational use case, and the no-touch framing rules out alternatives for probing index state.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-signal-argumentsGet the argument types a signal fires withA
Read-onlyIdempotent

Report the value TYPES that an RBXScriptSignal fires its connected handlers with, using the executor's getsignalarguments. This answers "what shape is the payload of this event?" without having to connect a handler and wait for a fire — invaluable when reverse-engineering a remote-driven or engine signal so you know how to construct a fire-signal / replicate-signal call. Pass the instance that owns the signal plus the signal member name (or leave signalName empty if instancePath already resolves to the signal). Requires the getsignalarguments executor function; degrades with a clear { error } if unavailable. Returns { Signal, Instance?, Arguments } where Arguments is the type/scalar info mapped through a safe serializer. Signature: { instancePath: string, signalName: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: getconnections. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
signalNameNoName of the RBXScriptSignal member on the instance (e.g. 'OnClientEvent', 'Touched', 'Changed'). Leave empty/omit if instancePath already evaluates to the signal.
instancePathYesLua expression resolving to the Instance that owns the signal (e.g. 'game.ReplicatedStorage.MyRemote', 'game.Workspace.Door'). Evaluated as `return <instancePath>`. Leave signalName empty if this expression already resolves to the RBXScriptSignal itself.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint, idempotentHint, non-destructive), and the description adds real context beyond them: dependency on the getsignalarguments executor function, graceful degradation to a { error } when unavailable, prerequisite state (active-client, resolved-target), cost level, and the return shape with a safe serializer. That is substantial disclosure not derivable from the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and value are front-loaded, but the tail is a dense metadata block that partly duplicates structured data: 'Safety: read-only' and 'idempotency=read-only' repeat the annotations, 'Signature: {...}' repeats the schema, and the closing 'inspect tool-schema' line is generic boilerplate. Useful information is present, but not every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, yet the description supplies the return shape ({ Signal, Instance?, Arguments }) and the serializer caveat. Combined with prerequisites, cost, failure behavior and the executor dependency, an agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents instancePath, signalName and threadContext, including the 'leave signalName empty' condition and the `return <instancePath>` evaluation semantics. The description's 'Signature: { instancePath, signalName?, threadContext? }' merely restates the schema rather than adding new meaning, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb + resource: 'Report the value TYPES that an RBXScriptSignal fires its connected handlers with'. It frames the exact question it answers ('what shape is the payload of this event?') and implicitly separates it from the connect-and-wait approach and from sibling tools like count-signal-connections or list-signal-connections.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear scenario — reverse-engineering a remote-driven or engine signal to know how to construct a fire-signal / replicate-signal call — which names downstream siblings. It does not state explicit exclusions (e.g. when an alternative signal-inspection tool is preferable), so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-signal-arguments-infoGet detailed per-argument info for a signalA
Read-onlyIdempotent

Report RICH, per-argument metadata about what an RBXScriptSignal fires with, using the executor's getsignalargumentsinfo. This is the more detailed companion to get-signal-arguments: where that tool reports just the argument types/values, this returns the executor's fuller per-argument descriptor table (e.g. type tags, optionality, and per-entry detail) so you can precisely reconstruct a signal's payload when crafting a fire-signal / replicate-signal call. Pass the instance that owns the signal plus the signal member name (or leave signalName empty if instancePath already resolves to the signal). Requires the getsignalargumentsinfo executor function; degrades with a clear { error } if unavailable. Returns { Signal, Instance?, ArgumentsInfo } with nested values mapped through a safe serializer. Signature: { instancePath: string, signalName: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: getconnections. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
signalNameNoName of the RBXScriptSignal member on the instance (e.g. 'OnClientEvent', 'Touched', 'Changed'). Leave empty/omit if instancePath already evaluates to the signal.
instancePathYesLua expression resolving to the Instance that owns the signal (e.g. 'game.ReplicatedStorage.MyRemote', 'game.Workspace.Door'). Evaluated as `return <instancePath>`. Leave signalName empty if this expression already resolves to the RBXScriptSignal itself.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and non-destructive, so the safety profile is covered. The description adds genuinely new behavior: it requires the getsignalargumentsinfo executor function and 'degrades with a clear { error } if unavailable', plus the returned payload is passed through a safe serializer. It does not mention rate limits or permission prerequisites, but the executor-function dependency and degradation path are solid added context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded and the sibling comparison is useful, but the text is padded with metadata that largely restates the schema or annotations (Signature line, Phase/cost/idempotency, Safety: read-only, Produces). Several sentences do not earn their place against the structured fields already provided.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description usefully declares the return shape ({ Signal, Instance?, ArgumentsInfo } with nested values mapped through a safe serializer) and the failure mode. Combined with the 100%-covered input schema and the annotations, an agent has enough to call it correctly; only the executor-availability fallback detail could be richer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so instancePath, signalName, and threadContext are already fully documented in the schema, including the 'leave signalName empty' rule the description repeats. The description adds no syntax or format detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (report) and resource (per-argument metadata about what an RBXScriptSignal fires with) and explicitly positions itself against the sibling get-signal-arguments, describing what that tool returns versus this one. An agent can distinguish the two without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the alternative (get-signal-arguments) and the condition that selects this one ('more detailed companion... so you can precisely reconstruct a signal's payload when crafting a fire-signal / replicate-signal call'). The when-to-use is tied to a concrete downstream task, leaving little to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-signal-whitelistList the engine's replicatable-signal whitelistA
Read-onlyIdempotent

Enumerate the full set of signals the Roblox engine allows to be replicated to the server, using the executor's getsignalwhitelist. This is the global allow-list that can-signal-replicate / replicate-signal check against, so it answers "which signals can I drive server-side?" up front, without probing them one at a time. The whitelist is typically large (180+ entries), so each entry is mapped through a safe serializer and the output is capped at limit. Requires the getsignalwhitelist executor function; degrades with a clear { error } if unavailable. Returns { Count, Truncated, Signals }. Signature: { limit: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Capabilities: getconnections. Produces: bounded-candidates, structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of whitelist entries to return (default 300). The list can be large, so this caps the response; Count always reflects the true total and Truncated indicates whether entries were dropped.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the read-only/idempotent/non-destructive safety profile, yet the description adds real context beyond them: the list is typically large (180+ entries), output is capped at limit, it requires the getsignalwhitelist executor function, and it degrades with a clear { error } if that function is unavailable. These are non-obvious operational traits an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening sentences are front-loaded and purposeful, but the definition is heavy with taxonomy tags ('Phase: observe; cost=medium... Capabilities: getconnections. Produces: bounded-candidates...') that partly duplicate the annotations (safety=read-only mirrors readOnlyHint) and a closing boilerplate pointer to tool-schema. These lines dilute rather than earn their place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fills the gap by naming the return shape ({ Count, Truncated, Signals }) and clarifying that Count reflects the true total while Truncated flags dropped entries. Combined with the stated requirements, failure behavior, and cap semantics, an agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents both parameters including the cap/Count/Truncated behavior. The description's restatement of the limit cap and a terse signature line add marginal value over that baseline, so a 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Enumerate the full set of signals the Roblox engine allows to be replicated to the server') and explicitly positions it against sibling tools ('the global allow-list that can-signal-replicate / replicate-signal check against'). An agent can distinguish it from those per-signal probes without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Frames the use case clearly ('which signals can I drive server-side? up front, without probing them one at a time'), implicitly routing the agent away from one-at-a-time checks like can-signal-replicate. However, it never states an explicit exclusion or 'use can-signal-replicate instead when...' condition, so it stops short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-stackdebug.getstack — read live Luau stack slots at a levelA
Read-onlyIdempotent

Read the raw values currently sitting on the Luau stack at a given call level via the executor's debug.getstack. Without 'index' it returns every live stack slot at that level as an encoded list (Instances/EnumItems/tables are flattened to a JSON-friendly shape); with 'index' it returns just that one slot's encoded value. This exposes the in-flight locals and temporaries of a running function — the values a frame is actively working with — letting you snapshot what a callback or hooked function holds at the moment your code runs. Requires debug.getstack — type-guarded and pcall-wrapped, returning { error } when missing or on failure. Returns { level, index?, value } / { level, count, values } or { error }. Signature: { level: number, index: number?, threadContext: number?, timeoutMs: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
indexNoOptional 1-based stack slot index. When given, only that single slot is read and returned; when omitted, the whole stack table at 'level' is returned.
levelYesThe call-stack level to read (1 = the function calling getstack, i.e. the executing chunk; higher = further up the stack).
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only/idempotent/non-destructive, but the description adds real behavioral detail beyond them: it is type-guarded and pcall-wrapped, returns { error } when debug.getstack is missing or the call fails, and documents the encoded/flattened output shape. Cost=medium and Requires: active-client are also surfaced.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core behavior and modes, then structured metadata (Phase/cost/Requires/Produces/Safety/failure pointer). It is dense for a 4-param tool and restates the full signature, which duplicates the schema, but every block still earns its place given the absence of an output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully carries the return contract ({ level, index?, value } / { level, count, values } / { error }), documents the failure mode, prerequisite (active-client), and cost. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all four parameters and the baseline is 3. The description adds value by explaining the observable effect of 'index' on output (whole-level encoded list vs. single slot's encoded value) and framing 'level' as in-flight locals/temporaries, which is more than the schema states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (read) and resource (raw values on the Luau stack at a call level), names the underlying API debug.getstack, and distinguishes the two operating modes (whole-level list vs. single indexed slot). An agent can immediately tell this apart from introspection siblings such as get-call-stack or dump-function-env.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for when this applies — 'snapshot what a callback or hooked function holds at the moment your code runs' — but names no explicit alternatives or exclusions against siblings like get-thread-stack or get-call-stack. Clear context without routing guidance, so a 4.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-thread-stackWalk a thread/coroutine's call stackA
Read-onlyIdempotent

Walk the call stack of a Luau thread/coroutine frame-by-frame using debug.info. For each level it reports the function Name, Source, and Line (debug.info(thread, level, 'nsl')), stopping when the stack is exhausted. If threadPath is omitted it traces the CURRENT injected thread instead, using debug.info(level, 'nsl'). A debug.traceback() string is also included when available. Use this to see exactly where a suspended coroutine or an event handler's thread is currently executing — e.g. pass a connection's .Thread, a stored coroutine, or leave it blank to inspect your own execution context. WARNING: this ACTS ON THE LIVE GAME — it evaluates your threadPath expression and introspects live thread state in the running client. Returns { Thread?, FrameCount, Frames:[{ Level, Name, Source, Line }], Traceback? } or { error }. Signature: { threadPath: string?, maxLevels: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
maxLevelsNoMaximum number of stack levels to walk before stopping (default 20, max 60). Walking stops early once debug.info returns nil for a level (top of stack reached).
threadPathNoOptional Luau expression resolving to a thread/coroutine, e.g. 'getconnections(game.Workspace.Part.Touched)[1].Thread', 'getgenv().myCoroutine', or 'coroutine.running()'. Evaluated as `return <threadPath>`. If omitted, the current injected thread is traced instead.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly/idempotent annotations, it discloses that the call evaluates the threadPath expression against live game state, warns that it acts on the live client, explains the omission behavior (traces the current injected thread), and includes the return shape and optional traceback.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and usage are front-loaded well, but the tail ('Phase: observe; cost=medium; ... Requires: active-client ... On failure: inspect tool-schema ...') largely restates annotation metadata and adds boilerplate that does not earn its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, but the description supplies the return shape ({ Thread?, FrameCount, Frames, Traceback? } or { error }) and covers safety, prerequisites, and failure handling, so an agent has enough to call it correctly; only minor edge behavior is left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents threadPath, maxLevels, and threadContext. The description reinforces the omitted-threadPath behavior but adds little syntax or constraint detail beyond the schema, matching the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Walk the call stack of a Luau thread/coroutine frame-by-frame using debug.info') and scopes it precisely to thread/coroutine introspection, distinguishing it from generic stack siblings like get-call-stack and get-stack.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete usage context ('see exactly where a suspended coroutine or an event handler's thread is currently executing') plus worked examples (pass a connection's .Thread, a stored coroutine, or leave blank). It does not explicitly name an alternative tool or a when-not-to-use condition, so it stops short of 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hook-and-log-functionHook a function and log every call (turnkey instrumentation, MUTATES STATE)A
Destructive

DANGER — INSTALLS A PERSISTENT GLOBAL HOOK. Turnkey call-tracing for any function: hook a target, automatically record every invocation (stringified arguments + return values + a timestamp), then fetch the captured call log and restore the original — all from three actions of this one tool. This is the fastest way to answer 'what is this function actually called with, how often, and what does it return?' without hand-writing a hook. WORKFLOW (stateful — survives across tool calls via getgenv().__mcp_fnlogs, keyed by functionPath): 1. action='start' with functionPath — resolves the target, captures the original, installs a logging hook that transparently calls the original and records up to maxCalls invocations. Returns { started, key }. 2. action='fetch' with the same functionPath — reads the accumulated call log so far WITHOUT stopping it. Returns { count, max, calls } where each call is { args[], returns[], t }. Call repeatedly to watch live. 3. action='stop' with the same functionPath — restores the original function and removes the registry entry. Returns { stopped }. ALWAYS stop when done. CAVEATS: The hook is GLOBAL and PERSISTS until you stop it (or the client restarts). It adds overhead on every call to the target and CAN TRIP ANTICHEAT or destabilize the game, especially on hot paths — prefer specific, low-frequency targets and keep maxCalls modest. Arguments/returns are captured by tostring (Instances become GetFullName()) and both arrays are capped at 8 entries each. Logging stops accumulating once maxCalls is reached, but the hook stays installed (and keeps calling the original) until you stop it. Requires hookfunction, newcclosure, and getgenv; restoration uses hookfunction(target, original) with a restorefunction fallback. Returns { error } with a clear message if a capability is missing, the target cannot be resolved, or there is no active log for fetch/stop. Signature: { action: "start" | "fetch" | "stop", functionPath: string?, maxCalls: any?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Produces: structured-result. Verify with: is-function-hooked. Safety: MUTATING; changes persistent executor-side observer or hook state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesWhat to do: 'start' installs the logging hook on functionPath; 'fetch' returns the call log captured so far (hook stays live); 'stop' restores the original function and clears the log. Use the SAME functionPath for all three so they address the same registry entry.
maxCallsNoMaximum number of calls to record before logging stops accumulating (default 100). On 'start' this sizes the ring of captured calls; on 'fetch' it caps how many entries are returned in this response. Keep modest on hot paths to limit overhead and output size.
functionPathNoLuau expression resolving to the function to instrument, e.g. 'getsenv(game.Players.LocalPlayer.PlayerScripts.Main).validate', 'getrawmetatable(game).__namecall', or 'getconnections(game.Workspace.Part.Touched)[1].Function'. Evaluated as `return <functionPath>` and must resolve to a function. REQUIRED for 'start'. For 'fetch'/'stop' it is the registry key identifying which running log to act on, so it must match the string used at start (defaults to the start expression).
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructive/mutating and non-idempotent, but the description goes far beyond them: persistence via getgenv().__mcp_fnlogs, global lifetime until stop/restart, per-call overhead, anticheat/destabilization risk on hot paths, tostring-based capture with arrays capped at 8 entries, maxCalls semantics, required capabilities (hookfunction/newcclosure/getgenv), the restore fallback, and error-return behavior. This is exactly the extra context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with a DANGER warning and then the workflow, so the critical risk and sequence come first. It is long, and some content (action semantics, signature) is repeated in both prose and signature form, but the length is largely justified by the tool's stateful, mutating nature.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex, stateful mutation tool with no output schema, the description specifies the return shapes ({started,key}, {count,max,calls}, {stopped}, {error}), failure modes, capability requirements, and lifecycle rules. Nothing an agent needs to invoke it correctly and clean up afterward is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents action, functionPath, maxCalls, and threadContext thoroughly. The description's signature and registry-key notes largely restate the schema rather than adding new semantic detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise compound action: install a persistent global hook, record invocations (args/returns/timestamp), fetch the log, and restore the original — all in one tool. It clearly distinguishes itself from hand-writing a hook and from single-purpose siblings like hook-function, list-hooks, and restore-hook by framing itself as the 'turnkey' three-action alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides an explicit numbered workflow (start/fetch/stop) with the conditions that select each action, states the use case ('fastest way to answer what is this function called with'), and gives when-not-to-use guidance ('prefer specific, low-frequency targets', 'ALWAYS stop when done'). It does not explicitly name sibling alternatives (hook-function/restore-hook), so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hook-functionHook a live function with a replacement (MUTATES STATE)A
Destructive

WRITES LIVE GAME STATE. DANGER — MUTATES STATE PERSISTENTLY. Resolve a target function and replace it with your own function via hookfunction. After hooking, every call to the target — from anywhere in the game — runs your replacement instead. This is the core primitive for intercepting/altering game behavior: log or rewrite arguments, spoof return values, or no-op a check. The hook is GLOBAL and PERSISTS until undone, so it can easily destabilize the game or trip anticheat. The original function is captured and stored in getgenv().__mcp_hooks keyed by the target expression so you can recover it; you (or your replacement) can call the original, and you can fully undo the hook with restorefunction(target). Requires hookfunction. Pass confirm=true to proceed. Returns { Target, Hooked, OriginalStored } or { error }. Signature: { targetPath: string, hookFunction: string, confirm: boolean, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Produces: structured-result. Verify with: is-function-hooked. Safety: MUTATING; changes persistent executor-side observer or hook state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmYesMust be true to actually install the hook (global, persistent mutation). When omitted or false, the tool refuses and changes nothing.
targetPathYesLuau expression resolving to the function to hook (the one whose calls you want to intercept), e.g. 'getsenv(game.Players.LocalPlayer.PlayerScripts.Main).validate' or 'getrawmetatable(game).__namecall'. Evaluated as `return <targetPath>`.
hookFunctionYesRaw Luau expression that evaluates to the REPLACEMENT function. Typically a function literal, e.g. 'function(...) print("called", ...) return getgenv().__mcp_hooks[<target>](...) end' or 'newcclosure(function(...) return true end)'. To call the original from inside your hook, read it from getgenv().__mcp_hooks. Evaluated as `return <hookFunction>` and must resolve to a function.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true and idempotentHint=false, yet the description adds substantial behavior: the hook is GLOBAL and PERSISTS until undone, the original is captured into getgenv().__mcp_hooks keyed by target, it can destabilize/trip anticheat, it requires hookfunction and explicit-mutation-approval, and it can be fully undone via restorefunction. This is well beyond what structured fields convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The destructive warning is correctly front-loaded, but the mutation caveat is repeated three times ('WRITES LIVE GAME STATE', 'MUTATES STATE PERSISTENTLY', 'MUTATING') and the trailing metadata block (Phase/cost/idempotency/Produces/Verify) is templated filler rather than agent-critical content. Dense but padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description still discloses the return shape ({ Target, Hooked, OriginalStored } or { error }), the recovery path, verification via is-function-hooked, prerequisites, and failure guidance. Nothing an agent needs to invoke and recover this mutation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so targetPath, hookFunction, confirm, and threadContext are already fully documented with evaluation rules and examples. The description's 'Pass confirm=true to proceed' and inline signature only restate what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Resolve a target function and replace it with your own function via hookfunction') and describes the exact mechanism and scope (global interception). An agent can distinguish this core primitive from narrower siblings like hook-and-log-function or spoof-function-return without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear when-to-use context ('core primitive for intercepting/altering game behavior: log or rewrite arguments, spoof return values, or no-op a check') plus the confirm=true prerequisite and undo path. However it never names or routes to the close alternatives (hook-and-log-function, block-function, spoof-function-return), so an agent with only logging intent gets no explicit redirection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hook-metamethodHook a metamethod on an object (MUTATES live state)A
Destructive

WRITES LIVE GAME STATE. DANGER — MUTATES STATE PERSISTENTLY. Resolve a Luau expression to an object and replace one of its metamethods (e.g. __namecall, __index, __newindex) with your own function via hookmetamethod. The hook stays active and INTERCEPTS EVERY call routed through that metamethod (for __namecall that is essentially every method call in the game), so a slow, throwing, or mis-behaving hook can hang or crash the client and is a strong anticheat signal. The original metamethod is stored in getgenv().__mcp_hooks under a descriptive key so you can later restore it with restorefunction or by re-hooking the saved original. Your hook function typically wraps the stored original (call it for unhandled cases) and should be a C closure (wrap it in newcclosure) to look native. Requires hookmetamethod. Because it mutates state you MUST pass confirm=true; otherwise the tool refuses and does nothing. Returns { Target, Method, Hooked, OriginalStored } or { error }. Signature: { objectPath: string, method: string, hookFunction: string, confirm: boolean, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Capabilities: getrawmetatable. Produces: structured-result. Verify with: is-function-hooked. Safety: MUTATING; changes persistent executor-side observer or hook state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
methodYesThe metamethod name to hook, e.g. '__namecall', '__index', '__newindex'. __namecall is the usual target for intercepting method calls (RemoteEvent:FireServer, etc.).
confirmYesSafety gate. Must be exactly true to apply the hook. If omitted or false the tool refuses and does nothing, because a persistent metamethod hook intercepts every call through it and can crash the game or trip anticheat.
objectPathYesLuau expression resolving to the object whose metamethod you want to hook, e.g. 'game', 'game.Players.LocalPlayer', or 'getgenv().SomeTable'. Evaluated as `return <objectPath>`.
hookFunctionYesRaw Luau expression evaluating to the replacement function, MUST resolve to a function. Usually wrapped in newcclosure so it appears as a native C closure, e.g. 'newcclosure(function(self, ...) local m = getnamecallmethod(); if m == "FireServer" then return end; return getgenv().__mcp_hooks["..."](self, ...) end)'. Evaluated as `return <hookFunction>`.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=false, but the description adds substantial behavioral context beyond them: the hook persists and intercepts every call (crash/hang/anticheat risk), the original is stored under getgenv().__mcp_hooks, the confirm gate, the C-closure wrapping guidance, and the exact return shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the danger warning and reasonably ordered, but it is dense and repeats the mutation warning multiple times ('WRITES LIVE GAME STATE', 'MUTATES STATE PERSISTENTLY', 'Safety: MUTATING') and includes a signature line that duplicates the schema. Some trimming would improve signal density.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description documents the return object ({ Target, Method, Hooked, OriginalStored } or { error }), the verification path (is-function-hooked), prerequisites, and failure guidance. An agent has everything needed to invoke and validate correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all five parameters thoroughly (method names, confirm gate, objectPath evaluation, hookFunction wrapping, threadContext default). The description restates the signature and confirm behavior but adds little syntax or format detail beyond what the schema provides, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (hook/replace) and resource (a metamethod on a resolved object) and names the underlying primitive hookmetamethod. It is clearly distinguishable from siblings like hook-function, list-hooks, restore-hook, and restore-function, which are referenced explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives strong context: when to hook (to intercept calls routed through a metamethod), the requirement to pass confirm=true, and the recovery path via restorefunction or re-hooking the saved original. It does not, however, explicitly contrast itself against the closest sibling hook-function, leaving that selection to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

http-requestMake an outbound HTTP request from the client (sUNC request)A
Destructive

SENDS A REAL OUTBOUND HTTP REQUEST from the Roblox client via the executor's request({ Url, Method, Headers, Body }) (falling back to http_request, then syn.request). Unlike Roblox's HttpService this can hit arbitrary hosts and set custom headers. The response body is capped at ~100 KB. Requires one of these functions to be available. The call is type-guarded and pcall-wrapped: if none is present you get { error = 'request is not available in this executor.' }, and a transport failure returns { error }. Returns { statusCode, success, headers, body, statusMessage } or { error }. Signature: { url: string, method: any?, headers: {[string]: string}?, body: string?, threadContext: number?, timeoutMs: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval. Capabilities: request. Produces: structured-result. Verify with: assert-state. Safety: MUTATING; performs external network or socket I/O. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe absolute request URL, e.g. 'https://api.example.com/v1/thing'.
bodyNoOptional request body (for POST/PUT/PATCH).
methodNoHTTP method (default GET).GET
headersNoOptional request headers as a string->string map.
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare the safety envelope (destructive, open-world, non-idempotent); the description adds substantial behavior beyond that: the ~100 KB response body cap, the internal fallback order, the type-guard/pcall wrapping, and both failure shapes ('request is not available in this executor.' vs transport {error}) and the success shape. This is exactly the extra context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the essential capability and executor details in the first quoted sentence, then layers operational metadata. It is dense and mostly earns its length, though the trailing boilerplate ('On failure: inspect tool-schema...') and the Phase/cost/idempotency line read as generic metadata rather than tool-specific value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully inlines the return shape ({statusCode, success, headers, body, statusMessage} or {error}) and covers error and fallback paths. It stops short of mentioning rate limits, header restrictions, or timeout defaults beyond what the schema says, but for a 6-param network tool with full annotation coverage this is close to complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all six parameters (url, method, headers, body, threadContext, timeoutMs) are already documented with defaults and enum values. The signature line in the description restates them without adding format, constraint, or interaction semantics, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('SENDS A REAL OUTBOUND HTTP REQUEST from the Roblox client') and immediately names the mechanism and fallback chain (request -> http_request -> syn.request). This distinguishes it cleanly from siblings like ws-connect or scan-network-endpoints, which are different I/O surfaces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It draws a contrast with Roblox's HttpService (arbitrary hosts, custom headers) and lists prerequisites ('Requires: active-client, explicit-mutation-approval'), which implies when it is legal to call. However, it names no sibling alternative and gives no explicit when-to-use / when-not-to-use routing, so guidance remains implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ignore-remoteSet a remote's ignore stateA
Destructive

WRITES LIVE GAME STATE. Reversibly ignores captures for an observed remoteId or Luau remotePath in the selected engine/direction. ignored=false undoes it. Calls continue executing. Ketamine suppresses MCP captures; Cobalt suppresses its engine/MCP history. Signature: { engine: "cobalt" | "ketamine"?, remotePath: string?, remoteId: string?, direction: any?, ignored: any?, threadContext: number? }. Phase: act; cost=medium; idempotency=idempotent-write. Requires: active-client, resolved-target, explicit-mutation-approval. Produces: structured-result. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
engineNoRemote-spy backend. Only one engine runs per client; starting another stops the current engine. Defaults to cobalt.cobalt
ignoredNoOptional validated input for ignored.
remoteIdNoOptional text value for remote id.
directionNoCapture/control direction; incoming includes client callbacks.Outgoing
remotePathNoOptional dotted Roblox instance/value path resolved in the active client.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.5/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description explicitly states 'idempotency=idempotent-write', but the annotations declare idempotentHint=false. These directly conflict, leaving an agent unsure whether repeating the call is safe. Despite the otherwise rich context (reversibility, engine-specific suppression semantics), the explicit contradiction with structured metadata forces a 1.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the critical fact ('WRITES LIVE GAME STATE') before mechanics, and the sentences are mostly information-dense. The duplicated signature block and metadata line ('Phase/cost/produces') are somewhat boilerplate but not wasteful enough to hurt much.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description covers produces ('structured-result'), verification, and preconditions, giving an agent enough to invoke safely. The idempotency conflict and the absence of guidance on the sibling block-remote are the remaining gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and two parameters have enums with defaults, so the schema already carries parameter meaning. The 'Signature' line merely restates names and types, adding little beyond the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+scope: 'Reversibly ignores captures for an observed remoteId or Luau remotePath in the selected engine/direction.' It also contrasts with blocking behavior ('Calls continue executing'), which lets an agent separate it from siblings like block-remote and monitor-remote without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit preconditions are given ('Requires: active-client, resolved-target, explicit-mutation-approval'), plus a verification path ('Verify with: assert-state') and a fallback ('On failure: inspect tool-schema'). It does not name the alternative tool to use instead (e.g., block-remote), so it stops short of a full when/when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inspect-callbacksInspect RemoteFunction / BindableFunction invoke callbacksA
Read-onlyIdempotent

Find and disassemble the invoke callbacks bound to RemoteFunctions and BindableFunctions — these are where a game's request/response logic lives (the server-authoritative answer to a client question, or a cross-script RPC), so they are frequently the single most valuable functions to reverse. Unlike signal events (OnClientEvent/OnServerEvent), invoke callbacks are stored as a hidden property on the instance and are NOT visible to getconnections; the ONLY way to retrieve them is getcallbackvalue. This tool walks a subtree (default the whole DataModel), and for every RemoteFunction reads its OnClientInvoke and OnServerInvoke slots, and for every BindableFunction reads its OnInvoke slot. For each slot that actually holds a function it captures debug.info (source / line-defined / name) so you can immediately pivot to inspect-closure, get-closure-constants, get-closure-upvalues, scan-proto-functions, or hook-function on the exact callback. Use the source/line to locate the defining script and the name to understand intent. Requires getcallbackvalue (returns a clean { error } if the executor lacks it). The scan is fully pcall-guarded (locked/parented-out/dead instances never abort it), capped by maxScan, and the output is capped by limit with a truncated flag. Read-only: it inspects callbacks but never invokes or modifies them. Returns { count, scanned, truncated, root, callbacks } where each entry is { remote, class, slot, callback = { source, line, name, pointer, isC } }. Signature: { root: any?, limit: any?, maxScan: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-observation, operation-receipt. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
rootNoLuau expression for the subtree root whose descendants are scanned, evaluated as `return <root>`. Defaults to 'game' (the whole DataModel). Narrow it to cut scan time and noise, e.g. 'game.ReplicatedStorage', 'game:GetService("ReplicatedStorage").Remotes', or 'game.Players.LocalPlayer.PlayerGui'.game
limitNoMaximum number of callback entries to return (default 150). Once reached the scan stops early and `truncated` is set true.
maxScanNoMaximum number of descendant instances to visit while scanning (default 8000). Caps cost on huge DataModels; if hit, `truncated` is set true. Clamped to 100..60000.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the read-only/idempotent profile, but the description adds substantial context beyond them: requires getcallbackvalue (clean { error } otherwise), fully pcall-guarded so dead/parented-out instances never abort, capped by maxScan, output capped by limit with a truncated flag, and an explicit 'never invokes or modifies' guarantee. This is rich behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose and the key 'why it matters' clause before mechanics. It is dense and long, and the trailing 'Signature/Phase/cost/idempotency' line partially duplicates annotations, but nearly every sentence carries actionable detail and little is truly wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by spelling out the return shape { count, scanned, truncated, root, callbacks } and each entry's fields (remote, class, slot, callback = { source, line, name, pointer, isC }), plus failure behavior. Nothing an agent needs to call and interpret it is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents root, limit, maxScan, and threadContext with defaults and clamps. The description only restates root's default (whole DataModel) and the limit/maxScan caps, adding marginal meaning over the schema; baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Find and disassemble the invoke callbacks bound to RemoteFunctions and BindableFunctions') and immediately distinguishes them from signal events, which an agent can tell apart from siblings like get-connection-info or find-instances-with-connections without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly explains why and when to reach for it: invoke callbacks hold request/response logic and are 'frequently the single most valuable functions to reverse', and it notes they are NOT visible to getconnections so getcallbackvalue is the only retrieval path. It also names the downstream pivots (inspect-closure, get-closure-constants, hook-function, scan-proto-functions).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inspect-closureInspect a function/closure by referenceA
Read-onlyIdempotent

Resolve a Luau expression to a function and dump everything about it in one call: whether it's a Lua or C closure, its name/source/line/param count/upvalue count (via getinfo), its constants, its upvalues (captured values), its nested-proto count, and its function hash. This is the by-reference counterpart to the gc-scan tools (scan-closures-by-*) — use it when you already have a handle on the function (e.g. a remote's OnClientEvent handler, a metamethod, or getsenv(script).someFunc). Requires getinfo/getconstants/getupvalues as available; missing capabilities are simply omitted. Returns { Target, Info, Constants, Upvalues, ProtoCount, FunctionHash } or { error }. Signature: { functionPath: string, includeConstants: any?, includeUpvalues: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: debug closure primitives. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
functionPathYesLuau expression resolving to a function, e.g. 'getsenv(game.Players.LocalPlayer.PlayerScripts.Main).update', 'getrawmetatable(game).__namecall', or 'getconnections(game.Workspace.Part.Touched)[1].Function'. Evaluated as `return <functionPath>`.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
includeUpvaluesNoInclude the function's upvalues — its captured variables (default true).
includeConstantsNoInclude the function's constants table (default true).

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only/idempotent/non-destructive, so the safety profile is covered. The description adds real behavior beyond that: it depends on getinfo/getconstants/getupvalues being available and silently omits missing capabilities, and it names the exact return payload fields. It stops short of detailing failure modes beyond 'inspect tool-schema', hence not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and what gets dumped, then usage routing, then return shape. The trailing metadata block (phase/cost/idempotency/requires/capabilities/produces/safety/on-failure) is partly redundant with annotations and the schema, but the payload itself is dense and earns its space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by enumerating the return object ({ Target, Info, Constants, Upvalues, ProtoCount, FunctionHash } or { error }), the required context (active-client, resolved-target), and capability dependencies. An agent has everything needed to call and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (functionPath, includeConstants, includeUpvalues, threadContext) are already documented in the schema, including defaults and the `return <functionPath>` evaluation semantics. The description's signature line largely restates the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (resolve + dump) on a specific resource (a function/closure), enumerating exactly what is reported: closure kind, getinfo fields, constants, upvalues, proto count, hash. It explicitly positions itself as the by-reference counterpart to the gc-scan/scan-closures-by-* siblings, so an agent can distinguish it without opening another schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives the precise precondition ('use it when you already have a handle on the function') and concrete examples (remote OnClientEvent handler, metamethod, getsenv(script).someFunc), and names the alternative family (scan-closures-by-*) for the case where you don't have a reference. When-to-use vs. when-not is explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inspect-instance-metatableInspect instance metatableA
Read-onlyIdempotent

Get metatable keys from a target Instance path (e.g., game.Players.LocalPlayer). Signature: { instancePath: string, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: getrawmetatable. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
instancePathYesdotted Roblox instance/value path resolved in the active client.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds value beyond that by naming the required capability (getrawmetatable), the preconditions, cost=medium, and that it produces a structured-observation, though it says nothing about the shape or size of the returned keys.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose well, but the back half is a boilerplate metadata string (Phase/cost/idempotency/Requires/Capabilities/Produces/Safety/On failure) that partially duplicates the annotations and restates the signature. Functional but not tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only inspection tool with full schema coverage and annotations covering the safety profile, the description supplies the remaining needed context: preconditions, required capability, and output type. Absent an output schema, a note on the return shape would be the only meaningful addition.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already fully documented in the schema. The description only restates the signature ({ instancePath: string, threadContext: number? }) without adding format, constraint, or default detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Get metatable keys from a target Instance path') with an inline example path. The focus on 'keys' distinguishes it somewhat from siblings like get-metatable and get-metamethod, but it never explicitly names or contrasts those alternatives, so the agent must infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides preconditions ('Requires: active-client, resolved-target') and a phase tag ('observe'), which implies when the tool fits. However, it never states when to prefer this over get-metatable, get-metamethod, or the set-metatable siblings, so alternative selection is left implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

invoke-closureInvoke a closure by reference with typed argumentsA
Destructive

WRITES LIVE GAME STATE. Invoke an arbitrary live function expression with up to 24 typed arguments and return up to 20 encoded results. Function calls may have side effects, so confirm=true is mandatory. Signature: { functionPath: string, threadContext: number?, arguments: any?, confirm: boolean? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Capabilities: debug closure primitives. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; executes caller-selected behavior in the live client. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmNoExplicit safety acknowledgement; must be true before this state-changing operation runs.
argumentsNoOptional ordered typed arguments forwarded to the selected operation.
functionPathYesLuau expression resolving to the target function.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructive/not-readonly/not-idempotent, and the description adds substantial context on top: confirm=true is mandatory, calls may have side effects, the write is idempotency=contextual-write, and caller-selected behavior executes in the live client. This is meaningful behavioral disclosure beyond the structured hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the critical mutation warning and otherwise organized as dense semicolon-separated metadata where each clause carries a distinct fact. Slightly telegraphic, but no filler sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive tool with no output schema, the description covers safety, mandatory confirmation, prerequisites, result cardinality (up to 20 encoded results), the produced artifact (operation-receipt), and a verification path (assert-state). Nothing essential an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters including types, defaults and the enum for argument kinds. The description's inline signature restates parameter names and optionality but adds no syntax or format detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (Invoke) plus a precise resource (an arbitrary live function expression by reference), with the argument and result cardinality quantified. It is clearly distinguishable from generic execution tools like execute or run-luau, though it never explicitly names near-identical siblings such as call-closure or invoke-method.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States the phase (act) and the preconditions (active-client, resolved-target, explicit-mutation-approval), which tells the agent when this tool is callable. It stops short of naming alternatives or exclusions, so the agent gets clear context but no routing guidance versus call-closure.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

invoke-methodCall a method on a live InstanceA
Destructive

ACTS ON LIVE GAME STATE. Resolve a Luau expression to an Instance and call one of its methods as a colon-call (inst:Method(args...)), returning whatever the method returns. Useful while debugging for :Destroy(), :GetChildren(), :FindFirstChild(name), :Clone(), :GetAttribute(name), :SetAttribute(name, value), :WaitForChild(name), Humanoid:TakeDamage(n), Humanoid:MoveTo(pos), Tool:Activate(), etc. Each argument is a typed value; use kind='raw' for non-primitive arguments (Vector3, Enum, Instance, ...). The call is pcall-guarded. WARNING: many methods MUTATE the game (e.g. :Destroy(), :SetAttribute, :TakeDamage) and the effect is immediate and may replicate — only call methods you understand. Returns { Path, Method, ok, ReturnValues } or { error }. Signature: { instancePath: string, methodName: string, args: {{ kind: "string" | "number" | "boolean" | "nil" | "raw", value: string | number | boolean? }}?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; executes caller-selected behavior in the live client. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
argsNoOrdered list of arguments passed to the method after `self`. Omit or pass [] for a no-argument call like :Destroy() or :GetChildren().
methodNameYesThe exact method name to call (case-sensitive), e.g. 'Destroy', 'GetChildren', 'FindFirstChild', 'Clone', 'GetAttribute', 'SetAttribute', 'TakeDamage'. Invoked as a colon-call so `self` is the instance.
instancePathYesLuau expression resolving to the Instance to call the method on, e.g. 'game.Workspace.Part', 'game.Players.LocalPlayer.Character.Humanoid', or 'game:GetService("Lighting")'. Evaluated as `return <instancePath>`.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructive/non-idempotent, and the description adds materially: the call is pcall-guarded, mutations are immediate and may replicate, explicit-mutation-approval/active-client/resolved-target are required, and it returns {Path, Method, ok, ReturnValues} or {error}. It also warns the agent to only call methods it understands.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the crucial 'ACTS ON LIVE GAME STATE' and the mutation warning, and the example method list aids recognition. It is on the verbose side, with a trailing phase/cost/requires/produces/verify metadata block that partly duplicates annotations and adds length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating live-state tool with no output schema, the description covers prerequisites, return shape, failure behavior, and the safety contract, and points to tool-schema for exact fields. An agent has everything needed to decide and invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents kind, value, methodName, instancePath, and threadContext thoroughly. The description's signature recap and kind='raw' guidance largely restate what the schema says, adding only marginal meaning, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: resolve a Luau expression to an Instance and call one of its methods as a colon-call, returning the method's result. This clearly separates it from generic execution siblings like execute, run-luau, eval-expression, and closure-invocation tools such as invoke-closure.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives strong context ('useful while debugging') and concrete example methods, plus the raw-kind rule for non-primitive args. It does not, however, route the agent to more specific siblings like destroy-instance, set-instance-property, or set-attribute, which cover some of the same mutations with tighter contracts.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

is-c-closureCheck whether a function is a native or wrapped C closureB
Read-onlyIdempotent

Resolve a function and call iscclosure with guarded metadata. Signature: { functionPath: string, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: debug closure primitives. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
functionPathYesLuau expression resolving to the target function.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the description's 'Safety: read-only' is redundant. It does add non-annotation context (cost=medium, requires=active-client/resolved-target, produces structured-observation), which is genuinely useful. It never states the return is a boolean predicate, which is the one behavioral fact an agent would want.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The predicate verb is front-loaded, but the body is a dense metadata dump (phase/cost/idempotency/requires/capabilities/produces/safety) that re-states the annotation set, and the trailing 'On failure: inspect tool-schema...' is a deflection rather than a fact. It could be materially shorter and still carry the same information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a low-complexity read-only predicate with 100% schema coverage and no output schema, so the description needn't explain return values. Preconditions and phase are covered, leaving only the boolean return nature and sibling selection unstated. Complete enough for a tool of this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are fully documented in the schema. The description merely restates the signature ({ functionPath: string, threadContext: number? }) without adding format, semantics, or examples beyond it. Baseline 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The title states a precise predicate (native/wrapped C closure) and the description puts a verb on it ('Resolve a function and call iscclosure'). An agent knows exactly what it checks. However it never differentiates the sibling predicates is-l-closure, is-executor-closure, and is-new-c-closure, so it doesn't tell you which of the family to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives preconditions ('Requires: active-client, resolved-target') and a phase tag ('observe'), which implies when it fits in a workflow. But it names no alternative and gives no when-not guidance, so the agent must infer selection among the several is-*-closure tools itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

is-executor-closureCheck whether a closure originates from the executorB
Read-onlyIdempotent

Call isexecutorclosure with checkclosure/isourclosure aliases to distinguish executor-created closures from game closures. Signature: { functionPath: string, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: debug closure primitives. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
functionPathYesLuau expression resolving to the target function.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so safety is covered. The description adds some non-redundant context (cost=medium, prerequisites active-client/resolved-target, 'structured-observation' output), but 'Safety: read-only' merely restates the annotations and no failure mode or return shape is explained.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loading the aliases is useful, but the description is a dense metadata block padded with fields that duplicate structured data (Safety/read-only, Capabilities, Produces) and a generic 'inspect tool-schema on failure' fallback. Roughly half the sentences restate annotations or schema rather than adding value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity, two-parameter predicate with full schema coverage and rich annotations, the description supplies the needed extras: aliases, prerequisites, phase, and cost. No output schema exists so return values need not be documented, though the relation to sibling closure-classification tools is the one notable gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both parameters carry their own descriptions, so the schema does the heavy lifting. The description only repeats the signature ({ functionPath, threadContext? }) without adding format, semantics, or examples beyond what the schema already states, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific predicate: distinguishing executor-created closures from game closures, with the aliases isexecutorclosure/checkclosure/isourclosure. It is clear what the tool answers, but it never distinguishes itself from close siblings like is-c-closure, is-l-closure, or is-new-c-closure, which an agent selecting among closure-inspection tools would need.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Requires: active-client, resolved-target' line implies a precondition context, and 'Phase: observe' hints at where it fits in a workflow. However, there is no explicit when-to-use/when-not-to-use statement or comparison to the alternative closure-classification tools, leaving selection largely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

is-function-hookedCheck whether a function currently has an executor hookB
Read-onlyIdempotent

Resolve a function and call isfunctionhooked before hooking/restoring it. Signature: { functionPath: string, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Produces: structured-observation. Verify with: is-function-hooked. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
functionPathYesLuau expression resolving to the target function.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

B3.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only/idempotent/non-destructive/closed-world, so safety is covered; the description adds genuine context the annotations do not, namely prerequisites (active-client, resolved-target) and cost=medium. It does not describe the return shape, leaving that partly opaque, but the prerequisites are real added value beyond the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The definition is compact but reads as a metadata field-block; 'Signature' duplicates the schema and 'Safety: read-only' duplicates readOnlyHint, while 'Verify with: is-function-hooked' is self-referential dead weight. The genuinely useful bits (prerequisites, cost, phase) are present but not sharply front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity read-only check with no output schema and full annotation coverage, the essentials are largely present (prerequisites, phase, safety). Still, the description never states what the check returns (boolean vs structured observation), and the self-referential 'Verify with' line leaves the verification story unclear.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters with constraints and meanings. The description's 'Signature: { functionPath: string, threadContext: number? }' merely repeats the schema without adding format or default nuance beyond it. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description's body leans on mechanism ('Resolve a function and call isfunctionhooked') which essentially restates the tool name rather than stating the outcome (a boolean/observation of whether the function is currently hooked). The title clarifies, but a reader relying on the description alone gets a vague sense of purpose. It does gesture at the hooking/restoring workflow, which lightly disambiguates it from siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'before hooking/restoring it' implies the intended usage context relative to hook-function/restore-function, giving implied guidance. However there are no explicit when-to-use/when-not conditions and no named alternatives; 'Verify with: is-function-hooked' unhelpfully points at itself rather than a distinct verification tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

is-l-closureCheck whether a function is a Luau closureB
Read-onlyIdempotent

Resolve a function and call islclosure with guarded metadata. Signature: { functionPath: string, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: debug closure primitives. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
functionPathYesLuau expression resolving to the target function.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so 'Safety: read-only' and 'idempotency=read-only' are redundant restatements. The description does add genuinely new context — cost=medium, phase=observe, active-client prerequisite, and a fallback pointer to tool-schema on failure — which supports a 3 rather than lower.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded, but the body is a semicolon-separated metadata dump with visible filler ('Safety: read-only' duplicates the annotation) and internal jargon ('call islclosure', 'guarded metadata'). Adequate but not tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-param, no-output-schema, heavily-annotated read-only predicate, the description covers prerequisites, phase, and cost. But 'Produces: structured-observation' is vague about the return (presumably a boolean) and nothing distinguishes this predicate from its siblings, leaving a real completeness gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both parameters are already documented in the schema. The description merely re-lists the signature without adding format, semantics, or validation detail beyond what the schema provides; this is the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The title states a specific verb+resource ('Check whether a function is a Luau closure') and the description confirms the mechanism ('resolve a function and call islclosure'). However, it never differentiates from the many sibling predicates (is-c-closure, is-new-c-closure, is-executor-closure), so an agent cannot tell which closure check applies without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It lists prerequisites ('Requires: active-client, resolved-target') but gives no guidance on when to choose this over is-c-closure, is-new-c-closure, or closure-capabilities. No exclusions, no alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

is-new-c-closureDistinguish newcclosure wrappers from native C closuresC
Read-onlyIdempotent

Call isnewcclosure with the iscustomcclosure alias when available. Signature: { functionPath: string, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: debug closure primitives. Produces: structured-observation, created-handle. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
functionPathYesLuau expression resolving to the target function.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

C2.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds genuinely new context (cost=medium, phase=observe, output kinds) that annotations do not carry. However 'Produces: created-handle' sits in tension with the read-only/idempotent hints, since creating a handle implies state mutation, and this is left unexplained.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The content is short but poorly front-loaded: the leading sentence is an alias note rather than the tool's purpose, and the body is a cryptic key=value metadata dump (phase, cost, idempotency, capabilities, produces). The agent has to parse boilerplate before reaching anything actionable, and the 'On failure' sentence just redirects to another tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must convey the return value, but it only says 'Produces: structured-observation, created-handle', never stating that it returns a boolean-style classification or how the wrapper verdict is expressed. It also omits any differentiation from the cluster of sibling is-* closure predicates, leaving the agent without the context needed to pick this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters (functionPath, threadContext) are fully documented in the schema, including the Luau-expression meaning of functionPath and the default behavior of threadContext. The description merely restates the signature types and adds nothing beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description never actually states what the tool determines (whether a function is a newcclosure wrapper); that meaning lives only in the title. It opens with invocation mechanics ('Call isnewcclosure with the iscustomcclosure alias') and 'Capabilities: debug closure primitives', which hint at purpose but do not distinguish it from the many sibling predicates (is-c-closure, is-l-closure, is-executor-closure, new-c-closure). An agent must fall back on the title to understand the tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives preconditions ('Requires: active-client, resolved-target') and a phase tag ('observe'), which is implied usage context. However it offers no guidance on when to choose this predicate over its close siblings like is-c-closure or is-l-closure, and no when-not conditions. Usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

is-parallel-contextCheck whether the current executor thread is parallelB
Read-onlyIdempotent

Call isparallel/is_parallel and return a boolean without changing scheduler state. Signature: { threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered structurally. The description adds value with "without changing scheduler state," "Requires: active-client," and "cost=medium," but these are terse and formulaic rather than rich behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose first, which is good. But it is fragmented into boilerplate key-value fragments (Phase/cost/idempotency/Safety) and closes with a generic "inspect tool-schema" instruction that adds little; a tighter sentence would carry the same information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only boolean check with one optional parameter and full annotations, the description covers return type, non-mutation, prerequisites, and failure handling. No output schema is needed since it returns a boolean. It is nearly complete, with only sibling routing left out.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the single threadContext parameter is fully documented by the schema (optional, server default on omission). The description merely restates the signature ({ threadContext: number? }) without adding meaning, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: calls isparallel/is_parallel and returns a boolean without changing scheduler state. An agent can tell what it does and that it is non-mutating. However, it does not differentiate itself from any sibling, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Phase: observe" and "Requires: active-client" give implied context for when the tool is applicable. There is no explicit when-not or named alternative among the many sibling introspection tools (e.g. is-c-closure, is-readonly), leaving selection largely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

is-readonlyCheck whether a table/metatable is read-onlyA
Read-onlyIdempotent

Resolve a Luau expression to a table (or a metatable expression) and report whether it is locked read-only via isreadonly. Roblox marks core metatables (e.g. getrawmetatable(game)) and many security tables read-only so __index/__namecall can't be swapped; this tells you whether you'd need setreadonly(t, false) before any mutation would take effect. Use it as a safe, non-mutating pre-check before attempting metatable hooks or constant/upvalue edits — it changes nothing. Typical targets: a plain table from getgenv(), or 'getrawmetatable(game)' to confirm the global instance metatable is frozen. Requires isreadonly; returns { Target, TargetType, ReadOnly } or { error } when the value isn't a table or isreadonly is unavailable. Signature: { targetPath: string, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetPathYesLuau expression resolving to the table or metatable to test, e.g. 'getgenv().SomeTable', 'getrawmetatable(game)', or 'getrawmetatable(game.Players.LocalPlayer)'. Evaluated as `return <targetPath>`. isreadonly only applies to tables — userdata/Instances will report an error.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds strong behavioral context beyond annotations: it changes nothing, requires isreadonly, returns { Target, TargetType, ReadOnly } or an error, and explains why Roblox marks metatables read-only. Annotations already declare read-only/idempotent, but the description usefully adds the failure case and the reason the check matters.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose and use-case, but the tail (signature, phase, cost, idempotency, requires, produces, safety, on-failure) is metadata-ish and somewhat repetitive of annotations. It is still readable and mostly earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-param read-only tool with no output schema, the description covers when to use it, error behavior, and return shape. It is close to complete, though it could mention pagination/timeout or threadContext semantics more explicitly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds semantic value by giving concrete examples of targetPath (getgenv().SomeTable, getrawmetatable(game)) and warning that isreadonly only applies to tables, which enriches how the agent should populate the field.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource: resolves a Luau expression to a table/metatable and reports whether it is read-only via isreadonly. It distinguishes itself from siblings like set-metatable-readonly (which mutates) and inspect-instance-metatable (which inspects metatables rather than lock state).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use it: as a safe pre-check before attempting metatable hooks or constant/upvalue edits. It names the alternative action (setreadonly(t, false)) that would be needed if locked, and gives typical targets, which routes the agent correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-actorsList Actor scripts (parallel-Luau VMs)A
Read-onlyIdempotent

Enumerate every Actor instance in the game (getactors). Actors run code in isolated parallel-Luau VMs, so scripts inside them execute outside the normal serial scheduler and are a common place to hide logic. For each Actor this returns its name, full path, where it lives (in tree / nil-parented / detached / CoreGui), how many LuaSourceContainer descendants it has, and — when includeScripts is true — a capped list (30 per Actor) of those scripts with their name, class, and full path. Requires getactors; degrades with a clear error otherwise. Output is capped at 200 actors with a truncated flag. Signature: { includeScripts: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Capabilities: getactors. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
includeScriptsNoIf true (default), list each Actor's descendant LuaSourceContainers (scripts), capped at 30 per actor. Set false for a lighter scan that only reports counts.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safe-read profile (readOnlyHint, idempotentHint, destructiveHint=false), but the description adds real operational traits: a 200-actor output cap with a truncated flag, a 30-scripts-per-actor cap, and a hard getactors prerequisite that degrades with an error. It does partially restate read-only/idempotency which annotations already carry, keeping it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core operation and return shape before trailing metadata (Phase, cost, capabilities, safety), so the most important content comes first. It is somewhat dense with boilerplate meta-fields and closes with a generic 'inspect tool-schema' redirect, but nothing is egregiously wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-format burden well, spelling out per-Actor name, path, location classification, descendant counts, and the conditional script list. It is essentially complete for calling the tool, with only minor overlap against the 100%-covered schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3: both parameters are already fully documented in the schema, including includeScripts' default and 30-per-actor cap. The description's signature line and includeScripts note largely repeat that, adding little new semantic detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Enumerate every Actor instance in the game') and even names the underlying capability (getactors), plus explains what Actors are (parallel-Luau VMs, outside normal scheduler). It stops short of distinguishing itself from near-neighbors like list-script-actors or get-actor-details, so the agent must infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides useful context ('a common place to hide logic', 'Phase: observe; cost=medium') that implies when this is worth running, and notes the getactors dependency and degradation behavior. However it names no explicit alternative or when-not condition despite several overlapping siblings, leaving selection to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-attributesList custom attributes used under a rootA
Read-onlyIdempotent

Walk the descendants of a root instance and collect every unique custom attribute name (via GetAttributes()), with its observed value type, a sample value, and how many instances carry it. Use this when debugging gameplay data that is stored as instance attributes (e.g. QuestId, Health, OwnerUserId) and you want to discover which attribute keys exist in a place without grepping scripts. Returns [{ Name, ValueType, SampleValue, InstanceCount }] sorted by frequency. Work is capped at limit instances scanned. Signature: { root: any?, limit: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
rootNoDotted root path to scan descendants of (e.g. 'game.Workspace', 'game.ReplicatedStorage'). Defaults to 'game' (the whole DataModel). Narrow this to keep scans fast on large places.game
limitNoMaximum number of instances to scan before stopping (default: 1000). Lower this for a quick sample; raise it for exhaustive coverage at the cost of speed.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only/idempotent/destructive, but the description adds real behavioral context: work is capped at `limit` instances scanned, cost=medium, and it 'Requires: active-client'. Some tags (Safety: read-only, idempotency=read-only) merely restate annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Core purpose and scope are front-loaded, and the return shape is stated compactly. The trailing metadata block (Phase/cost/idempotency/Safety/On failure) is verbose and partly redundant, but each clause carries some information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, but the description supplies the return shape [{ Name, ValueType, SampleValue, InstanceCount }], sort order, scan cap, prerequisites, and failure guidance. Everything needed to call it correctly is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (root, limit, threadContext) are already documented with defaults and constraints. The description repeats the signature but adds no format or semantics beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('walk the descendants of a root instance and collect every unique custom attribute name') plus the returned fields. An agent can distinguish this from set-attribute or get-instance-properties without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear when-to-use ('debugging gameplay data stored as instance attributes... discover which attribute keys exist without grepping scripts') and implies the alternative is script grep. No explicit when-not or named sibling alternative, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-clientsList connected Roblox clientsA
Read-onlyIdempotent

List every connected Roblox executor client with its clientId, account, place, and executor. Also reports which client THIS session currently resolves to (after applying its selection). Call this before select-client when more than one client is connected or the target is unknown. Signature: {}. Phase: observe; cost=low; idempotency=read-only. Requires: none. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds genuine value by disclosing that it reports which client the current session resolves to, plus produces=bounded-candidates; the phase/cost/idempotency tail partly restates the annotations rather than extending them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose and when-to-use guidance are front-loaded and efficient. The trailing 'Signature/Phase/Requires/Safety' block is largely boilerplate that partly duplicates the annotations, so it does not fully earn its space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of explaining return content, and it does so by naming the returned fields and the bounded-candidates output. For a zero-param read tool this is essentially complete, though pagination or result-size limits are not addressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4 and the description has no parameter semantics to describe. There is nothing the schema leaves ambiguous that the description needs to compensate for.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List every connected Roblox executor client') and enumerates the returned fields (clientId, account, place, executor). It is clearly distinguishable from select-client, which changes the selection rather than enumerating clients.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to call it: 'before select-client when more than one client is connected or the target is unknown.' It names the alternative tool and the exact triggering condition, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-closure-referencesList retained clone/newclosure handlesA
Read-onlyIdempotent

Read the bounded getgenv().__mcp_closure_refs registry created by clone-function/new-c-closure/new-l-closure and return reusable expressions plus function metadata. Signature: { limit: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Capabilities: debug closure primitives. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoOptional hard result/work budget used to keep output and runtime bounded.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive/non-open-world, so the safety profile is covered. The description adds genuinely useful context beyond that: the result is bounded (limit budget), phase=observe, cost=medium, it requires an active client session, and it emits 'bounded-candidates'. It does not describe the exact shape of returned expressions or how the registry is truncated, but it is well above annotation-only coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose first, then compact labelled facets (Phase/Cost/Requires/Capabilities/Produces/Safety) that are easy to scan. The final 'On failure' sentence is near-boilerplate filler, but overall the text is tight and each label carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must carry return-value burden; it does state the tool returns reusable expressions plus function metadata and flags bounded output. It also names the precondition (active-client) and cost. Missing only finer detail on result shape and truncation behaviour for a tool with a large sibling surface.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters are documented in the schema with types, defaults and min/max bounds. The description only restates the signature, and its 'limit: any?' is looser than the schema's integer 1..200 default 100, so it adds no meaning beyond structured data. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: reads the bounded getgenv().__mcp_closure_refs registry and returns reusable expressions plus function metadata. It also names the sibling tools that populate that registry (clone-function/new-c-closure/new-l-closure), which helps an agent place it among the many closure/gc siblings. The prose is dense executor jargon, but an agent can tell what it returns.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied, via provenance ('registry created by clone-function/new-c-closure/new-l-closure') and the 'Requires: active-client' constraint. There is no explicit when-not guidance and no alternative tool is named (e.g. get-closure-protos or list-gc-functions for other retrieval paths). The 'On failure: inspect tool-schema' line is a fallback, not routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-drawingsList all MCP-managed Drawing objectsA
Read-onlyIdempotent

Read-only inventory of the Drawing overlay this server manages: walks getgenv().__mcp_drawings and returns each registered object's { id, type, visible } so you can see what is currently on screen before updating or removing it. 'visible' reflects the live handle.Visible property (pcall-read; null if it could not be read). Does NOT touch the screen or any handle. Requires the Drawing table (type-guarded). On an executor without it, it returns { error = "Drawing is not available in this executor." }. The list is capped by 'limit'. Returns { count, truncated, drawings[] } or { error }. Signature: { limit: any?, threadContext: number?, timeoutMs: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Capabilities: Drawing. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax number of drawings to return (default 200). Hitting it sets truncated=true.
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the readOnlyHint/destructiveHint annotations by disclosing that 'visible' is a live handle.Visible read via pcall (null on failure), that the tool does not touch the screen or handle, that it requires the type-guarded Drawing table, and the exact error return on an executor lacking it. This is substantial behavioral context for a read-only tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded and the failure/return info is useful, but the trailing boilerplate ('Phase: observe; cost=medium; idempotency=read-only... Safety: read-only... On failure: inspect tool-schema...') partly duplicates the annotations and schema and dilutes the payload.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description supplies the return shape ({ count, truncated, drawings[] } or { error }), the truncation behavior, the executor requirement, and the failure message, leaving nothing an agent needs to invoke or interpret the call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents limit, timeoutMs, and threadContext. The description only restates that the list is capped by 'limit' (and truncated semantics, which the schema also states) and lists the signature; it adds no meaning for threadContext or timeoutMs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read-only inventory of the Drawing overlay') and names exactly what it walks (getgenv().__mcp_drawings) and returns ({ id, type, visible }). This clearly separates it from the mutation siblings draw-create, draw-update, and draw-remove.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit use context: 'so you can see what is currently on screen before updating or removing it,' which implicitly routes the agent to the mutation siblings afterward. It stops short of naming the alternatives directly or stating when-not to use it, so it is clear context rather than full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-filesList entries in an executor workspace folder (UNC listfiles)A
Read-onlyIdempotent

List the files and subfolders directly inside a folder in the executor's workspace, returning their paths. NOTE: this is executor-side file I/O — the connector runs INSIDE the executor, so the path is relative to the executor's workspace directory on the host machine, NOT the Roblox game. Pass an empty string to list the workspace root. Requires the UNC function listfiles(path) -> { string }. The call is type-guarded and pcall-wrapped: if listfiles is missing you get { error = 'listfiles is not available in this executor.' }, and any failure (missing folder) returns { error = }. The returned list is capped at 1000 entries with a 'truncated' flag. Returns { path, files, count, truncated } or { error }. Signature: { path: string, threadContext: number?, timeoutMs: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: readfile. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesPath to the folder within the executor workspace, e.g. '' (root), 'logs', or 'data/snapshots'.
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the readOnly/idempotent annotations: it discloses the type-guard and pcall wrapper, the exact error object when listfiles is missing, the error-on-failure behavior, and the 1000-entry truncation cap with a 'truncated' flag. These are precisely the failure modes an agent needs before invoking.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the key semantic distinction, but the tail is boilerplate-heavy and redundant: 'Safety: read-only' restates destructiveHint=false and readOnlyHint=true, and the Signature/Phase/Capabilities block largely duplicates schema and annotation data.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, and the description compensates fully by specifying the return shape { path, files, count, truncated } and both error shapes. For a read-only listing tool this covers everything needed to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents path, threadContext, and timeoutMs. The description repeats the signature and the empty-string-root tip, which is already in the schema description, so it adds little beyond the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+scope: list files and subfolders directly inside a folder in the executor's workspace, returning paths. It explicitly disambiguates from the Roblox game side and from sibling file tools, so an agent can identify it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives useful context (pass empty string for the workspace root, requires the UNC listfiles function, phase=observe) but never states when to choose this over siblings like read-file, ws-list, or verify-path-exists, nor any when-not condition. Usage is implied rather than directed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-gc-functionsList GC functionsA
Read-onlyIdempotent

Enumerate live Lua closures from getgc() and return debug metadata (name/source/line/upvalues). Useful for live reversing when script paths are unknown. Signature: { nameQuery: string?, sourceQuery: string?, includeCClosures: any?, limit: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Capabilities: getgc. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of functions to return (default: 100).
nameQueryNoOptional case-insensitive substring match against function names.
sourceQueryNoOptional case-insensitive substring match against debug source/short_src.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
includeCClosuresNoInclude C closures from getgc (default: false).

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, non-destructive, so the safety profile is covered. The description adds genuinely new context: it requires an active-client, has cost=medium, depends on the getgc capability, produces bounded-candidates, and points to tool-schema on failure. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose and the 'useful for' clause are front-loaded, but the trailing Signature/Phase/Cost/Capabilities/Produces/Safety/On-failure tag block is listy and partly duplicates structured fields (the Signature repeats the schema). It is information-dense but not maximally economical.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description steps in by naming the returned fields (name/source/line/upvalues) and adds prerequisites (active-client), cost, and failure routing. That covers most of what an agent needs, though it omits result ordering/pagination behavior for what is implicitly a bounded candidate list.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and each parameter has its own description, so the schema carries the semantic load. The description only re-lists the parameter names/types in a Signature line without adding format, matching behavior, or default details beyond the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (enumerate), a specific resource (live Lua closures from getgc()), and the return content (debug metadata: name/source/line/upvalues). This clearly separates it from siblings like list-gc-tables and list-gc-threads, which enumerate different object classes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete use case: 'live reversing when script paths are unknown.' However, it never names an alternative or a when-not condition, despite close siblings such as filter-gc, search-gc-value, or scan-closures-by-name that an agent might weigh against this.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-gc-tablesList GC tablesB
Read-onlyIdempotent

Enumerate table objects from getgc(true) with optional key/value query matching. Signature: { query: any?, limit: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Capabilities: getgc. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoOptional hard result/work budget used to keep output and runtime bounded.
queryNoOptional search text used to filter and rank bounded results.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, non-open-world, non-destructive, so the safety profile is covered. The description adds modestly useful context beyond that: cost=medium, capability getgc, and requires active-client. But 'Produces: bounded-candidates' is vague and no output/pagination behavior is described.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded, which is good, but the body is dense template boilerplate ('Signature', 'Phase', 'Capabilities', 'Produces', 'Safety') and the Signature block redundantly restates the schema. The 'On failure: inspect tool-schema' pointer is meta rather than useful content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, idempotent list tool with 100% schema coverage and full annotations, the description is largely sufficient: it names the source, the precondition, and the safety profile. It lacks any note on result volume/ordering, but nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3 and the schema already documents limit, query, and threadContext. The description adds little, and its phrase 'key/value query matching' slightly mischaracterizes the schema's 'search text' string parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (enumerate) and resource (table objects from getgc(true)), plus the filtering behavior. It is distinguishable from list-gc-functions and find-tables-by-key, but it never explicitly names a sibling or contrast condition, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a precondition (Requires: active-client) and a phase marker (observe), which implies when the tool fits. However, it never states when NOT to use it or which alternative (find-tables-by-key, search-gc-value, filter-gc) to pick, leaving routing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-gc-threadsList GC threadsA
Read-onlyIdempotent

Enumerate thread objects from getgc(true), including coroutine status where available. Signature: { limit: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Capabilities: getgc. Produces: bounded-candidates, structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoOptional hard result/work budget used to keep output and runtime bounded.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/idempotentHint/destructiveHint=false, so the safety profile is covered. The description adds genuine context beyond that: the active-client prerequisite, cost=medium, and the bounded-candidates / structured-observation output characteristics. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then compact metadata tags. Efficient overall, though the signature and Safety: read-only lines partly duplicate structured data (schema and annotations), which is slight redundancy rather than bloat.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, no-output-schema tool, the definition covers prerequisites, cost, phase, and failure handling (inspect tool-schema). Nothing critical is missing, though a brief note on pagination/bounding behavior of the limit budget would make it fully self-sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already fully documented in the schema. The signature line ({ limit: any?, threadContext: number? }) merely restates what the schema provides and adds no extra syntax or default meaning, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Enumerate) and resource (thread objects from getgc(true)), and even notes the coroutine-status enrichment. This clearly distinguishes it from close siblings like list-gc-tables, list-gc-functions, and list-registry-objects without opening their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides prerequisites (Requires: active-client) and a phase (observe), which gives implied usage context. However, it never names an alternative tool or states when NOT to use this versus filter-gc or search-gc-value, leaving selection to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-global-env-keysList global env keysA
Read-onlyIdempotent

List keys from getgenv/_G with type/value previews. Signature: { query: any?, limit: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoOptional hard result/work budget used to keep output and runtime bounded.
queryNoOptional search text used to filter and rank bounded results.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only/idempotent/non-destructive, and the description adds genuinely new traits: cost=medium, the active-client prerequisite, bounded-candidates output, and an on-failure recovery path. That is meaningful context beyond the structured hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose, then terse clause-per-fact metadata. Slight redundancy between the signature line and the schema, but no filler prose and everything is scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-optional-param read tool with no output schema and rich annotations, the description covers purpose, prerequisite, cost, and failure handling adequately. Only the absence of alternative-routing guidance keeps it short of complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents query, limit, and threadContext with defaults. The description's signature line restates parameter names/types without adding filtering or ranking semantics, so it earns only the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('List keys from getgenv/_G') and adds the return preview ('type/value previews'), so the agent knows what it retrieves. It doesn't differentiate itself from nearby siblings like find-global-xrefs or dump-function-env, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides contextual markers (Phase: observe; Requires: active-client) that imply when the tool applies, and a failure fallback to tool-schema. However, it never states when to prefer this over alternatives such as find-global-xrefs, so usage selection is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-gui-elementsList GUI elements under a rootA
Read-onlyIdempotent

Enumerate the live GUI tree under a root Instance (defaults to the LocalPlayer's PlayerGui) and return a flat list of every GuiObject it contains. This is the fastest way to discover what UI is actually on screen — the exact paths, classes and current Text — so you can then read or drive a specific element with get-gui-text, set-gui-text, click-button or type-text-box. Walks root:GetDescendants() with each property access pcall-guarded so a single hostile element never aborts the scan. For every descendant that is (or, with classFilter, exactly matches) a GuiObject it records { path = GetFullName(), class = ClassName, name = Name, Visible, Text } where Visible and Text are only present when readable. Output is capped at limit; when more elements exist than the cap, truncated is true. Returns { count, truncated, elements } or { error }. Signature: { root: any?, classFilter: string?, limit: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
rootNoLuau expression resolving to the Instance whose descendant GUI tree to list. Defaults to 'game:GetService("Players").LocalPlayer.PlayerGui'. Pass a deeper expression such as 'game:GetService("CoreGui")' or 'game.Players.LocalPlayer.PlayerGui.MainMenu' to scope the walk. Evaluated as `return <root>`.game:GetService("Players").LocalPlayer.PlayerGui
limitNoMaximum number of elements to return (default 200). The walk stops adding once this many matches are collected and sets `truncated` to true so you know to scope the root or filter more tightly.
classFilterNoOptional exact ClassName to keep, e.g. 'TextButton', 'TextLabel', 'TextBox', 'ImageButton', 'Frame'. When set, only descendants whose ClassName equals this string are returned. When omitted, every GuiObject (IsA('GuiObject')) is returned. Case-sensitive.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish read-only/idempotent/non-destructive, and the description adds substantial behavioral context beyond them: pcall-guarded property access so one hostile element cannot abort the scan, conditional presence of Visible/Text fields, the limit cap with a truncated flag, and the { count, truncated, elements } or { error } return contract.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose and workflow before metadata, and most sentences earn their place. However the 'Signature: { root, classFilter, limit, threadContext }' line duplicates the input schema verbatim, and the trailing failure boilerplate adds little, so there is minor waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description fully specifies the return shape (count/truncated/elements or error), the default root, and truncation semantics. For a 4-parameter, zero-required observation tool, an agent has everything needed to invoke and interpret it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so root, limit, classFilter and threadContext are already fully documented in the schema. The description mostly restates this via the 'Signature' line and reiterates classFilter exact-matching and the limit/truncated behavior, adding little beyond the structured fields. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (enumerate), resource (live GUI tree under a root Instance), and output shape (flat list of every GuiObject). It also distinguishes itself from siblings by naming the downstream tools it feeds (get-gui-text, set-gui-text, click-button, type-text-box), so an agent can place it in the workflow without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly frames the tool as 'the fastest way to discover what UI is actually on screen' and routes the agent to the correct follow-up tools for reading or driving a specific element. Strong positive guidance, but it never states when a sibling (e.g. find-hidden-guis, get-hidden-ui, list-rendered-instances) should be preferred instead, so no explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-hooksList active hooks installed via this MCPA
Read-onlyIdempotent

List every function/metamethod hook installed through hook-function / hook-metamethod in this session. Each entry shows the key (pass it to restore-hook to undo), the kind (function or metamethod), the target expression, the metamethod name (if any), and the type of the stored original. Use this to audit what you've hooked before restoring — important because hooks are global and persistent. Reads getgenv().__mcp_hooks. Signature: { threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Verify with: is-function-hooked. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only/idempotent/non-destructive, and the description adds real behavioral context beyond that: it reveals the backing store (getgenv().__mcp_hooks), warns hooks are global and persistent, notes it requires an active client, and names is-function-hooked for verification. Cost and produce-class are minor boilerplate but the underlying-mechanism disclosure is genuinely useful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and the audit-before-restore workflow are front-loaded in the first two sentences, and the entry-field enumeration is informative. The trailing metadata block repeats annotation-derived facts ('idempotency=read-only', 'Safety: read-only') and the schema signature, adding minor redundancy but staying structured and skimmable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully enumerates what each entry contains (key, kind, target, metamethod name, original type) and the storage source, so an agent knows what it will receive. Requires/produces/verify fields round it out; nothing critical for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Single optional parameter with 100% schema description coverage, so the schema already fully documents threadContext. The description only restates the signature ({ threadContext: number? }) without adding format or default semantics beyond what the schema provides; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('List every function/metamethod hook installed') and names the exact origin tools (hook-function / hook-metamethod). An agent can distinguish it from restore-hook and is-function-hooked without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Use this to audit what you've hooked before restoring', which ties the tool to a concrete workflow and names restore-hook as the follow-up. It also justifies the need ('hooks are global and persistent') but does not state an explicit when-not or list alternative audit tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-instance-signalsList instance signalsA
Read-onlyIdempotent

Enumerate the RBXScriptSignal members (events) of a Roblox Instance and report how many connections each has. Connection counts require an executor exposing getconnections(signal); if it is unavailable the signals are still listed but ConnectionCount is reported as null with a note. Returns { Instance, GetConnectionsAvailable, Signals: [{ SignalName, ConnectionCount, Note? }] } or { error } when the instance cannot be resolved. Signature: { instancePath: string, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: getconnections. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
instancePathYesLua expression resolving to the target Instance, e.g. 'game.Players.LocalPlayer' or 'game.Workspace.Part'. Evaluated as `return <instancePath>`.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only/idempotent/non-destructive, and the description goes well beyond them: it discloses the getconnections dependency, the degraded fallback behavior (signals still listed, ConnectionCount=null with a Note), the full return shape, and the error case. This is unusually rich behavioral context for a read tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose followed by fallback behavior and return shape, which is well ordered. There is minor redundancy (read-only appears as both idempotency=read-only and Safety: read-only, and the Signature line duplicates the schema), but no sentence is truly wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a medium-complexity read tool with no output schema, the description supplies the return shape, the degraded-mode semantics, prerequisites, and the error case, so an agent can call it correctly. Only marginal gaps remain (e.g., what the Note text contains).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents instancePath (with Lua-expression semantics) and threadContext. The description restates the signature but adds no format or constraint detail beyond what the schema provides, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Enumerate) and resource (RBXScriptSignal members/events of a Roblox Instance) with the scope of also reporting per-signal connection counts. This clearly separates it from a pure connection-lister, though it never names the relevant siblings (list-signal-connections, count-signal-connections) to route the agent explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides usage context via metadata (Phase: observe, cost=medium, Requires: active-client, resolved-target) and a fallback condition when getconnections is unavailable. However, it never states when to prefer this over alternatives like list-signal-connections or find-instances-with-connections, so the when-to-use guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-lua-statesList every active LuaStateProxyA
Read-onlyIdempotent

Call getactorstates(), cap output at 256 states, and return compact state/actor metadata plus reusable References. Signature: { includeActors: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Capabilities: getluastate. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
includeActorsNoWhether to include actors in the bounded result.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only/idempotent/non-destructive, so the bar is lower, and the description still adds real value: the 256-state output bound, the active-client precondition, and the failure path (inspect tool-schema). Bounded-result behavior is exactly the kind of trait annotations cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded and compact, but several clauses merely restate structured data: 'Safety: read-only' duplicates readOnlyHint and 'idempotency=read-only' duplicates idempotentHint, while the signature duplicates the schema. That label-value padding dilutes rather than wastes away.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with no output schema, it covers the essentials an agent needs: return shape ('compact state/actor metadata plus reusable References'), the result bound, the precondition, and where to look on failure. Adequate and nearly complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters, making 3 the baseline. The inline signature repeats the schema and even renders includeActors as 'any?' which is less precise than the schema's boolean, so it adds no meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: it calls getactorstates(), caps output at 256, and returns state/actor metadata plus References. That is clearly distinct from singular siblings like get-lua-state. It does not, however, explicitly differentiate itself from get-lua-state-actors, which sounds like the same territory.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives preconditions (Requires: active-client) and a phase marker (Phase: observe), which implies when to reach for it, but names no alternative tools and no when-not condition. Guidance is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-registry-objectsList registry objectsA
Read-onlyIdempotent

Inspect debug.getregistry() and summarize object types for runtime reversing. Signature: { threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so 'Safety: read-only' and 'idempotency=read-only' mostly repeat structured data. The description does add genuinely new context — cost=medium and the active-client prerequisite — but 'Produces: bounded-candidates' is too vague to count as real behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded and the metadata clauses are compact and information-dense rather than padded. The trailing 'On failure: inspect tool-schema...' sentence is boilerplate that would apply to nearly any tool, but it does not bloat the description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, one-optional-param observe tool with annotations covering the safety profile, this is close to complete: it names the API, the phase, the prerequisite, and the cost. The only gap is that 'bounded-candidates' does not really tell the agent what shape the summarized output takes, and no output schema exists to fill that in.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single optional parameter is fully documented in the schema ('omit it to use the server default'). The description only restates the signature '{ threadContext: number? }', adding no meaning beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — 'Inspect debug.getregistry() and summarize object types' — naming the exact underlying API and the goal (runtime reversing). It is distinguishable from GC-focused siblings like list-gc-tables or list-global-env-keys, though it never explicitly contrasts itself with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Phase: observe' plus 'Requires: active-client' gives a situational frame and a prerequisite, which is more than nothing. However, there is no when-not guidance and no mention of which sibling to use instead when the agent wants GC tables, threads, or globals rather than the registry.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-remotesInventory every remote in a subtree of the DataModelA
Read-onlyIdempotent

Read-only inventory of every remote/bindable object in a part of the game tree. Resolves root (a Luau expression, default the whole game), walks its :GetDescendants() pcall-guarded, and collects every instance whose ClassName is RemoteEvent, RemoteFunction, UnreliableRemoteEvent, BindableEvent, or BindableFunction. For each it records its full path (GetFullName), class, and name. This is the map you build BEFORE spying: use it to discover which remotes exist and where, then feed an interesting path into get-remote-signature (to learn its shape) or monitor-remote (to watch one remote's traffic). Complements get-remote-spy-logs, which shows calls that have already happened; this shows the static set of remotes regardless of whether they have fired. Scanning the entire DataModel can be large, so results are capped at limit (a truncated flag is set when the cap is hit) and a per-class tally is always returned. Uses only :GetDescendants and reflection — no special executor functions required. Returns { ok, root, total, scanned, byClass, truncated, remotes } where remotes is a list of { path, class, name }, or { error }. Signature: { root: any?, limit: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
rootNoLuau expression resolving to the Instance to scan under, e.g. 'game', 'game:GetService("ReplicatedStorage")', or 'game.Players.LocalPlayer.PlayerGui'. Its :GetDescendants() is walked. Evaluated as `return <root>`. Defaults to 'game' (the whole DataModel).game
limitNoMaximum number of remote/bindable entries to return (default 400). When the subtree contains more matching instances than this, the list is truncated to `limit` and `truncated` is set true (the byClass tally and `total` still reflect everything that was scanned up to the cap).
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, but the description adds substantial context beyond them: pcall-guarded traversal, truncation at `limit` with a truncated flag, always-returned per-class tally, no executor functions required, requires active-client, and phase/cost metadata. This is rich behavioral disclosure, not restatement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense and somewhat long, but front-loaded with the core purpose and organized as purpose -> mechanics -> routing -> return shape -> signature/meta. Nearly every sentence adds usable information, though some metadata (phase/cost/idempotency) is redundant with annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although no output schema exists, the description enumerates the return shape ({ ok, root, total, scanned, byClass, truncated, remotes } with per-entry fields) and error case, plus the truncation/failure guidance. For a complex, tree-scanning tool this is complete enough to call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema fully documents root, its default, limit's cap/truncation, and threadContext. The description largely restates these and adds the signature list; it doesn't add format or constraint meaning beyond what the schema already carries, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (read-only inventory of remotes/bindables) and precisely defines scope: walks a subtree's :GetDescendants() and collects instances of five named classes. It explicitly distinguishes itself from get-remote-spy-logs (static set vs already-fired calls) and names the downstream consumers, so an agent can place it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit workflow: 'the map you build BEFORE spying', then feed a path into get-remote-signature or monitor-remote. It also states the exclusion case against get-remote-spy-logs, so when-to-use and when-to-use-something-else are both covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-rendered-instancesgetrendered — list instances currently being renderedA
Read-onlyIdempotent

Enumerate the instances the engine is currently rendering via getrendered() and return { count, samples }, where each sample is { class, name, path }. This is the set of objects actually on screen this frame — useful for ESP/render auditing, spotting which parts/GUIs are visible, or correlating a render spike with specific instances. The full count is reported even though only a capped sample of entries is returned. Requires getrendered — type-guarded and pcall-wrapped, returning { error } when missing or on failure. Returns { count, truncated, samples } or { error }. Signature: { limit: number?, threadContext: number?, timeoutMs: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of sample entries to return (default 100). The full count is always reported.
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Even with annotations covering read-only, idempotent, non-destructive behavior, the description adds substantial context: the getrendered dependency is type-guarded and pcall-wrapped, failures return { error }, the result shape is { count, truncated, samples }, and the full count is reported while samples are capped. It also notes prerequisites like active-client and failure guidance.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded and the content is generally structured, but the description is dense and somewhat repetitive — for example, return details and read-only/safety metadata are restated. Most sentences still earn their place for a complex runtime inspection tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description adequately explains return values and failure behavior, and the annotations cover safety. Prerequisites, cost, phase, and fallback guidance are all present, making it complete enough for an agent to call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, so the three optional parameters are already fully documented in the schema. The description includes a signature summary but does not add meaningful semantics beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: it enumerates instances currently being rendered by the engine and returns a structured sample. It also clarifies the scope ('set of objects actually on screen this frame'), but it does not explicitly differentiate itself from nearby siblings such as get-render-stats or list-gui-elements.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives concrete use cases — ESP/render auditing, spotting visible parts/GUIs, and correlating render spikes with specific instances — which makes the intended context clear. It does not, however, name alternative tools or state when not to use this one.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-roblox-windowsList Roblox OS windowsA
Read-onlyIdempotent

List the Roblox player processes on the SERVER host that own a visible OS window, with their process id and window title. Windows-only: runs a PowerShell Get-Process RobloxPlayerBeta query and returns the windows that have a non-empty MainWindowTitle. Use this before screenshot-window when more than one Roblox window may be open, to pick the right processId. Returns { windows: [{ pid, title }] } (empty if none are open). Signature: {}. Phase: observe; cost=low; idempotency=read-only. Requires: none. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is partly redundant. The description still adds real context the annotations cannot: Windows-only execution, the underlying PowerShell query, that the result is empty when no windows are open, cost=low, and a failure route to tool-schema. Missing only rate/timeout behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The what/why is front-loaded in the first two sentences and the return shape follows immediately. The metadata footer (Phase, cost, idempotency, Requires, Produces, Safety, On failure) is somewhat boilerplate and restates the annotations, but it is compact and consistently formatted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, but the description supplies the return contract ({ windows: [{ pid, title }] }, empty if none). Combined with the Windows constraint, failure guidance, and read-only safety, an agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters, so the baseline of 4 applies; the explicit 'Signature: {}' confirms no arguments are accepted. There is no parameter semantics to describe beyond that.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (list), resource (Roblox player processes on the SERVER host), and scope constraint (own a visible OS window), plus the exact implementation path (PowerShell Get-Process RobloxPlayerBeta with non-empty MainWindowTitle). This is precise enough to distinguish it from screenshot-window and every other listing sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger: 'Use this before screenshot-window when more than one Roblox window may be open, to pick the right processId.' That names the alternative tool and the exact condition that selects it, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-runtime-modulesList runtime modulesA
Read-onlyIdempotent

Find ModuleScripts in game + nil instances and report metadata for reversing. Signature: { includeNil: any?, limit: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates, operation-receipt. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoOptional hard result/work budget used to keep output and runtime bounded.
includeNilNoWhether to include nil in the bounded result.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only/idempotent/non-destructive/non-openWorld. The description goes beyond them with phase, cost tier, the active-client precondition, and the produced artifacts (bounded-candidates, operation-receipt), which materially help the agent reason about calling it. Minor gap: it doesn't mention rate or result-size behavior beyond the limit param.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded and the metadata is terse, but restating the signature '{ includeNil: any?, limit: any?, threadContext: number? }' duplicates the schema without adding meaning and bloats an otherwise compact definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description should carry the return-value burden; 'Produces: bounded-candidates, operation-receipt' is only a vague tag and doesn't explain what metadata is actually reported. Safety and preconditions are covered, so it is adequate but incomplete for a read tool returning structured metadata.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and all three parameters are documented in the schema, so the description's inline signature adds no semantics beyond restating names/types. Correct baseline of 3 when structured fields do the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

'Find ModuleScripts in game + nil instances and report metadata for reversing' gives a specific verb (find/list) and resource (ModuleScripts) plus the distinguishing scope of including nil instances. However, it does not name the obvious sibling find-module-scripts, so an agent must infer the boundary between the two.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Phase: observe; cost=medium; Requires: active-client' supplies usable invocation context (precondition and phase). But there is no when-to-use-vs-alternatives guidance, and neither find-module-scripts nor get-module-source is referenced as a related or preferred tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-script-actorsList script actorsA
Read-onlyIdempotent

List Actor instances and contained LuaSourceContainer scripts for actor-based reverse analysis. Signature: { limit: any?, threadContext: number? }. Phase: orchestrate; cost=medium; idempotency=read-only. Requires: active-client. Capabilities: getactors. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoOptional hard result/work budget used to keep output and runtime bounded.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent and non-destructive, so the safety profile is covered. The description adds non-redundant context: cost=medium, requires an active client, the getactors capability, and bounded output. These are useful operational facts beyond the structured annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded, which is good, but roughly half the text is metadata labels that partly duplicate the annotations ('idempotency=read-only', 'Safety: read-only') or the schema (the signature line). It is compact but padded with low-value fields rather than earning every sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description should carry more of the return-shape burden; 'Produces: bounded-candidates' is vague about what a caller actually gets back. The active-client prerequisite and budget framing are present, but the result format and how it differs from list-actors remain underspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both 'limit' (hard result/work budget, default 250) and 'threadContext' (optional Roblox thread identity). The description only restates the signature as 'limit: any?, threadContext: number?' and adds no new meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (Actor instances and contained LuaSourceContainer scripts) plus a use case (actor-based reverse analysis). It does not, however, differentiate itself from the very similar sibling 'list-actors', and an agent cannot tell from the text alone which of the two to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for actor-based reverse analysis' and 'Phase: orchestrate' imply a usage context, and 'Requires: active-client' is a real prerequisite. But there is no explicit when-to-use guidance, no when-not, and no routing against alternatives like list-actors, get-actor-details, or actor-capabilities.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-signal-connectionsList all connections on a signalA
Read-onlyIdempotent

Enumerate every connection on an RBXScriptSignal and, for each, report its full Connection metadata: Index, Enabled, LuaConnection, ForeignState, whether it has a Function/Thread, and (for Lua connections) the connected function's Source script, Name, LineDefined and NumParams. This is the main tool for answering "what is listening to this event and where is each handler defined?". Requires getconnections; degrades with a clear { error } if unavailable. Returns { Signal, Instance?, ConnectionCount, Connections: [...] }. Signature: { instancePath: string, signalName: any?, includeFunctionInfo: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: getconnections. Produces: bounded-candidates, created-handle. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
signalNameNoName of the RBXScriptSignal member on the instance (e.g. 'Touched', 'Changed', 'OnClientEvent'). Leave empty/omit if instancePath already evaluates to the signal.
instancePathYesLua expression resolving to the Instance that owns the signal (e.g. 'game.Workspace.Door', 'game.Players.PlayerAdded' parent). Evaluated as `return <instancePath>`. Leave signalName empty if this expression already resolves to the signal itself.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
includeFunctionInfoNoWhen true (default), resolve each Lua connection's function source/line/name via debug.info. Set false for a faster summary that only reports counts and connection flags.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the readOnly/idempotent annotations by disclosing the getconnections dependency, the graceful { error } degradation path, phase=observe, cost=medium, required preconditions (active-client, resolved-target), and produced artifacts. This is exactly the behavioral context annotations can't carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose and answer-value, then layers structured metadata lines. It is dense but each line is informative, though the signature/phase/cost block runs slightly long relative to the core explanation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only enumeration tool with no output schema, the description supplies the full return shape ({ Signal, Instance?, ConnectionCount, Connections }), the failure behavior, and preconditions, leaving nothing an agent needs unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the four parameters (including defaults and the includeFunctionInfo=true behavior) are already documented. The description's signature line mostly restates the schema, adding little semantic meaning beyond it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Enumerate every connection on an RBXScriptSignal') and enumerates exactly what metadata it reports per connection. It is clearly distinguishable from siblings like count-signal-connections or get-connection-info.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It self-identifies as 'the main tool for answering "what is listening to this event and where is each handler defined?"', giving a clear usage context and a prerequisite (requires getconnections). It does not, however, name an alternative tool or state a when-not condition, so a 4 rather than a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-stringsList all string constants (IDA Strings window)A
Read-onlyIdempotent

Walk every Luau function in the GC and collect their string constants — the runtime equivalent of IDA's Strings window. Reports each unique string, how many functions reference it (xref count), and a sample owning script, sorted by frequency. Filter by substring and minimum length to cut noise. Great for finding interesting literals (remote names, URLs, error messages, anti-cheat tags) and then pivoting with find-string-xrefs. Requires getgc + getconstants; caps the scan and flags truncation. Signature: { filter: any?, minLength: any?, limit: any?, maxScan: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax unique strings to return, most-referenced first (default 200).
filterNoOnly include strings containing this substring (case-sensitive). Empty = all strings.
maxScanNoMax GC functions to scan (default 9000).
minLengthNoIgnore strings shorter than this (default 4) to cut noise.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover safety (readOnly, idempotent, non-destructive), and the description adds real context beyond them: prerequisite capabilities (getgc + getconstants), that the scan is capped and truncation is flagged, that results are sorted by reference frequency, and the failure path (inspect tool-schema). The 'idempotency=read-only' line merely restates the annotations, keeping this below a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening purpose and pivot sentence are strong and front-loaded, but the trailing metadata block (Signature, Phase, cost, idempotency, Requires, Produces, Safety, On failure) pads the definition and partly restates the schema and annotations. The Signature line in particular is redundant and less accurate than the schema it summarizes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-shape burden and does so reasonably: it states each result is a unique string with an xref count and a sample owning script, sorted by frequency, and that truncation is flagged. Prerequisites and failure handling are noted, leaving little an agent needs missing to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all five parameters including defaults and meanings. The description's 'Filter by substring and minimum length to cut noise' and its Signature line duplicate that content, with the signature actually being less precise (any? vs the schema's integer types), so it adds no meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource with exact scope: 'Walk every Luau function in the GC and collect their string constants.' It gives a memorable analogy (IDA's Strings window) and names the natural next-step sibling (find-string-xrefs), so an agent can place it against list-gc-functions, search-gc-value, and find-functions-by-constant without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly frames the use case (finding literals like remote names, URLs, anti-cheat tags) and the follow-up pivot to find-string-xrefs, plus how to trim noise via filter/minLength. It stops short of explicit when-not-to-use guidance or contrasting itself with the closest alternatives such as find-functions-by-constant or find-string-in-tables.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-toolsBrowse/search this server's tools by categoryA
Read-only

Discover the right tool without scanning all of them. Call with NO arguments to see every category and its count plus the total. Pass { category } to list that category's tools, or { search } to rank tools against a natural-language goal. Results include signatures, required inputs, definition quality, execution phase/cost, capabilities, outputs, contracts, and safety flags. Reads the server's own tool catalog (no game client required).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum detailed tool entries when category/search is used; default 50.
searchNoNatural-language goal or keyword matched with intent aliases across the tool catalog.
categoryNoRestrict to one category (use the exact name shown in the no-argument overview).

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true; the description adds meaningful context beyond that — it reads the server's own catalog with 'no game client required' and discloses what each result entry contains (quality, execution phase/cost, capabilities, safety flags). No mutation, auth, or rate-limit caveats are given, but none are expected for a read-only catalog browse.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the value proposition, then the invocation modes, then the result contents. Slightly dense in the trailing enumeration of return fields, but every sentence carries information an agent needs to pick the right mode.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description carries the return-value burden and does so by enumerating the fields each entry yields. Combined with the no-arg overview behavior and the no-game-client note, an agent has everything needed to call this correctly in any of its three modes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and already documents the three params, but the description adds the key behavioral semantics the schema cannot: that calling with NO arguments produces the category overview, which is the prerequisite for discovering valid category names. It also clarifies that search ranks against a natural-language goal.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb (browse/search) applied to a specific resource (this server's own tool catalog), and the opening line 'Discover the right tool without scanning all of them' frames the tool's role precisely as a discovery/meta tool among query-oriented siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly maps invocation modes to intent: no-argument overview, { category } for a category listing, { search } for goal-ranked results. This is strong when-to-use guidance, though it never routes the agent to alternative siblings like suggest-tools or tool-schema for adjacent needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

load-fileCompile a workspace Luau file without running it (UNC loadfile)A
Read-onlyIdempotent

Compile a Luau file from the executor's workspace into a function WITHOUT executing it, to verify that it parses and to surface any syntax error. NOTE: this is executor-side file I/O — the connector runs INSIDE the executor, so the path is relative to the executor's workspace directory on the host machine, NOT the Roblox game. This tool intentionally does NOT call the compiled function; it only reports whether compilation succeeded. Requires the UNC function loadfile(path) -> (fn?, err?). The call is type-guarded and pcall-wrapped: if loadfile is missing you get { error = 'loadfile is not available in this executor.' }. On a compile error loadfile returns nil plus an error string, surfaced as { compiled = false, error }. Returns { path, compiled, error? } or { error }. Signature: { path: string, threadContext: number?, timeoutMs: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: readfile. Produces: structured-result. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesPath to the Luau file within the executor workspace to compile-check, e.g. 'scripts/main.lua'.
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial value beyond annotations: the path is host-side executor I/O relative to the workspace, not the Roblox game; it requires UNC loadfile and is type-guarded/pcall-wrapped; missing loadfile yields a specific error object; compile failure yields { compiled=false, error }. Annotations only cover the read-only safety profile, so this extra context is genuinely informative.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core is front-loaded and well structured, but the definition is long and its metadata tail ('Phase: observe; cost=medium; idempotency=read-only... Safety: read-only') partly duplicates annotations and the schema signature, and the 'On failure: inspect tool-schema' line is boilerplate. Efficient in places, padded in others.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description supplies return shapes ({ path, compiled, error? } / { error }) and both failure modes. Prerequisites (active-client, resolved-target) and the readfile capability are noted. An agent has everything needed to invoke and interpret the call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description clarifies path semantics well beyond the schema ('relative to the executor's workspace directory on the host machine, NOT the Roblox game'). It also lists type/default notes for threadContext and timeoutMs, adding modest value over the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Compile a Luau file from the executor's workspace into a function') plus a scope qualifier ('WITHOUT executing it'). This clearly separates it from execute-file/run-luau, which would actually run the code.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear use condition ('to verify that it parses and to surface any syntax error') and an implicit exclusion ('intentionally does NOT call the compiled function'). It stops short of naming the sibling to use when you DO want execution, but the context is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lookup-functionLookup function in getgcB
Read-onlyIdempotent

Search live functions by name/source/constant text and return enriched debug metadata for reverse engineering. Signature: { query: string, mode: any?, limit: any?, includeCClosures: any?, includeConstantsPreview: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoWhere to match the query (default: both).both
limitNoMaximum number of functions to return (default: 25).
queryYesCase-insensitive query to match against name/source/constants.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
includeCClosuresNoInclude C closures from getgc (default: false).
includeConstantsPreviewNoInclude a small constants preview if available (default: true).

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/idempotentHint/destructiveHint, so the safety profile is covered. The description usefully adds a prerequisite (active-client) and cost signal, but its "Signature" line merely restates the schema and it says nothing about result size, truncation, or what "enriched debug metadata" contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose sentence is front-loaded and there is no filler prose. But the inline "Signature" block duplicates the input schema verbatim and the "On failure: inspect tool-schema" line offloads rather than informs, slightly diluting structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With annotations covering the safety profile and no output schema required, the description is close to adequate. It still leaves gaps for a medium-cost observation tool: no note on result shape, default limits, or how results relate to siblings like list-gc-functions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description's signature block only repeats parameter names with placeholder types ("mode: any?") and adds no semantics beyond the schema's own enum, defaults, and descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ("Search") and resource ("live functions") and clarifies the match dimensions (name/source/constant text) plus the return value ("enriched debug metadata"). It is clearly distinguishable from generic tools, though it never names the very close siblings like find-functions-by-constant, scan-closures-by-name, or list-gc-functions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides workflow context ("Phase: observe; cost=medium") and a prerequisite ("Requires: active-client"), which implies when the tool fits. However, it offers no explicit when-to-use-vs-alternative routing despite a dense sibling set of function-search tools (scan-closures-by-name, find-functions-by-constant, find-function-xrefs).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lua-state-event-monitorMonitor a LuaStateProxy's generic cross-state EventA
Destructive

WRITES LIVE GAME STATE when starting/stopping. Connect to LuaStateProxy.Event, retain a bounded 200-event buffer, and poll it by key. Start/stop require confirm=true. Signature: { action: "start" | "poll" | "stop", key: any?, limit: any?, clear: any?, confirm: any?, threadContext: number?, state: any?, stateExpression: any? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval. Capabilities: getluastate. Produces: bounded-event-snapshot. Verify with: assert-state. Safety: MUTATING; changes persistent executor-side observer or hook state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyNoOptional validated input for key.default
clearNoOptional validated input for clear.
limitNoOptional hard result/work budget used to keep output and runtime bounded.
stateNoOptional state selector or reusable state reference returned by a discovery tool.current
actionYesoperation selector; use one of the schema's allowed values.
confirmNoExplicit safety acknowledgement; must be true before this state-changing operation runs.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
stateExpressionNoFor state='expression', a Luau expression resolving to a LuaStateProxy, Actor, or BaseScript.

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false, so the safety profile is known. The description still adds material context beyond them: the specific side effect ("changes persistent executor-side observer or hook state"), the buffer bound, and the confirm=true requirement for start/stop. It stops short of describing failure/partial-write behavior, so 4 rather than 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The critical fact is front-loaded, but the body is a telegraphic metadata dump where several clauses duplicate structured data ("idempotency=contextual-write" vs idempotentHint=false, the full signature vs the schema, "Phase: act; cost=medium"). It is compact but not every line earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 8 parameters, one required, no output schema, and annotations present, the description covers most of what an agent needs: the lifecycle, buffer bounds, the confirm gate, the mutated state, and a suggested verification step (assert-state). It omits return/poll-result shape and what happens on the read path, leaving a small gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description's inline signature restates the fields already in the schema and adds only two semantic points: polling is by key and confirm must be true for start/stop. Undocumented-in-prose parameters like clear and limit gain nothing, so no lift above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb and effect ("WRITES LIVE GAME STATE when starting/stopping") and names the exact resource (LuaStateProxy.Event), plus the retention model (bounded 200-event buffer polled by key). It does not name its near-identical siblings (actor-event-monitor, comm-channel-monitor, watch-value), so an agent must infer differentiation itself, keeping it below 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The action lifecycle (start/poll/stop) and the confirm=true gate imply how to use the tool, and "Requires: active-client, explicit-mutation-approval" gives a precondition. However, it never states when to choose this over actor-event-monitor or comm-channel-monitor, so the when/when-not guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

make-folderCreate a folder in the executor workspace (UNC makefolder)A
Destructive

Create a folder (and any required parent folders) inside the executor's workspace. NOTE: this is executor-side file I/O — the connector runs INSIDE the executor, so the path is relative to the executor's workspace directory on the host machine, NOT the Roblox game. Requires the UNC function makefolder(path). The call is type-guarded and pcall-wrapped: if makefolder is missing you get { error = 'makefolder is not available in this executor.' }, and any failure returns { error = }. Creating a folder that already exists is a no-op on most executors. Returns { path, ok = true } or { error }. Signature: { path: string, threadContext: number?, timeoutMs: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Capabilities: executor filesystem. Produces: structured-result. Verify with: file-exists. Safety: MUTATING; writes executor workspace filesystem. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesPath to the folder to create within the executor workspace, e.g. 'data' or 'data/snapshots'.
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well past the annotations (readOnlyHint=false, destructiveHint=true, openWorldHint=true) by disclosing exact error payloads, that calls are type-guarded and pcall-wrapped, that a missing makefolder yields a specific error object, and that creating an existing folder is a no-op on most executors. This is rich behavioral context an agent cannot get from the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose and the key workspace-scope caveat, then packs prerequisites and error behavior efficiently. It is dense with metadata-style fields (Phase, cost, idempotency, Capabilities, Produces) that read as boilerplate, costing some crispness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return contract itself ({ path, ok = true } or { error }) and specifies failure modes and approval requirements. An agent has everything needed to invoke and interpret the call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so path, timeoutMs, and threadContext are already documented. The description restates the signature and clarifies the path is relative to the executor workspace, which adds marginal value but no syntax or format detail beyond the schema; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create a folder') with explicit scope: the executor's workspace, not the Roblox game, and it creates required parent folders. This distinguishes it cleanly from siblings like create-instance or write-file, which target different resources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides concrete prerequisites ('Requires: active-client, resolved-target, explicit-mutation-approval'), a phase ('act'), and a verification step ('Verify with: file-exists'). It does not, however, name sibling alternatives (e.g., write-file vs. make-folder) or exclusions, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

measure-memoryMeasure Lua heap growth caused by running a snippetA
Destructive

Measure how much Lua heap memory a snippet allocates: read gcinfo() (the live Lua heap size in KB) immediately before running your code, run it once (compiled via loadstring, executed inside a pcall), then read gcinfo() again and report the delta. Use this to spot leaks or unexpectedly heavy allocations — e.g. confirm a function frees what it creates, see how big a table a constructor builds, or detect a snippet that balloons the heap. Caveats: gcinfo reports the WHOLE Lua heap, so concurrent game activity and garbage collection between the two samples add noise; a negative deltaKB means a GC cycle ran (it does NOT mean your code freed memory). For a cleaner reading, allocate enough to dwarf the noise or run a tight loop inside code. A compile error or a runtime error is reported in error with ok=false (the before/after samples are still returned). Requires gcinfo and loadstring (both guarded). Returns { beforeKB, afterKB, deltaKB, ok, error? } or { error }. Signature: { code: string, threadContext: number? }. Phase: act; cost=high; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval, validated-source. Produces: structured-result. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYesThe Luau snippet to measure. Compiled with loadstring and run once between two gcinfo() samples. Whatever it returns is ignored; only the change in heap size is reported. Allocate enough (or loop) so the delta stands out above background GC noise.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: explains that gcinfo measures the WHOLE heap, that concurrent activity and GC add noise, that a negative deltaKB means a GC cycle rather than freeing, and that compile/runtime errors return ok=false with samples still present. It also states the gcinfo/loadstring dependency and mutation intent, matching the destructive/mutating annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Core purpose and caveats are front-loaded and each sentence carries real information. However the trailing template boilerplate (Phase, cost, Requires, Produces, 'On failure: inspect tool-schema...') adds operational noise that dilutes the tight core.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully specifies the return shape ({ beforeKB, afterKB, deltaKB, ok, error? } or { error }), the error semantics, and the two-sample measurement methodology, so an agent has everything needed to invoke and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents code and threadContext fully. The description mirrors the schema's advice about allocating enough or looping to stand out from noise, so it adds little param-level meaning beyond the structured fields. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — measure Lua heap allocation of a snippet — and describes the exact mechanism (gcinfo before/after around a loadstring+pcall run). This clearly distinguishes it from generic execution siblings like execute or profile-code.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete when-to-use scenarios: spot leaks, confirm a function frees what it creates, size a table constructor, detect heap ballooning. It lacks explicit exclusion language naming alternatives (e.g. get-memory-stats or compare-gc-snapshots), keeping it short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

message-boxShow a native host message box (sUNC messagebox)A
Destructive

WRITES HOST STATE — pops up a NATIVE OS dialog on the host machine via messagebox(text, caption, flags). The flags select the dialog's button/icon style (the same bitfield as the Win32 MessageBox), and the call blocks until the user dismisses it, returning the numeric code of the button pressed. Requires messagebox. The call is type-guarded and pcall-wrapped: if messagebox is missing you get { error = 'messagebox is not available in this executor.' }. Returns { result } or { error }. Signature: { text: string, caption: any?, flags: any?, threadContext: number?, timeoutMs: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval. Produces: structured-result. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesThe body text shown in the dialog.
flagsNoThe dialog style bitfield (default 0 = OK button only).
captionNoThe dialog title/caption (default 'MCP').MCP
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial context beyond the annotations: it discloses that the call blocks until dismissal, returns the numeric button code, is type-guarded and pcall-wrapped, and yields { error = 'messagebox is not available...' } when the executor lacks the function. Annotations only flag destructive=true/readOnly=false; the description carries the real behavioral burden.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with 'WRITES HOST STATE' and dense but non-repetitive across the clause about blocking/returns, the failure path, and the safety line. It is long, but nearly every sentence carries distinct information; only the Phase/cost/idempotency metadata feels like padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description still documents the return shapes ({ result } or { error }), the failure message, the verification tool (assert-state), and the mutation safety profile. Nothing an agent needs in order to invoke or interpret the call is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds genuine meaning: it explains flags selects the button/icon style and is the same bitfield as the Win32 MessageBox, and it frames caption and threadContext within a signature. That is real value beyond the schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'pops up a NATIVE OS dialog on the host machine via messagebox(text, caption, flags).' The signature, the blocking behavior, and the returned button code leave no ambiguity about what the tool does, distinguishing it clearly from the execute/eval family of siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives preconditions ('Requires messagebox', 'Requires: active-client, explicit-mutation-approval') but never states when to prefer this over an alternative such as logging or a non-blocking notification. Usage is implied rather than routed against siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

monitor-remoteMonitor one remoteA
Destructive

WRITES LIVE GAME STATE on start. Starts the selected engine if needed and opens a view for one remote instance. fetch reads retained captures since the view started; stop closes the view while shared capture continues. Starting a different engine stops the previous spy on this client. No additional game hooks are installed for views. Signature: { engine: "cobalt" | "ketamine"?, action: "start" | "fetch" | "stop", remotePath: string?, direction: any?, limit: any?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Produces: bounded-event-snapshot. Verify with: assert-state. Safety: MUTATING; changes persistent executor-side observer or hook state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoOptional hard result/work budget used to keep output and runtime bounded.
actionYesoperation selector; use one of the schema's allowed values.
engineNoRemote-spy backend. Only one engine runs per client; starting another stops the current engine. Defaults to cobalt.cobalt
directionNoCapture/control direction; incoming includes client callbacks.Both
remotePathNoOptional dotted Roblox instance/value path resolved in the active client.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (destructiveHint=true, readOnlyHint=false), the description discloses that start WRITES LIVE GAME STATE, that a different engine start silently stops the previous spy on this client, that shared capture continues after stop, that no extra hooks are installed, that explicit-mutation-approval is required, and that it mutates persistent executor-side observer state. This is unusually rich disclosure that matches, and does not contradict, the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The most consequential fact, 'WRITES LIVE GAME STATE on start', is front-loaded first, followed by operational modes and then metadata. The signature, phase, requires/produces, and safety lines are formulaic density but each carries usable information, so it is efficient without being wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and a complex mutating tool, the description supplies preconditions, mutation scope, idempotency character (contextual-write), a verification path (assert-state), and failure remediation (inspect tool-schema). It stops short of describing the shape of fetched captures, but otherwise covers what an agent needs to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all six parameters (including enum values, defaults, and the engine-exclusivity note) are already documented in the schema. The description's signature block largely restates those types and enums rather than adding new semantics, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource set: it starts an engine, opens a view for one remote instance, and exposes start/fetch/stop operations on a remote spy. An agent can tell it monitors a single remote, but it never names or contrasts against close siblings such as remote-spy, trace-remote-traffic, or ensure-remote-spy, so sibling differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It delineates the three action modes (start writes live state, fetch reads retained captures since the view started, stop closes the view) and lists preconditions (active-client, resolved-target, explicit-mutation-approval). However, it gives no explicit when-to-use-this-vs-alternative guidance against the many other remote-spy and traffic-tracing tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

new-c-closureWrap a Luau function as a C closure and retain itA
Read-onlyIdempotent

Call newcclosure(function, debugName?) and store the wrapper under getgenv().__mcp_closure_refs, returning a reusable Reference expression and closure metadata. Signature: { functionPath: string, threadContext: number?, key: any?, debugName: any? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: debug closure primitives. Produces: created-handle. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyNoOptional stable key for getgenv().__mcp_closure_refs. Generated when omitted.
debugNameNoOptional validated input for debug name.
functionPathYesLuau expression resolving to the target function.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is structured. The description adds real value beyond that: it discloses the persistent side effect of registering the wrapper under getgenv().__mcp_closure_refs, the produced handle, and medium cost. No contradiction with the read-only annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the call form and storage effect, which is good. But it carries redundancy: 'idempotency=read-only' and 'Safety: read-only' restate the same fact, and the 'Signature:' block duplicates the input schema verbatim rather than adding meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully explains what is produced (reusable Reference expression, closure metadata, created-handle) and points to tool-schema for exact fields, defaults and an invocation example. Requirements, cost and failure guidance are covered, so an agent has enough to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description restates the signature with types (functionPath, threadContext, key, debugName) but adds nothing the schema does not already document via per-parameter descriptions, including the note about threadContext defaulting to the server thread.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (wrap), resource (Luau function as a C closure), the storage side effect (getgenv().__mcp_closure_refs) and the return shape (Reference expression + closure metadata). An agent can distinguish it from siblings like new-l-closure, clone-function, and is-new-c-closure without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives prerequisites ('Requires: active-client, resolved-target') and a phase/cost framing, which provides context for when it fits. However, it never names an alternative (e.g., new-l-closure for L closures) or states when NOT to use it, leaving the routing decision to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

new-l-closureWrap a C closure as a Luau closure and retain itA
Read-onlyIdempotent

Call newlclosure(function) and store the wrapper under getgenv().__mcp_closure_refs, returning a reusable Reference expression and closure metadata. Signature: { functionPath: string, threadContext: number?, key: any? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: debug closure primitives. Produces: created-handle. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyNoOptional stable key for getgenv().__mcp_closure_refs. Generated when omitted.
functionPathYesLuau expression resolving to the target function.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnly/idempotent/non-destructive, so the bar is lower, and the description still adds substantive behavior: the wrapper is stored under getgenv().__mcp_closure_refs, it yields a reusable Reference expression plus metadata, and it produces a created-handle. It does not say how that handle is later released (a sibling, release-closure-reference, exists), but the key side effect and return shape are disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core action and storage location are front-loaded in the first sentence, followed by compact labeled facts (Phase, Requires, Capabilities, Produces, Safety). The trailing 'On failure: inspect tool-schema...' line is boilerplate-ish but brief, and nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of describing returns and it does (Reference expression plus closure metadata). Preconditions, side effects, and safety are covered; the only real gap is lifecycle guidance on releasing the retained reference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and all three parameters are documented in the schema, so the baseline is 3. The description only echoes the signature ({functionPath, threadContext, key}) without adding semantics such as what threadContext defaults to or what makes a good key, leaving the schema to carry the load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the concrete action (call newlclosure, store the wrapper, return a reusable Reference expression plus closure metadata), which is specific and distinguishes it from analysis-only siblings like inspect-closure. It never states the C-closure-to-Luau-closure direction or the target-creation nature that the title carries, so an agent reading only the description could confuse it with the is-c-closure / new-c-closure family.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It supplies preconditions (active-client, resolved-target, phase=observe, cost=medium) and a fallback pointer ('inspect tool-schema for exact fields'), which is real usage context. It never says when to prefer this over siblings such as new-c-closure, call-closure, or list-closure-references, nor when not to use it, so the guidance remains implied rather than routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

new-lua-state-proxyConstruct a LuaStateProxy for the current stateA
Read-onlyIdempotent

Call LuaStateProxy.new() and return serializable state metadata plus a retained Reference; no raw proxy crosses JSON. Signature: { threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Capabilities: getluastate. Produces: created-handle. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered and 'Safety: read-only' is largely redundant. The description does add genuine context beyond the annotations: a retained Reference is produced and 'no raw proxy crosses JSON.' It omits how/when to release that retained reference, which a construct-and-retain tool should mention.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is compact and front-loads the core action before the machine-oriented metadata. There is minor redundancy ('idempotency=read-only' alongside 'Safety: read-only') and a meta-sentence deferring to tool-schema, but overall it is tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description partially compensates by stating it returns serializable state metadata plus a retained Reference. Given the medium complexity of a state-handle tool, this is nearly sufficient, though it never explains what the retained Reference is for or its lifecycle.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single threadContext parameter, so the schema carries the meaning. The description only restates the signature ('{ threadContext: number? }') without explaining what thread context does or when to supply it, adding no value beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource: it constructs a LuaStateProxy and returns serializable state metadata plus a retained Reference. This distinguishes it from read-only siblings like get-lua-state or list-lua-states, though it never explicitly names those alternatives. The purpose is clear on its own.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides context signals ('Phase: observe; cost=medium; Requires: active-client') that imply when it fits, but it never states when to prefer this over get-lua-state or list-lua-states, nor when not to use it. Usage is implied rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

observe-worldObserve the live world and build semantic entity handlesA
Read-onlyIdempotent

Perform one bounded, read-only Luau observation that fuses the active or custom character rig, current camera, visible PlayerGui objects, nearby 3D parts, ClickDetectors, ProximityPrompts, TouchTransmitters, and Tool instances in the Backpack or character. It returns compact evidence, exact GetFullName paths plus executable bracket-safe Luau expressions, screen position and distance where available, exact scanned/truncation counts, and stable session-local handles. Handles are retained in getgenv().__mcp_world_brain using weak references when supported and structural fingerprints for later resolve-entity calls. The walk uses GetChildren with hard instance/result caps; it never starts frame loops or performs an unbounded GetDescendants scan. Signature: { radius: any?, roots: any?, features: any?, maxInstances: any?, maxResults: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: grounded-evidence, bounded-world-model, semantic-entity-handles, actionable-instance-expressions, custom-character-resolution, camera-state, visible-gui, nearby-3d-entities, interaction-targets, inventory-tools. Verify with: resolve-entity. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
rootsNoBounded traversal roots. Omit a root to avoid scanning that domain; custom rigs generally require workspace.
radiusNoMaximum distance in studs for nearby 3D objects and interactables. Distance uses the resolved character root, then the character pivot, then the camera position.
featuresNoObservation features to include. Character and camera probes are cheap; GUI, nearby, interactables, and tools are collected only from the selected bounded roots.
maxResultsNoMaximum ranked semantic entities returned. Exact candidate counts and result truncation are still reported.
maxInstancesNoHard maximum number of unique Instances examined across all selected roots.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite annotations already covering the safety profile, the description adds substantial behavioral context: it discloses that handles are retained in getgenv().__mcp_world_brain via weak references and structural fingerprints, that the walk uses GetChildren with hard caps, and that it 'never starts frame loops or performs an unbounded GetDescendants scan.' This goes well beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose well, but the body is dense jargon and the 'Signature' restatement plus the long 'Produces:' enumeration duplicate what the schema and annotations already convey, adding bulk without proportional value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden and does so reasonably: compact evidence, GetFullName paths, bracket-safe expressions, positions/distances, scan/truncation counts, and session handles. It also covers failure guidance ('inspect tool-schema') and a verification path, making it nearly complete, though the return shape remains somewhat terse.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so each parameter is already documented in the schema with defaults, ranges, and meaning. The description merely restates the signature without adding new semantics beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and scope ('Perform one bounded, read-only Luau observation') and enumerates exactly what it fuses: character rig, camera, PlayerGui, nearby parts, ClickDetectors, ProximityPrompts, TouchTransmitters, and Tools. This is clearly distinguishable from siblings like get-instance-tree or search-instances because it fuses multiple domains into a session handle model.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives prerequisites ('Requires: active-client'), phase ('Phase: observe; cost=medium'), and a follow-up route ('Verify with: resolve-entity'). It does not, however, explicitly contrast when to prefer this over siblings such as get-instance-tree or discover-character, so a 4 rather than a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

packet-spySpy on outgoing RakNet packetsA
Destructive

WRITES LIVE GAME STATE — installs a RakNet send hook. Captures every OUTGOING low-level packet the client sends (RakNet only exposes outgoing traffic). For each packet it records Size, Priority, Reliability, OrderingChannel, and a hex preview of the payload. action='start' installs the hook (idempotent), 'fetch' returns the captured packets newest-first without stopping, 'stop' removes the hook and clears the buffer. Requires the raknet library. WARNING: packet-level interception is risky and can disconnect or flag the client — stop when done. Signature: { action: "start" | "fetch" | "stop", limit: any?, previewBytes: any?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval. Capabilities: RakNet packet APIs. Produces: bounded-event-snapshot. Verify with: assert-state. Safety: MUTATING; performs external network or socket I/O. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax packets to return on fetch (default 100, newest-first).
actionYes'start' installs the send hook; 'fetch' returns captured packets; 'stop' removes the hook.
previewBytesNoHow many payload bytes to include as a hex preview per packet (default 64).
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Far exceeds the annotations: it discloses that start installs a hook, fetch is non-destructive to the buffer, stop removes the hook and CLEARS the buffer, requires the `raknet` library, and that interception is risky and can disconnect/flag the client. This is exactly the kind of consequence-level detail the destructiveHint annotation cannot convey. The only nuance is that 'start ... (idempotent)' sits in mild tension with tool-level idempotentHint=false, but the qualifier is action-specific and the overall non-idempotent reading holds.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the most important fact ('WRITES LIVE GAME STATE') and stays dense, but the trailing metadata boilerplate (Signature, Phase, Capabilities, Produces, Verify, Safety) partially duplicates the schema and annotations. Efficient overall, with a little redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter mutating tool with no output schema, it covers everything an agent needs: the three actions and their side effects, buffer-clearing behavior, the returned packet fields, the library dependency, and the risk warning. Nothing material to correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents action, limit, previewBytes, and threadContext. The description's signature block largely restates the schema and adds no format, default, or constraint detail beyond it, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: installs a RakNet send hook and captures every OUTGOING low-level packet, enumerating exactly what is recorded (Size, Priority, Reliability, OrderingChannel, hex preview). This cleanly distinguishes it from higher-level siblings like get-remote-spy-logs, trace-remote-traffic, send-packet, and block-packets without needing to open any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context: 'RakNet only exposes outgoing traffic' scopes what it can see, and the WARNING plus 'stop when done' tells the agent the right lifecycle. However it never names an alternative tool (e.g. remote-level spies) or an explicit when-not condition, so routing between it and its siblings is left partly to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

playbook-deleteDelete a Saved Luau PlaybookA
Destructive

Remove a persisted playbook from ~/.executor-mcp/playbooks/. Returns { ok: true, removed: true } when the file was deleted, or { ok: true, removed: false } when it didn't exist. Does not touch any running scripts that loaded this playbook earlier. Signature: { name: string }. Phase: orchestrate; cost=low; idempotency=contextual-write. Requires: explicit-mutation-approval. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; changes server-side persisted/session state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesPlaybook name to delete.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations (destructiveHint=true, non-idempotent) by disclosing exact return payloads for both hit and miss cases, the guarantee that running scripts loaded from the playbook are unaffected, and the approval/verification contract. This is exactly the extra behavioral context annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence and the metadata segments are short and scannable. Some trailing items — notably 'On failure: inspect tool-schema...' — read as template boilerplate rather than tool-specific value, but the overall structure is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description supplies the return shape ({ ok, removed }) for both outcomes, plus approval requirements and post-conditions. For a one-parameter delete tool this covers everything an agent needs to invoke and interpret it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'name' parameter is already documented as 'Playbook name to delete.' The description restates the signature as { name: string } but adds no format, naming-convention, or resolution detail beyond what the schema provides, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — 'Remove a persisted playbook from ~/.executor-mcp/playbooks/' — including the exact storage location. This clearly distinguishes it from siblings like playbook-save, playbook-list, and playbook-run, so an agent can select it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear operating context: it is a persisted-state mutation requiring explicit-mutation-approval and should be verified with assert-state. However, it never explicitly names an alternative (e.g., playbook-list to confirm existence first) or states when NOT to use it, so it stops short of full when/when-not/alternative guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

playbook-listList Saved Luau PlaybooksA
Read-onlyIdempotent

List every persisted playbook in ~/.executor-mcp/playbooks/, optionally filtered to one tag. Returns summary metadata only (name, description, tags, params, timestamps) — fetch the source via playbook-run or by reading the file. Newest-first by updatedAt. Signature: { tag: string?, includeSource: boolean? }. Phase: orchestrate; cost=low; idempotency=read-only. Requires: none. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagNoRestrict to playbooks carrying this tag.
includeSourceNoWhen true, include the full Luau source per entry (default false: metadata only).

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive/openWorld=false, so the safety profile is covered. The description still adds real value beyond them: storage location, metadata-only return (not source), newest-first ordering by updatedAt, and failure guidance. Minor gap: no pagination or empty-result behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and return shape are front-loaded, and sentences are efficient. However, the trailing meta-block (Phase, cost, idempotency, Requires, Produces, Safety) is boilerplate that repeats annotation-derived facts and adds length without much new information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly compensates by enumerating the returned metadata fields and noting source is excluded. Complete for a read-only list tool; only pagination/empty-result handling is unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both params are already documented in the schema. The description's 'Signature: { tag, includeSource }' and 'optionally filtered to one tag' largely restate the schema rather than adding format/syntax detail. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('List every persisted playbook'), scopes it to a directory, and explicitly distinguishes itself from siblings by noting it returns metadata only and routing source retrieval to playbook-run. An agent can tell it apart from playbook-save/run/delete without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context ('optionally filtered to one tag', metadata-only) and names an alternative path for getting source (playbook-run or reading the file). It stops short of explicit when-not-use guidance, but the routing to playbook-run is genuinely useful.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

playbook-runRun a Saved Luau Playbook by NameA
Destructive

Load a saved playbook from ~/.executor-mcp/playbooks/.json, replace any ${param} placeholders with the values from params, then run through the same scripting surface as the script tool — mcp.* calls, mcp.all(), the persistent VM, per-script RPC budget, output capture, and pre-flight ALL apply identically. Returns { result, output } or { error, output } from the run. Signature: { name: string, params: {[string]: string}?, persistent: boolean?, timeoutMs: number?, rpcBudget: number?, threadContext: number? }. Phase: orchestrate; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; changes server-side persisted/session state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesThe saved playbook's name.
paramsNoKey/value substitutions for ${param} placeholders in the source.
rpcBudgetNoOptional numeric value for rpc budget.
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
persistentNoUse the persistent VM (default true). Same semantics as script.persistent.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true and readOnly=false, and the description reinforces this ('Safety: MUTATING; changes server-side persisted/session state'). It goes well beyond the annotations by disclosing the load path, placeholder substitution, execution surface details (mcp.* calls, persistent VM, RPC budget, output capture, pre-flight), the return shapes, and failure handling guidance.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then layered detail; sentences carry distinct information. The metadata block (Phase/cost/idempotency/Requires/Produces/Verify) is dense boilerplate that slightly inflates length, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating, no-output-schema, 6-param tool this is thorough: it names the return shape ({result, output} / {error, output}), the execution surface, prerequisites, verification tool (assert-state), and where to look on failure (tool-schema). An agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema coverage the baseline is 3, but the description adds real meaning: it explains that `params` keys substitute ${param} placeholders in the source, that `persistent` defaults to true with the same semantics as script.persistent, and enumerates the full signature. Still, several params (timeoutMs, rpcBudget, threadContext) get no extra context beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb + resource ('Run a Saved Luau Playbook by Name') and immediately clarifies scope: loads from a named path, substitutes placeholders, and executes through the same surface as the `script` tool. This lets an agent distinguish it from playbook-save/list/delete and from script/execute without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete preconditions ('Requires: active-client, explicit-mutation-approval') and routes execution semantics to the sibling `script` tool. It does not explicitly state when to prefer this over `script`/`execute` (i.e., only for reused, parameterized playbooks), so the exclusion is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

playbook-saveSave a Named Reusable Luau PlaybookB
Read-onlyIdempotent

Persist a named, optionally parameterized Luau snippet to ~/.executor-mcp/playbooks/.json so it can be re-run later via playbook-run. Names must match /^[a-zA-Z0-9][a-zA-Z0-9_-]{0,63}$/. The source is stored verbatim; ${param} placeholders are replaced at run-time by playbook-run. Use tags to group related playbooks (e.g. 'recon', 'farming'). Upserts on existing names. Signature: { name: string, source: string, description: string?, tags: {string}?, params: {string}? }. Phase: orchestrate; cost=medium; idempotency=read-only. Requires: validated-source. Produces: structured-result. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesFilename-safe identifier (used as the JSON file basename).
tagsNoFree-form tags for grouping (e.g. ['recon','farming']).
paramsNoNames of ${param} placeholders the source uses (for documentation + UI hints).
sourceYesThe Luau source to save. Use ${param} placeholders for runtime substitution.
descriptionNoOne-line summary of what this playbook does.

TDQS

B3.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description describes a mutating operation — 'Persist... to ~/.executor-mcp/playbooks/<name>.json' and 'Upserts on existing names' — while the annotations declare readOnlyHint=true. Persisting/upserting writes state, so the description directly contradicts the read-only annotation. This is an Annotation Contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded and the path/constraint are stated early, but the trailing 'Signature:' line duplicates the input schema and the Phase/cost/requires/produces metadata plus the boilerplate 'inspect tool-schema' sentence add bulk without helping invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

It covers storage location, naming rules, upsert behavior, and placeholder handling, which is reasonably complete for a save tool. But with no output schema, the description should say what the call returns (e.g. confirmation/path), and it is silent on that; combined with the mislabeled read-only safety, it is only adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all five parameters including the name pattern and ${param} placeholder usage. The description's 'source is stored verbatim; ${param} placeholders are replaced at run-time by playbook-run' adds cross-tool runtime semantics, but the Signature line merely restates the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Persist) and resource (named Luau snippet/playbook), names the exact storage path, and ties it to the sibling playbook-run for later re-execution. An agent can distinguish this from playbook-list/playbook-run/playbook-delete without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context ('so it can be re-run later via playbook-run') and notes tags are for grouping related playbooks, which implies when to use it. However it never explicitly states when NOT to use it versus playbook-run or playbook-delete, so no exclusions are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

press-keySimulate a keyboard key pressA
Destructive

WRITES LIVE GAME STATE. Simulate a real keyboard key press/release in the running client via VirtualInputManager, so anything listening for keyboard input runs: UserInputService.InputBegan/InputEnded, ContextActionService bindings, default movement, and keybound abilities. The key string is resolved against Enum.KeyCode (e.g. 'E', 'Space', 'W', 'LeftShift', 'F'); an unknown name returns a clean error listing what you passed. Sends KeyDown immediately, optionally holds for holdSec seconds (via task.wait) to emulate a held key, then sends KeyUp. NOTE: VirtualInputManager:SendKeyEvent only works from an exploit/elevated context (injected executor thread) — in an ordinary game script it is locked and will error. Getting the service and each SendKeyEvent are pcall-guarded. WARNING: this drives real input into the game and may move the character or trigger abilities. Returns { key, ok, error? }. Signature: { key: string, holdSec: any?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval. Capabilities: VirtualInputManager. Produces: structured-result. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYesThe Enum.KeyCode name to press, e.g. 'E', 'Space', 'W', 'A', 'S', 'D', 'LeftShift', 'F', 'Return', 'One'. Case-sensitive and must match an Enum.KeyCode member exactly. Resolved as Enum.KeyCode[key].
holdSecNoHow long (seconds) to hold the key down before releasing it. Default 0 = a quick tap (KeyDown then immediate KeyUp). Use a small positive value (e.g. 0.5) to emulate a held key for charge/hold mechanics. Capped at 10 seconds to avoid blocking the executor thread.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well past the annotations: it discloses that VirtualInputManager is locked outside an elevated/injected context and will error, that service acquisition and each SendKeyEvent are pcall-guarded, the KeyDown/holdSec/KeyUp sequencing via task.wait, the 10s cap rationale, and the return shape { key, ok, error? }. The mutating/destructive claims are consistent with destructiveHint=true and idempotentHint=false.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loading the state-mutation warning is good, but the body is padded with generated metadata that repeats itself — the signature line duplicates the schema, and 'Safety: MUTATING; writes live game/client state' restates the opening warning. High signal-to-noise is undermined by this redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex mutating tool with no output schema, the description covers the critical operational facts: elevated-context prerequisite, pcall guarding, hold timing, and the returned fields. Nothing essential to calling it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already documents all three parameters, so the baseline is 3. The description adds genuinely useful semantics beyond the schema: an unknown Enum.KeyCode name yields a clean error listing what was passed, and holdSec emulates a physically held key for charge/hold mechanics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Simulate a real keyboard key press/release in the running client via VirtualInputManager') and enumerates exactly what fires downstream (UserInputService.InputBegan/InputEnded, ContextActionService, default movement, keybound abilities). That specificity separates it from text-entry and pointer siblings such as type-text-box and click-button.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description conveys strong context for when this is usable ('only works from an exploit/elevated context'), but it never routes the agent between this and near-neighbors like virtual-input, click-button, or type-text-box. Usage is implied by the described effect chain rather than stated as an explicit when/when-not choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

profile-codeTime a Luau snippet over repeated runs (micro-benchmark)A
Destructive

Micro-benchmark a Luau snippet by running it runs times and timing each run with os.clock(), returning the wall-clock statistics so you can measure how fast (or how variable) a piece of code is. Use this to compare two implementations, to find out how expensive a function call / loop / property access really is, or to confirm a fix actually made something faster. The code is COMPILED ONCE via loadstring (a syntax error is returned cleanly as { error } and nothing is run); then each of the runs invocations is timed individually inside its own pcall so a runtime error in one run is counted but does not abort the benchmark. Timing measures the run only — compile time is excluded. Note os.clock() resolution is coarse, so for very cheap snippets raise runs or wrap a loop inside your code. Requires loadstring and os.clock (both guarded). Returns { runs, totalMs, avgMs, minMs, maxMs, errorCount, firstError? } or { error }. Signature: { code: string, runs: any?, threadContext: number? }. Phase: act; cost=high; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval, validated-source. Produces: structured-result. Verify with: assert-state. Safety: MUTATING; writes executor workspace filesystem. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYesThe Luau snippet to benchmark. Compiled once with loadstring, then executed `runs` times. It may do anything (call a function, run a loop, read properties); any value it returns is ignored — only the elapsed time per run is measured. Wrap an inner loop here if a single execution is too cheap to time accurately.
runsNoHow many times to execute the compiled snippet (default 1, clamped 1..100000). More runs give a more stable average for cheap code but take longer. Each run is timed and pcall-guarded independently.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial behavior beyond the annotations: loadstring compilation happens once and syntax errors return { error } without running anything, each run is individually pcall-guarded so a runtime error is counted but does not abort, compile time is excluded from timing, and os.clock() resolution is coarse so cheap snippets need more runs or an inner loop. It also discloses the dependency on guarded loadstring/os.clock and the mutating filesystem side effect, consistent with destructiveHint=true.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose and usage scenarios are front-loaded and dense with useful detail. The trailing metadata block (Phase/cost/idempotency/Requires/Produces/Verify with/On failure) is boilerplate padding that does not add tool-selection value, but the body itself earns its length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, and the description fully compensates by naming the return shape ({ runs, totalMs, avgMs, minMs, maxMs, errorCount, firstError? } or { error }) and the error path. Combined with the mutation/safety disclosure, an agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so baseline is 3, but the description adds real value: it explains that `runs` is clamped 1..100000 (the schema only shows a loose min/max), that returned values are ignored, and that an inner loop can be wrapped in `code` when a single run is too cheap to time. threadContext is left unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: micro-benchmark a Luau snippet by running it `runs` times and timing each run with os.clock(). The scoping detail (compile once, time each run, exclude compile time) makes it clearly distinct from execution siblings like run-luau or execute, which run code but do not measure it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete when-to-use scenarios: compare two implementations, find the cost of a function call/loop/property access, confirm a fix made something faster. It does not name an alternative tool or state exclusions (e.g., memory profiling belongs to measure-memory), so the routing guidance stops short of fully explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

queue-on-teleportQueue code to run after the next teleport (sUNC queueonteleport)A
Destructive

WRITES EXECUTOR STATE — queues a chunk of Luau source via queueonteleport(code) so it runs automatically after the NEXT teleport completes, in the destination place. This is how scripts survive a Roblox teleport: the queued code persists across the place change and executes once the new place loads. Requires queueonteleport. The call is type-guarded and pcall-wrapped: if queueonteleport is missing you get { error = 'queueonteleport is not available in this executor.' }. Returns { queued } or { error }. Signature: { code: string, threadContext: number?, timeoutMs: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval, validated-source. Produces: structured-result. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYesThe Luau source to run automatically after the next teleport.
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond annotations: discloses the exact error payload when queueonteleport is missing, the return shape ({ queued } or { error }), the pcall/type-guard behavior, the 'contextual-write' idempotency class, and an explicit MUTATING safety warning that it writes live game/client state. Annotations already mark destructiveHint=true, but the description adds concrete failure modes and preconditions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the mutation warning and the core behavior, then layers requirements and failure handling efficiently. The trailing metadata block (Phase/cost/idempotency/Produces/Verify) and the generic 'On failure: inspect tool-schema' line are slightly boilerplate-heavy, keeping it just under top marks.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return contract itself ({ queued } or { error }), plus preconditions, error semantics, and safety. Nothing an agent needs to invoke and interpret this mutation correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents code, threadContext, and timeoutMs with meaningful descriptions. The description restates the signature and optionality but adds no new syntax, format, or constraint detail. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('queues a chunk of Luau source via queueonteleport(code) so it runs automatically after the NEXT teleport completes, in the destination place') and explains the mechanism distinguishing it from the sibling clear-queue-on-teleport. An agent can immediately tell what this does and its exact effect.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear scenario ('This is how scripts survive a Roblox teleport') that implies when to use it, and the requirement list (requires queueonteleport, active-client, explicit-mutation-approval) provides actionable preconditions. However, it never explicitly contrasts with adjacent executors like execute/run-luau or the sibling clear-queue-on-teleport, so routing is left partly to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read-fileRead a file from the executor workspace (UNC readfile)A
Read-onlyIdempotent

Read the entire contents of a file from the executor's workspace folder and return it as a string. NOTE: this is executor-side file I/O — the connector runs INSIDE the executor, so paths are relative to the executor's workspace directory on the host machine, NOT anything in the Roblox game/DataModel. Requires the UNC function readfile(path). The call is type-guarded and pcall-wrapped: if readfile is missing you get { error = 'readfile is not available in this executor.' }, and any read failure (missing file, permission) returns { error = }. Returns { path, content } or { error }. Signature: { path: string, threadContext: number?, timeoutMs: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: readfile. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesPath to the file within the executor workspace folder, e.g. 'config.json' or 'logs/session.txt'. Resolved by the executor relative to its own workspace directory.
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations (which only cover read-only/idempotent safety) by disclosing the type-guard and pcall-wrapping, the exact error shapes for missing readfile and read failures, and the return shape { path, content } or { error }. This is exactly the behavioral context an agent needs to handle failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and scope are front-loaded, then failure behavior and signature follow. It is dense but nearly every clause earns its place, though the trailing metadata (idempotency=read-only, Safety: read-only) partly duplicates the annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden and does so explicitly ({ path, content } or { error }), and it covers failure modes, prerequisites, and the capability required. An agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters are already documented, and the description merely restates the signature. It adds no syntax, format, or default nuance beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (read entire contents of a file) plus the exact scope (executor's workspace folder, returned as a string). It explicitly distinguishes itself from the Roblox game/DataModel reads, which separates it from in-game siblings like get-script-content and read-path-value.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clear context for when to use it — executor-side file I/O with paths relative to the executor workspace, plus the required capability readfile and active-client/resolved-target prerequisites. It does not name a specific alternative sibling (e.g. load-file or file-exists) to route away from, so it stops short of a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read-path-valueRead the value at a Luau path/expression (precise slot read)A
Read-onlyIdempotent

Evaluate a single Luau expression and report exactly what lives at that slot — without mutating anything. Use it to read a precise location after a heap scan points you at it (e.g. 'getgenv().PlayerData.Coins', 'game.Players.LocalPlayer.leaderstats.Cash.Value', 'require(game.ReplicatedStorage.Config).GodMode', 'getrawmetatable(game).__namecall'), or to spot-check any value/global/field during analysis. The expression is evaluated as return <expression> under a pcall, so a bad path returns a clean { error } instead of throwing. The result reports { type } (Roblox typeof), { value } (a safe encoded scalar — Instances become 'Instance: ', tables/functions become their address), and { isTable }. When the value is a table it also returns { length } (the array-part length via #) and { keys } (up to 50 string keys, so you can see the shape and pick the next field to read). Reading the keys/length is pcall-guarded. This is read-only and is the counterpart to write-path-value. Requires loadstring/load. Returns { type, value, isTable, length?, keys?, keysTruncated? } or { error }. Signature: { expression: string, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-observation. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
expressionYesLuau expression resolving to the value to read, e.g. 'getgenv().PlayerData.Coins', 'game.Players.LocalPlayer.leaderstats.Cash.Value', 'require(game.ReplicatedStorage.Config)', '_G.Settings', or any table slot / global / call result. Evaluated as `return <expression>`.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint, idempotentHint, destructiveHint=false), yet the description adds substantial behavioral detail beyond them: evaluation as `return <expression>` under pcall, clean { error } on a bad path, the exact result shape (type/value/isTable, length, keys capped at 50, keysTruncated), pcall-guarded key reads, and the loadstring/load and active-client requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and a dense, useful body. It is long and the trailing boilerplate ('Phase: observe; cost=medium... On failure: inspect tool-schema...') partly restates annotations and metadata, but the bulk of the text carries real informational weight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully specifies the return contract ({ type, value, isTable, length?, keys?, keysTruncated? } or { error }), the failure mode, and runtime prerequisites. Nothing an agent needs to invoke or interpret this tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both parameters (expression, threadContext) are already documented in the schema with examples. The description reinforces rather than extends them — it clarifies the `return <expression>` evaluation form but adds little syntax or constraint meaning beyond the schema, matching the baseline for full coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Evaluate a single Luau expression and report exactly what lives at that slot') and explicitly positions itself against a sibling ('the counterpart to write-path-value'). An agent can distinguish it from eval-expression, run-luau, and write-path-value from the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete usage context: read a precise location after a heap scan points you at it, or spot-check any value/global/field during analysis, and names its mutation counterpart write-path-value. It stops short of explicit when-not or naming read-style alternatives like eval-expression, so it is clear context without full routing rules.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

release-closure-referenceRelease a retained closure handleA
Destructive

WRITES LIVE GAME STATE. Remove one key from getgenv().__mcp_closure_refs so cloned/wrapped closures can be garbage-collected. This invalidates its returned Reference and requires confirm=true. Signature: { key: string, confirm: boolean?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval. Capabilities: debug closure primitives. Produces: structured-result. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYestext value for key.
confirmNoExplicit safety acknowledgement; must be true before this state-changing operation runs.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=false, but the description adds value beyond them: invalidating the returned Reference, the confirm=true requirement, explicit-mutation-approval, and idempotency=contextual-write. It describes what state is mutated, though it doesn't detail failure modes or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The critical warning ('WRITES LIVE GAME STATE') and the core action are front-loaded, and the labeled segments (Phase, Requires, Capabilities, Produces, Verify, Safety, On failure) are scannable. It is dense but each clause carries distinct information, with only mild redundancy between the signature block and the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the description covers safety, prerequisites, verification, and failure handling ('inspect tool-schema for exact fields'). It is largely complete, though 'Produces: structured-result' is vague without an output schema to define the actual return.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both key and confirm are fully documented in the schema. The description only restates the signature ('{ key: string, confirm: boolean?, threadContext: number? }') without adding semantic meaning, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise verb and resource: 'Remove one key from getgenv().__mcp_closure_refs so cloned/wrapped closures can be garbage-collected.' This is a specific action distinguishable from siblings like list-closure-references or closure-capabilities, and it explains the downstream effect (enabling GC).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear preconditions (requires confirm=true, active-client, explicit-mutation-approval), a phase (act), and a verification path ('Verify with: assert-state'). It does not name a specific alternative sibling or explicitly state when NOT to use it, so it falls just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

remote-spyRemote-spy controls and diagnosticsA
Destructive

WRITES LIVE GAME STATE. Controls the selected spy: lifecycle, configure, pause/resume, clear, block/unblock, ignore/unignore, and reset-controls. Select engine=cobalt (default) or ketamine. status reports effective capture/GUI settings and capabilities; controls lists active block/ignore rules with limit/offset and direction filters. reset-controls clears block/ignore rules in the selected direction, preserving history and capture settings; Ketamine GUI ignore rules span Both and reset only with direction=Both. logs/list support query filters; code generates call code without replay. Configuration is a patch; capture filters affect future MCP records only. Pausing recording preserves blocking and GUI logging. Starting another engine stops the previous after preflight; read/configuration operations never load spies. Block/ignore require observed remotes. Ketamine cannot block incoming RemoteEvents. restart expires IDs/views and resets configuration; stop unloads the selected GUI/hooks. Signature: { engine: "cobalt" | "ketamine"?, operation: "status" | "start" | "restart" | "stop" | "list" | "logs" | "clear" | "block" | "unblock" | "ignore" | "unignore" | "code" | "configure" | "pause" | "resume" | "controls" | "reset-controls", mode: any?, max: number?, capture: { enabled: boolean?, direction: "Incoming" | "Outgoing" | "Both"?, nameFilter: string?, method: "FireServer" | "InvokeServer" | "OnClientEvent" | "OnClientInvoke"?, classFilter: "RemoteEvent" | "UnreliableRemoteEvent" | "RemoteFunction"?, blockedOnly: boolean? }?, resetFilters: boolean?, guiVisible: boolean?, guiLogging: boolean?, limit: any?, offset: number?, resetControl: "block" | "ignore" | "Both"?, direction: any?, remotePath: string?, remoteId: string?, nameFilter: string?, method: "FireServer" | "InvokeServer" | "OnClientEvent" | "OnClientInvoke"?, classFilter: "RemoteEvent" | "UnreliableRemoteEvent" | "RemoteFunction"?, blockedOnly: boolean?, afterId: number?, summaryOnly: any?, callId: number?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Produces: bounded-event-snapshot. Verify with: assert-state. Safety: MUTATING; changes persistent executor-side observer or hook state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
maxNoRetained MCP calls; resizing preserves newest history and IDs.
modeNoCobalt auto prefers supported RakNet; luau selects standard hooks. Ketamine accepts auto/luau only. Changing an active mode requires restart.auto
limitNoOptional hard result/work budget used to keep output and runtime bounded.
callIdNoOptional numeric value for call id.
engineNoRemote-spy backend. Only one engine runs per client; starting another stops the current engine. Defaults to cobalt.cobalt
methodNoOptional validated input for method.
offsetNoPagination offset for controls only.
afterIdNoOptional numeric value for after id.
captureNoPersistent capture patch; omitted fields keep their current values. Filters apply only to future MCP captures.
remoteIdNoOptional text value for remote id.
directionNoCapture/control direction; incoming includes client callbacks.Both
operationYesoperation selector; use one of the schema's allowed values.
guiLoggingNoKetamine only: enable/disable new GUI log entries independently of MCP capture; disabling also clears queued GUI entries.
guiVisibleNoKetamine only: show/hide its GUI without unloading the spy.
nameFilterNoOptional text value for name filter.
remotePathNoOptional dotted Roblox instance/value path resolved in the active client.
blockedOnlyNoWhether to enable blocked only.
classFilterNoOptional validated input for class filter.
summaryOnlyNoOptional validated input for summary only.
resetControlNoFor reset-controls: which rules to clear; defaults to Both.
resetFiltersNoClear persistent capture filters before applying this patch; preserve pause state, history, and network controls.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only give the safety profile (destructive/non-idempotent/write); the description adds substantial behavior beyond them: configuration is a patch affecting only future captures, pausing preserves blocking/GUI logging, restarting expires IDs/views, stop unloads GUI/hooks, starting another engine stops the previous, and Ketamine cannot block incoming RemoteEvents and its ignore rules only reset with direction=Both. This is exactly the kind of context an agent needs before invoking a mutating tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the critical 'WRITES LIVE GAME STATE' warning, and most sentences carry actionable information. It is dense and run-on in places (the signature block and behavioral notes could be structured as bullets), but there is little dead weight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 22-parameter mutating tool with no output schema, the description covers operations, key parameter relationships, engine semantics, and safety/state effects well. It partially covers returns (status/controls descriptions) but not the other operations' output, leaving a modest gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3; the description adds value by mapping parameters to operations (limit/offset/direction filters apply to controls, direction=Both governs Ketamine reset-controls) and disclosing engine default and switching behavior. It goes beyond enumerating the schema but does not fully cover all 22 parameter interactions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Controls the selected spy') and enumerates the operation families (lifecycle, configure, pause/resume, block/unblock, ignore/unignore, reset-controls), so an agent knows the tool's scope. However it doesn't differentiate itself from closely related siblings like configure-remote-spy, ensure-remote-spy, block-remote, ignore-remote, or monitor-remote, which overlap heavily with its operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains what each operation does (status reports settings, controls lists rules, reset-controls clears rules) which implies when to pick them, but there is no explicit when-to-use-this-vs-alternatives guidance despite multiple overlapping siblings. Usage is implied rather than stated, matching a 3.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

replicate-signalReplicate a signal to the SERVERA
Destructive

MUTATES SERVER STATE — DANGEROUS. Invokes replicatesignal(signal, ...) to fire the event with full network REPLICATION, so the event is delivered to the SERVER and can have real, authoritative server-side effects (unlike fire-signal, which is local only). Only some signals are replicable; this tool first checks cansignalreplicate(signal) and refuses to fire if it returns false. Use this ONLY for authorized testing of a game you own/control — firing replicated signals on other games may violate their rules. Requires the executor's replicatesignal; degrades with a clear { error } if unavailable. Returns { Signal, Instance?, Replicated, ArgCount } or { error }. Signature: { instancePath: string, signalName: any?, args: {{ kind: "string" | "number" | "boolean" | "nil" | "instance" | "raw", value: string | number | boolean? }}?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Capabilities: getconnections, firesignal. Produces: structured-result. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
argsNoOrdered list of arguments to replicate with the event, mirroring the real arguments the server expects for this signal. Omit or pass [] to replicate with no arguments.
signalNameNoName of the RBXScriptSignal member on the instance (e.g. 'OnClientEvent'). Leave empty/omit if instancePath already evaluates to the signal.
instancePathYesLua expression resolving to the Instance that owns the signal (e.g. 'game.ReplicatedStorage.RemoteEvent'). Evaluated as `return <instancePath>`. Leave signalName empty if this expression already resolves to the signal itself.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructive=true and idempotent=false, but the description adds substantial extra context: it requires the executor's replicatesignal, degrades to a clear { error } if unavailable, refuses when cansignalreplicate is false, and has authoritative server-side effects. This goes well beyond the safety profile the annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The dangerous-mutation warning and the fire-signal contrast are front-loaded, which is exactly right. However, the trailing template metadata (Phase/cost/idempotency/Requires/Capabilities/Produces/Verify/Safety) partially duplicates the opening 'MUTATES SERVER STATE — DANGEROUS' warning, adding some bulk.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although there is no output schema, the description discloses the return shape ({ Signal, Instance?, Replicated, ArgCount } or { error }), the failure mode, prerequisites, and invocation signature. Nothing an agent needs to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters in detail. The inline signature and notes (e.g. omitting args, threadContext default) largely restate what the schema provides rather than adding new semantics. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (replicate/fire with network replication) and resource (signal to the SERVER), and explicitly distinguishes itself from the sibling fire-signal ('which is local only'). An agent can tell exactly what this does and how it differs from the closest alternative without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use it, the prerequisite (checks cansignalreplicate and refuses if false), and the exclusions ('ONLY for authorized testing of a game you own/control'). It also names the alternative fire-signal and the gating sibling can-signal-replicate, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resolve-entityResolve or rediscover a semantic world handleA
Read-onlyIdempotent

Resolve a session-local handle returned by observe-world back to its live Roblox Instance. A live weak reference resolves immediately. If it is stale or destroyed, optional bounded rediscovery scores candidates against the stored structural fingerprint (class, name, parent, grandparent, original path, and root) and reattaches the same handle when confidence clears the requested threshold. Returns an exact path and executable bracket-safe expression, class, confidence, staleness state, match evidence, runner-up ambiguity, scan counts, and truncation. Read-only; uses GetChildren with a hard cap and never performs GetDescendants or a frame loop. Signature: { handle: string, rediscover: any?, roots: any?, maxInstances: any?, minConfidence: any?, threadContext: number? }. Phase: verify; cost=medium; idempotency=read-only. Requires: active-client, semantic handle previously returned by observe-world. Produces: grounded-evidence, live-instance-path, actionable-instance-expression, staleness-status, structural-match-confidence, resolution-evidence. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
rootsNoRoots searched when rediscovering a stale handle.
handleYesSession-local entity handle returned by observe-world, for example 'wb:12'.
rediscoverNoWhen true, a stale weak reference triggers bounded structural rediscovery. False only reports staleness.
maxInstancesNoHard maximum number of unique instances examined during rediscovery.
minConfidenceNoMinimum adjusted fingerprint confidence required to reattach a stale handle. Ambiguous runner-up matches reduce confidence.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish read-only/idempotent/non-destructive, yet the description adds real behavioral detail beyond them: weak reference resolves immediately, stale handles trigger bounded structural scoring with a hard instance cap, and it explicitly never calls GetDescendants or runs a frame loop. It also routes failure inspection to tool-schema, so the agent knows the cost and limits of the call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded well with purpose first, but the block is dense and contains redundancy: the inline Signature duplicates the input schema, and 'Safety: read-only' plus 'idempotency=read-only' repeat the annotations. The Phase/Requires/Produces meta lines are useful but inflate the size for a single tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully enumerates the return payload (path, expression, class, confidence, staleness, match evidence, runner-up ambiguity, scan counts, truncation) and names prerequisites and failure guidance. It is close to complete, missing only explicit guidance on what confidence values imply for the caller.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents every parameter, giving a baseline of 3. The description mostly restates the same facts (minConfidence reattachment threshold, rediscover=false only reporting staleness, maxInstances cap) and its inline Signature even degrades types to 'any?', adding little conceptual value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource: resolve a handle returned by observe-world back to its live Roblox Instance. It also distinguishes itself from the producer sibling (observe-world) and from the staleness/rediscovery case, so an agent can pick it without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Use condition is explicit: you must hold a handle previously returned by observe-world and an active client, and rediscover=true vs false is spelled out. However it never names alternative tools (e.g. verify-path-exists, search-instances) that might be chosen instead, so it stops short of full when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

restore-functionRestore any currently hooked functionA
Destructive

WRITES LIVE GAME STATE. Resolve a target and call restorefunction/restorefunc directly. Unlike restore-hook, this also restores hooks not installed through the MCP registry. Requires confirm=true. Signature: { functionPath: string, threadContext: number?, confirm: boolean? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Produces: structured-result. Verify with: is-function-hooked. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmNoExplicit safety acknowledgement; must be true before this state-changing operation runs.
functionPathYesLuau expression resolving to the target function.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, but the description adds real context beyond them: it writes LIVE GAME STATE, requires explicit-mutation-approval and an active/resolved client, has contextual-write idempotency, and produces a structured result. This is well past what the structured fields convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the critical warning ('WRITES LIVE GAME STATE') and the sibling differentiator. The metadata tail (phase/cost/idempotency/preconditions) is dense but each item is actionable; only the generic 'On failure: inspect tool-schema' sentence is near-filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the description covers preconditions, the required confirmation, the verification sibling, and the failure fallback. It stops short of describing the structured result's shape, but that is a minor gap given the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents functionPath, threadContext, and confirm including defaults and the Luau-expression semantics. The description's signature line largely restates the schema, adding only the emphasis that confirm is mandatory; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (restore) and resource (a currently hooked function) and explicitly differentiates itself from the sibling restore-hook: 'Unlike restore-hook, this also restores hooks not installed through the MCP registry.' An agent can pick between the two without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the alternative (restore-hook) and the exact condition that selects this tool (hooks not installed through the MCP registry), plus a verification path ('Verify with: is-function-hooked') and a hard precondition (confirm=true). Routing is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

restore-hookRestore (undo) a hook installed via this MCPA
Destructive

WRITES LIVE GAME STATE. Undo a hook created by hook-function / hook-metamethod, restoring the captured original. Pass a key from list-hooks to restore one, or all: true to restore every tracked hook. For function hooks it re-resolves the target and calls restorefunction (falling back to re-hooking with the original); for metamethod hooks it re-installs the original metamethod. Successfully restored hooks are removed from the registry. Use this to clean up after debugging so your hooks don't linger and destabilize the game. Signature: { key: any?, all: any?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval. Produces: structured-result. Verify with: is-function-hooked. Safety: MUTATING; changes persistent executor-side observer or hook state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
allNoRestore ALL tracked hooks (ignores key). Default false.
keyNoThe hook key to restore (from list-hooks). Ignored when all=true.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations (readOnlyHint=false, destructiveHint=true) by disclosing the actual mechanism: function hooks re-resolve the target and call restorefunction with a fallback to re-hooking, metamethod hooks re-install the original, and restored hooks are removed from the registry. It also flags the approval requirement (explicit-mutation-approval), the active-client prerequisite, and that executor-side hook state is persistently mutated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The critical warning and purpose are front-loaded, and the trailing metadata (phase, cost, requires, verify, safety) is terse and scannable. There is some redundancy in the Signature line duplicating the input schema and in the closing tool-schema pointer, keeping it off a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating tool with no output schema, the description covers everything needed to call it safely: what state changes, prerequisites, verification path (is-function-hooked), and a recovery hint on failure (inspect tool-schema). Nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents `all`, `key`, and `threadContext`. The description restates the key/all interplay (key from list-hooks, all ignores key) but adds no format, default, or constraint detail beyond the schema; baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Undo a hook created by hook-function / hook-metamethod, restoring the captured original') and names the two sibling tools that created the hook. An agent can distinguish this from list-hooks, is-function-hooked, or restore-function without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to pass a `key` from list-hooks for one hook or `all: true` for every tracked hook, and gives the motivating scenario ('clean up after debugging so your hooks don't linger'). It does not state when not to use it (e.g., versus restore-function or hook-and-log-function), so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run-deferredSchedule a Luau snippet to run in the background (fire-and-forget)A
Destructive

Schedule a Luau snippet to run on its OWN thread and return immediately WITHOUT waiting for it to finish. Use this to start something that runs in the background while you keep issuing other tool calls — e.g. kick off a long watcher loop, an auto-farm, a background poller, or any snippet you want to run later or off the current thread so it cannot block the round trip. The code is COMPILED FIRST via loadstring; a syntax error is returned cleanly as { error } and nothing is scheduled. On success the compiled function is handed to the chosen scheduler and the tool returns at once. Modes (Luau task library): - 'spawn' -> task.spawn(fn): start running on a fresh thread immediately. - 'defer' -> task.defer(fn): run on a fresh thread at the end of the current resumption cycle. - 'delay' -> task.delay(delaySec, fn): run on a fresh thread after delaySec seconds. IMPORTANT: because this does not wait, you will NOT see the snippet's return value, errors, or print output here — the scheduled thread runs independently and any error inside it is swallowed by task. If the snippet installs state or a loop, YOU are responsible for stopping it (e.g. have it watch a getgenv() flag). Requires loadstring and the task library (both guarded). Returns { scheduled = true, mode, delaySec? } or { error }. Signature: { code: string, mode: any?, delaySec: any?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval, validated-source. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; executes caller-selected behavior in the live client. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYesThe Luau snippet to run in the background. Compiled with loadstring, then scheduled on its own thread. Since nothing is awaited, build in your own stop condition for long loops (e.g. `while getgenv().__myFlag do ... task.wait() end`) so you can shut it down later.
modeNoHow to schedule the thread: 'spawn' starts it immediately on a new thread; 'defer' runs it at the end of the current resumption cycle; 'delay' runs it after `delaySec` seconds. Default 'spawn'.spawn
delaySecNoSeconds to wait before the snippet runs. Only used when mode='delay' (ignored for spawn/defer). Default 0.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial behavior beyond the annotations: compile-first via loadstring with syntax errors returned as {error} before anything is scheduled, errors/returns/prints being invisible because the thread runs independently, errors swallowed by task, and the caller's responsibility to build a stop condition. It also states prerequisites (loadstring, task library) and the return shape, none of which the schema or annotations convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose is front-loaded and the mode list is scannable, but the fire-and-forget guarantee is restated several times and the trailing metadata footer (Phase/cost/Requires/Produces) adds boilerplate weight. Efficient enough, slightly verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description discloses the return value ({scheduled=true, mode, delaySec?} or {error}), the safety posture, prerequisites, and the lifecycle caveat about stopping my own loops. Nothing needed to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents code, mode, delaySec, and threadContext. The description's restatement of the spawn/defer/delay modes and delaySec behavior largely duplicates the schema, so it earns the baseline 3 with only marginal added meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Schedule a Luau snippet to run') and its defining trait ('on its OWN thread and return immediately WITHOUT waiting'). This cleanly separates it from waiting/blocking siblings like execute-and-wait and batch-execute, so an agent can distinguish it without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete when-to-use guidance with examples ('kick off a long watcher loop, an auto-farm, a background poller') and frames it as the non-blocking counterpart to synchronous execution. It stops short of explicitly naming a sibling alternative for the blocking case, so the exclusion is implied rather than a named route.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run-loopRun a Luau snippet repeatedly with a delay, collecting resultsA
Destructive

Run a Luau snippet iterations times in a single round trip, waiting delayMs between iterations, and collect what each run produced. Use this to poll a value over time (e.g. read a Humanoid's Health every 200ms for 20 iterations to watch it change), to retry a flaky operation, or to drive a short repeated action and see the outcome of each pass — all without paying a tool round trip per iteration. The code is COMPILED ONCE (a syntax error returns { error } and nothing runs); each iteration is executed inside its own pcall so a runtime error in one pass is recorded and the loop continues. When collectReturns is true the FIRST return value of each run is captured via __encVal (Instances become their full path, tables/functions become a stable string), giving an ordered series you can compare across iterations. The whole loop runs synchronously in-client, so total time is roughly iterations * delayMs — keep that under the tool timeout. Requires loadstring; uses task.wait for the delay (falls back to wait). Returns { iterations, results, errorCount, errors } or { error }. Signature: { code: string, iterations: any?, delayMs: any?, collectReturns: any?, threadContext: number? }. Phase: act; cost=high; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval, validated-source. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; executes caller-selected behavior in the live client. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYesThe Luau snippet to run each iteration. Compiled once, then executed every pass. To capture a value per iteration, `return` it (only the first return value is collected when collectReturns is true).
delayMsNoMilliseconds to wait between iterations via task.wait (default 0 = yield minimally each pass). Use e.g. 200 to sample a value five times a second. There is no wait after the final iteration.
iterationsNoHow many times to run the snippet (default 5, clamped 1..1000). Combined with delayMs this determines total wall-clock time, so keep iterations * delayMs comfortably under the request timeout.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
collectReturnsNoWhen true (default), capture the FIRST return value of each iteration into `results` (encoded as a serializable scalar/string). Set false to skip capture when the snippet returns nothing useful or large values you do not want echoed back.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: compile-once semantics, syntax errors returning {error} with nothing run, per-iteration pcall isolation so runtime errors are recorded but the loop continues, the __encVal encoding of returned values, synchronous in-client execution with total time ≈ iterations * delayMs, and the loadstring requirement. The mutating/live-client nature is stated explicitly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Core purpose is front-loaded in sentence one, and most sentences carry distinct information (compile behavior, error isolation, timing budget, return shape). It is dense and long, and the trailing Phase/cost/idempotency/Requires metadata block is boilerplate that could be trimmed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a high-complexity mutating tool with no output schema, the description covers return shape ({iterations, results, errorCount, errors} or {error}), failure modes, timing constraints, and prerequisites. An agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real value: the collectReturns encoding behavior (Instances become full paths, tables/functions become stable strings) and the first-return-value-only capture rule go beyond the schema's 'serializable scalar/string' phrasing. The signature line is somewhat redundant with the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (run a Luau snippet N times with a delay, collecting each run's output) and the loop semantics immediately separate it from single-shot siblings like run-luau or execute. The title and first sentence align exactly with what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives three concrete usage scenarios (polling a value over time, retrying a flaky operation, driving a short repeated action) and explains the benefit over per-iteration round trips. It stops short of naming alternatives (e.g., watch-value for monitoring, run-with-timeout for bounded execution) or stating when NOT to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run-luauRun Luau on the active clientA
Destructive

Execute arbitrary Luau in the active Roblox client and return its first returned value (decoded from JSON). This is the core PURE-LUAU execution tool — no access to this server's other tools from inside the script. If you want to use any other tool's data (get-players, search-instances, discover-player-values, anything from list-tools) inside your Luau, STOP and use the script tool instead: it binds a live mcp table so you can write local p = mcp.getPlayers() / mcp.searchInstances({...}) / mcp.parallel({...}) and use the results directly — one call instead of dozens of round-trips. Use run-luau only when your Luau is fully self-contained (reading workspace, looping over a part, returning a value). return the value(s) you want back; do NOT call JSONEncode yourself, the connector serializes automatically. A chunk that returns nothing yields null. Use eval-expression for a single expression, or the higher-level inspection tools for structured reads. Signature: { source: string, threadContext: number?, timeoutMs: number?, client: string?, agent: string? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval, validated-source. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; executes caller-selected behavior in the live client. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNoOptional. A stable label for WHICH agent is calling when several share this MCP session (e.g. 'researcher'). Gives that agent its own fair scheduling lane, its own persistent VM on each game, and its own queue budget, so co-tenant agents don't starve or clobber each other.
clientNoOptional. Run on a specific connected client — its clientId OR username — for THIS call only, overriding your session's select-client binding without changing it. Lets multiple agents drive different games at the same time; omit to use your session's selected client.
sourceYesLuau source to execute. Use `return <value>` to get data back.
timeoutMsNoPer-call deadline in milliseconds. Server default if omitted.
threadContextNoRoblox thread identity to run under (e.g. 2 = game scripts, 8 = elevated). Server default if omitted.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructive/non-idempotent/non-readonly, but the description adds real behavior: no cross-tool binding inside the script, automatic JSON serialization ('do NOT call JSONEncode yourself'), null for a chunk that returns nothing, and prerequisites (active-client, explicit-mutation-approval, validated-source) plus a verification path (assert-state). That is materially more than the annotations carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and the critical `script`-vs-`run-luau` routing decision, and the serialization caveat is placed well. The trailing metadata block (Phase/cost/idempotency/Requires/Produces/Verify/On-failure) is boilerplate-heavy and slightly dilutes an otherwise tight description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a high-risk mutation tool with no output schema, the description covers prerequisites, return-value handling, failure recovery ('inspect tool-schema for exact fields'), and the verification step. An agent has everything needed to invoke it correctly without opening other artifacts.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, and the per-parameter descriptions (agent scheduling lane, client override, threadContext identity) are already thorough. The description adds modest extra meaning by restating the signature and clarifying the `source` return contract, but it does not deepen the optional-parameter semantics beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Execute arbitrary Luau in the active Roblox client') plus the exact return contract ('first returned value, decoded from JSON'). It actively distinguishes itself from siblings by naming `script` and `eval-expression` and the conditions that separate them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-not guidance: 'no access to this server's other tools from inside the script' and a hard redirect to the `script` tool when other tools' data is needed, including the exact `mcp.getPlayers()` / `mcp.parallel({...})` pattern to use instead. It also routes single expressions to `eval-expression` and structured reads to higher-level inspection tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run-on-actorSchedule Luau source on an Actor's isolated stateA
Destructive

WRITES LIVE GAME STATE. Resolve an Actor expression and call run_on_actor(actor, source, ...arguments). The operation is asynchronous and returns scheduling metadata, not the actor script's return value. Requires confirm=true. Signature: { actorPath: string, source: string, arguments: any?, confirm: boolean?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval, validated-source. Capabilities: getactors. Produces: operation-receipt. Verify with: get-lua-state. Safety: MUTATING; executes caller-selected behavior in the live client. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
sourceYesLuau source text processed by this tool; it is never inferred or guessed.
confirmNoExplicit safety acknowledgement; must be true before this state-changing operation runs.
actorPathYesdotted Roblox instance/value path resolved in the active client.
argumentsNoOptional ordered typed arguments forwarded to the selected operation.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial behavior beyond the annotations: asynchronous execution that returns scheduling metadata rather than the script's return value, the confirm=true gate, phase/cost/idempotency metadata, produced operation-receipt, and a safety note that it runs caller-selected behavior in the live client. Annotations cover the destructive/idempotent profile, and the description enriches it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the critical 'WRITES LIVE GAME STATE' warning, then tightly structured metadata blocks. Some redundancy (MUTATING repeats destructiveHint) but the density is justified for a high-risk mutation tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly explains that the return is scheduling metadata, not the script result. Prerequisites, verification path, and failure guidance are all present, so nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents every parameter; the description's signature largely restates it. It adds minor value by noting the actor path is resolved in the active client and that arguments are forwarded. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: resolve an Actor expression and call run_on_actor(actor, source, ...arguments), scoped to an Actor's isolated state. This clearly separates it from sibling executors like run-luau, execute, and execute-lua-state.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Supplies strong context: requires active-client, resolved-target, explicit-mutation-approval and confirm=true, and points to get-lua-state for verification. It does not explicitly say when to prefer a sibling executors (e.g. run-luau for non-Actor code), so the when-not is left implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run-with-timeoutRun a Luau snippet with a watchdog timeoutA
Destructive

Run a Luau snippet under a WATCHDOG: it executes on its own thread and this tool polls a done flag against an os.clock() deadline, so a snippet that hangs on a YIELD (a yield that never resumes, a :Wait() on a signal that never fires) cannot block the round trip forever — after timeoutSec you get a clean { timedOut = true } instead of waiting out the whole connector timeout. Use this any time you are about to run code that MIGHT hang or take an unknown amount of time. The snippet is COMPILED FIRST via loadstring (a syntax error returns { error } and nothing runs), then started with task.spawn; the watchdog waits with task.wait until either the thread sets its done flag or the deadline passes. LIMITATION: Luau is cooperatively scheduled, so the watchdog can only fire when the snippet YIELDS. A snippet that never yields (e.g. while true do end or a tight CPU loop with no task.wait) runs synchronously inside task.spawn, so the watchdog loop never gets to run and the round trip still blocks until the connector timeout. Put a task.wait() inside long loops if you want the watchdog to be able to interrupt them. IMPORTANT: a timeout REPORTS that the snippet did not finish in time, but it does NOT kill the runaway thread — the executor keeps running it in the background (and it may keep consuming CPU). Prefer snippets that can finish, and keep timeoutSec sane. If the snippet completes in time, its FIRST return value is encoded via __encVal and returned; a runtime error inside it is captured in error. Requires loadstring, the task library, and os.clock (all guarded). Returns { completed, timedOut, elapsedMs, result?, error? } or { error }. Signature: { code: string, timeoutSec: any?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval, validated-source. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; executes caller-selected behavior in the live client. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYesThe Luau snippet to run under the watchdog. Compiled with loadstring and run on its own thread. To report a value back, `return` it (only the first return value is captured, encoded as a serializable scalar/string).
timeoutSecNoHow long to let the snippet run before giving up and reporting timedOut, in seconds (default 5, clamped 0.1..60). On timeout the snippet is NOT killed — it keeps running in the background; this only stops THIS tool from waiting.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Far exceeds the annotations: discloses the cooperative-scheduling limitation (watchdog only fires on yield), that a timeout does NOT kill the runaway thread, that compile happens first via loadstring, and that only the first return value is encoded. These are exactly the behavioral traits an agent needs beyond destructiveHint=false/idempotentHint=false.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core behavior, but very long and repetitive — the 'timeout does not kill the snippet' caveat is stated twice, and ALL-CAPS emphasis plus the signature/phase/safety trailer add length without new information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description compensates fully by spelling out both return shapes ({ completed, timedOut, elapsedMs, result?, error? } and { error }) and the runtime-error capture behavior. Complete for a mutating, high-complexity tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds some timeout semantics, but largely repeats the signature and the schema's own notes (timeout clamps, thread not killed); it does not add new meaning for threadContext.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('run a Luau snippet under a WATCHDOG') plus the exact mechanism (own thread, polled done flag vs os.clock() deadline). This clearly separates it from siblings like execute, run-luau, and execute-and-wait, which lack the timeout watchdog.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use it: 'any time you are about to run code that MIGHT hang or take an unknown amount of time,' and explains the failure mode it prevents. It does not name the alternative siblings (execute-and-wait, run-luau) or when to prefer them instead, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save-instanceDump the game or a subtree to the workspace (sUNC saveinstance)A
Destructive

WRITES HOST FILES — serializes the entire game (or a chosen instance subtree) to a file in the executor's workspace folder via saveinstance. By default it dumps the whole DataModel; pass instancePath to dump just that subtree. This is a heavy operation that can take a while and produce a large file. Because executors differ on the exact signature, this tries several forms in order: saveinstance({ FilePath = fileName }), then saveinstance(instance, fileName) / saveinstance(instance), then saveinstance(). Requires saveinstance. The call is type-guarded and pcall-wrapped: if saveinstance is missing you get { error = 'saveinstance is not available in this executor.' }. Returns { saved, note } or { error }. Signature: { fileName: string?, instancePath: string?, threadContext: number?, timeoutMs: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Produces: structured-result. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
fileNameNoOptional output file name within the executor workspace (e.g. 'place.rbxm'). Omit to let the executor pick a default.
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
instancePathNoOptional Luau expression for the instance to dump (e.g. 'game.Workspace'). Defaults to the whole game.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructive/non-idempotent/not-read-only, but the description adds substantial context beyond them: the multi-signature fallback strategy, pcall/type-guard wrapping, the exact error object returned when saveinstance is missing, the { saved, note } success shape, and the heavy/large-file caveat.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The critical warning (WRITES HOST FILES) is front-loaded and each sentence carries distinct information. It runs long with the metadata tail (Phase/cost/Verify with), but nothing is redundant padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating tool with no output schema, the description covers prerequisites, fallback invocation behavior, error and success payload shapes, and failure recovery (inspect tool-schema). An agent has everything needed to call it and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all four parameters. The description restates defaults (whole game when instancePath omitted, executor default fileName) but adds little syntax or constraint detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource: 'serializes the entire game (or a chosen instance subtree) to a file in the executor's workspace folder via saveinstance.' This is clearly distinguished from siblings like write-file (arbitrary content) and get-instance-tree (inspection only).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States the default behavior (whole DataModel) and the alternative trigger (pass instancePath for a subtree), plus the prerequisites (saveinstance availability, active-client, explicit-mutation-approval) and a warning that it is heavy/slow. It does not explicitly route away from alternative snapshot tools (e.g., diff-instance-snapshot or get-instance-tree), so it falls just short of 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scan-closures-by-nameScan closures by nameB
Read-onlyIdempotent

Find getgc closures whose debug name contains a target substring. Signature: { nameQuery: string, limit: any?, includeCClosures: any?, threadContext: number? }. Phase: observe; cost=high; idempotency=read-only. Requires: active-client. Capabilities: getgc, debug closure primitives. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoOptional hard result/work budget used to keep output and runtime bounded.
nameQueryYessearch text used to filter and rank bounded results.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
includeCClosuresNoWhether to include cclosures in the bounded result.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, and 'Safety: read-only' merely repeats that. The description does add non-annotation context (cost=high, requires active-client, produces bounded-candidates, fallback to tool-schema on failure), which is genuinely useful, but the additions are terse labels rather than explanatory behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded, which is good, but the text is a dense run-on of colon-delimited metadata and the 'Signature: {...}' block simply duplicates the input schema, wasting space without adding information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only observe tool with no output schema, the description covers purpose, preconditions and cost, but leaves the output shape vague ('bounded-candidates') and says nothing about ranking/ordering or what a bounded result actually contains, deferring everything to tool-schema on failure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents every parameter fully. The description only re-lists the signature with types ('nameQuery: string, limit: any?...'), which adds no meaning beyond the schema; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Find getgc closures') and the exact filter mechanism ('debug name contains a target substring'), which self-differentiates it from the sibling scan-closures-by-source. It stops short of explicitly naming that sibling, so it is clear but not fully sibling-routing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides partial context ('Phase: observe; cost=high', 'Requires: active-client') that implies when the tool is appropriate, but gives no explicit when-not guidance and never names an alternative tool to use instead (e.g., scan-closures-by-source or search-gc-value).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scan-closures-by-sourceScan closures by sourceA
Read-onlyIdempotent

Find getgc closures whose debug source contains a target substring. Signature: { sourceQuery: string, limit: any?, includeCClosures: any?, threadContext: number? }. Phase: observe; cost=high; idempotency=read-only. Requires: active-client. Capabilities: getgc, debug closure primitives. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoOptional hard result/work budget used to keep output and runtime bounded.
sourceQueryYessearch text used to filter and rank bounded results.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
includeCClosuresNoWhether to include cclosures in the bounded result.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, non-destructive, closed-world, so the bar is lower. The description adds genuinely new context: cost=high, requires an active client, phase=observe, and that output is bounded-candidates. It does not detail pagination or result shape, but that is a minor gap given the annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence and the metadata is compact and structured. Minor redundancy: 'Safety: read-only' and 'idempotency=read-only' partly duplicate the readOnlyHint/idempotentHint annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description carries return-value burden, and 'Produces: bounded-candidates' plus cost=high gives an adequate sense of output and expense. The active-client prerequisite is stated. Only minor gaps (exact result shape) remain for a read-only scan tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters including defaults and constraints. The description only restates the signature with types and adds no semantic meaning beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Find), resource (getgc closures), and a precise filter scope (debug source contains a target substring), which inherently separates it from the by-name sibling. It does not explicitly name scan-closures-by-name as the alternative, so the differentiation is left implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides context markers ('Phase: observe', 'Requires: active-client') that imply when the tool is applicable, but offers no explicit when-to-use versus the by-name/by-source sibling or any exclusion conditions. Guidance is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scan-connections-by-sourceScan connection handlers by defining script sourceA
Read-onlyIdempotent

Sweep a hierarchy for event handlers and pinpoint which script each handler is defined in. Walks descendants of root, probes a set of signal names on each instance, and for every Lua connection describes its Function via debug.info. If the function's Source contains sourcePattern (case-insensitive substring; empty matches ALL handlers) it records { Instance, Signal, ConnectionIndex, Source, LineDefined, Name }. This answers "find every event handler defined in ", "where are all the handlers backed by this module?", or simply "enumerate all connected handlers under here and tell me their source". Requires the executor's getconnections and debug.info; degrades to a clear { error } if getconnections is unavailable (a missing debug.info merely leaves Source blank). Bounded by maxScan (instances) and maxResults (matches). Returns { Matches, MatchCount, ScannedInstances, Truncated }. Signature: { root: any?, sourcePattern: any?, signalNames: any?, maxResults: any?, maxScan: any?, threadContext: number? }. Phase: observe; cost=high; idempotency=read-only. Requires: active-client. Capabilities: getconnections. Produces: bounded-candidates, created-handle. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
rootNoLua expression for the root instance whose descendants are scanned (e.g. 'game', 'game.Workspace', 'game.StarterGui'). Evaluated as `return <root>`. Defaults to 'game'. Narrow this to speed up the scan and reduce noise.game
maxScanNoMaximum number of descendant instances to examine. Defaults to 4000. Raise to cover larger games (slower); lower for a quick sample. If this cap is hit, Truncated is true.
maxResultsNoMaximum number of matching connections to collect before stopping the scan. Defaults to 200. If this cap is hit, Truncated is true.
signalNamesNoSignal member names to probe on each instance. When omitted, a sensible default list of commonly connected signals is used (Touched, Changed, ChildAdded, Activated, InputBegan, OnClientEvent, Heartbeat, etc.). Provide your own list to target specific events and make the scan faster.
sourcePatternNoCase-insensitive substring matched against each Lua connection's function Source (the chunk/script path reported by debug.info, e.g. a script name or asset path). When empty (default), EVERY Lua handler is matched, giving a full inventory. Set it to a script name to find only handlers defined in that script.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly/idempotent/non-destructive annotations, it discloses the hard dependency on getconnections and debug.info, the graceful degradation to { error } when getconnections is missing, that a missing debug.info merely blanks Source, and that results are bounded by maxScan/maxResults. That is exactly the extra behavioral context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose and the sourcePattern/maxScan semantics are front-loaded and each sentence carries information. The trailing metadata footer (Signature, Phase, cost, idempotency, capabilities, safety) partly duplicates the annotations and schema, adding some redundancy without derailing the read.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a high-cost, read-only scan with no output schema, the description supplies the return shape ({ Matches, MatchCount, ScannedInstances, Truncated }), truncation behavior, the required capabilities, and failure modes. An agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents every parameter (including root expression semantics, defaults, and the sourcePattern substring rule). The description largely restates sourcePattern behavior and the caps rather than adding new syntax or constraints, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('sweep a hierarchy for event handlers and pinpoint which script each handler is defined in') and differentiates itself from adjacent tools like scan-closures-by-source and find-event-connections by scoping to connected handlers and their defining script. An agent can identify it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete triggering questions ('find every event handler defined in <script>', 'where are all the handlers backed by this module?'), which clarifies when to reach for it. It does not, however, explicitly name competing siblings such as find-event-connections or list-signal-connections and when to prefer them, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scan-hook-surfacesScan hook surfacesA
Read-onlyIdempotent

Check availability of common exploit/debug hook APIs and return capability map. Signature: { threadContext: number? }. Phase: observe; cost=high; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Verify with: is-function-hooked. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnly/idempotent/non-destructive, so the bar is lower. The description still adds real context beyond them: cost=high, the active-client requirement, 'Produces: bounded-candidates', and a verification path via is-function-hooked, plus a failure fallback to tool-schema. It does not explain the shape of the capability map, but adds meaningful behavioral signal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then structured metadata fields that are quick to scan. There is minor redundancy between 'idempotency=read-only', 'Safety: read-only', and the readOnlyHint annotation, but overall it is dense and earns its space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only scan with no output schema, the description covers purpose, prerequisites, cost, verification, and a failure fallback. The only gap is the exact structure of the returned capability map, which an agent would have to discover at call time.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With a single optional parameter at 100% schema coverage, the schema already documents threadContext fully. The description restates the signature as '{ threadContext: number? }' without adding format or default nuance beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Check availability of common exploit/debug hook APIs') plus the output ('return capability map'), which lets an agent distinguish it from sibling scanners like scan-connections-by-source or list-hooks. It stops short of naming a competing alternative to choose between, but the purpose itself is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives operational context: 'Phase: observe; cost=high' and the prerequisite 'Requires: active-client', plus 'Verify with: is-function-hooked' as a follow-up. There is no explicit when-not or list of alternatives, but the context is clear enough to route the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scan-network-endpointsScan network endpointsA
Read-onlyIdempotent

Find URL-like strings in function constants across getgc closures. Signature: { limit: any?, threadContext: number? }. Phase: observe; cost=high; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoOptional hard result/work budget used to keep output and runtime bounded.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnly/idempotent/non-destructive/closed-world, so the bar is lower, and the description still adds real context: 'Requires: active-client', 'Produces: bounded-candidates', cost=high, and failure-handling guidance pointing to tool-schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose in the first sentence and then uses compact labeled segments (Phase/cost/Requires/Produces/Safety/On failure). Nearly every clause earns its place, with only the redundant Signature line as minor waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-required-param, no-output-schema scan tool, the description covers preconditions (active-client), cost profile, output shape hint (bounded-candidates), and failure recovery, leaving little an agent needs missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents limit as a work budget and threadContext as a Roblox thread identity. The description's 'Signature' line merely restates the parameter names/types without adding semantics, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource with scope: 'Find URL-like strings in function constants across getgc closures.' This lets an agent distinguish it from generic string search siblings like find-string-in-tables or search-gc-value, though it never names an alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No direct when/when-not statement or named alternatives among the many scan-*/find-* siblings. Usage is only implied via 'Phase: observe' and 'cost=high', which signal reconnaissance use but leave the choice between this and neighboring scanners to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scan-number-rangeScan GC tables for numeric values in a [min, max] range (CE range scan)A
Read-onlyIdempotent

Cheat-Engine-style RANGE scan: walk every live Luau GC table and report each numeric value v where min <= v <= max. Use this when you know a stat is within a window but not its exact value — e.g. coins between 100 and 200, health under 50, a timer in [0, 10], a damage multiplier in [1, 5]. It is the inexact complement to search-gc-value (which needs the precise number). Each hit is recorded as { table, key, value } where 'table' is the container address (tostring), 'key' is the field the number sits under (tostring), and 'value' is the raw number. To narrow many hits down to one, run successive scans with tightening bounds after the stat changes in-game (the Cheat-Engine 'next scan' technique), then read or flip the survivor with read-path-value / write-path-value. Each table's pairs() iteration is pcall-guarded so a locked/proxy table never aborts the scan; GC objects examined are capped by maxScan and results by limit, with a 'truncated' flag. NaN values are skipped. Requires getgc (falls back from getgc(true) to getgc()). Returns { min, max, matchCount, scannedObjects, truncated, matches } or { error }. Signature: { min: number, max: number, limit: any?, maxScan: any?, threadContext: number? }. Phase: observe; cost=high; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
maxYesInclusive upper bound of the numeric window to match (e.g. 200). Must be >= min for any hits. Combine with min to bracket a stat you don't know exactly (e.g. min=100, max=200).
minYesInclusive lower bound of the numeric window to match (e.g. 100 to find a coin total at least 100). Values v with v >= min and v <= max are recorded.
limitNoMaximum number of matching numbers to return (default 150). Hitting this sets truncated=true.
maxScanNoMaximum number of GC objects to examine before stopping (default 40000). Hitting this sets truncated=true. Raise for a deeper sweep at the cost of time; lower if scans are slow.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the read-only/idempotent/non-destructive profile, and the description adds substantial extra behavior: pcall-guarded iteration so locked/proxy tables never abort, NaN skipping, maxScan/limit caps with a 'truncated' flag, and a getgc(true)->getgc() fallback. These are exactly the operational traits an agent needs beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core operation and the key alternative, and most sentences carry real content (bounds semantics, next-scan technique, return shape). However, the trailing tag block ('Phase: observe; cost=high; idempotency=read-only', etc.) partly restates the annotations and adds some padding against an otherwise dense description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description supplies the return shape ({ min, max, matchCount, scannedObjects, truncated, matches } or { error }), the resource caveat (requires getgc), and the follow-up tooling. For a high-cost, complex scan tool this is sufficient to invoke and interpret correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all five parameters including defaults and bounds, which sets the baseline at 3. The description reinforces limit/maxScan capping and the truncated behavior but adds little syntax or meaning beyond what the schema descriptions already state, and gives only a bare signature with types.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource with scope: 'walk every live Luau GC table and report each numeric value v where min <= v <= max'. It explicitly distinguishes itself from the sibling search-gc-value by framing itself as the 'inexact complement' that does not need the precise number. An agent can tell it apart from find-string-in-tables and search-gc-value without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete when-to-use conditions ('when you know a stat is within a window but not its exact value') with examples, names the alternative it complements (search-gc-value), and describes the follow-up workflow (successive tightening scans, then read-path-value / write-path-value). When and why are fully specified with no inference required.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scan-proto-functionsScan proto functionsA
Read-onlyIdempotent

Use debug.getprotos over matched functions and return proto summaries. Signature: { query: any?, limit: any?, threadContext: number? }. Phase: observe; cost=high; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoOptional hard result/work budget used to keep output and runtime bounded.
queryNoOptional search text used to filter and rank bounded results.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, and the description adds genuinely useful context: cost=high, requires an active client, and output is bounded-candidates. The failure-path guidance (consult tool-schema) is a nice extra, though return shape detail is thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose before the metadata and is compact. Some redundancy exists ('idempotency=read-only' and 'Safety: read-only' both duplicate the readOnlyHint annotation), but the structured phrasing keeps it scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-optional-param read-only scan with no output schema, the description covers prerequisites, cost, and the nature of the output (bounded-candidates). It is largely complete, missing only an explicit distinction from sibling proto-scanning tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents query, limit, and threadContext with defaults and constraints. The description only restates the signature with types and adds no syntax or format detail beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (scan) and resource (proto functions) and names the underlying mechanism (debug.getprotos over matched functions), so the agent knows what it does. It does not distinguish itself from close siblings like get-function-protos or get-closure-protos, so it falls short of the top score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides contextual signals ('Phase: observe', 'Requires: active-client', 'cost=high') that imply when it fits, and the failure note routes to tool-schema. However, it never states when to choose this over sibling proto tools, leaving the alternative selection to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scan-remote-listenersScan remote listenersA
Read-onlyIdempotent

List RemoteEvent/RemoteFunction objects and listener counts from connection APIs. Signature: { limit: any?, threadContext: number? }. Phase: observe; cost=high; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoOptional hard result/work budget used to keep output and runtime bounded.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description nonetheless adds real behavioral context beyond that: cost=high, a required active-client precondition, and a bounded-candidates output characteristic, which warns the agent about expense and result bounds.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The telegraphic phase/cost/safety tag structure is compact and front-loads the resource being listed. It is slightly padded by the redundant signature restatement that duplicates the schema, but nothing is wasted overall.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-required-parameter, read-only enumeration tool with no output schema, the description covers the key operational facts an agent needs: phase, cost, precondition, output boundedness, and a pointer to tool-schema for exact field details. Only the sibling-disambiguation and return-shape aspects are left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both parameters fully documented in-schema (limit as hard result/work budget with default 300, threadContext as a Roblox thread identity). The description only restates the signature as '{ limit: any?, threadContext: number? }', adding no meaning beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (list) and resource (RemoteEvent/RemoteFunction objects plus listener counts from connection APIs), which is more specific than a generic 'scan'. However it does not distinguish itself from close siblings such as list-remotes, find-hidden-remotes, or scan-connections-by-source, so an agent cannot route between them from the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Phase: observe' and 'Requires: active-client' tags imply when the tool is applicable and what precondition must hold, which is genuine usage context. But there is no explicit statement of when to prefer this over sibling remote/connection enumeration tools, so the guidance remains implied rather than actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

screenshot-windowScreenshot a Roblox windowA
Read-onlyIdempotent

Capture an OS screenshot of a Roblox window on the SERVER host and save it as a PNG, returning the saved file path and the captured width/height. Windows-only: locates the RobloxPlayerBeta window by processId or by a windowTitle substring (the single Roblox window if neither is given), reads its rectangle via the user32 GetWindowRect API, and copies the pixels with System.Drawing.CopyFromScreen. Provide savePath to choose the output file, otherwise a PNG is written to the system temp directory. Use list-roblox-windows first when several windows are open. Returns { savedPath, width, height }. Signature: { windowTitle: string?, processId: number?, savePath: string? }. Phase: observe; cost=low; idempotency=read-only. Requires: resolved-target. Produces: structured-result. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
savePathNoDestination .png path. Defaults to a file in the system temp directory.
processIdNoExact process id of the target Roblox window.
windowTitleNoCase-insensitive substring of the target window's title.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint, so the safety profile is covered. The description adds genuinely new behavioral context: execution happens on the SERVER host, it is Windows-only, it uses user32 GetWindowRect plus System.Drawing.CopyFromScreen, and a file is written to the temp directory when no savePath is supplied. What it does not add is failure behavior (it defers to tool-schema), which is explicitly pushed elsewhere.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action, result, and platform constraint in the first sentence, followed by targeting and implementation detail. It is dense but well ordered; the trailing boilerplate 'On failure: inspect tool-schema...' and the restated Signature line are redundant padding against the compact JSON schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter read-only tool with no output schema, the description covers everything an agent needs: targeting semantics, platform limitation, side effect (file written), default destination, and the returned fields { savedPath, width, height }.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters and the baseline is 3. The description adds meaning beyond the schema by specifying the resolution order between processId and windowTitle, including the default of 'the single Roblox window if neither is given,' which the schema does not state.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (capture an OS screenshot of a Roblox window, save as PNG, return path and dimensions) and distinguishes it from siblings like list-roblox-windows and observe-world. The scope (server host, Windows-only) and the exact return shape are named, so an agent can identify it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete routing rule: 'Use list-roblox-windows first when several windows are open,' naming the alternative and the condition that selects it. It also explains the no-argument default (single Roblox window). It does not discuss when this tool is inappropriate (e.g., vs. in-game capture tools like observe-world), so it stops short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scriptRun a Luau Script with the Whole Tool Surface (mcp.*) + Persistent VMA
Destructive

Run a Luau program in the active Roblox client that can ALSO call any other tool inline through a live mcp table, and use the results in the same script — one call instead of dozens of round-trips. Inside the script: game, workspace, and all in-game globals are available, PLUS mcp.<tool>(args) invokes any of this server's tools and RETURNS its data. Tool names are camelCase of the tool id (e.g. mcp.getPlayers(), mcp.searchInstances({ className = 'RemoteEvent' }), mcp.findFunctionsByName({ name = 'buy' })), or use mcp.call('kebab-tool-name', { ... }). For N independent calls use mcp.all({ players = { 'get-players' }, remotes = { 'search-instances', { className = 'RemoteEvent' } } }) — runs them all in parallel server-side with a single round-trip and returns a table keyed identically. print/warn are captured and returned as output (and still stream to the dashboard Output console). By default the script runs in a PERSISTENT VM: globals and functions you define survive across script calls (a REPL-like session) — set persistent:false for a clean one-shot, or call vm-reset to wipe the VM. Returns { result = <return value>, output = [lines] } or { error, output }. mcp.script is disabled (no recursion). Signature: { source: string, client: string?, agent: string?, persistent: boolean?, timeoutMs: number?, rpcBudget: number?, threadContext: number? }. Phase: orchestrate; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval, validated-source. Produces: structured-result. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNoOptional. A stable label for WHICH agent is calling when several share this MCP session (e.g. 'researcher'). Gives that agent its own fair scheduling lane, its own persistent VM on each game, and its own queue budget, so co-tenant agents don't starve or clobber each other.
clientNoOptional. Run on a specific connected client — its clientId OR username — for THIS call only, overriding your session's select-client binding WITHOUT changing it. Lets multiple agents drive different games at the same time; omit to use your session's selected client.
sourceYesThe Luau script to run. Has `game`/`workspace`/all in-game globals, plus `mcp.<tool>(args)` to call any tool and use its returned data inline. `print`/`warn` are captured. `return <value>` to hand a value back.
rpcBudgetNoMax number of `mcp.*` tool calls this script can make through the bridge (default 500). Further calls reject with BUDGET_EXCEEDED so a runaway loop can't saturate the connection.
timeoutMsNoOverall timeout for the whole script including nested tool calls (default 120000).
persistentNoRun in the persistent VM so defined globals/functions survive across calls (default true). false = a fresh, isolated environment each run.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only mark it non-read-only, destructive, and non-idempotent; the description adds substantial behavior beyond that: persistent-VM semantics (globals survive across calls, REPL-like), recursion disabled for mcp.script, print/warn capture and streaming, the exact return shapes, cost/idempotency class, mutation-approval and validated-source requirements, and the MUTATING-writes-live-state warning.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core value proposition and then organized into mechanics, return shape, signature, and safety. It is long and repeats some material that also lives in the schema (the trailing Signature line, default values), but most sentences earn their place for such a complex tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return shape itself ({ result, output } / { error, output }) and covers requirements, failure guidance ('inspect tool-schema'), verification (assert-state), and the recursion/budget limits. Nothing an agent needs to call it correctly appears missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all seven parameters thoroughly, giving a baseline of 3. The description adds value by summarizing the full signature and clarifying behavioral interplay — persistent:false for a one-shot, resetting via vm-reset, and the mcp.all parallel pattern — beyond the raw field definitions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource (runs a Luau program in the active Roblox client) with a sharp differentiator from siblings like execute/run-luau: it can call any other tool inline via the live `mcp` table and uses a persistent VM. The opening sentence lets an agent distinguish it from the other execution tools without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Strong context for when to reach for it — 'one call instead of dozens of round-trips' for orchestration, and `mcp.all` for N independent calls in a single round-trip. It names vm-reset and persistent:false behaviors, but does not explicitly contrast against the many sibling execution tools (execute, execute-and-wait, batch-execute, eval-expression), nor state a when-not-to-use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

script-fanoutRun a Luau script on N clients in parallelA
Destructive

Run ONE Luau program across multiple connected clients in parallel and return per-client results. Targets are chosen by either passing clients: ['<clientId>', ...] (explicit ids from list-clients) or clients: 'all' (every connected client). Each client gets its own scriptToken / RPC budget, and the same mcp.* surface (including mcp.all()) is available inside. Concurrency is capped to 8 in flight so a 50-client fanout doesn't saturate the bridge. Returns { results: [{ clientId, displayName, ok, result?, error?, output, durationMs }], summary: { total, ok, failed, totalMs } }. Pre-flight catches typo'd mcp.* calls before dispatch, so a typo doesn't fail N times. Signature: { clients: {string} | any, source: string, persistent: boolean?, timeoutMs: number?, rpcBudget: number?, threadContext: number? }. Phase: orchestrate; cost=high; idempotency=contextual-write. Requires: explicit-mutation-approval, validated-source. Produces: structured-result. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
sourceYesThe Luau script to run on every target client. Has the same `mcp.*` surface as the regular script tool.
clientsYesEither an array of clientIds (from list-clients) or the string 'all'.
rpcBudgetNoPer-client mcp.* RPC cap (default 500).
timeoutMsNoPer-client timeout in ms (default 60000).
persistentNoEach client's persistent VM is independent (default true = use VM env).
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: discloses the 8-in-flight concurrency cap, per-client scriptToken/RPC budget isolation, pre-flight validation that catches typo'd mcp.* calls before fanout, and the MUTATING/no-idempotency nature of writing live client state. Annotations already flag destructive/non-readOnly, and the description adds the operational constraints an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose and target selection before the return shape and metadata tail. The trailing 'Phase/cost/Requires/Produces/Verify' line and the signature block partly duplicate the schema and annotations, costing a little tightness, but every section is scannable and purposeful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description fully specifies the return shape ({results: [...], summary: {...}}), the failure mode, and the prerequisites. For a high-cost mutating fanout tool with 6 params, nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents source, clients, rpcBudget, timeoutMs, persistent, and threadContext with defaults and bounds. The description's signature block largely restates the schema rather than adding new semantics, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Run), resource (ONE Luau program), and scope (across multiple connected clients in parallel) with the return shape named. It is clearly separable from single-target siblings like execute, run-luau, and batch-execute.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly explains how targets are selected (explicit clientId array from list-clients vs the 'all' sentinel) and gives the failure path (inspect tool-schema, verify with assert-state). It does not, however, contrast itself against near-neighbors such as batch-execute or run-on-actor, so the when-not guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

script-grepGrep across all scripts in the gameA
Read-onlyIdempotent

Search decompiled Roblox scripts for a pattern, line by line, with surrounding context. Enumerates every client-readable LuaSourceContainer reachable from the active client (via QueryDescendants plus nil-parented scripts), decompiles each, and reports matching lines grouped per script. Matching uses Luau string.find: with literal=true the query is matched as a plain substring; otherwise it is treated as a Luau string pattern (note: Luau patterns, not JavaScript regex). Use exact identifiers or simple patterns; use semantic-search-scripts when behavior is known but names are not. Decompilation is best-effort and can be slow on large places, so the scan is capped by maxScripts. Signature: { query: string, root: any?, limit: any?, contextLines: any?, maxMatchesPerScript: any?, maxScripts: any?, literal: any?, caseSensitive: any?, threadContext: number? }. Phase: orchestrate; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-result. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
rootNoRoot instance to enumerate scripts under (e.g. 'game', 'game.ReplicatedStorage'). Defaults to 'game'. Narrow this to keep the scan fast.game
limitNoMaximum number of scripts to return results from (default: 50).
queryYesThe search pattern. With literal=false it is interpreted as a Luau string pattern (%d, %w, %s, character classes [a-z], anchors, etc.). Use the literal flag for exact substring matching.
literalNoWhen true, treats the query as a plain literal substring - no pattern interpretation (string.find with plain=true). (default: false)
maxScriptsNoMaximum number of scripts to decompile and scan before stopping (default: 400).
contextLinesNoNumber of lines of context to show before and after each match (default: 2).
caseSensitiveNoWhen false, matches case-insensitively. (default: true)
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
maxMatchesPerScriptNoMaximum number of matches to return per script (default: 20).

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only/idempotent/non-destructive, yet the description adds substantial context beyond them: enumeration path (QueryDescendants plus nil-parented scripts), best-effort decompilation, the maxScripts cap, the active-client prerequisite, and the Luau-pattern (not JavaScript regex) matching semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening is well front-loaded and dense with useful behavior, but it then restates the full parameter signature (mostly typed as 'any') that the schema already provides, and repeats annotation-derived metadata lines like idempotency=read-only. Those blocks are padding against an otherwise efficient description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter, no-output-schema tool, the description covers the return shape (matching lines grouped per script), performance caveats, prerequisites, and a failure path (inspect tool-schema), leaving nothing an agent needs to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so baseline is 3, but the description adds genuine meaning by clarifying that literal=true means plain substring and otherwise the query is a Luau string pattern rather than a JS regex, which is the most likely source of misuse.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+scope: grep decompiled Roblox scripts line by line with context, and describes what gets enumerated. It clearly separates itself from semantic-search-scripts by naming that sibling in the same breath.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to use exact identifiers or simple patterns and to switch to semantic-search-scripts when behavior is known but names are not. It also warns that decompilation is slow on large places and advises narrowing root, giving concrete when-and-how guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search-bytecodeSearch compiled bytecode for a signature (IDA byte search)A
Read-onlyIdempotent

Scan every script returned by getscripts(), dump each one's compiled bytecode with getscriptbytecode, and find scripts whose bytecode contains a given hex byte pattern — the runtime equivalent of an IDA binary/byte search. Use it to locate scripts carrying a known opcode signature, constant blob, or fingerprint. Provide the pattern as a hex byte string (e.g. '1a2b3c' or '1a 2b 3c'); it is normalized (spaces stripped, lowercased) and must be an even number of hex digits. Each match reports the script's full name, class, and the byte offset of the first hit. Requires getscripts + getscriptbytecode; caps the scan and flags truncation. Signature: { hexPattern: string, limit: any?, maxScan: any?, threadContext: number? }. Phase: observe; cost=high; idempotency=read-only. Requires: active-client. Capabilities: getscriptbytecode. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax matching scripts to return (default 50).
maxScanNoMax scripts to scan (default 6000).
hexPatternYesHex byte pattern to search for in compiled bytecode, e.g. '1a2b3c' or '1a 2b 3c'. Spaces are stripped and it is lowercased; must contain an even number of hex digits.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish read-only/idempotent/non-destructive, so the bar is lower, and the description adds real context beyond them: scan caps with truncation flagging, the required dependency tools, and what each match reports (full name, class, first-hit byte offset). Could disclose more about scan cost behavior, but it is solid.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and use case, then constraints, then return shape. However, the trailing signature block and Phase/cost/idempotency boilerplate duplicate structured data and add length without new information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description competently explains the return shape (script name, class, byte offset of first hit) and the truncation behavior. Dependencies and pattern format are covered, leaving little an agent needs missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so constraints and defaults are already documented. The description restates the hexPattern format and normalization, adding no meaning beyond the schema fields; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (scan scripts, dump compiled bytecode, find a hex byte pattern) and even names the underlying pipeline (getscripts + getscriptbytecode). The 'runtime equivalent of an IDA binary/byte search' framing clearly distinguishes it from siblings like get-script-bytecode or find-string-in-tables.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit use cases ('locate scripts carrying a known opcode signature, constant blob, or fingerprint') and states its prerequisite dependency (getscripts + getscriptbytecode). It does not name a specific alternative tool for the when-not case, which keeps it just below a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search-gc-valueSearch the GC heap for where a value is stored (Cheat-Engine scan)A
Read-onlyIdempotent

Cheat-Engine-style heap scanner: find WHERE a specific value lives across the entire Luau garbage collector. Resolve a target from one of five value types, then walk every live object via getgc(true) and report each place that holds it. For GC TABLES the tool pcall-iterates pairs() and records a hit when a KEY or a VALUE matches (string matches support exact OR 'contains' substring search). For Lua CLOSURES (when scanFunctions is on) it scans the function's constants and upvalues. Each match reports { container, where, keyText? } where 'where' is one of value/key/constant/upvalue and 'container' is the table address or the closure's source:line. Use it to locate a coin/HP total, a remote or flag name, a boolean toggle, or a Part/Instance reference, then pivot with inspect-closure / dump-table / set-closure-upvalue to read or mutate it. Requires getgc; closure scanning additionally requires getconstants/getupvalues (each is type-guarded and simply skipped if the executor lacks it). Everything is pcall-guarded so locked/dead objects never abort the scan; the object count is capped by maxScan and the result list by limit, with a 'truncated' flag. Returns { valueType, matchCount, truncated, matches } or { error }. Signature: { valueType: "number" | "string" | "boolean" | "instance" | "raw", value: string | number | boolean?, match: any?, scanFunctions: any?, limit: any?, maxScan: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Capabilities: getgc. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax number of match locations to return (default 150). Hitting this sets truncated=true.
matchNoMatch mode (default 'exact'). 'contains' performs a plain (non-pattern) substring search and applies ONLY to string searches; for every other value type it is ignored and exact equality is used.exact
valueNoThe value to search for. For number/string/boolean this is the literal value. For 'instance' it is a Luau path expression resolving to an Instance. For 'raw' it is a Luau expression that is evaluated and whose result is the search target. Required for instance/raw; for number/string/boolean it may be omitted only if you really mean to search for the empty string / 0 / false (prefer always supplying it).
maxScanNoMax number of GC objects (tables + functions) to examine before stopping (default 40000). Hitting this sets truncated=true. Raise for a more thorough sweep at the cost of time; lower it if scans are slow.
valueTypeYesHow to interpret 'value' and what to search for: - 'number': search for a numeric value (e.g. a coin/HP total). - 'string': search for a string (supports match='contains' for substrings — e.g. a remote name). - 'boolean': search for true/false (e.g. a god-mode flag). - 'instance': 'value' is a Luau path/expression resolving to an Instance (e.g. 'game.Workspace.Boss'). - 'raw': 'value' is an arbitrary Luau expression; the tool searches for whatever it evaluates to (e.g. 'Enum.KeyCode.E', 'Vector3.new(0,0,0)', 'game:GetService("Players").LocalPlayer').
scanFunctionsNoAlso scan Lua closures' constants and upvalues for the target (default true). Disable to scan tables only, which is faster and avoids getconstants/getupvalues overhead.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the read-only/idempotent annotations it discloses a great deal of behavioral detail: pcall-guarded traversal so dead/locked objects never abort, type-guarded getgc/getconstants/getupvalues that skip silently, table pair-iteration vs closure constant/upvalue scanning, and truncation via limit/maxScan. This is well past what the annotations convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose and scan mechanics before usage and prerequisites, and most sentences carry information. It is somewhat long and the tail (Signature restatement, phase/cost/idempotency, capabilities, safety) partly duplicates the schema and annotations rather than adding to them.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description supplies the return shape ({ valueType, matchCount, truncated, matches } or { error }), the getgc dependency plus optional getconstants/getupvalues, the active-client requirement, and the truncation behavior. An agent can call and interpret it correctly without opening other artifacts.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so the baseline is 3, but the description adds meaningful semantics: match='contains' applies only to strings and is ignored otherwise, and value semantics per valueType (Luau path for instance, evaluated expression for raw). It also describes the match record shape { container, where, keyText? }, which is not in any schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: find WHERE a value lives across the Luau garbage collector, with five value types. Distinguishes itself from siblings like find-string-in-tables and find-tables-by-key by its value-type-driven, heap-wide scan scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete use cases (coin/HP total, remote or flag name, boolean toggle, Part/Instance reference) and names the pivot tools (inspect-closure / dump-table / set-closure-upvalue) for follow-up. It does not explicitly contrast itself with nearby scanners like find-string-in-tables or scan-number-range, leaving that to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search-instancesSearch for instances in the gameB
Read-onlyIdempotent

Search Roblox instances with QueryDescendants selector syntax. Use for class, name, tag, property, and attribute queries against a chosen root. Signature: { selector: string, root: any?, limit: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
rootNoThe root instance to search from (e.g., 'game.Workspace', 'game.ReplicatedStorage'). Defaults to 'game' if not specified.game
limitNoMaximum number of results to return (default: 50, to avoid overwhelming output)
selectorYesSelector string to filter instances. Supports classes (Part), tags (.Tagged), names (#HumanoidRootPart), properties ([CanCollide = false]), attributes ([$QuestId] or [$Health = 100]), child/descendant combinators (> and >>), OR selectors (,), :not(), and :has(); chain selectors for AND logic, e.g. Part.Tagged[Anchored = false].
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so safety is covered structurally. The description reinforces with 'Phase: observe', 'idempotency=read-only', and 'Safety: read-only' — redundant rather than additive. 'Produces: bounded-candidates' hints at return shape (tied to the limit param) but lacks rate-limit, pagination, or failure-mode detail beyond a pointer to tool-schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose well, but the 'Signature/Phase/cost/idempotency/Requires/Produces/Safety/On failure' block is dense metadata that largely restates annotations and schema. Several statements ('read-only' twice, annotations restated) don't earn their place, though the signature is useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, no annotations gap for safety, but the description does not describe return structure beyond 'bounded-candidates', nor does it cover error conditions beyond directing the agent to tool-schema. Selector grammar coverage is good, but the read-only/observation tool is under-specified on result shape and pagination behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, giving a baseline of 3, but the description adds value: it surfaces the full selector grammar (combinators, :not(), :has(), attribute syntax) and the required/optional split via the 'Signature' line, which the schema only partially conveys. It also frames the limit's purpose ('bounded-candidates').

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Search Roblox instances') and narrows the query capability (class, name, tag, property, attribute) against a chosen root. It's clear, but it never explicitly distinguishes from sibling finders like find-instances-with-connections or get-instance-tree, leaving the agent to infer routing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description says what kinds of queries it supports ('class, name, tag, property, and attribute queries'), which implies usage, but never states when to prefer this over siblings like get-instance-tree or find-instances-with-connections. No explicit exclusions or alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

select-clientBind this session to a Roblox client (account-sticky)A
Read-onlyIdempotent

Bind a connected Roblox client to THIS session so your tool calls target that game. Select by clientId OR by username — selecting by username is recommended because it is account-sticky: a client gets a new clientId every reconnect, but a username binding follows the account across rejoins without re-selecting. Each session keeps its own selection, so two sessions can drive two games at once. If several agents SHARE one session, don't fight over this binding — instead pass client (clientId or username) directly to script/execute/run-luau to target a game for a single call without changing anyone's selection. Provide exactly one of clientId or username; the value must match a currently-connected client. Signature: { clientId: string?, username: string? }. Phase: observe; cost=low; idempotency=read-only. Requires: none. Produces: structured-result. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
clientIdNoOptional connected Roblox client id returned by list-clients.
usernameNoOptional Roblox username used to resolve a connected client.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnly/idempotent/non-destructive, and the description adds genuinely non-obvious behavior: per-session scoping, account-stickiness of username bindings, and the requirement that the value match a currently-connected client. It also restates the safety/idempotency metadata, which is redundant, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core guidance is front-loaded and every substantive sentence earns its place, but the trailing 'Phase/cost/Requires/Produces/Safety/On failure' block is boilerplate that largely duplicates the annotations and adds length without new information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-param, no-output-schema, read-only binding tool, the description covers selection semantics, scoping, alternatives, failure handling, and validation rules. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds constraints the schema does not encode: exactly one of clientId/username must be provided, and the username form is preferred for stickiness. That is real added meaning beyond the per-field descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (bind) and resource (connected Roblox client to THIS session) and explains the resulting effect on subsequent tool calls. It is clearly distinguishable from siblings like list-clients, get-active-client, and clear-selection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to select by clientId OR username, recommends username with a concrete reason (account-sticky across reconnects), and tells the agent what to do instead when several agents share a session (pass `client` to script/execute/run-luau). This is a full when/when-not/alternative treatment.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

semantic-search-scriptsSemantic search over loaded scriptsA
Read-onlyIdempotent

Rank the active client's loaded/GC scripts by semantic relevance to a natural-language query. NOTE: full decompiled source is NOT available in the clean protocol, so each script is indexed over its GetFullName() path + Name + ClassName + the string constants reachable from its closure (capped per script) — think of it as 'find the script most likely about X', not a full-text code search. The first call embeds and caches every script (locally or via the configured embeddings endpoint); later calls reuse the cache for unchanged scripts. Returns { hits: [{ path, score, snippet }], model } sorted by descending cosine similarity. Signature: { query: string, limit: any?, maxScripts: any? }. Phase: orchestrate; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of ranked hits to return.
queryYesNatural-language description of the script you are looking for.
maxScriptsNoCap on how many scripts to harvest and index from the client.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only/idempotent, and the description adds substantial behavior: the first call embeds and caches every script (locally or via the embeddings endpoint), later calls reuse the cache for unchanged scripts, and the index is capped/approximate because decompiled source is unavailable. It also notes cost=medium and the active-client requirement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose is front-loaded and every clause about indexing and caching earns its place. However, the trailing metadata block (Phase/cost/idempotency/Requires/Produces/Safety) partially repeats the annotations (Safety: read-only, idempotency=read-only), diluting density.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only search tool, the description is complete: it explains the indexing substrate and its limits, caching cost model, requirement for an active client, the return shape ({ hits, model } sorted by cosine similarity), and even failure guidance. No output schema exists, yet return values are described.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already carries parameter meaning; the description's signature line ({ query, limit: any?, maxScripts: any? }) restates the parameters without adding semantics beyond what the schema documents. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+mechanism: it ranks the active client's loaded/GC scripts by semantic relevance to a natural-language query. It also clarifies the indexing substrate (GetFullName path + Name + ClassName + reachable string constants), which no sibling tool does. An agent can tell this apart from grep/name-based scanners without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a clear framing for when to reach for it ('find the script most likely about X') and an explicit when-not ('not a full-text code search'). It stops short of naming a concrete alternative sibling such as script-grep or scan-closures-by-name, so the routing is implied rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

send-packetSend a custom RakNet packetA
Destructive

WRITES LIVE GAME STATE — transmits a raw packet. Sends a custom OUTGOING low-level packet via raknet.send with the given payload (a hex string, converted to a byte array), priority, reliability, and ordering channel. Requires the raknet library. WARNING: malformed packets or wrong metadata can disconnect the client or break protocol behavior — only send payloads you understand. Signature: { dataHex: string, priority: any?, reliability: any?, orderingChannel: any?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval. Capabilities: RakNet packet APIs. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; performs external network or socket I/O. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
dataHexYesPayload as a hex string (e.g. '01020304' or '01 02 03 04'); whitespace is ignored. Must be an even number of hex digits.
priorityNoRakNet send priority (default 0).
reliabilityNoRakNet reliability mode (default 0).
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
orderingChannelNoOrdering channel for ordered traffic (default 0).

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructive/openWorld/non-idempotent, but the description goes further: it discloses the disconnect/protocol-break risk of malformed payloads or wrong metadata, the approval requirement, the medium cost, that it performs external socket I/O, and that it produces an operation-receipt verifiable with assert-state. That is real behavioral context beyond structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the critical 'WRITES LIVE GAME STATE' warning and the core verb before the metadata tags. The trailing phase/cost/idempotency/requires/verify tag block is dense but each token is terse and distinct, though 'On failure: inspect tool-schema...' is generic filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by stating what is produced (operation-receipt) and how to verify it (assert-state), plus prerequisites, approval and failure guidance. For a destructive, open-world socket-writing tool this is complete enough to call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents each parameter, setting the baseline at 3. The description only adds the byte-array conversion detail for dataHex and restates the field list; priority/reliability/orderingChannel semantics come entirely from the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb, resource, direction and mechanism: sends a custom OUTGOING low-level RakNet packet via raknet.send with payload/priority/reliability/ordering channel. The 'OUTGOING raw packet' framing clearly separates it from observation siblings like packet-spy or block-packets without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear preconditions (requires the `raknet` library, active-client, explicit-mutation-approval) and a use constraint ('only send payloads you understand'). However, it never names a higher-level alternative such as fire-remote for normal remote traffic, so the agent gets context but no explicit routing away from this low-level tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session-listList Recorded Tool-Call SessionsA
Read-onlyIdempotent

Enumerate every recorded session in ~/.executor-mcp/sessions/.jsonl. Each entry shows sessionId, sessionLabel (e.g. 'live'), startedAt, endedAt, the number of recorded calls, and file size. Newest-first. Use the returned sessionId with session-show to read records or session-replay to plan a re-issue. Signature: {}. Phase: observe; cost=low; idempotency=read-only. Requires: none. Produces: bounded-candidates, diagnostic-report. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered structurally. The description adds real value beyond that: the file layout, the exact fields returned per entry, and the newest-first ordering. It stops short of noting limits like pagination or how many sessions can accumulate, so a 4 rather than 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and storage location, then return fields, ordering, and routing. The trailing 'Phase: observe; cost=low; idempotency=read-only. Requires: none. Produces: ...' block is somewhat boilerplate and partially duplicates the annotations, but it is short and scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of describing the return shape and does so (sessionId, sessionLabel, startedAt, endedAt, call count, file size). It also covers location, ordering, and downstream usage, so nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters (signature {}), so there is nothing to disambiguate; baseline for an empty schema is 4. The description correctly confirms the empty signature rather than leaving the agent to wonder whether filters exist.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Enumerate every recorded session') plus the exact storage location (~/.executor-mcp/sessions/<id>.jsonl), which pins down what is being listed. It also names the two sibling tools it feeds into (session-show, session-replay), so an agent can distinguish it from them immediately.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent onward: 'Use the returned sessionId with session-show to read records or session-replay to plan a re-issue.' Each alternative is paired with the condition that selects it, leaving no inference required for the obvious next step.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session-replayPlan or Replay a Recorded Session's Tool CallsA
Destructive

Read one session's recorded trace and (by default) return a structured plan; pass dryRun: false to actually re-issue each plannable step through the invoker. Tools that mutate game/host state are flagged blockedByDefault: true and are SKIPPED unless includeMutating: true is also set. Steps that originally failed are never replayed. The replayed steps themselves get recorded into the CURRENT session's trace, so you can chain saves into a playbook via playbook-save. session-replay refuses to call itself. Signature: { sessionId: string, from: number?, to: number?, includeMutating: boolean?, dryRun: boolean? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: explicit-mutation-approval. Produces: diagnostic-report. Verify with: assert-state. Safety: MUTATING; changes server-side persisted/session state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
toNoOptional numeric value for to.
fromNoOptional numeric value for from.
dryRunNoWhen true (default) only returns the plan. False actually re-issues each step in order.
sessionIdYesSession UUID from session-list.
includeMutatingNoAllow flagged-mutating tools to be planned/executed. Default false.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructive=true / readOnly=false, and the description richly extends this: mutating tools are flagged blockedByDefault and skipped, failed steps are never replayed, replayed steps are recorded into the CURRENT session, and explicit-mutation-approval is required. This is real behavioral context beyond the annotation flags, with no contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Core semantics are front-loaded and tight. However, the appended metadata block (Phase/cost/idempotency/Produces/Verify/On-failure) and the signature line that restates the schema params are templated padding that dilutes an otherwise crisp description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex, state-changing tool with no output schema, it covers safety, approval requirements, and chain-to-playbook flow well. It does not describe the shape of the returned 'structured plan', which is the main piece an agent might still need.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both `dryRun` and `includeMutating` defaults are already documented in the schema. The description adds a useful signature recap and reinforces the defaults, but adds no syntax or format detail beyond what the schema provides — baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb pair (read/plan vs re-issue) on a specific resource (a recorded session's tool-call trace), and makes the default behavior explicit. It is clearly distinguishable from siblings like playbook-run or execute, since it operates on an existing session trace.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when/when-not: default is planning, `dryRun: false` to actually replay, `includeMutating: true` needed for side-effecting steps, and failed steps are never replayed. It names `playbook-save` as the chaining alternative and states session-replay refuses to call itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session-showRead Recorded Tool Calls from a SessionA
Read-onlyIdempotent

Read a window of recorded tool calls from one session's JSONL trace. Pass sessionId (from session-list) and optional from/to to bound the seq range (1-indexed, inclusive). Each record has { seq, at, tool, input, result?|error?, elapsedMs, clientId?, sessionId } — exactly what the invoker saw. Useful for auditing what was run and feeding session-replay. Signature: { sessionId: string, from: number?, to: number?, limit: number? }. Phase: observe; cost=low; idempotency=read-only. Requires: none. Produces: diagnostic-report. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
toNoStop at this seq (default end-of-file).
fromNoStart at this seq (default 1).
limitNoCap returned records (default 200).
sessionIdYesSession UUID from session-list.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so safety is covered; the description adds value beyond them by disclosing the record shape ({ seq, at, tool, input, result?|error?, elapsedMs, clientId?, sessionId }), that it reflects 'exactly what the invoker saw', and phase/cost/output semantics. It still omits capacity/pagination behavior at the window edge, keeping it below a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then the required arg, then the record shape and routing. It is dense but every sentence carries information; the tail metadata ('Phase: observe; cost=low; idempotency=read-only ... Safety: read-only') slightly duplicates the annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully supplies the return record shape and error alternation, plus a failure fallback pointing at tool-schema. Only the boundary behavior of from/to/limit (e.g. what happens when the range exceeds limit) is left unexplained, so it is nearly complete rather than fully so.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds semantics the schema does not: the range is seq-based, 1-indexed and inclusive, and from/to bound that range while limit caps output. This goes beyond restating field names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+scope: 'Read a window of recorded tool calls from one session's JSONL trace', and immediately scopes it to a single session with a seq range. This distinguishes it from session-list (which produces the id) and session-replay (which consumes the output), so an agent can route without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete context: use it for auditing what was run and for feeding session-replay, and tells the caller where sessionId comes from ('from session-list'). It stops short of stating exclusions (e.g. when to prefer session-replay instead), so it is clear but not a full when/when-not/alternative matrix.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set-attributeSet (or clear) a custom attribute on a live InstanceA
Destructive

WRITES LIVE GAME STATE. Resolve a Luau expression to an Instance and write one of its custom attributes via inst:SetAttribute(name, value), returning both the OLD and NEW value so the change is auditable. Attributes are the named, typed key/value pairs games store on instances (visible in the Studio Attributes panel and read with :GetAttribute) — distinct from engine properties (use set-instance-property for those). Common uses while debugging: flip a 'IsAdmin'/'Frozen' boolean attribute, bump a 'Cooldown'/'Damage' number, or set a 'State' string. For non-primitive attribute types use value.kind='raw'; use kind='nil' to DELETE the attribute. The read of the old value and the write are each pcall-guarded. WARNING: this mutates the running game on the client — the change takes effect immediately and may replicate, and game scripts listening on GetAttributeChangedSignal will fire. Returns { Path, Attribute, OldValue, NewValue, ok } or { error }. Signature: { instancePath: string, attributeName: string, value: { kind: "string" | "number" | "boolean" | "nil" | "raw", value: string | number | boolean? }, threadContext: number? }. Phase: act; cost=medium; idempotency=idempotent-write. Requires: active-client, resolved-target, explicit-mutation-approval. Produces: operation-receipt. Verify with: get-instance-properties. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
valueYesThe new attribute value to write, expressed as a typed argument.
instancePathYesLuau expression resolving to the Instance whose attribute you want to set, e.g. 'game.Workspace.Part', 'game.Players.LocalPlayer', or 'game:GetService("Workspace").Boss'. Evaluated as `return <instancePath>`.
attributeNameYesThe exact attribute name to write, e.g. 'IsAdmin', 'Cooldown', 'State'. Case-sensitive. If the attribute does not exist yet it will be created; passing value.kind='nil' deletes it.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only supply destructive/not-read-only/not-idempotent hints; the description goes well past them, disclosing that the write takes effect immediately on the client and may replicate, that GetAttributeChangedSignal listeners will fire, that both the old-value read and the write are pcall-guarded, and that approval plus an operation receipt are required. It also gives the failure envelope ({ error }) and the UserWarning-style 'MUTATING' framing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the highest-value fact (this writes live game state) and organized from purpose to mechanics to warnings. It is dense and nearly every sentence earns its place, but the trailing 'Signature', 'Phase/cost/idempotency' and 'Safety: MUTATING' lines largely restate the schema and annotations, so it is slightly longer than needed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description supplies the return shape ({ Path, Attribute, OldValue, NewValue, ok } or { error }), the idempotency/approval prerequisites, the verification tool, and a fallback instruction ('inspect tool-schema for exact fields'). For a mutating tool with nested value object and 4 parameters, nothing an agent needs to invoke it safely is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds integrative meaning by tying value.kind to intent ('raw' for the non-primitive types Roblox permits, 'nil' to DELETE the attribute) and by giving realistic attributeName examples (IsAdmin, Frozen, Cooldown, Damage, State). The 'Signature' restatement of the schema is redundant, so it does not reach a full 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ('WRITES LIVE GAME STATE', sets a custom attribute on a live Instance) and explicitly differentiates the resource from the sibling concept: 'distinct from engine properties (use set-instance-property for those)'. An agent can select this over list-attributes, set-instance-property, or set-properties-bulk without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the alternative tool and the condition that routes to it (set-instance-property for engine properties), gives concrete debugging scenarios, and states the branch conditions for the call itself: kind='raw' for non-primitive attribute types, kind='nil' to delete. Prerequisites (active-client, explicit-mutation-approval) and a verification step (get-instance-properties) are spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set-clipboardCopy text to the host clipboard (sUNC setclipboard / toclipboard)A
Destructive

WRITES HOST STATE — copies the given text to the host machine's OS clipboard via setclipboard(text), falling back to toclipboard(text) on executors that use that name. This overwrites whatever is currently on the clipboard. Requires one of these functions. The call is type-guarded and pcall-wrapped: if neither is present you get { error = 'setclipboard is not available in this executor.' }. Returns { ok, length } or { error }. Signature: { text: string, threadContext: number?, timeoutMs: number? }. Phase: act; cost=medium; idempotency=idempotent-write. Requires: active-client, explicit-mutation-approval. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesThe text to place on the host clipboard.
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Rich disclosure beyond annotations: the overwrite of existing clipboard contents, the fallback function, the type-guard/pcall wrapping, exact error object and return shape, and required preconditions. However it declares 'idempotency=idempotent-write' while the annotations set idempotentHint=false, a direct conflict on a retry-relevant trait that an agent could act on.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the critical fact ('WRITES HOST STATE') and the overwrite warning, which is good. But it is bloated with metadata tags (Phase, cost, idempotency, Produces) and a redundant signature dump, plus a trailing 'On failure: inspect tool-schema...' sentence that is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by specifying the return shape ({ ok, length } or { error }), the exact failure object, prerequisites, and a verification tool. An agent has everything needed to invoke and check the call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and all three parameters are documented in the schema; the description merely restates the signature, adding no new meaning. Baseline 3 is appropriate when the schema carries the semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('copies the given text to the host machine's OS clipboard') and even names the underlying primitives (setclipboard/toclipboard). No sibling tool touches the clipboard, so it is trivially distinguishable from the rest of the catalog.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear preconditions (requires an available setclipboard/toclipboard, active-client, explicit-mutation-approval) and a follow-up (verify with assert-state). It does not name a when-not or an alternative tool, but no plausible alternative exists, so context is sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set-closure-constantSet a constant of a live function (MUTATES STATE)A
Destructive

WRITES LIVE GAME STATE. DANGER — Resolve a Luau expression to a function and overwrite one of its bytecode constants via setconstant. Constants are the literal values baked into a function's bytecode (numbers, strings, the names of globals/methods it calls). Patching one changes the function's behavior the next time it runs — e.g. rewrite a magic number, swap a hardcoded string, or redirect a method call by changing its name constant. This persists for the exact closure and can destabilize the game or trip anticheat; some executors only allow same-type replacement. Index is 1-based and must be within the function's constant count (inspect-closure reports ConstantCount). Requires setconstant. Pass confirm=true to proceed. Returns { Target, Index, ok } or { error }. Signature: { functionPath: string, index: number, value: { kind: "string" | "number" | "boolean" | "nil" | "raw", value: string | number | boolean? }, confirm: boolean, threadContext: number? }. Phase: act; cost=medium; idempotency=idempotent-write. Requires: active-client, resolved-target, explicit-mutation-approval. Capabilities: debug closure primitives. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
indexYes1-based index of the constant to overwrite. Must be within the function's constant count (see ConstantCount from inspect-closure).
valueYesNew value for the constant. kind 'string'|'number'|'boolean' uses value literally; kind 'nil' sets nil; kind 'raw' treats value as a Luau expression evaluated for its result. Note: many executors require the replacement to be the same type as the original constant.
confirmYesMust be true to actually mutate the live function. When omitted or false, the tool refuses and changes nothing.
functionPathYesLuau expression resolving to the function whose constant you want to change, e.g. 'getsenv(game.Players.LocalPlayer.PlayerScripts.Main).fire' or a function found via scan-closures-by-source. Evaluated as `return <functionPath>`.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true/readOnlyHint=false, but the description adds substantial context beyond them: scope of persistence ('persists for the exact closure'), risk ('can destabilize the game or trip anticheat'), an executor-specific constraint ('some executors only allow same-type replacement'), and the confirm-gating behavior. The only wrinkle is the self-declared 'idempotency=idempotent-write' label sitting alongside idempotentHint=false, but this is a taxonomy field rather than a claim about the tool's effect, and the safety-relevant annotations are fully consistent with the text.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The DANGER warning and the explanation of constants are front-loaded, which is correct for a destructive mutation tool. It is long, however, and the trailing metadata block (phase/cost/signature) partially duplicates the schema rather than earning its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no output schema, the description specifies the return shape ({ Target, Index, ok } or { error }) and failure handling ('inspect tool-schema for exact fields, defaults, constraints'). It also covers prerequisites, verification (assert-state), and safety, so an agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents index bounds, value.kind semantics, confirm behavior, and functionPath evaluation. The description restates the same facts (1-based index, same-type constraint) and its Signature line simply mirrors the schema, adding little beyond the baseline. The 'requires setconstant' note is the only genuinely additive element.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a precise verb+resource ('overwrite one of its bytecode constants via setconstant') and defines what a constant is (literal values baked into bytecode), which distinguishes it from sibling tools like set-closure-upvalue, get-closure-constants, and inspect-closure. It even enumerates concrete outcomes (rewrite a magic number, swap a string, redirect a method call).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear preconditions ('Requires setconstant', 'Pass confirm=true to proceed', index must be within ConstantCount from inspect-closure) and use cases, and points to assert-state for verification. It stops short of an explicit when-not-to-use or a named alternative (e.g., set-closure-upvalue for upvalue edits), so it is strong but not complete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set-closure-upvalueSet an upvalue of a live function (MUTATES STATE)A
Destructive

WRITES LIVE GAME STATE. DANGER — Resolve a Luau expression to a function and overwrite one of its upvalues (a variable captured by the closure) via setupvalue. Use this to patch a function's behavior in place without rehooking it — e.g. flip a captured enabled boolean, swap a captured config table, or zero out a captured cooldown. The change persists for every future call of that exact closure and is shared by all closures that captured the same upvalue, so it can destabilize the game or trip anticheat. Index is 1-based and must be within the function's upvalue count (inspect-closure reports UpvalueCount). Requires setupvalue. Pass confirm=true to proceed. Returns { Target, Index, ok } or { error }. Signature: { functionPath: string, index: number, value: { kind: "string" | "number" | "boolean" | "nil" | "raw", value: string | number | boolean? }, confirm: boolean, threadContext: number? }. Phase: act; cost=medium; idempotency=idempotent-write. Requires: active-client, resolved-target, explicit-mutation-approval. Capabilities: debug closure primitives. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
indexYes1-based index of the upvalue to overwrite. Must be within the function's upvalue count (see UpvalueCount from inspect-closure).
valueYesNew value for the upvalue. kind 'string'|'number'|'boolean' uses value literally; kind 'nil' sets nil; kind 'raw' treats value as a Luau expression evaluated for its result (e.g. value='Vector3.new(0,0,0)' or 'game.Workspace').
confirmYesMust be true to actually mutate the live function. When omitted or false, the tool refuses and changes nothing.
functionPathYesLuau expression resolving to the function whose upvalue you want to change, e.g. 'getsenv(game.Players.LocalPlayer.PlayerScripts.Main).update' or 'getrawmetatable(game).__namecall'. Evaluated as `return <functionPath>`.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds excellent context beyond the annotations (persistence across all future calls, sharing across closures that captured the same upvalue, anticheat risk, confirm gating). However it declares 'idempotency=idempotent-write', which directly conflicts with the annotation idempotentHint=false, so the description contradicts structured metadata on a key behavioral trait.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the highest-severity fact ('WRITES LIVE GAME STATE. DANGER') and then organized into use case, constraints, signature, and failure handling. It is long and the signature block duplicates the schema, but nearly every sentence carries actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex, destructive mutation tool with no output schema, the description supplies the return shape ({ Target, Index, ok } / { error }), prerequisites, verification route (assert-state), confirm gating, and failure guidance. Nothing critical to correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so parameters are already fully documented; the description mostly restates the signature. It still adds value by tying the index bound to a cross-tool workflow ('inspect-closure reports UpvalueCount') and clarifying the raw kind's evaluation semantics, which helps correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource: resolve a Luau expression to a function and overwrite one of its upvalues via setupvalue. It cleanly distinguishes itself from siblings like get-closure-upvalues, set-closure-constant, and set-function-env, and the danger/mutation scope is stated up front.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete when-to-use examples ('flip a captured enabled boolean, swap a captured config table'), an implied alternative ('patch in place without rehooking it' vs hook-function), a prerequisite ('Requires setupvalue') and a gate ('Pass confirm=true'). It stops short of explicitly naming an alternative sibling tool to use instead, so it is strong but not fully routable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set-connection-stateDisable / enable / disconnect signal connectionsA
Destructive

MUTATES CLIENT STATE. Toggles or removes connections on an RBXScriptSignal: 'disable' temporarily stops a connection from firing (reversible with 'enable'), 'enable' re-activates a disabled connection, and 'disconnect' permanently removes it — IRREVERSIBLE for that connection (you would have to recreate it). scope='one' targets a single connection by index; scope='all' applies to every connection on the signal. Because they are destructive, scope='all' and action='disconnect' BOTH require confirm=true or the tool refuses without changing anything. Useful for silencing or surgically removing event handlers (e.g. anti-cheat or input listeners) while debugging. Requires getconnections plus the connection's Disable/Enable/Disconnect methods; degrades with a clear { error } if unavailable. Returns { Signal, Instance?, Action, Scope, Affected } or { error }. Signature: { instancePath: string, signalName: any?, action: "disable" | "enable" | "disconnect", scope: any?, connectionIndex: any?, confirm: any?, threadContext: number? }. Phase: act; cost=medium; idempotency=idempotent-write. Requires: active-client, resolved-target, explicit-mutation-approval. Capabilities: getconnections. Produces: created-handle, operation-receipt. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
scopeNo'one' (default) affects only the single connection at connectionIndex; 'all' affects every connection on the signal and REQUIRES confirm=true.one
actionYesWhat to do to the matched connection(s): 'disable' = stop firing but keep it (reversible via 'enable'); 'enable' = re-activate a disabled connection; 'disconnect' = permanently remove it (IRREVERSIBLE, requires confirm=true).
confirmNoSafety gate for destructive operations. Must be true to run when scope='all' OR action='disconnect'; otherwise the tool refuses without making changes. Not required for disable/enable on a single connection.
signalNameNoName of the RBXScriptSignal member on the instance (e.g. 'Touched', 'Changed'). Leave empty/omit if instancePath already evaluates to the signal.
instancePathYesLua expression resolving to the Instance that owns the signal (e.g. 'game.Workspace.Door'). Evaluated as `return <instancePath>`. Leave signalName empty if this expression already resolves to the signal itself.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
connectionIndexNoZero-based index of the target connection within getconnections(signal) when scope='one', matching the Index from list-signal-connections. Ignored when scope='all'. Defaults to 0.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: distinguishes reversible disable/enable from irreversible disconnect, explains the confirm gate for scope='all' and action='disconnect', notes dependency on getconnections and graceful { error } degradation, and describes the return payload. This is materially more than destructiveHint/readOnlyHint convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the critical 'MUTATES CLIENT STATE' warning and the irreversible-action caveat, and the prose is dense but readable. It loses a point for restating the full signature and appending metadata boilerplate (Phase/cost/capabilities/produces) that duplicates the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, 7-parameter tool with no output schema, the description covers reversibility, safety gating, prerequisites, failure mode, and the return shape. Nothing an agent needs to invoke it safely is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds cross-parameter meaning: the interaction of scope='all' with confirm=true, and that connectionIndex matches the Index from list-signal-connections. That is real semantics beyond field-level text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb set (toggles/removes) on a specific resource (connections on an RBXScriptSignal), and defines each action's effect. It is clearly distinguishable from siblings like list-signal-connections or fire-connection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete use context ('silencing or surgically removing event handlers while debugging') and spells out when each action/scope applies, plus the confirm=true precondition. It references list-signal-connections for indices, but does not name a true alternative tool for the same job.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set-fast-flagOverride a Roblox FastFlag value (sUNC setfflag)A
Destructive

WRITES CLIENT STATE — overrides a Roblox engine FastFlag for the running client via setfflag(name, value). The value is supplied as a string and coerced by the executor to the flag's native type (bool/int/string). This changes engine behavior at runtime and can destabilize the client if a flag is set to an invalid value. Requires setfflag. The call is type-guarded and pcall-wrapped: if setfflag is missing you get { error = 'setfflag is not available in this executor.' }. Returns { name, value, ok } or { error }. Signature: { name: string, value: string, threadContext: number?, timeoutMs: number? }. Phase: act; cost=medium; idempotency=idempotent-write. Requires: active-client, explicit-mutation-approval. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesThe FastFlag name to override, e.g. 'DFIntTaskSchedulerTargetFps'.
valueYesThe new value as a string (e.g. '120', 'true', or a text value).
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description is rich beyond the annotations: value coercion to the flag's native type, pcall/type-guard wrapping, the exact error shape, and the { name, value, ok } return shape. However, it labels the operation 'idempotency=idempotent-write' while the annotations declare idempotentHint=false, a genuine conflict that undercuts the otherwise strong behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the critical 'WRITES CLIENT STATE' warning, then error/return behavior, then signature. It is long and somewhat repetitive ('Requires setfflag' is restated, trailing 'On failure: inspect tool-schema' is boilerplate), but each block is useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description supplies the return shape and error envelope, and there are no nested objects or enums to cover. It is nearly complete for a mutating 4-param tool; the only slip is the inconsistent idempotency claim.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and documents each parameter, so the baseline is 3. The description adds real meaning: the value is a string coerced by the executor to the flag's native type (bool/int/string), and it repeats the full signature with optionality, exceeding the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (overrides) and resource (a Roblox engine FastFlag for the running client) with the underlying call setfflag(name, value). It is clearly distinguishable from siblings like get-fast-flag and set-fps-cap. An agent knows exactly what this does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear invocation context: 'Phase: act', 'Requires: active-client, explicit-mutation-approval', 'Requires setfflag', and 'Verify with: assert-state'. It does not explicitly name alternatives or when-not-to-use (e.g., 'use get-fast-flag to read'), so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set-fps-capSet the render FPS cap (sUNC setfpscap)A
Destructive

WRITES CLIENT STATE — sets the executor's render frame-rate cap via setfpscap(cap). Pass 0 to uncap the frame rate. Requires setfpscap. The call is type-guarded and pcall-wrapped: if setfpscap is missing you get { error = 'setfpscap is not available in this executor.' }. Returns { cap, ok } or { error }. Signature: { cap: number, threadContext: number?, timeoutMs: number? }. Phase: act; cost=medium; idempotency=idempotent-write. Requires: active-client, explicit-mutation-approval. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
capYesThe target frame-rate cap; 0 means uncapped (no limit).
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds genuinely useful behavior beyond annotations: type-guarding, pcall wrapping, the exact error payload for a missing setfpscap, the return shape { cap, ok }/{ error }, and prerequisites. However, it declares 'idempotency=idempotent-write' while the annotations set idempotentHint=false, a direct conflict that misleads an agent deciding whether retrying is safe.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the most important fact ('WRITES CLIENT STATE') and is organized into scannable clauses. It is somewhat bloated with metadata boilerplate (Phase, cost, Produces, Verify with) that dilutes the key guidance, but each part is on-task.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly documents return values and error cases, states prerequisites and a verification path, and warns that the operation is mutating. Only the internal idempotency inconsistency and missing guidance on unsetting/restoring state leave a small gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents cap/timeoutMs/threadContext. The description restates the signature and reinforces the 0=uncapped rule (already in the schema) but adds no new syntax or constraints, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (set) and resource (render FPS cap) plus the underlying primitive setfpscap(cap), and says explicitly it writes client state. It is immediately distinguishable from the sibling get-fps-cap (read) and from unrelated executors/setters.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete call context: 'Pass 0 to uncap the frame rate', 'Requires setfpscap', plus gates (active-client, explicit-mutation-approval) and a verification step (assert-state). It does not name alternatives or when NOT to use it, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set-function-envReplace a closure's environment tableA
Destructive

WRITES LIVE GAME STATE. Resolve a function and an environment-table expression, then call setfenv. This changes global lookup behavior for the live closure and requires confirm=true. Signature: { functionPath: string, threadContext: number?, environmentExpression: string, confirm: boolean? }. Phase: act; cost=medium; idempotency=idempotent-write. Requires: active-client, resolved-target, explicit-mutation-approval. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmNoExplicit safety acknowledgement; must be true before this state-changing operation runs.
functionPathYesLuau expression resolving to the target function.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
environmentExpressionYestext value for environment expression.

TDQS

A3.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds real behavioral data beyond the annotations: it mutates live game state, alters global lookup semantics for the live closure, requires confirm=true, and requires explicit mutation approval. However, it labels the operation 'idempotency=idempotent-write' while the annotations declare idempotentHint=false, which conflicts on retry-safety for a destructive operation and risks an agent re-running it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The most important fact ('WRITES LIVE GAME STATE') is front-loaded before any mechanics. The remaining sentences are dense and mostly earn their place, though the templated Phase/cost/idempotency/receipt metadata reads as boilerplate that dilutes the core message.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, no-output-schema mutation, the description covers what is changed, approval preconditions, the produced operation-receipt, verification via assert-state, and failure guidance pointing to tool-schema. It omits reversibility (there is no undo for a replaced environment) and the exact effect on the previous environment table.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four fields, including confirm's safety meaning and threadContext's default behavior. The description's signature line merely restates the schema without adding format, constraints, or examples beyond what is already structured, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a concrete verb+resource: resolve a function and an environment-table expression, then call setfenv, changing global lookup for the live closure. This clearly separates it from read-only siblings like dump-function-env and get-function-env, and from the adjacent set-closure-upvalue/constant tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It names preconditions (active-client, resolved-target, explicit-mutation-approval), the mandatory confirm=true acknowledgement, the acting phase, and a follow-up verification step with assert-state. It stops short of explicitly naming alternatives (e.g., 'inspect first with dump-function-env') or stating when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set-gui-textSet the .Text of a GUI elementA
Destructive

WRITES LIVE GAME STATE. Resolve a Luau expression to a GUI Instance and overwrite its .Text property, returning both the OLD and NEW text so the change is auditable. Use this to directly poke a value into a TextLabel, TextButton or TextBox without simulating keystrokes — e.g. to pre-fill a login/search box or change a label while debugging UI logic. This sets the raw .Text directly and does NOT fire FocusLost / text-changed input events, so scripts that react only to player typing may not run; use type-text-box with useKeyPress when you need real keystroke side-effects. The read of the old text and the write are each pcall-guarded. WARNING: this mutates the running client UI immediately. Returns { Path, OldText, NewText, ok } or { error }. Signature: { path: string, text: string, threadContext: number? }. Phase: act; cost=medium; idempotency=idempotent-write. Requires: active-client, resolved-target, explicit-mutation-approval. Produces: operation-receipt. Verify with: get-gui-text. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesLuau expression resolving to the GUI Instance whose .Text to set, e.g. 'game.Players.LocalPlayer.PlayerGui.Login.UsernameBox'. Evaluated as `return <path>`. The resolved Instance must have a writable .Text property (TextLabel / TextButton / TextBox).
textYesThe new string to assign to the element's .Text property. Written verbatim.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: discloses that FocusLost/text-changed events do NOT fire, that reads/writes are pcall-guarded, and returns { Path, OldText, NewText, ok } for auditability. Safety traits (mutating, destructive, non-readOnly) agree with the annotations. Minor tension: the metadata line says 'idempotency=idempotent-write' while the annotation sets idempotentHint=false, though the practical effect (setting a fixed string) is harmless to retry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the critical warning 'WRITES LIVE GAME STATE' and organized into purpose, caveats, return shape, and metadata. Slightly long due to the boilerplate trailing 'On failure: inspect tool-schema...' line, but nearly every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description fully specifies the return shape ({ Path, OldText, NewText, ok } or { error }), the mutation warning, prerequisites (active-client, explicit-mutation-approval), and verification (get-gui-text). Nothing an agent needs to invoke it safely is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so path/text/threadContext are already fully documented in the schema, including the Luau expression form and 'written verbatim'. The description's 'Signature:' line only restates the schema, adding no new parameter meaning, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource: 'Resolve a Luau expression to a GUI Instance and overwrite its .Text property.' It is clearly distinguishable from siblings get-gui-text (read), type-text-box/click-button (input simulation), and set-instance-property (generic property write).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use ('pre-fill a login/search box or change a label while debugging UI logic') and when-not, naming the alternative: 'use type-text-box with useKeyPress when you need real keystroke side-effects.' This gives an agent a clean routing rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set-instance-propertySet a property on a live InstanceA
Destructive

WRITES LIVE GAME STATE. Resolve a Luau expression to an Instance and assign one of its properties, returning both the OLD and NEW value so the change is auditable. Common uses while debugging: toggle a GUI's Visible, bump a character's Humanoid WalkSpeed/JumpPower, set a part's Transparency/Anchored/CanCollide, or write an IntValue/StringValue's Value. For non-primitive property types (Vector3, CFrame, Color3, UDim2, Enum, Instance references) use value.kind='raw' and pass a Luau expression. The read of the old value and the write are each pcall-guarded. WARNING: this mutates the running game on the client — the change takes effect immediately and may replicate. Returns { Path, Property, OldValue, NewValue, ok } or { error }. Signature: { instancePath: string, propertyName: string, value: { kind: "string" | "number" | "boolean" | "nil" | "raw", value: string | number | boolean? }, threadContext: number? }. Phase: act; cost=medium; idempotency=idempotent-write. Requires: active-client, resolved-target, explicit-mutation-approval. Produces: operation-receipt. Verify with: get-instance-properties. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
valueYesThe new value to write, expressed as a typed argument.
instancePathYesLuau expression resolving to the Instance to modify, e.g. 'game.Players.LocalPlayer.Character.Humanoid', 'game.Workspace.Part', or 'game:GetService("StarterGui")'. Evaluated as `return <instancePath>`.
propertyNameYesThe exact property name to write, e.g. 'WalkSpeed', 'Visible', 'Transparency', 'Anchored', 'Value', 'Position', 'BrickColor'. Case-sensitive.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.6/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description embeds 'idempotency=idempotent-write' while the annotations declare idempotentHint=false, directly contradicting the structured safety metadata an agent relies on to decide whether repeated calls are safe. This is an Annotation Contradiction, so per the rubric the score collapses to 1 despite the otherwise useful disclosure that reads/writes are pcall-guarded, that the mutation replicates, and that old/new values are returned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The critical safety warning is front-loaded, which is good, but the body then mixes return format, signature, use cases, and a trailing metadata block (Phase/cost/idempotency/Requires/Produces/Verify/Safety) that adds bulk and repeats information already in the schema. Information-rich but not tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description helpfully specifies the return shape ({ Path, Property, OldValue, NewValue, ok } or { error }) and states prerequisites and a failure-recovery path via tool-schema. Coverage is strong; only the self-contradictory idempotency claim undercuts it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so the baseline is 3, but the description adds practical property-value color not present in the schema ('toggle a GUI's Visible', 'bump WalkSpeed/JumpPower', 'set Transparency/Anchored/CanCollide') that helps the agent choose valid property/value combinations. The raw-expression guidance largely restates the schema, keeping it below 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('assign one of its properties' on an Instance) plus the exact scope ('WRITES LIVE GAME STATE'), which cleanly separates it from read-oriented siblings like get-instance-properties and watch-instance-property. The mechanism (resolve a Luau expression to an Instance, then write) is spelled out, so an agent understands what it does before opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit debugging use cases, tells the agent to use value.kind='raw' for non-primitive types, and routes verification to get-instance-properties. It does not, however, say when to prefer this over adjacent mutators such as set-attribute, set-properties-bulk, or write-path-value, so the when-not/alternative coverage is incomplete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set-metatable-readonlyToggle a metatable's read-only flag (MUTATES live state)A
Destructive

WRITES LIVE GAME STATE. DANGER — Resolve a Luau expression to a metatable (usually getrawmetatable(obj)) and flip its read-only flag via setreadonly. Most game metatables (e.g. getrawmetatable(game)) are locked read-only so __index/__namecall cannot be overwritten. Pass readonly=false to temporarily unlock a metatable so you can edit a metamethod (or run hook-metamethod), then call this again with readonly=true to RE-LOCK it — leaving a core metatable writable is a classic anticheat tripwire and can crash or destabilize the game. This changes the live runtime; it does not create a copy. Requires setreadonly. Because it mutates state you MUST pass confirm=true; otherwise the tool refuses and does nothing. Returns { Target, ReadOnly, ok } or { error }. Signature: { targetPath: string, readonly: boolean, confirm: boolean, threadContext: number? }. Phase: act; cost=medium; idempotency=idempotent-write. Requires: active-client, resolved-target, explicit-mutation-approval. Capabilities: getrawmetatable. Produces: structured-observation, operation-receipt. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmYesSafety gate. Must be exactly true to apply the change. If omitted or false the tool refuses and does nothing, because toggling read-only on a live metatable can destabilize the game or trip anticheat.
readonlyYesThe new read-only state. false = unlock the metatable so its metamethods can be edited/hooked; true = re-lock it. Always re-lock when you are done to avoid leaving the game in an unprotected, anticheat-suspicious state.
targetPathYesLuau expression resolving to the metatable (a table) whose read-only flag you want to change, e.g. 'getrawmetatable(game)', 'getrawmetatable(game.Players.LocalPlayer)', or 'getgenv().SomeLockedTable'. Evaluated as `return <targetPath>`. Should resolve to a table — non-tables will error.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=true, and the description adds real value beyond them: it writes the live runtime rather than a copy, requires confirm=true or refuses, returns {Target, ReadOnly, ok} or {error}, and carries crash/anticheat consequences. The one blemish is that its own metadata line claims 'idempotency=idempotent-write' while the annotations declare idempotentHint=false — an internal inconsistency that keeps this from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The critical warning ('WRITES LIVE GAME STATE. DANGER') is correctly front-loaded and the workflow paragraph earns its place, but the tail is appended machine metadata ('Phase: act; cost=medium; ... Requires: ... Capabilities: ... Produces: ... Verify with: ... On failure: inspect tool-schema...') that pads the definition without helping selection or invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, no-output-schema mutation tool this is complete: it covers purpose, required approval, prerequisites, failure behavior, verification path (assert-state), and even the return shape. An agent has everything needed to call it correctly and to know what to do afterward.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and every parameter (confirm, readonly, targetPath, threadContext) is already fully documented in the schema, including the confirm safety gate and the Luau-expression evaluation of targetPath. The description restates the signature and the confirm rule but adds no semantics the schema lacks, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+mechanism: 'Resolve a Luau expression to a metatable (usually getrawmetatable(obj)) and flip its read-only flag via setreadonly.' It also distinguishes itself from nearby siblings (get-metatable, is-readonly, set-rawmetatable) by naming the exact target form and the live-state mutation, so an agent can pick it without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear when-to-use context and workflow: unlock with readonly=false to edit a metamethod or run hook-metamethod, then re-lock with readonly=true, and it warns that leaving a core metatable writable is an anticheat tripwire. It also lists prerequisites (active-client, resolved-target, explicit-mutation-approval). It stops short of explicit 'use X instead of this' routing to siblings, so 4 rather than 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set-properties-bulkSet many properties on one live Instance in a single callA
Destructive

WRITES LIVE GAME STATE. Resolve a Luau expression to a single Instance ONCE, then apply a list of property writes to it in order. For each property the OLD value is read, the new value is written, and the NEW value is read back — each step pcall-guarded so one bad property never aborts the others. This is the efficient way to reconfigure an instance with several changes at once (e.g. make a Part Anchored + CanCollide=false + Transparency=0.5 + a new Size in one round-trip) instead of issuing many set-instance-property calls. The writes happen sequentially within the same execution so the result is effectively atomic from the game's perspective for that frame. WARNING: this mutates the running game on the client — changes take effect immediately and may replicate. Returns { Path, results:[{ name, OldValue, NewValue, ok, error? }], okCount, failCount }, or { error } if the instance itself cannot be resolved. Signature: { instancePath: string, properties: {{ name: string, value: { kind: "string" | "number" | "boolean" | "nil" | "raw", value: string | number | boolean? } }}, threadContext: number? }. Phase: act; cost=medium; idempotency=idempotent-write. Requires: active-client, resolved-target, explicit-mutation-approval. Produces: operation-receipt. Verify with: get-instance-properties. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
propertiesYesNon-empty, ordered list of property assignments to apply, e.g. [{ name: 'Anchored', value: { kind: 'boolean', value: true } }, { name: 'Transparency', value: { kind: 'number', value: 0.5 } }, { name: 'Size', value: { kind: 'raw', value: 'Vector3.new(4,1,4)' } }]. Applied top-to-bottom.
instancePathYesLuau expression resolving to the single Instance to modify, e.g. 'game.Workspace.Part', 'game.Players.LocalPlayer.Character.Humanoid', or 'game:GetService("Lighting")'. Evaluated once as `return <instancePath>`.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes far beyond the annotations: it explains the read-old/write-new/read-back sequence, pcall-guarding so one failure doesn't abort others, sequential effectively-atomic execution within a frame, immediate client mutation and replication, and the full return shape with per-property ok/error. It also names prerequisites (active-client, resolved-target, explicit-mutation-approval) and a verification tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the critical warning ('WRITES LIVE GAME STATE') and orders the explanation logically: resolution, per-property lifecycle, efficiency rationale, atomicity, warning, return shape, signature, and metadata. Some redundancy between the prose return-shape and the signature block keeps it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description fully documents the return object (Path, results array with name/OldValue/NewValue/ok/error, okCount/failCount, error on unresolvable instance), the failure mode, required preconditions, and verification path. An agent has everything needed to call and interpret the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds real value beyond the schema by showing an ordered property list example with the kind:'raw' Vector3 case and by restating the signature with optional threadContext, clarifying the shape of the properties array and the resolution-once semantics of instancePath.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (apply property writes), resource (a single live Instance resolved from a Luau expression), and scope (many properties in one call). It explicitly distinguishes itself from the sibling set-instance-property by presenting itself as the batch alternative: 'instead of issuing many set-instance-property calls'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear condition for use: reconfiguring an instance with several changes at once, with a concrete Part example. It names the alternative (set-instance-property) and why this is preferred, but does not state when NOT to use it (e.g. single property, or when instance is unresolvable).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set-rawmetatableReplace an object's raw metatable (MUTATES live state)A
Destructive

WRITES LIVE GAME STATE. DANGER — Resolve a Luau expression to an object (table/Instance/userdata) and REPLACE its entire metatable via setrawmetatable, bypassing any __metatable lock. This swaps out __index/__namecall/__newindex etc. wholesale, so it can completely change how the object behaves — overwriting the game's core metatable can break the client, sever security routing, and is a strong anticheat signal. Typical RE use: clone the existing metatable, modify a metamethod, then set it back. Note the target metatable may need to be writable (see set-metatable-readonly) before this succeeds. This changes the live runtime in place. Requires setrawmetatable. Because it mutates state you MUST pass confirm=true; otherwise the tool refuses and does nothing. Returns { Target, ok } or { error }. Signature: { objectPath: string, metatableExpr: string, confirm: boolean, threadContext: number? }. Phase: act; cost=medium; idempotency=idempotent-write. Requires: active-client, resolved-target, explicit-mutation-approval. Capabilities: getrawmetatable. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmYesSafety gate. Must be exactly true to apply the change. If omitted or false the tool refuses and does nothing, because replacing a live object's metatable can break the game or trip anticheat.
objectPathYesLuau expression resolving to the object whose metatable you want to replace, e.g. 'game', 'game.Players.LocalPlayer', or 'getgenv().SomeProxy'. Evaluated as `return <objectPath>`.
metatableExprYesRaw Luau expression evaluating to the NEW metatable (a table) to install, e.g. '{ __index = function() return nil end }', 'setmetatable({}, nil)', or a previously cloned/edited table referenced from getgenv(). Evaluated as `return <metatableExpr>`; must resolve to a table or nil.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond annotations (destructiveHint=true, readOnlyHint=false, idempotentHint=false) by naming the concrete mechanics: bypasses __metatable lock, swaps __index/__namecall/__newindex wholesale, can break the client, severs security routing, is an anticheat signal, and requires confirm=true or it refuses. Also discloses the return shape { Target, ok } / { error }. Minimal tension: the trailing 'idempotency=idempotent-write' sits awkwardly against idempotentHint=false, but this is a labeling nuance, not a safety contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the danger and the core action, and most sentences carry distinct information (mechanics, risk, RE workflow, confirm gate, return values). The trailing metadata block (Phase/cost/idempotency/Requires/Capabilities/Produces/Verify/On failure) is dense but partly redundant with structured fields, keeping it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description states the return shape, the mandatory confirm gate, prerequisites (setrawmetatable, writable target, resolved target), and failure behavior. For a destructive live-state mutation tool, an agent has everything needed to invoke it correctly and safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters with examples and the confirm safety gate. The description's inline signature and confirm restatement add little beyond the schema. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: resolve a Luau expression to an object and REPLACE its entire metatable via setrawmetatable, bypassing any __metatable lock. This is clearly distinguishable from siblings like get-metatable, inspect-instance-metatable, hook-metamethod, and set-metatable-readonly. An agent knows exactly what this does versus neighbors.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete when-to-use pattern ('clone the existing metatable, modify a metamethod, then set it back') and a prerequisite chain (target metatable may need to be writable — see set-metatable-readonly; requires setrawmetatable). It stops short of naming an explicit alternative to prefer for lighter-touch cases (e.g. hook-metamethod), so it is strong context without full when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set-stack-hiddenSet a closure's stack-trace visibilityB
Destructive

WRITES LIVE GAME STATE. Call setstackhidden(function, hidden) to alter runtime stack/debug visibility. This can make diagnostics incomplete, so it is confirmation-gated and should be restored with hidden=false after use. Signature: { functionPath: string, threadContext: number?, hidden: boolean, confirm: boolean? }. Phase: act; cost=medium; idempotency=idempotent-write. Requires: active-client, resolved-target, explicit-mutation-approval. Produces: operation-receipt. Verify with: is-function-hooked. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
hiddenYesTrue hides the function; false restores normal visibility.
confirmNoExplicit safety acknowledgement; must be true before this state-changing operation runs.
functionPathYesLuau expression resolving to the target function.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

B3.4/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations declare idempotentHint=false, but the description asserts 'idempotency=idempotent-write', a direct contradiction of the structured metadata. Although the text otherwise adds useful context (confirmation-gated, requires explicit-mutation-approval, produces operation-receipt), the conflict with annotations is disqualifying.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loading the mutation warning is good, but the text is templated and redundant, repeating 'writes live game/client state' twice and padding with Phase/cost metadata that adds little for the caller.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating tool with no output schema, the description covers safety, prerequisites, a verification path, and the produced operation-receipt, which is close to sufficient. The idempotency mismatch is a completeness defect rather than a missing detail.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters including defaults and the confirm gate. The prose signature repeats types and optionality but adds no syntax, format, or constraint detail beyond what is already structured.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource (alter runtime stack/debug visibility for a function) and immediately distinguishes it from sibling reads such as is-function-hooked and list-hooks by declaring it a live-state write. An agent can tell what this does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear context: use it to hide stack traces, it is confirmation-gated, and it 'should be restored with hidden=false after use', plus a verification sibling (is-function-hooked). It stops short of naming when a different tool is preferable beyond the verify step.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

smart-taskPlan or execute an adaptive verified taskA
Destructive

DETERMINISTIC ADAPTIVE ORCHESTRATOR. Turn a goal into a schema-aware plan, preview an explicit typed workflow, or execute it one tool at a time under hard step, tool-call, and wall-clock budgets. Inputs may safely reference prior outputs with exact $steps.. references. Mutations require allowMutations=true and an identical mutating action is never retried. Semantic postconditions are delegated to the read-only assert-state tool and count as verified only when it returns explicit boolean truth. Handled failures may be diagnosed by the read-only explain-failure tool and can activate only named recoverWith branches supplied by the caller. The result includes an evidence timeline, consumed budgets, confidence, unresolved assertions, recovery advice, and a continuation plan. This tool contains no LLM: omitted steps produce deterministic rankTools/matchWorkflows suggestions and leave required arguments blank instead of guessing them. Signature: { goal: string, mode: any?, steps: {{ type: any?, id: string, phase: "observe" | "act" | "verify"?, tool: string, input: any?, assertions: any?, recoverWith: any?, onFailure: any? }}?, fallbacks: any?, successAssertions: any?, finalRecoverWith: any?, allowMutations: any?, budgets: any? }. Phase: orchestrate; cost=medium; idempotency=contextual-write. Requires: explicit-mutation-approval. Produces: grounded-evidence, schema-aware plan, evidence timeline, assertion truth, completion confidence, continuation plan. Verify with: assert-state. Safety: MUTATING; writes live game/client state, may invoke multiple registered tools, mutating nested tools require allowMutations=true. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
goalYesConcrete outcome the workflow must prove.
modeNoPlan ranks tools, preview validates explicit steps, and execute invokes tools.plan
stepsNoExplicit ordered workflow. Omit it for deterministic schema-aware planning.
budgetsNoOptional validated input for budgets.
fallbacksNoNamed recovery branches; only a step's recoverWith list can activate one.
allowMutationsNoUser approval gate for every tool whose contract mutates state.
finalRecoverWithNoExplicit branches available if goal-level assertions fail.
successAssertionsNoGoal-level assertions evaluated after the main workflow.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations (destructiveHint/readOnlyHint/idempotentHint) by disclosing that mutations require allowMutations=true, that an identical mutating action is never retried, that verification is delegated to read-only assert-state and only counts on explicit boolean truth, that recovery runs only through named recoverWith branches, and that no LLM guesses omitted args. This is exactly the mutation/retry/permission context the annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded and the flow is logical, but the definition is very dense and includes a full Signature block that restates the input schema, plus metadata lines (Phase:, cost=, Produces:, Safety:) that partly echo structured fields. Every sentence is informative, yet the size exceeds what is needed and hurts scannability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a highly complex nested-input tool with no output schema, it covers modes, budgets, recovery, mutation gating, and even summarizes the return payload (evidence timeline, consumed budgets, confidence, continuation plan). Near complete, though the absence of an output schema means return-shape detail is only sketched rather than specified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already carries field-level semantics, but the description adds real value beyond it: $steps.<id>.<path> reference syntax, the no-guessing omission behavior, and the caller-supplied recovery branch model. The embedded signature largely duplicates the schema, limiting it below 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb-set: turn a goal into a plan, preview an explicit typed workflow, or execute it one tool at a time. It names the resource (schema-aware workflow) and its three modes, letting an agent distinguish it from raw executors like execute or batch-execute.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the mode semantics ('Plan ranks tools, preview validates explicit steps, execute invokes tools') and cross-references assert-state for verification and explain-failure for diagnosis. However it never explicitly says when to prefer this over siblings such as run-loop, execute-and-wait, or playbook-run, so routing guidance is implied rather than excluded.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

spoof-function-returnForce a function to always return a chosen value (MUTATES STATE via hookfunction)A
Destructive

WRITES LIVE GAME STATE — INSTALLS A PERSISTENT GLOBAL HOOK. Replaces a target function with a stub that IGNORES its arguments and ALWAYS returns a value you choose, without ever calling the original. This is the canonical anticheat/validation bypass: make a check like isValid() return true, force a server-config getter to return your value, or stub a paywall test to return false. Distinct from block-function (which makes the target a no-op returning nothing) because here you control the exact return value. WORKFLOW (stateful — survives across tool calls via getgenv().__mcp_spoofReturns, keyed by functionPath): 1. action='start' with functionPath + returnValue — resolves the target, captures the original, installs a stub that returns your value. Returns { started, key, returns }. 2. action='stop' with the same functionPath — restores the original function. Returns { stopped }. CAVEATS: the hook is GLOBAL and PERSISTS until you stop it (or the client restarts). The original is NEVER called while spoofed, so any side effects the real function had will not happen — this can desync state or destabilize the game, and a live function hook CAN TRIP ANTICHEAT. Always stop when done. Requires hookfunction, newcclosure, and getgenv; restoration uses hookfunction(target, original) with a restorefunction fallback. Returns { error } if a capability is missing, the target cannot be resolved, or there is already an active spoof for fetch/stop. Signature: { action: "start" | "stop", functionPath: string?, returnValue: { kind: "string" | "number" | "boolean" | "nil" | "raw", value: string | number | boolean? }?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Produces: structured-result. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYes'start' installs the return-spoofing stub on functionPath (requires returnValue); 'stop' restores the original function. Use the SAME functionPath for both so they address the same registry entry.
returnValueNoThe value the spoofed function should always return, expressed as a typed argument.
functionPathNoLuau expression resolving to the function to spoof, e.g. 'getsenv(game.Players.LocalPlayer.PlayerScripts.AntiCheat).isValid' or 'getrawmetatable(game).__index'. Evaluated as `return <functionPath>` and must resolve to a function. REQUIRED for 'start'. For 'stop' it is the registry key identifying which spoof to restore, so it must match the string used at start.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and non-idempotent, but the description adds substantial context beyond them: the hook is GLOBAL and persists until stopped, the original is never called so side effects are lost, state can desync, and live hooks can trip anticheat. It also documents the registry mechanism (getgenv().__mcp_spoofReturns keyed by functionPath) and error return shapes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core mutating purpose and caveats before the workflow and metadata. It is dense and the signature/metadata tail partially duplicates the schema, but the ordering is logical and most sentences carry real information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description explicitly documents the return shapes ({ started, key, returns }, { stopped }, { error }) and failure conditions. Combined with the safety, workflow, and capability requirements, an agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents action, returnValue/kind/value, functionPath, and threadContext in detail. The description reinforces the functionPath-as-registry-key semantics and repeats the signature, but adds little that the schema does not already state, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (installs a persistent return-spoofing hook on a target function) and explicitly contrasts with the sibling block-function, explaining the exact behavioral difference (controlled return value vs no-op). An agent can distinguish this from related hook tools without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit step-by-step workflow for action='start' vs 'stop', names the sibling alternative it is not, and instructs 'Always stop when done.' Prerequisites (active-client, resolved-target, explicit-mutation-approval) and the correct usage sequence are spelled out rather than inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

state-transactionCapture, rollback, and clean up bounded live state transactionsA
Destructive

STATE TRANSACTION JOURNAL. Begin a bounded getgenv-backed transaction, capture explicitly requested Instance properties/attributes and camera fields, register known cleanup resources, inspect status, commit without restoration, or rollback in reverse journal order with pcall-isolated per-item results. Cleanup can rollback or explicitly discard expired and cross-place/job orphaned transactions. Registered resources support MCP Drawing ids, held virtual input releases, connection state restoration or cleanup-only disconnection, and canonical __mcp_hooks/__mcp_hook_meta entries when their metadata is safely understood. This is a best-effort client-state journal, not a general undo system: destroyed Instances, fired remotes, server-side changes, arbitrary script side effects, and connections destroyed before capture cannot be reconstructed. Commit intentionally discards all snapshots and does not clean registered resources. Signature: { action: "begin" | "capture" | "status" | "commit" | "rollback" | "cleanup", transactionId: string?, name: string?, targets: any?, captureCamera: any?, cleanupItems: any?, maxItems: number?, expirySeconds: number?, limit: any?, cleanupMode: any?, includeOrphans: any?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval, begin a transaction before mutating reversible client state, capture every property/attribute before changing it. Capabilities: getgenv. Produces: grounded-evidence, bounded state journal, per-item rollback evidence, expired/orphaned cleanup report. Verify with: assert-state. Safety: MUTATING; writes live game/client state, capture stores raw client references in getgenv until commit, rollback, cleanup, expiry, or executor shutdown, rollback writes captured values and runs registered cleanup actions in reverse order, connection disconnect cleanup is irreversible and is reported separately from reversible restoration. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoHuman label; later actions may use it only when exactly one active transaction matches.
limitNoMaximum transactions listed or processed by status/cleanup.
actionYesTransaction lifecycle action.
targetsNoExplicit property/attribute snapshot targets for begin or capture.
maxItemsNoPer-transaction journal cap (default 128, hard maximum 256); for capture it may only lower the call cap.
cleanupModeNoFor cleanup, rollback attempts every journal item first; discard explicitly removes journals without restoration.rollback
cleanupItemsNoResources to release/restore during reverse-order rollback.
captureCameraNoSnapshot CurrentCamera CameraType, CameraSubject, CFrame, Focus, and FieldOfView.
expirySecondsNoSeconds until cleanup considers the transaction expired (default 900 on begin). Capture may refresh it.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
transactionIdNoStable transaction id. Optional for begin (one is generated); required by id for unambiguous later actions.
includeOrphansNoFor cleanup, include transactions created in another PlaceId/JobId, not only expired ones.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Far exceeds the destructive=true annotation: it discloses that capture stores raw client references in getgenv until commit/rollback/cleanup/expiry/shutdown, that rollback writes in reverse order, that connection disconnect cleanup is irreversible and reported separately, and that commit intentionally discards snapshots without cleaning registered resources.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded and organized into labeled sections, but bulky; the inline { action: ... } signature duplicates information already fully covered by the schema, and the Safety sentence restates points made earlier about rollback ordering and disconnect irreversibility.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a 12-param mutating tool with no output schema, the description covers lifecycle semantics, best-effort limitations, failure handling ('inspect tool-schema'), and verification path. An agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all 12 params (baseline 3). The description adds real meaning beyond field text by tying params to actions, e.g. 'capture it may only lower the call cap' for maxItems and the per-action semantics of cleanupMode/rollbackAction, though the inline signature largely echoes the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource framing ('STATE TRANSACTION JOURNAL') and enumerates the six lifecycle actions (begin/capture/status/commit/rollback/cleanup) with what each does. This is unmistakably distinct from siblings like assert-state or world-delta, which it explicitly cross-references.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States when to use it ('begin a transaction before mutating reversible client state, capture every property/attribute before changing it'), when not to ('not a general undo system') with a concrete exclusion list, and routes verification to 'assert-state'. Explicit alternatives and prerequisites are all present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

suggest-toolsSuggest tools for a keyword, ranked by past successA
Read-only

Surface the right tool faster from a natural-language goal. Uses intent aliases and field-aware ranking, then adds a small past-success bias so tools that have worked in this server session float above equally relevant untouched tools. Returns exact signatures, required inputs, definition quality, phase/cost, capabilities, outputs, mutation/client flags, match reasons, and usage stats. Pass { keyword } (required), { limit } (default 10), and optional { category }.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum tools to return, default 10.
keywordYesNatural-language goal or keyword, e.g. 'find the player's money' or 'click a UI button'.
categoryNoOptional: restrict to one category.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true; the description adds real behavioral context beyond that, disclosing the intent-alias/field-aware ranking and the session-local past-success bias that reorders equally relevant results. It also enumerates the returned fields, which is valuable given no output schema exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the purpose before the ranking mechanics and the return-field list. The long enumeration of return fields is dense but earns its place given the absence of an output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only discovery tool with no output schema, the description covers purpose, ranking behavior, and the shape of the response, which is what an agent needs to call it correctly. Only the relationship to the closest sibling (list-tools) is left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents keyword, limit (default 10), and category. The description restates the same three parameters with defaults and required/optional status, adding no syntax or semantic detail beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('surface the right tool') scoped to a natural-language goal, and the title adds the keyword/ranking angle. It is clearly distinct from a plain listing tool like list-tools, though it never names that sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Surface the right tool faster from a natural-language goal' implies the use case (goal-based discovery rather than enumeration), but there is no explicit when-not guidance and no alternative sibling is named for comparison against list-tools or tool-schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

summarize-hidden-surfacesSummarize hidden surfaces (what is hiding)A
Read-onlyIdempotent

One-call high-level overview answering 'what is hiding in this game?'. Safely (each executor call guarded + pcall'd) counts: Actors (parallel-Luau VMs), nil-parented instances (with a small top-classes breakdown), currently-running scripts, loaded modules, and a quick count of scripts that are NOT sitting normally in the tree (hiddenScripts). Any executor function that is unavailable in this executor is listed under unavailable instead of aborting the whole summary. Use this first to decide which deeper hidden-* tool to run next. Signature: { threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: structured-result. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, and the description adds substantial context beyond them: each executor call is guarded and pcall'd, unavailable functions are surfaced in `unavailable` instead of aborting the whole summary, cost=medium, and it requires an active-client. Partial-failure behavior and prerequisite are exactly the kind of traits structured fields cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose and enumerated counts, but the trailing block is padded and partly redundant: 'idempotency=read-only' and 'Safety: read-only' duplicate each other and the annotations, and 'On failure: inspect tool-schema...' is boilerplate. Those sentences do not fully earn their place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does the work of describing the return shape (counts per category, top-classes breakdown, `unavailable` list) plus prerequisites and failure handling. It is complete enough for a read-only aggregate, though it could note output field naming conventions more precisely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one optional param (threadContext) with 100% schema coverage; the schema already explains it fully. The description merely repeats the signature '{ threadContext: number? }' without adding new semantics, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('One-call high-level overview answering what is hiding in this game?') and then enumerates exactly what it counts: Actors, nil-parented instances, running scripts, loaded modules, and hiddenScripts. It clearly positions itself as the aggregate entry point that feeds deeper 'hidden-*' siblings, so an agent can distinguish its role from individual finders.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use this first to decide which deeper hidden-* tool to run next' gives explicit ordering guidance and names the sibling family it routes to. It lacks named exclusions (e.g., when NOT to bother with the summary vs. a direct finder), but the entry-point framing is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

summarize-runtime-surfacesSummarize runtime surfacesA
Read-onlyIdempotent

Quick high-level reverse summary: counts for scripts/modules/remotes/actors/functions/tables/threads. Signature: { threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: operation-receipt. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, non-destructive. The description adds genuinely new context: the 'active-client' prerequisite, cost=medium, phase=observe, and 'Produces: operation-receipt'. It also gives a failure-recovery pointer. It stops short of describing the receipt's contents or pagination, but it meaningfully exceeds the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose and enumerates surfaces before pivoting to compact key-value metadata. Slightly telegraphic and crowded, but every line carries information and nothing is redundant filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers purpose, signature, phase, cost, prerequisite, output artifact, safety, and failure guidance for a read-only count tool with full schema coverage and annotations. Only the return shape (the 'operation-receipt' contents) is left implicit, which is a minor gap given no output schema exists.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single optional threadContext parameter is fully documented in the schema. The description only restates the signature ('{ threadContext: number? }'), adding no syntax or default behavior beyond what the schema already says. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (summarize) and resource (runtime surfaces), then enumerates exactly what it counts: scripts/modules/remotes/actors/functions/tables/threads. This distinguishes it from generic list tools, though it never explicitly contrasts with the close sibling summarize-hidden-surfaces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides operational context ('Phase: observe', 'Requires: active-client') that implies when the tool applies, but never says when to prefer it over alternatives like summarize-hidden-surfaces, get-instance-counts, or the many list-* siblings. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

teach-modeRecord a user demonstration and draft a reusable playbookA
Destructive

Event-driven demonstration recorder with start, poll, stop, and cancel actions. It uses bounded client-side ring buffers and temporary Roblox signal connections to observe keyboard, mouse, touch, throttled movement, GuiButton activation, meaningful GUI appearance/disappearance, ProximityPrompt triggers, character respawns, and Tool equip/backpack transitions. Optional remote capture is started and read through trace-remote-traffic via the normal nested-tool invoker when that tool and executor capabilities are available. stop disconnects every owned listener, returns the retained chronological timeline, and builds a conservative review-required playbook draft with semantic selectors, path evidence, virtual-input/click-button/fire-proximity-prompt and remote candidates, inferred waits/guards, placeholders, uncertainty, and manual-review flags. It does not claim perfect intent inference. cancel disconnects and discards. Idle sessions self-expire and a three-session cap evicts the oldest recorder to prevent leaked listeners. Signature: { action: "start" | "poll" | "stop" | "cancel", sessionId: string?, maxEvents: any?, movementThrottleMs: any?, expirySeconds: any?, maxGuiWatch: any?, sinceSeq: any?, limit: any?, includeRemoteSpy: any?, remoteLimit: any?, sinceRemoteTime: any?, threadContext: number? }. Phase: orchestrate; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval, An active Roblox client, The user is ready to demonstrate the workflow after action=start. Capabilities: UserInputService and Roblox RBXScriptSignal connections, getgenv for state across calls, Optional hookmetamethod/getnamecallmethod/newcclosure for remote capture. Produces: grounded-evidence, Bounded chronological event timeline, Semantic instance selectors and path evidence, Conservative reusable playbook draft with uncertainty and review flags. Verify with: assert-state, teach-mode action=poll to confirm events are arriving, Manual review of selectors, guards, timing, and server-reaching candidates, Run the reviewed playbook in a disposable/test game state and verify outcomes. Safety: MUTATING; writes live game/client state, Temporarily installs bounded signal listeners in the active client, May temporarily install a remote-spy metamethod hook when explicitly requested, Does not execute the generated playbook automatically. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNopoll only: maximum local timeline events returned in this page.
actionYesstart creates the recorder; poll returns retained events after sinceSeq; stop disconnects and returns the timeline plus a conservative playbook draft; cancel disconnects and discards the recording.
sinceSeqNopoll only: return retained local events whose sequence is greater than this cursor.
maxEventsNostart only: circular timeline capacity; oldest events are overwritten and counted.
sessionIdNoOptional stable recording id. start generates one when omitted; later actions default to the newest active session.
maxGuiWatchNostart only: cap on watched GuiObjects/ScreenGuis to keep listener overhead bounded.
remoteLimitNoMaximum remote-spy entries merged into poll/stop results.
expirySecondsNostart only: idle time before automatic disconnection and session removal.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
sinceRemoteTimeNopoll only: omit remote events at or before this many seconds since recording start.
includeRemoteSpyNostart only: request outgoing remote capture through trace-remote-traffic. Failure is non-fatal and reported.
movementThrottleMsNostart only: minimum interval for mouse/touch/gamepad movement observations.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations only marking it non-read-only/destructive/non-idempotent, the description carries the behavioral load and delivers: bounded ring buffers, temporary signal listeners, a three-session cap that evicts the oldest recorder, idle self-expiry, cancel discarding data, optional metamethod hook installation, and the explicit note that the generated playbook is not auto-executed. This is far richer than the annotations alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose is front-loaded, but the passage is dense and long, and the full signature enumeration duplicates the input schema, adding length without new information. Much of the remaining content earns its place, but the overall size is excessive for a description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex, 12-parameter mutation tool with no output schema, the description covers behavior, safety, failure handling ('inspect tool-schema for exact fields...'), and verification steps. An agent has enough to call it correctly and interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 12 parameters, and the description's signature block largely duplicates it. It adds some meaning (linking includeRemoteSpy to trace-remote-traffic, the self-expiry/cap context) but no syntax or format detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource ('Event-driven demonstration recorder with start, poll, stop, and cancel actions') and details exactly what it observes and produces (a timeline plus a review-required playbook draft). An agent can distinguish it from siblings like playbook-save/playbook-run/session-replay without opening a schema, since it records live demonstrations rather than persisting or replaying them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Each action's role is spelled out (start creates, poll reads after sinceSeq, stop disconnects and returns, cancel discards) and prerequisites are named ('active-client, explicit-mutation-approval, the user is ready to demonstrate the workflow after action=start'). It also routes optional remote capture through trace-remote-traffic. However, it never states when NOT to use this tool versus alternatives like session-replay or manual scripting.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

test-capabilitiesExecutor capability matrix (closure, Actor/state, and core primitives)A
Read-onlyIdempotent

In-game capability matrix: probes a curated executor/runtime function list and reports which are present (callable) versus missing. For each name it checks the global environment (getgenv() first, then the thread's environment / _G) and, for dotted names like 'debug.getinfo', walks the parent table — classifying the leaf as available only when it is type=='function'. The probe NEVER calls the functions, so it is completely safe and side-effect-free. Covers reflection/closures (getgc, constants/upvalues/protos, clone/wrapper/type/hook/hash functions), script access (getscriptbytecode/closure/caller/hash, getscripts, getrunningscripts, getloadedmodules, getsenv/getfenv), instance/actor discovery (getactors, Actor Lua-state execution, communication channels, getnilinstances, getinstances), environments (getgenv, getrenv, getreg), hooking (hookfunction, hookmetamethod, restorefunction, newcclosure), metatables (getrawmetatable, setrawmetatable, setreadonly, isreadonly), namecall/signals (getnamecallmethod, getconnections, firesignal, replicatesignal, getcallbackvalue, getsignalarguments), threads (setthreadidentity, getthreadidentity), and IO/misc (loadstring, fireproximityprompt, fireclickdetector, firetouchinterest, virtual-input globals, getcustomasset, request/http_request). Use this to decide up-front whether a workflow is supported, or to compare two executors. Returns { total, availableCount, missingCount, available[], missing[] } (both lists sorted) or { error }. Signature: { threadContext: number? }. Phase: verify; cost=medium; idempotency=read-only. Requires: active-client. Produces: diagnostic-report. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent and non-destructive, so the bar is met, yet the description adds substantive mechanics: the lookup order (getgenv() first, then thread env/_G), the dotted-name parent-table walk, and the classification rule (leaf counted available only when type=='function'). It also states the strong safety guarantee that functions are never invoked, and covers failure handling via tool-schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and the safety guarantee are front-loaded, and the return shape is stated plainly. The long middle enumeration of ~60 probed functions is verbose, but it does define the probe's coverage boundary, which is genuinely useful for deciding whether a workflow is supported.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description supplies the return contract ({ total, availableCount, missingCount, available[], missing[] }, both sorted, or { error }), plus phase, cost, idempotency, required context, and a failure fallback. Combined with the annotations, an agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single optional parameter and schema description coverage is 100%, so the schema already carries its meaning. The description only restates the signature ({ threadContext: number? }) without adding format, range, or effect details beyond the schema, which is the expected baseline-3 case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'probes a curated executor/runtime function list and reports which are present (callable) versus missing.' The enumerated coverage areas (reflection/closures, script access, actor discovery, hooking, etc.) make the scope concrete, but it never names the near-overlapping siblings (closure-capabilities, actor-capabilities, get-executor-info), leaving the agent to infer which capability tool to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear triggering context: 'Use this to decide up-front whether a workflow is supported, or to compare two executors.' That is a real when-to-use statement, but there is no when-not guidance and no explicit routing to alternatives when the agent only needs closure or actor coverage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tool-planTurn a natural-language goal into a tool workflowA
Read-onlyIdempotent

READ-ONLY. Convert a vague Roblox goal into a ranked, schema-aware workflow. Give this tool the user's natural-language objective, not a guessed tool name. It combines intent aliases, the live tool catalog, available schemas, mutation flags, and curated discover→act→verify recipes. Use the returned workflow as a starting point, then inspect the exact schema of the selected tool before calling it. This is especially useful for goals involving UI/input, remotes, player values, instance inspection, or reverse engineering. Signature: { goal: string, limit: any?, includeMutating: any?, capabilityAware: any? }. Phase: observe; cost=low; idempotency=read-only. Requires: none. Produces: agent-guidance. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
goalYesNatural-language objective, e.g. 'find the player's cash and verify its current value'.
limitNoMaximum ranked alternatives to return.
capabilityAwareNoUse the active client's advertised executor capabilities to annotate and rank candidates.
includeMutatingNoInclude tools that write game state in the ranked alternatives and workflow.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only/idempotent/non-destructive, so the repeated 'READ-ONLY'/'Safety: read-only' lines add little. However, the description adds real behavioral context: what signals it fuses (intent aliases, live catalog, schemas, mutation flags, recipes), that the output is a starting point rather than a final answer, and the failure path.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded, but the trailing templated metadata (Signature, Phase, cost, idempotency, Requires, Produces, Safety, On failure) is boilerplate that repeats the annotations and the READ-ONLY claim multiple times. Some sentences earn their place, several do not.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description carries the return-value burden, yet it only says 'Produces: agent-guidance' and 'ranked, schema-aware workflow' without describing the workflow's shape or fields. For a planning tool this leaves a meaningful gap, though the failure guidance partly compensates.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all four parameters. The Signature line merely restates them with looser types ('any?') and adds no syntax or default detail beyond the schema. The only genuine value-add is the instruction to pass a natural-language goal, not a guessed tool name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('convert a vague Roblox goal into a ranked, schema-aware workflow') and differentiates from siblings by contrasting its role with tool-schema (which is for exact field inspection). An agent can immediately tell it apart from list-tools/suggest-tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit context ('especially useful for goals involving UI/input, remotes, player values, instance inspection, or reverse engineering') and names the follow-up alternative (inspect the exact schema of the selected tool; on failure use tool-schema). No explicit when-NOT-to-use, but strong routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tool-quality-auditAudit definition quality across the complete tool catalogA
Read-onlyIdempotent

Read-only catalog self-audit for tool titles, descriptions, input documentation, schema examples/defaults/constraints, AI data-flow contracts, prerequisites, capability requirements, mutation side effects, verification paths, and recovery guidance. It evaluates the same centrally compiled metadata exposed to MCP clients, so every current and future tool can be checked against one measurable standard. Filter by exact name/category or return only entries below a score threshold; no Roblox client is required. Signature: { name: string?, category: string?, minimumScore: any?, includePassing: any?, limit: any? }. Phase: verify; cost=medium; idempotency=read-only. Requires: none. Produces: structured-result. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoOptional exact kebab-case tool name to audit.
limitNoMaximum detailed tool reports returned; aggregate counts always cover the full filter.
categoryNoOptional exact category name to audit.
minimumScoreNoScore below which a definition is reported as needing attention.
includePassingNoInclude definitions meeting or exceeding minimumScore.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description still adds value beyond them: cost=medium, 'Produces: structured-result', that it audits the centrally compiled metadata exposed to MCP clients, and a failure path pointing to tool-schema for exact fields. This is useful behavioral context on top of the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening sentence is dense but front-loaded and information-rich. However, the trailing Signature block restates the input schema verbatim and 'Safety: read-only' duplicates the readOnlyHint annotation, so several clauses do not earn their place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries return-format burden and does so ('Produces: structured-result'), and it supplies a failure-recovery path and cost/phase metadata. For a zero-required-param, client-independent audit tool this is close to complete; only the exact shape of the structured result is left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all five parameters, including defaults and bounds. The description's 'filter by exact name/category or return only entries below a score threshold' adds modest semantic context about the minimumScore/includePassing interplay, but the verbatim Signature line contributes nothing beyond the schema, keeping this at the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (read-only catalog self-audit of tool definition quality) and enumerates exactly what is evaluated: titles, descriptions, input docs, schema examples, data-flow contracts, prerequisites, side effects, verification and recovery. This is clearly distinct from siblings like tool-schema, list-tools, and suggest-tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete usage context: filter by exact name/category, or surface only entries below a score threshold, and notes no Roblox client is required, which helps the agent know it can run this standalone. It does not explicitly name an alternative (e.g., tool-schema) or state when NOT to use it, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tool-schemaGet a tool's input schema and Luau signatureA
Read-onlyIdempotent

Return the input schema for any tool on this server in a compact, Luau-friendly form. Pass { name } for one tool: returns its title, compiled description, quality grade, safety/execution/success/recovery guidance, category, mutatesState, requiresClient, a one-line Luau signature (e.g. { limit: number?, includeBots: boolean? }), a per-field list with type + description + optional flag, and a runnable mcp.<camelCase>({...}) example. Pass { search } to match a keyword across names/titles and get back a compact list of name + signature. With no args, returns every tool's name + signature (one line each) — bulky but the fastest way for a script to discover the whole surface. Use this from inside script to learn the right args BEFORE calling a tool, instead of guessing and failing. Signature: { name: string?, search: string?, category: string? }. Phase: observe; cost=low; idempotency=read-only. Requires: none. Produces: agent-guidance. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoExact kebab-case tool name (e.g. 'get-players'). If unknown, returns near-miss suggestions.
searchNoKeyword matched across each tool's name, title, and description.
categoryNoRestrict to one category (case-sensitive exact match).

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnly/idempotent/non-destructive, so safety is redundant but consistent. The description adds real value beyond annotations: return shape, phase=observe, cost=low, produces agent-guidance, and on-failure guidance to inspect for exact fields/defaults. No contradiction, and several behavioral traits disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but well-organized and front-loaded: the core purpose is in sentence one, modes follow, then a signature/metadata footer. The footer ('Phase: observe; cost=low; idempotency=read-only. Requires: none...') partially restates the annotations, adding minor redundancy, but every sentence largely earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description carries the full burden of describing returns, and it does so thoroughly: title, compiled description, quality grade, guidance fields, category, mutatesState, Luau signature, per-field list, and a runnable example. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is already 100%, establishing a baseline of 3. The description adds meaning beyond the schema by explaining the mode interaction: {name} returns near-miss suggestions when unknown, {search} matches across name/title, and no-args returns the whole surface. It clarifies how parameter presence changes output, which the schema alone does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (return) and resource (input schema for any tool) with a clear modifier (compact, Luau-friendly form). Distinguishes itself from siblings like list-tools and suggest-tools by describing its exact output shape (title, signature, per-field list, runnable example). An agent can tell it apart without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use it: 'Use this from inside script to learn the right args BEFORE calling a tool, instead of guessing and failing.' Also enumerates the three invocation modes (name, search, no-args) with what each returns, giving the agent clear selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

trace-call-durationsProfile how long a function takes per call (MUTATES STATE via hookfunction)A
Destructive

WRITES LIVE GAME STATE — INSTALLS A PERSISTENT GLOBAL HOOK. Per-function profiler: hook a target so that every invocation is timed with os.clock(), accumulating call count plus total/min/max time, then read the aggregated stats, then restore the original. This is the fastest way to answer 'how expensive is this function and how often does it run?' — ideal for finding the hot path in an anticheat loop, a render step, or a remote handler. Unlike count-function-calls (count only) it also measures duration; unlike hook-and-log-function it stores only aggregates (no per-call args), so it is cheap enough for hot paths. WORKFLOW (stateful — survives across tool calls via getgenv().__mcp_callTimings, keyed by functionPath): 1. action='start' with functionPath — resolves the target, captures the original, installs a timing wrapper that transparently calls the original and records elapsed time. Returns { started, key }. 2. action='fetch' with the same functionPath — returns { count, totalMs, avgMs, minMs, maxMs } so far WITHOUT stopping. Poll to watch live. 3. action='stop' with the same functionPath — restores the original and clears the entry. Returns final stats. CAVEATS: the hook is GLOBAL and PERSISTS until you stop it (or the client restarts), adds (small) timing overhead on every call, and a live function hook CAN TRIP ANTICHEAT — always stop when done. The original is called through real-time (its return values are passed back unchanged); timing/aggregation is pcall-isolated. Requires hookfunction, newcclosure, and getgenv; restoration uses hookfunction(target, original) with a restorefunction fallback. Returns { error } if a capability is missing, the target cannot be resolved, or there is no active profile for fetch/stop. Signature: { action: "start" | "fetch" | "stop", functionPath: string?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Produces: bounded-event-snapshot, operation-receipt. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYes'start' installs the timing wrapper on functionPath; 'fetch' returns the aggregated timing stats so far (hook stays live); 'stop' restores the original function and clears the stats. Use the SAME functionPath for all three so they address the same registry entry.
functionPathNoLuau expression resolving to the function to profile, e.g. 'getsenv(game.Players.LocalPlayer.PlayerScripts.Main).update' or 'getrawmetatable(game).__namecall'. Evaluated as `return <functionPath>` and must resolve to a function. REQUIRED for 'start'. For 'fetch'/'stop' it is the registry key identifying which running profile to act on, so it must match the string used at start.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the destructiveHint/readOnlyHint/non-idempotent annotations: discloses the hook is GLOBAL and persists until stopped, adds timing overhead on every call, can trip anticheat, requires hookfunction/newcclosure/getgenv, uses pcall-isolation, and specifies fallback restoration. This is exactly the mutation-risk context an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Long but well-organized, front-loading the critical WRITES LIVE GAME STATE warning and structuring the WORKFLOW and CAVEATS clearly. The trailing Phase/cost/idempotency/Requires/Produces metadata block somewhat duplicates the annotations and could be trimmed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a stateful, mutating tool with no output schema, the description is complete: it explains the return shape at each step ({started,key}, {count,totalMs,avgMs,minMs,maxMs}, {error}) and the failure conditions. Nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents each parameter thoroughly; baseline is 3. The description adds the combined signature and reinforces that functionPath must match across calls, but largely restates what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: a per-function profiler that hooks a target and times every invocation with os.clock(), accumulating count plus total/min/max. It explicitly distinguishes itself from siblings, noting it measures duration unlike count-function-calls and stores only aggregates unlike hook-and-log-function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names concrete use cases (finding the hot path in an anticheat loop, render step, or remote handler) and provides an explicit stateful workflow with the three actions and the condition for each. It also routes against alternatives by stating their tradeoffs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

trace-connection-functionTrace connection functionA
Read-onlyIdempotent

Inspect a single connection on an Instance's RBXScriptSignal and return debug metadata about the connected Luau function. Uses getconnections(inst[signalName]) to enumerate connections and debug.info/debug.getinfo to resolve the function's name, source, line and parameter count. REQUIRES an executor exposing getconnections and debug.info (or debug.getinfo); if either capability is missing, or the signal/connection/function cannot be resolved, returns a clear { error } describing the missing capability. Returns { Signal, ConnectionIndex, ConnectionCount, Function: { Name, Source, ShortSource, LineDefined, NumParams, IsVararg, What, Pointer } }. Signature: { instancePath: string, signalName: string, connectionIndex: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Capabilities: getconnections. Produces: bounded-event-snapshot, created-handle. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
signalNameYesName of the RBXScriptSignal member on the instance to inspect, e.g. 'Touched', 'Changed', 'OnClientEvent'.
instancePathYesLua expression resolving to the target Instance, e.g. 'game.Players.LocalPlayer.Character.Humanoid'. Evaluated as `return <instancePath>`.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
connectionIndexNoZero-based index of the connection to trace within getconnections() results (default: 0 = first connection).

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only/idempotent, and the description goes well beyond that: it names the exact executor capabilities required (getconnections, debug.info), documents the graceful-degradation contract (returns a clear { error } describing the missing capability), and describes the produced artifacts (bounded-event-snapshot, created-handle) and cost.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose and mechanism before the returns/phase metadata, and every clause carries information. It is dense and slightly long due to the metadata tail, but nothing is truly redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description usefully enumerates the return shape ({ Signal, ConnectionIndex, ConnectionCount, Function: {...} }) and the failure shape, plus prerequisites and capabilities. An agent has everything needed to call it and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters are already documented with types, defaults, and examples. The description restates the signature but adds no semantics (format, resolution order, or threading implications) beyond what the schema provides. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource: 'Inspect a single connection on an Instance's RBXScriptSignal and return debug metadata about the connected Luau function.' This distinguishes it from siblings like get-connection-info, count-signal-connections, and scan-connections-by-source by scoping to a single indexed connection with debug resolution.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context ('Phase: observe; cost=medium') and prerequisites ('Requires: active-client, resolved-target', requires an executor exposing getconnections/debug.info). It does not, however, explicitly name the alternative siblings (e.g. get-connection-info) or state when-not-to-use, so routing between the many connection tools is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

trace-remote-trafficTrace remote trafficA
Destructive

WRITES LIVE GAME STATE on start. Starts/adopts the selected engine and opens a traffic view. fetch reads retained captures since start; stop closes only the view. Direction is selected on start. Shares the bounded buffer with all spy tools for this engine. Starting another engine stops the previous spy on this client. Signature: { engine: "cobalt" | "ketamine"?, action: "start" | "fetch" | "stop", direction: any?, limit: any?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval. Produces: bounded-event-snapshot. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoOptional hard result/work budget used to keep output and runtime bounded.
actionYesoperation selector; use one of the schema's allowed values.
engineNoRemote-spy backend. Only one engine runs per client; starting another stops the current engine. Defaults to cobalt.cobalt
directionNoCapture/control direction; incoming includes client callbacks.Both
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=true, and the description goes well beyond them: it warns that start WRITES LIVE GAME STATE, notes the buffer is shared and bounded, explains that stop only closes the view (captures persist), and flags that starting another engine kills the previous spy. It also states prerequisites (active-client, explicit-mutation-approval) and a verification path (assert-state).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The critical warning ('WRITES LIVE GAME STATE') is front-loaded and the lifecycle is described in tight clauses. The trailing metadata block (Phase/cost/idempotency/Requires/Produces/Verify/Safety) is somewhat boilerplate-heavy, but each line carries actionable content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a stateful, mutating, multi-action tool with no output schema, the description covers lifecycle, shared-state hazards, prerequisites, and verification. It does not describe the shape of a fetch result, but with a bounded-snapshot contract stated, that gap is minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds real semantic value: 'Direction is selected on start' tells the agent that direction is only meaningful on start, not fetch/stop, and the shared-buffer note explains what limit bounds. The engine conflict rule is restated from the schema but reinforced.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('trace remote traffic'), and unpacks the three actions (start/fetch/stop) so their distinct purposes are visible. The 'shares the bounded buffer with all spy tools' clause distinguishes it from sibling readers like get-remote-spy-logs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear per-action guidance: start opens the view and selects direction, fetch reads retained captures since start, stop closes only the view. It implies the spy-tool family relationship but never explicitly names an alternative tool to use instead when a different need arises, so it stops short of full when/when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

trace-require-callersTrace require callersA
Read-onlyIdempotent

Find functions that reference 'require' in constants/source via getgc scan. Signature: { limit: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-event-snapshot, operation-receipt. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoOptional hard result/work budget used to keep output and runtime bounded.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only/idempotent/non-destructive/openWorld, so the description earns credit for adding context beyond them: cost=medium, a required precondition (active-client), and declared outputs (bounded-event-snapshot, operation-receipt). These are useful traits the structured fields do not carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded, but the metadata block is template-y and partly redundant with annotations ('Safety: read-only' and 'idempotency=read-only' restate readOnlyHint/idempotentHint). The failure-handling sentence is boilerplate that adds little for an agent that already has tool-schema available.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-param, all-optional, read-only scan with no output schema, the description covers phase, precondition, cost, and output kinds—enough for an agent to invoke it confidently. Only the lack of sibling differentiation and parameter nuance keeps it from being fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented, making 3 the baseline. The description restates the signature ('{ limit: any?, threadContext: number? }') without adding syntax, format, or default detail beyond the schema—in fact its 'any?' for limit is less precise than the schema's number/default 80.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Find functions that reference require' plus the mechanism ('via getgc scan') and the location ('in constants/source'). This clearly distinguishes it from generic GC scans, though it does not name a specific sibling to route away from.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides contextual signals ('Phase: observe', 'Requires: active-client') that imply when it is callable, and a fallback pointer ('On failure: inspect tool-schema'). But it never states when to prefer this over alternatives like find-string-xrefs or find-constants-xref, so routing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

type-text-boxType into a TextBoxA
Destructive

WRITES LIVE GAME STATE. Enter text into a Roblox TextBox by path. Resolves the path to a TextBox, captures its focus, then either simulates real keystrokes (useKeyPress=true, via VirtualInputManager:SendTextInput with a keypress/keyrelease fallback) so text-changed and FocusLost handlers run, or directly sets the .Text property (useKeyPress=false). Optionally presses Enter afterwards and releases focus. Use the keystroke path when scripts react to player typing; use the direct path for a fast value poke. Returns { Path, ok } or { error }. Signature: { path: string, text: string, enter: any?, useKeyPress: any?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Capabilities: VirtualInputManager. Produces: structured-result. Verify with: get-gui-text. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesThe instance path to the TextBox
textYesThe string to type into the TextBox
enterNoWhether to press Enter after typing
useKeyPressNoIf true, simulates real keystrokes using VirtualInputManager / keypress. If false, directly sets the Text property.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=false, but the description goes well beyond them: it flags 'WRITES LIVE GAME STATE', discloses the VirtualInputManager mechanism and keypress/keyrelease fallback, explains the side effect that text-changed and FocusLost handlers fire, notes focus capture/release, and states the exact failure/verification path. This is unusually rich behavioral context for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the critical 'WRITES LIVE GAME STATE' warning and the core behavior, then organized into mode guidance, signature, phase/requires/capabilities, and safety. Some metadata (the full signature, 'Safety: MUTATING') restates schema or annotations, so it is slightly longer than strictly necessary, but every block is scannable and earns most of its space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description states the return shape ({ Path, ok } or { error }), names the verification tool (get-gui-text), lists required preconditions, and covers both operating modes. For a 5-parameter mutation tool, nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema: it explains *why* useKeyPress matters (handler firing vs. fast poke), mentions the fallback mechanism, and summarizes the full signature including optional enter/threadContext. Only threadContext gets no extra treatment beyond the schema's default note.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('Enter text into a Roblox TextBox by path') and then distinguishes its two operating modes, which separates it from siblings like set-gui-text (direct property write) and press-key/virtual-input (raw input). An agent can identify the tool's job without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit mode-selection guidance: 'Use the keystroke path when scripts react to player typing; use the direct path for a fast value poke.' It also lists prerequisites (active-client, resolved-target, explicit-mutation-approval) and a verification tool. It does not, however, name sibling alternatives such as set-gui-text or press-key for cases where this tool is the wrong choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify-path-existsVerify an instance path resolvesA
Read-onlyIdempotent

Cheaply check whether a dotted instance path currently resolves to a real Instance, without reading any properties. Use this as a pre-flight check before click-button / type-text-box / get-instance-properties so you fail fast with a clear message instead of acting on a stale or wrong path (UI gets created/destroyed dynamically). Returns { exists, className?, fullName? } — when exists is false, className/fullName are omitted. Signature: { path: string, threadContext: number? }. Phase: verify; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Produces: structured-result. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesDotted path to verify, starting at 'game' (e.g. 'game.Players.LocalPlayer.PlayerGui.Shop.BuyButton'). Resolution walks each '.'-separated segment via direct indexing then FindFirstChild.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already carry readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds real context beyond that: it declares it reads no properties, reports cost=medium, states prerequisites (active-client, resolved-target), and describes the return shape including the omitted-field behavior. Return format is well disclosed, so 4 is warranted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then usage, then return shape, then a compact metadata tail. Every sentence carries information, though the trailing structured block (phase/cost/idempotency/requires/produces/safety) is somewhat dense and partly overlaps annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description fully specifies the return object and its optional-field behavior, plus prerequisites, cost, and failure guidance pointing to tool-schema. For a simple two-parameter read tool, nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both parameters are already documented in the schema, including the dotted-path resolution semantics. The description restates the signature and echoes the path concept without adding syntax or format detail beyond the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (check/verify) and resource (a dotted instance path resolving to a real Instance), and immediately scopes it with 'without reading any properties'. This distinguishes it cleanly from property-reading siblings without needing to open their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the condition for use ('pre-flight check before click-button / type-text-box / get-instance-properties') and the payoff ('fail fast with a clear message instead of acting on a stale or wrong path'), with the concrete rationale that UI is created/destroyed dynamically. Naming specific sibling tools leaves little to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

virtual-inputSend a low-level virtual input eventA
Destructive

WRITES LIVE GAME STATE. Send keyboard, mouse, touch, or gamepad input into the active Roblox client. This is the broad low-level input surface: keyDown/keyUp/keyPress use Enum.KeyCode, mouseMove supports absolute or relative movement, mouseButton supports down/up/click, mouseWheel sends a wheel delta, touch sends Begin/Change/End, and gamepadButton/gamepadAxis target a gamepad. VirtualInputManager is preferred; common executor mouse/key fallbacks are used when available. Calls into VirtualInputManager are wrapped with newcclosure when the executor provides it. Unsupported executor APIs return a structured error rather than silently claiming success. Use press-key for the simpler keyboard-only case. Signature: { action: "keyDown" | "keyUp" | "keyPress" | "mouseMove" | "mouseButton" | "mouseWheel" | "touch" | "gamepadButton" | "gamepadAxis", key: string?, x: number?, y: number?, relative: any?, button: any?, buttonAction: any?, delta: any?, holdSec: any?, touchId: any?, touchState: any?, gamepad: any?, gamepadButton: string?, gamepadDown: any?, axis: string?, axisX: any?, axisY: any?, axisZ: any?, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval. Capabilities: VirtualInputManager. Produces: structured-result. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
xNoScreen X coordinate, or horizontal relative delta for mouseMove.
yNoScreen Y coordinate, or vertical relative delta for mouseMove.
keyNoExact Enum.KeyCode member name for keyDown, keyUp, or keyPress, such as W, Space, or LeftShift.
axisNoExact Enum.KeyCode member name for a gamepad axis, such as Thumbstick1 or ButtonL2.
axisXNoGamepad axis X/value component.
axisYNoGamepad axis Y/value component.
axisZNoGamepad axis Z/value component.
deltaNoWheel delta for mouseWheel.
actionYesThe virtual input operation to perform.
buttonNoMouse button for mouseButton.Left
gamepadNoGamepad input device.Gamepad1
holdSecNoSeconds to hold keyPress or a mouse click between down and up, capped at 10 seconds.
touchIdNoTouch identifier for touch.
relativeNoFor mouseMove, use executor relative movement when true; otherwise use absolute screen coordinates.
touchStateNoTouch phase.Begin
gamepadDownNoPressed state for gamepadButton.
buttonActionNoWhether mouseButton sends down, up, or a down/up click.click
gamepadButtonNoExact Enum.KeyCode member name for a gamepad button, such as ButtonA or ButtonStart.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true and non-idempotent, and the description adds substantial context on top: VirtualInputManager is preferred with executor fallbacks, calls are wrapped in newcclosure when available, and unsupported executor APIs return a structured error instead of a false success. It also names the phase, cost, idempotency class, and verification path.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The safety warning is front-loaded ('WRITES LIVE GAME STATE') and most sentences carry distinct information. The inline Signature block largely restates the enum and parameter list already present at 100% schema coverage, which is the one redundant element keeping this from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 19-parameter mutation tool with no output schema, the description covers what it writes, how it behaves on failure (structured error, inspect tool-schema), the preferred implementation path, required approvals, and how to verify. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds real mapping value by binding each action to its parameters (keyDown/keyUp/keyPress use Enum.KeyCode, mouseMove supports absolute or relative, touch sends Begin/Change/End, gamepadButton/gamepadAxis target a gamepad). This clarifies which fields apply to which action, which the flat schema does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Send keyboard, mouse, touch, or gamepad input into the active Roblox client') and immediately frames itself as 'the broad low-level input surface', distinguishing it from press-key and the higher-level GUI input siblings. An agent can place it precisely without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the simple case to press-key ('Use press-key for the simpler keyboard-only case') and lists prerequisites (active-client, explicit-mutation-approval) and a verification step (assert-state). It does not mention other plausible alternatives such as click-button or type-text-box for GUI targeting, so it falls just short of full when/when-not coverage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vm-resetReset the Persistent Script VMA
Read-onlyIdempotent

Wipe the persistent VM environment used by the script tool. Every global, function, and value defined by previous persistent script runs is discarded, giving you a clean session. This does NOT touch the game itself — only the VM's own sandboxed environment. Wipes ONLY your own scope: pass the same agent label (and client) you run scripts under so it clears that agent's VM on that game, never a co-tenant's. Returns { reset = true }. Signature: { client: string?, agent: string? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: operation-receipt. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNoOptional. Reset the persistent VM belonging to this agent lane — pass the SAME label you run your `script`/`execute` calls under, so you clear your own VM and not a co-tenant agent's.
clientNoOptional. Reset the VM on a specific connected client (its clientId or username) for THIS call, instead of your session's selected client.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false; the description adds real context beyond them — exactly what gets discarded (all globals/functions/values), that the game is untouched, scope isolation per agent/client, and the return shape { reset = true }. Some inline metadata (Phase/cost/Safety) restates annotation hints redundantly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and scope, then constraints. It is somewhat long and ends with a dense metadata tail (Phase/cost/idempotency/Requires/Produces) that is partly redundant with annotations, diluting conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, but the description supplies the return value, required active-client precondition, failure guidance, and safety scoping — enough for an agent to call it correctly. Remaining gap is minimal.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% (baseline 3), and the description adds meaning: `agent` must match the label scripts run under to avoid clearing a co-tenant, and `client` overrides the session's selected client for this call. This clarifies intent beyond the schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific verb and resource ('Wipe the persistent VM environment used by the `script` tool') and ties it to the relevant sibling (`script`). An agent immediately understands this resets persistent script state rather than touching the game.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear operational guidance: pass the same `agent`/`client` you run `script`/`execute` under so you clear your own lane, not a co-tenant's. It does not explicitly contrast with alternate VM tools (e.g. new-lua-state-proxy), but the when/how context is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

watch-instance-propertyWatch a property change over timeA
Read-onlyIdempotent

Poll a single property of an instance at a fixed interval for a bounded duration and record every sampled value plus whether it changed since the previous sample. Use this to debug timing/animation/state issues — e.g. confirm a Frame's Visible actually toggles when you click, watch a Humanoid's Health drop, or see whether a Value object updates. Returns { Path, Property, Samples = [{ t, value, changed }], changeCount }. The call blocks for roughly durationMs while sampling, so it is inherently synchronous; keep durations short. Signature: { instancePath: string, propertyName: string, checkIntervalMs: any?, durationMs: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Produces: bounded-event-snapshot. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
durationMsNoTotal time to watch in milliseconds (default: 3000, max: 30000). The tool blocks for about this long; perform the triggering action (click/type) shortly before or during the watch.
instancePathYesDotted path to the instance to watch, starting at 'game' (e.g. 'game.Players.LocalPlayer.PlayerGui.HUD.HealthBar').
propertyNameYesName of the single property to sample each interval (e.g. 'Visible', 'Text', 'Value', 'Health', 'Position').
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
checkIntervalMsNoPolling interval in milliseconds (default: 100, min: 10, max: 5000). Smaller catches fast transitions but produces more samples.

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly/idempotent annotations, it discloses the critical non-obvious trait that the call blocks for roughly durationMs and is inherently synchronous, plus prerequisites (active-client, resolved-target), cost class, and the exact return shape. That is exactly the behavioral context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is well front-loaded, but the trailing metadata block ('Phase: observe; cost=medium; idempotency=read-only... Safety: read-only... On failure: inspect tool-schema') largely duplicates the annotations (readOnlyHint, idempotentHint) and the schema signature, padding the text without adding selection or invocation value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly supplies the return shape ({ Path, Property, Samples, changeCount }) and the sampling semantics, so an agent knows what it will get back. It is nearly complete; the only gap is behavior on error (missing instance/property) beyond the generic 'inspect tool-schema' pointer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and each parameter already documents defaults, ranges and semantics, so the baseline is 3. The description restates the signature and repeats the blocking/duration guidance that the schema's durationMs description already provides, adding little parameter-level meaning beyond it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise verb+resource+mechanism: 'Poll a single property of an instance at a fixed interval for a bounded duration and record every sampled value plus whether it changed.' That is far more specific than the title. However, it never distinguishes itself from close siblings such as watch-value or watch-property-changes, so an agent cannot rule those out from this text alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage context is concrete and actionable: 'Use this to debug timing/animation/state issues' with three worked examples (Frame.Visible toggling on click, Humanoid.Health dropping, Value object updating). It also warns to keep durations short. It stops short of stating when NOT to use it or naming an alternative tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

watch-property-changesRecord EVERY property that changes on an Instance over a window (event-driven)A
Read-onlyIdempotent

Connects an Instance's Changed signal for a bounded window and records EVERY property that changes (its name plus the new value), then disconnects and returns the log. This is the discovery counterpart to watch-instance-property (inspection), which polls a SINGLE named property you already know: use watch-property-changes when you DON'T know which property to watch and want to see everything that moves — e.g. perform an action in-game and learn which of an Instance's properties the game actually mutates, or catch a property you didn't expect to change. It uses a signal CONNECTION (Instance.Changed), not a function hook, so it is low-risk; the connection is always disconnected at the end of the window. HOW IT WORKS: resolves the Instance, connects inst.Changed (which fires with the changed property NAME for plain Instances), and for each fire pcall-reads inst[prop] and appends { property, newValue, t }; then task.wait(s) for the duration and disconnects. The call BLOCKS for roughly durationMs while listening — perform the triggering action (click/move/etc.) shortly before or during the watch. NOTE: on some objects the Changed signal carries a Property-value-changed payload rather than a name (e.g. ValueBase objects fire with the value); this tool records the raw signal argument as the property identifier in that case. Returns { Path, ClassName, durationMs, changeCount, changes = [{ property, newValue, t }], truncated } or { error }. Signature: { instancePath: string, durationMs: any?, limit: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client, resolved-target. Produces: bounded-event-snapshot. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of change records to keep (default 500, max 2000). When more changes fire than this cap, the extra ones are dropped and `truncated` is set true.
durationMsNoHow long to listen, in milliseconds (default 4000, min 100, max 30000). The tool BLOCKS for about this long; trigger the change you want to observe shortly before or during this window.
instancePathYesLuau expression resolving to the Instance to watch, e.g. 'game.Players.LocalPlayer.Character.Humanoid', 'game.Workspace.Part', or 'game.Players.LocalPlayer.PlayerGui.HUD.Frame'. Evaluated as `return <instancePath>` and must resolve to an Instance.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnly/idempotent/non-destructive, but the description adds substantial behavior beyond them: signal CONNECTION vs function hook, guaranteed disconnect at window end, blocking for ~durationMs, the ValueBase payload edge case, and the truncation cap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and the sibling contrast, and every major section is useful. Some redundancy exists (read-only stated in annotations, idempotency line, and Safety line; the signature duplicates the schema), which keeps it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully documents the return shape ({ Path, ClassName, durationMs, changeCount, changes[], truncated } or { error }), plus prerequisites and failure handling, leaving nothing essential missing for a 4-param read tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters. The description mostly restates them (blocking for durationMs, signature line), adding little syntax or meaning beyond what the schema provides, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (connects/records/disconnects) and resource (an Instance's Changed signal over a bounded window), and explicitly distinguishes itself from the sibling watch-instance-property by naming the counterpart and contrasting discovery vs inspection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when-to-use rule ('when you DON'T know which property to watch'), names the alternative (watch-instance-property) and the condition that selects it, and adds operational guidance about triggering the action during the blocking window.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

watch-valueWatch any Luau expression over time (live value monitor)A
Read-onlyIdempotent

Sample ANY Luau expression repeatedly over a bounded window and report every value it took plus exactly when it changed — a live memory/state monitor for reverse-engineering and debugging. Give an expression that returns a value (no return keyword needed) and the tool compiles it once with loadstring, then polls it on a fixed interval until the duration elapses, recording the first sample and every subsequent change. Use it to confirm a value actually mutates when you trigger an action (watch leaderstats coins tick up, a Humanoid's Health drop, an anti-cheat flag flip, a getgenv() speed multiplier, a remote-cooldown timer, or a metatable hook's hit counter), to time how fast something updates, or to prove a suspected variable is the one that drives a behavior. EXAMPLES: 'game.Players.LocalPlayer.leaderstats.Coins.Value', 'getgenv().speed', 'workspace.CurrentCamera.CFrame.Position', '#getconnections(game.Workspace.Part.Touched)', 'tostring(getgenv().__mcp_remotespy and #getgenv().__mcp_remotespy.logs)'. The expression is evaluated as return <expression>, so it can be any value-producing Luau (indexing, function calls, arithmetic, concatenation). Each sample is stringified for transport (Instances become GetFullName()). Change detection compares the live value with the previous one via ~=, so it catches identity changes for tables/Instances and equality changes for scalars. IMPORTANT: this call BLOCKS in-client for roughly durationMs while it samples (it is inherently synchronous), so keep durations short and perform the triggering action shortly before or during the watch. Everything is fully pcall-guarded: a compile error returns { error } immediately, and a per-tick evaluation error is recorded as a sample value rather than aborting the loop. Returns { expression, intervalMs, durationMs, sampleCount, changeCount, truncated, samples = [{ t, value, changed, ok }] }. Signature: { expression: string, intervalMs: any?, durationMs: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Produces: bounded-event-snapshot. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
durationMsNoTotal time to watch in milliseconds (default: 5000, clamped 50..30000). The call blocks in-client for about this long; the network timeout is set to durationMs + 10000 so the loop finishes before the round-trip times out.
expressionYesAny Luau expression that returns a value (do NOT prefix with `return` — the tool adds it). Examples: 'game.Players.LocalPlayer.leaderstats.Coins.Value', 'getgenv().speed', 'workspace.CurrentCamera.FieldOfView', '#getconnections(game.Workspace.Part.Touched)'.
intervalMsNoPolling interval in milliseconds between samples (default: 200, clamped 20..5000). Smaller catches fast transitions but produces more samples and more in-client overhead.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses far more than the annotations: the call BLOCKS in-client for ~durationMs (synchronous), is fully pcall-guarded, returns { error } on compile failure, records per-tick evaluation errors as sample values, detects change via ~=, and stringifies values (Instances become GetFullName()). These are exactly the operational behaviors annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads what it does well, but is long and carries redundant tail metadata that restates the annotations (idempotency=read-only, Safety: read-only, Requires/Produces) alongside a repeated signature block. Informative but not tight; several lines could be trimmed without loss.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema is present, but the description fully specifies the return shape (expression, intervalMs, durationMs, sampleCount, changeCount, truncated, samples with t/value/changed/ok), the blocking constraint, and error semantics. An agent has everything needed to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (including clamps and defaults) are already documented. The description reinforces expression semantics ('evaluated as return <expression>') and the polling model, but this largely duplicates the schema rather than adding new meaning, matching the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (sample/watch) and resource (ANY Luau expression over a bounded window), plus the underlying mechanism (loadstring + fixed-interval polling). The emphasis on 'ANY expression' implicitly sets it apart from the property-specific watchers, but no sibling is named explicitly, which is what separates a 4 from a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Offers concrete when-to-use triggers: confirm a value mutates when you act, time how fast something updates, or prove a variable drives behavior, with realistic worked examples. It stops short of explicit exclusions or named alternatives (e.g. when to prefer watch-instance-property over this generic watcher), leaving that inference to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

world-deltaStream bounded event-driven world changesA
Destructive

WRITES LIVE CLIENT OBSERVER STATE — installs RBXScriptConnections and a bounded getgenv registry, but never modifies gameplay Instances. action='start' subscribes to Workspace and PlayerGui descendant additions/removals, character spawn/removal and equipped Tool changes, CurrentCamera replacement and selected camera properties, Backpack Tool changes, plus explicitly requested Instance properties. action='poll' returns cursor-ordered deltas; action='status' reports observer health; action='stop' disconnects every connection. Events are filtered before storage, repeated noise is coalesced/throttled, the ring buffer reports every capacity eviction, and a fixed TTL runs the same cleanup path automatically. It uses Roblox signals only: no RenderStepped, per-frame polling, or world scans. Signature: { action: "start" | "poll" | "status" | "stop", observerId: string?, cursor: any?, limit: any?, maxEvents: any?, ttlSeconds: any?, throttleMs: any?, coalesceWindowMs: any?, filters: any?, cameraProperties: any?, watchedProperties: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval, getgenv capability, a bounded observation goal and relevant source filters. Capabilities: getgenv. Produces: grounded-evidence, event-driven-world-deltas, monotonic-cursor, buffer-gap-and-drop-accounting, observer-health-and-expiry, coalescing-and-throttle-statistics. Verify with: assert-state, observe-world. Safety: MUTATING; writes live game/client state, installs bounded RBXScriptConnections in the active client, stores bounded observer metadata and events in getgenv.__mcp_world_delta, schedules one TTL cleanup callback; no gameplay Instance is written. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum events returned by one poll; hasMore indicates another page.
actionYesstart installs a bounded observer, poll reads events after cursor, status reports health without consuming events, and stop disconnects every observer connection.
cursorNoPoll returns events whose monotonically increasing cursor is greater than this value. Always persist the returned nextCursor for the next poll.
filtersNoStart-only server-side filters, applied before an event consumes buffer capacity.
maxEventsNoStart-only ring-buffer capacity. Older events are evicted at the cap and reported through droppedEvents, droppedThrough, and cursorGap.
observerIdNoObserver id returned by start. For poll/status/stop, omission selects the newest active observer owned by this MCP session.
throttleMsNoMinimum interval for storing repeated events with the same semantic key after their coalescing window.
ttlSecondsNoStart-only fixed lifetime. A delayed cleanup retires the observer and disconnects every connection even if the AI never calls stop.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.
cameraPropertiesNoCamera properties watched with GetPropertyChangedSignal. Add CFrame or Focus only when needed; throttling protects high-frequency properties.
coalesceWindowMsNoWindow in which identical source/kind/path/property events update one record and increment repeatCount.
watchedPropertiesNoAdditional explicit Instance/property subscriptions. At most 16 targets and 64 aggregate property connections are allowed.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, but the description adds substantial context beyond them: it clarifies exactly what is written (client observer state, getgenv.__mcp_world_delta) and what is NOT ('no gameplay Instance is written'), plus TTL cleanup, ring-buffer eviction accounting, and coalescing/throttling behavior. This is a rich, non-contradictory disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core behavior is front-loaded in the first sentence before the action enumeration, and the structured metadata (Phase/cost/idempotency/Requires/Capabilities/Safety) is scannable. It is dense and slightly redundant with the inline signature block, but every section carries usable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 12-parameter, nested-schema tool with no output schema, the description covers what it produces (grounded-evidence, deltas, cursor, health), verification tools (assert-state, observe-world), safety, and failure guidance ('inspect tool-schema for exact fields...'). Nothing an agent needs to call it correctly appears missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by spelling out what action='start' subscribes to, how poll returns 'cursor-ordered deltas', and how status/stop behave, giving the agent inline semantic context for the key action parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ('installs RBXScriptConnections and a bounded getgenv registry' that 'streams event-driven world changes') and enumerates the four actions. It hints at differentiation from scanning siblings ('no RenderStepped, per-frame polling, or world scans') but never names an alternative tool explicitly, so it falls just short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear context for invocation ('Requires: active-client, explicit-mutation-approval, getgenv capability, a bounded observation goal and relevant source filters') and explains each action's purpose. It does not, however, tell the agent when to choose this over observe-world, watch-value, or watch-property-changes, so exclusions are absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

write-fileWrite (overwrite) a file in the executor workspace (UNC writefile)A
Destructive

Write content to a file in the executor's workspace folder, creating it if needed and OVERWRITING any existing contents. NOTE: this is executor-side file I/O — the connector runs INSIDE the executor, so the path is relative to the executor's workspace directory on the host machine, NOT the Roblox game. Requires the UNC function writefile(path, content). The call is type-guarded and pcall-wrapped: if writefile is missing you get { error = 'writefile is not available in this executor.' }, and any write failure (bad path, denied extension, permission) returns { error = }. Returns { path, ok = true } or { error }. Signature: { path: string, content: string, threadContext: number?, timeoutMs: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, resolved-target, explicit-mutation-approval. Capabilities: executor filesystem. Produces: operation-receipt. Verify with: file-exists. Safety: MUTATING; writes executor workspace filesystem. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesPath to the file within the executor workspace folder, e.g. 'config.json'. Existing contents are replaced.
contentYesThe full text content to write to the file.
timeoutMsNoOptional per-call deadline in milliseconds; omit it to use the tool or server default.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructive/not-idempotent/openWorld, so the bar is lower, yet the description adds real substance: the UNC writefile dependency, the exact error shape for missing writefile and for write failures, and the success shape { path, ok = true }. That is material behavioral detail beyond what the structured fields convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and the overwrite warning, then reference detail. Most sentences earn their place, though the internal taxonomy line (Phase: act; cost=medium; idempotency=...) plus the generic 'inspect tool-schema' closer add some boilerplate weight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating file-write tool with no output schema, it covers prerequisites, side effects, verification, and both error and success return shapes, so an agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents path, content, timeoutMs and threadContext. The inline signature restates the same parameters without adding format, default, or constraint semantics beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource (write content to a file) with explicit scope: 'executor's workspace folder', overwriting existing contents. The clarification that this is executor-side I/O 'NOT the Roblox game' cleanly separates it from read-file, append-file, and the many in-game mutation siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives strong context: requires active-client, resolved-target and explicit-mutation-approval, and names file-exists as the verification step. It implies rather than states when to prefer append-file or create-instance; no explicit when-not/alternative routing sentence beyond the executor-vs-game distinction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

write-path-valueWrite a value into a live GC table slot (direct heap write, MUTATES STATE)A
Destructive

WRITES LIVE GAME STATE. Direct heap write into one slot of a Luau table: resolve a Luau expression to a TABLE container, then assign container[key] = , returning BOTH the OLD and NEW value so the change is auditable. This is the write counterpart to read-path-value and the natural action after a memscan locates a field — e.g. flip a config flag ('require(game.ReplicatedStorage.Config)', key='GodMode', value true), bump a cached stat ('getgenv().PlayerData', key='Coins', value=9999), or clear a slot (kind='nil'). The container expression is evaluated as return <containerExpr> and MUST resolve to a table (anything else returns a clean { error }). For non-primitive values (Vector3, Color3, Enum, an Instance, a table, …) use kind='raw' and pass a Luau expression. The read of the old value and the write are each pcall-guarded. WARNING: this mutates running client memory immediately; writing through a metatable-protected/readonly table may error (returned as { error }) and some game state replicates or is server-authoritative. Requires loadstring/load. Returns { Container, Key, OldValue, NewValue, ok } or { error }. Signature: { containerExpr: string, key: string, value: { kind: "string" | "number" | "boolean" | "nil" | "raw", value: string | number | boolean? }, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; writes live game/client state. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYesThe string key (field name) within the container table to write, e.g. 'Coins', 'GodMode', 'WalkSpeed'. Indexed as container[key]. (String keys only — for non-string keys, write via a raw containerExpr that already indexes to the parent.)
valueYestyped value consumed by this operation.
containerExprYesLuau expression resolving to the TABLE whose slot you want to write, e.g. 'getgenv().PlayerData', 'require(game.ReplicatedStorage.Config)', '_G.Settings', or any table reference found via a heap scan. Evaluated as `return <containerExpr>` and must resolve to a table.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false, and the description goes well beyond them: pcall-guarded old-read and write, error-as-{error} on metatable-protected/readonly tables, server-authoritative replication caveats, and the loadstring/load requirement. It also discloses that both OLD and NEW values are returned for auditability, which is genuinely useful context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The critical constraint ('WRITES LIVE GAME STATE') is front-loaded and the prose is dense with usable detail, but the trailing block ('Signature: ... Phase: act; cost=medium; idempotency=contextual-write. Requires: ... Produces: ...') restates schema and metadata in prose, which pads the definition without adding selection value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no output schema, the description fully documents the return contract ({ Container, Key, OldValue, NewValue, ok } or { error }), the failure modes, and the verification path. Nothing an agent needs to invoke or interpret this mutating call is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: containerExpr is evaluated as `return <containerExpr>` and MUST resolve to a table, and kind='raw' is the route for Vector3/Color3/Enum/Instance/table values. It stops short of documenting threadContext semantics beyond what the schema already says.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource+scope ('Direct heap write into one slot of a Luau table') and explicitly names its counterpart and predecessor ('the write counterpart to read-path-value and the natural action after a memscan locates a field'). An agent can distinguish it from read-path-value, set-instance-property, and set-attribute without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states when to use it (after a memscan locates a field), gives three concrete scenarios (flip a config flag, bump a cached stat, clear a slot with kind='nil'), and routes the agent explicitly: kind='raw' for non-primitive values, and verify with assert-state. Alternatives and the condition selecting them are spelled out rather than inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ws-closews:Close — close a WebSocket and free its slotA
Destructive

WRITES LIVE GAME STATE — closes a WebSocket previously opened with ws-connect, addressed by its registry id. Calls socket:Close() (pcall-wrapped), marks the entry closed, and removes it from the getgenv registry so the id no longer appears in ws-list. Requires getgenv and the entry from ws-connect — guarded, returning { error } when the id is unknown. Returns { id, closed = true } or { error }. Signature: { id: number, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval. Capabilities: WebSocket. Produces: structured-result. Verify with: ws-list. Safety: MUTATING; performs external network or socket I/O. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe WebSocket registry id returned by ws-connect.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=false, readOnlyHint=false, so the mutation profile is known. The description adds real value beyond that: the underlying call is pcall-wrapped, the entry is marked closed and removed from the getgenv registry, and unknown ids are guarded to return { error }. The 'MUTATING; external network/socket I/O' line largely restates annotations, keeping it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The operational content is front-loaded and useful, but the description is padded with metadata boilerplate (Phase, cost, idempotency, Capabilities, Produces, Safety) that duplicates annotation-provided facts. Several of these clauses do not earn their place, diluting an otherwise efficient opening.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the description supplies the return shapes ({ id, closed = true } or { error }), the failure/guard behavior, prerequisites (active-client, getgenv, ws-connect entry), and a verification path via ws-list. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both parameters (id, threadContext) are already fully documented in the schema. The description echoes the signature but adds no syntax or format meaning beyond what the schema provides, matching the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('closes a WebSocket') and immediately scopes it to sockets 'previously opened with ws-connect, addressed by its registry id'. An agent can distinguish it from ws-connect, ws-send, ws-receive, and ws-list without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear precondition (must be opened via ws-connect, entry comes from that call) and names ws-list as the verification step. It does not explicitly state when-not to use it or name a competing alternative, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ws-connectWebSocket.connect — open a live WebSocketA
Destructive

WRITES LIVE GAME STATE — opens a real outbound WebSocket from the client via WebSocket.connect(url). The live socket is parked in a getgenv registry keyed by an integer id (returned as id); use that id with ws-send, ws-receive, ws-close, and ws-list. On connect, ws.OnMessage is wired into a capped (200-frame) ring buffer and ws.OnClose flips the entry's open flag to false, so inbound frames accumulate server-side until you fetch them. Requires the WebSocket library (type(WebSocket)=='table' with WebSocket.connect) — both are type-guarded and the connect is pcall-wrapped, returning { error } when missing or on failure. Returns { id, url } or { error }. Signature: { url: string, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval. Capabilities: WebSocket. Produces: created-handle. Verify with: ws-list. Safety: MUTATING; performs external network or socket I/O. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe WebSocket URL to connect to, e.g. 'ws://127.0.0.1:8080' or 'wss://example.com/socket'.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well past the annotations (which only mark it non-readonly, destructive, open-world, non-idempotent) by disclosing the 200-frame capped ring buffer, OnMessage/OnClose wiring, the open-flag side effect, the type-guard plus pcall wrapping, and the { id, url } vs { error } return shape. It also flags the mutation approval requirement and external socket I/O, which an agent needs before invoking.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The critical facts (what it opens, the registry id, the downstream siblings) are front-loaded in the first two sentences, and the buffer/pcall behavior follows logically. The trailing block of templated metadata (phase, cost, capabilities, 'On failure: inspect tool-schema...') is partly boilerplate that duplicates annotations, keeping it short of a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, and the description compensates by spelling out the return shape ({ id, url } or { error }) and the id contract used by sibling tools. With annotations covering safety and the schema covering both parameters, nothing an agent needs to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both url and threadContext are already fully documented in the schema. The description's 'Signature: { url: string, threadContext: number? }' restates the schema without adding format, default, or constraint detail beyond it. Baseline 3 applies when the schema carries the load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — 'opens a real outbound WebSocket from the client via WebSocket.connect(url)' — and explicitly separates itself from siblings by naming ws-send, ws-receive, ws-close, and ws-list as the follow-on tools. An agent can place this as the entry point of the WebSocket workflow without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear operational context: the returned id is what the sibling tools consume, and 'Verify with: ws-list' tells the agent how to confirm success. It also states prerequisites (requires active-client and explicit-mutation-approval). It stops short of saying when NOT to open a socket or naming a competing alternative, so it is clear context rather than explicit routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ws-listList open WebSocketsA
Read-onlyIdempotent

List every WebSocket currently parked in the getgenv registry by ws-connect. For each entry it reports { id, url, open, messageCount } so you can see which sockets are still live and how many inbound frames are buffered. Read-only. Requires getgenv — guarded; when no registry exists it simply returns an empty list. Returns { count, sockets } or { error }. Signature: { threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Capabilities: WebSocket. Produces: bounded-candidates. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnly/idempotent/openWorld, yet the description adds real context beyond them: it reads from the getgenv registry, is guarded so it returns an empty list rather than erroring when no registry exists, and describes the failure mode. It stops short of discussing ordering or cost beyond a vague 'cost=medium'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening two sentences are excellent and front-loaded, but the trailing metadata block ('Phase: observe; cost=medium; idempotency=read-only... Safety: read-only') largely repeats what the annotations already declare, and the 'inspect tool-schema' fallback sentence is boilerplate.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden and does so ('Returns { count, sockets } or { error }'), plus per-entry fields. It is complete enough to invoke correctly, with only minor gaps in pagination/ordering.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single threadContext parameter is fully documented in the schema. The description only restates the signature, adding no format or default nuance beyond what the schema provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List every WebSocket currently parked in the getgenv registry by ws-connect') and pins the source of the data. This clearly distinguishes it from the sibling ws-connect/ws-send/ws-receive/ws-close family.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the registry framing ('see which sockets are still live'), but there is no explicit when-to-use or when-not-to-use guidance and no named alternative for related inspection tasks. An agent can infer intent but must guess at boundaries.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ws-receiveRead buffered WebSocket framesA
Read-onlyIdempotent

Read the inbound frames buffered for a WebSocket opened with ws-connect, addressed by its registry id. Returns the captured { text, t } records newest-first, capped at limit, plus whether the socket is still open and the total buffered count. With clear=true the ring is emptied after the snapshot is taken, so the next call only sees newer frames. Requires getgenv and the entry from ws-connect — guarded, returning { error } when the id is unknown. Returns { id, open, messageCount, messages } or { error }. Signature: { id: number, limit: any?, clear: any?, threadContext: number? }. Phase: observe; cost=medium; idempotency=read-only. Requires: active-client. Capabilities: WebSocket. Produces: structured-result. Safety: read-only. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe WebSocket registry id returned by ws-connect.
clearNoWhen true, empty the buffered frames after returning them (default false).
limitNoMax frames to return, newest-first (default 100).
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well past the annotations: newest-first ordering, limit capping, the buffer-emptying side effect of clear=true, the getgenv/registry guard, and the {error} fallback for unknown ids. The one unreconciled point is that clear=true mutates the ring while readOnlyHint=true and destructiveHint=false are declared; the description discloses this, so it is context rather than a contradiction, but the tension is left for the agent to resolve.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The prose opening is dense and front-loaded with the important facts, but the trailing block (Phase/cost/idempotency/Requires/Capabilities/Produces/Safety) duplicates annotation data and the Signature line repeats the schema, padding a description that was already complete.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return contract itself ({ id, open, messageCount, messages } or { error }) plus the guard and clearing semantics, which is sufficient for correct invocation. It defers exact field details to tool-schema, a minor gap rather than a missing requirement.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: it states that results are newest-first, that clear acts after the snapshot, and that threadContext is optional/omit-for-default. The signature line largely restates the schema rather than extending it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (read) and resource (inbound frames buffered for a WebSocket) and scopes it by registry id from ws-connect. An agent can distinguish it from ws-list, ws-connect, ws-send, and ws-close without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly implies the prerequisite (a socket opened with ws-connect) and explains the key usage decision — clear=true empties the ring so subsequent calls only see newer frames. No explicit when-not or named alternative, but the operating context is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ws-sendws:Send — send a frame over an open WebSocketA
Destructive

WRITES LIVE GAME STATE — sends a text frame over a WebSocket previously opened with ws-connect, addressed by its registry id. Looks up getgenv().__mcp_ws[id], verifies the socket is still open, and calls socket:Send(message). Requires getgenv and the live socket from ws-connect — both are guarded and the send is pcall-wrapped, returning { error } when the id is unknown, the socket is closed, or Send fails. Returns { id, sent = true } or { error }. Signature: { id: number, message: string, threadContext: number? }. Phase: act; cost=medium; idempotency=contextual-write. Requires: active-client, explicit-mutation-approval. Capabilities: WebSocket. Produces: operation-receipt. Verify with: assert-state. Safety: MUTATING; performs external network or socket I/O. On failure: inspect tool-schema for exact fields, defaults, constraints, and an invocation example.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe WebSocket registry id returned by ws-connect.
messageYesThe text frame to send over the socket.
threadContextNoOptional Roblox thread identity for this call; omit it to use the server default.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructive=true, idempotent=false and openWorld=true, but the description adds real behavioral detail beyond them: the registry lookup, an open-socket check, pcall wrapping, and the exact failure surface ({ error } on unknown id, closed socket, or Send failure). The 'MUTATING / external network I/O' line merely restates the annotations, which keeps it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded and information-dense, but the trailing metadata block (Phase, cost, idempotency, capabilities, Produces, Verify with, Safety, On failure) repeats concepts already stated in prose or annotations, and the closing 'inspect tool-schema' pointer is boilerplate rather than tool-specific value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description supplies the return contract ({ id, sent = true } or { error }) and the failure conditions, plus prerequisites and safety framing. For a 3-param mutating socket-write tool, nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so id (registry id from ws-connect), message (text frame) and optional threadContext are fully documented in the schema. The description's Signature line and registry-lookup narrative confirm but do not materially extend that per-parameter meaning; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource+scope: sends a text frame over a WebSocket opened by ws-connect, addressed by registry id. It also names the exact internal mechanism (getgenv().__mcp_ws[id], socket:Send), so an agent can distinguish it cleanly from ws-connect, ws-receive and ws-close.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the prerequisite tool explicitly ('previously opened with ws-connect') and lists gating requirements (active-client, explicit-mutation-approval), giving clear context for when the tool is callable. It does not, however, spell out when an agent should prefer this over alternatives or what to do if no socket exists.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 291 tool updatesv2.0.0-spies.2
    • First observedactor-capabilities
    • First observedactor-event-monitor
    • First observedagent-context
    • First observedagent-memory
    • First observedagent-run
    • First observedappend-file
    • First observedassert-state
    • First observedbatch-execute
    • First observedblock-function
    • First observedblock-packets
    • First observedblock-remote
    • First observedbridge-status
    • First observedbuild-call-graph
    • First observedcache-invalidate
    • First observedcache-is-cached
    • First observedcache-replace
    • First observedcall-closure
    • First observedcamera-control
    • First observedcan-signal-replicate
    • First observedcapture-log-output
    • First observedcheck-caller
    • First observedclear-queue-on-teleport
    • First observedclear-remote-spy-logs
    • First observedclear-selection
    • First observedclear-semantic-index
    • First observedclick-button
    • First observedclone-function
    • First observedclone-instance
    • First observedclosure-capabilities
    • First observedcomm-channel-monitor
    • First observedcompare-gc-snapshots
    • First observedcompare-instances
    • First observedconfigure-remote-spy
    • First observedcount-function-calls
    • First observedcount-signal-connections
    • First observedcreate-comm-channel
    • First observedcreate-instance
    • First observedcrypt-base64-decode
    • First observedcrypt-base64-encode
    • First observedcrypt-decrypt
    • First observedcrypt-encrypt
    • First observedcrypt-generate-bytes
    • First observedcrypt-generate-key
    • First observedcrypt-hash
    • First observeddelete-file
    • First observeddelete-folder
    • First observeddestroy-instance
    • First observeddiff-instance-snapshot
    • First observeddisassemble-function
    • First observeddiscover-character
    • First observeddiscover-player-values
    • First observeddraw-clear
    • First observeddraw-create
    • First observeddraw-remove
    • First observeddraw-update
    • First observeddump-function-env
    • First observeddump-table
    • First observedensure-remote-spy
    • First observedeval-expression
    • First observedexecute
    • First observedexecute-and-wait
    • First observedexecute-file
    • First observedexecute-lua-state
    • First observedexecution-footprint-audit
    • First observedexplain-failure
    • First observedfile-exists
    • First observedfilter-gc
    • First observedfind-bytecode-size-outliers
    • First observedfind-constants-xref
    • First observedfind-detached-instances
    • First observedfind-duplicate-functions
    • First observedfind-event-connections
    • First observedfind-function-xrefs
    • First observedfind-functions-by-complexity
    • First observedfind-functions-by-constant
    • First observedfind-global-xrefs
    • First observedfind-hidden-guis
    • First observedfind-hidden-instances
    • First observedfind-hidden-remotes
    • First observedfind-hidden-scripts
    • First observedfind-instance-xrefs
    • First observedfind-instances-with-connections
    • First observedfind-module-scripts
    • First observedfind-path-references
    • First observedfind-remote-xrefs
    • First observedfind-running-scripts
    • First observedfind-string-in-tables
    • First observedfind-string-xrefs
    • First observedfind-table-references
    • First observedfind-tables-by-key
    • First observedfind-upvalue-sharing
    • First observedfind-upvalue-xref
    • First observedfire-click-detector
    • First observedfire-comm-channel
    • First observedfire-connection
    • First observedfire-lua-state-event
    • First observedfire-proximity-prompt
    • First observedfire-remote
    • First observedfire-signal
    • First observedget-active-client
    • First observedget-actor-details
    • First observedget-anticheat-surfaces
    • First observedget-call-stack
    • First observedget-closure-constants
    • First observedget-closure-protos
    • First observedget-closure-upvalues
    • First observedget-comm-channel
    • First observedget-connection-constants
    • First observedget-connection-info
    • First observedget-connection-protos
    • First observedget-connection-upvalues
    • First observedget-connector-diagnostics
    • First observedget-console-output
    • First observedget-custom-asset
    • First observedget-executor-info
    • First observedget-fast-flag
    • First observedget-fps-cap
    • First observedget-function-env
    • First observedget-function-hash
    • First observedget-function-protos
    • First observedget-function-upvalues
    • First observedget-game-info
    • First observedget-game-state
    • First observedget-gui-text
    • First observedget-hidden-ui
    • First observedget-hwid
    • First observedget-instance-counts
    • First observedget-instance-properties
    • First observedget-instance-tree
    • First observedget-local-player-info
    • First observedget-lua-state
    • First observedget-lua-state-actors
    • First observedget-memory-stats
    • First observedget-metamethod
    • First observedget-metatable
    • First observedget-module-source
    • First observedget-nil-instances
    • First observedget-place-details
    • First observedget-players
    • First observedget-remote-signature
    • First observedget-remote-spy-logs
    • First observedget-render-stats
    • First observedget-script-bytecode
    • First observedget-script-closure
    • First observedget-script-content
    • First observedget-script-env
    • First observedget-script-hash
    • First observedget-semantic-index-stats
    • First observedget-signal-arguments
    • First observedget-signal-arguments-info
    • First observedget-signal-whitelist
    • First observedget-stack
    • First observedget-thread-stack
    • First observedhook-and-log-function
    • First observedhook-function
    • First observedhook-metamethod
    • First observedhttp-request
    • First observedignore-remote
    • First observedinspect-callbacks
    • First observedinspect-closure
    • First observedinspect-instance-metatable
    • First observedinvoke-closure
    • First observedinvoke-method
    • First observedis-c-closure
    • First observedis-executor-closure
    • First observedis-function-hooked
    • First observedis-l-closure
    • First observedis-new-c-closure
    • First observedis-parallel-context
    • First observedis-readonly
    • First observedlist-actors
    • First observedlist-attributes
    • First observedlist-clients
    • First observedlist-closure-references
    • First observedlist-drawings
    • First observedlist-files
    • First observedlist-gc-functions
    • First observedlist-gc-tables
    • First observedlist-gc-threads
    • First observedlist-global-env-keys
    • First observedlist-gui-elements
    • First observedlist-hooks
    • First observedlist-instance-signals
    • First observedlist-lua-states
    • First observedlist-registry-objects
    • First observedlist-remotes
    • First observedlist-rendered-instances
    • First observedlist-roblox-windows
    • First observedlist-runtime-modules
    • First observedlist-script-actors
    • First observedlist-signal-connections
    • First observedlist-strings
    • First observedlist-tools
    • First observedload-file
    • First observedlookup-function
    • First observedlua-state-event-monitor
    • First observedmake-folder
    • First observedmeasure-memory
    • First observedmessage-box
    • First observedmonitor-remote
    • First observednew-c-closure
    • First observednew-l-closure
    • First observednew-lua-state-proxy
    • First observedobserve-world
    • First observedpacket-spy
    • First observedplaybook-delete
    • First observedplaybook-list
    • First observedplaybook-run
    • First observedplaybook-save
    • First observedpress-key
    • First observedprofile-code
    • First observedqueue-on-teleport
    • First observedread-file
    • First observedread-path-value
    • First observedrelease-closure-reference
    • First observedremote-spy
    • First observedreplicate-signal
    • First observedresolve-entity
    • First observedrestore-function
    • First observedrestore-hook
    • First observedrun-deferred
    • First observedrun-loop
    • First observedrun-luau
    • First observedrun-on-actor
    • First observedrun-with-timeout
    • First observedsave-instance
    • First observedscan-closures-by-name
    • First observedscan-closures-by-source
    • First observedscan-connections-by-source
    • First observedscan-hook-surfaces
    • First observedscan-network-endpoints
    • First observedscan-number-range
    • First observedscan-proto-functions
    • First observedscan-remote-listeners
    • First observedscreenshot-window
    • First observedscript
    • First observedscript-fanout
    • First observedscript-grep
    • First observedsearch-bytecode
    • First observedsearch-gc-value
    • First observedsearch-instances
    • First observedselect-client
    • First observedsemantic-search-scripts
    • First observedsend-packet
    • First observedsession-list
    • First observedsession-replay
    • First observedsession-show
    • First observedset-attribute
    • First observedset-clipboard
    • First observedset-closure-constant
    • First observedset-closure-upvalue
    • First observedset-connection-state
    • First observedset-fast-flag
    • First observedset-fps-cap
    • First observedset-function-env
    • First observedset-gui-text
    • First observedset-instance-property
    • First observedset-metatable-readonly
    • First observedset-properties-bulk
    • First observedset-rawmetatable
    • First observedset-stack-hidden
    • First observedsmart-task
    • First observedspoof-function-return
    • First observedstate-transaction
    • First observedsuggest-tools
    • First observedsummarize-hidden-surfaces
    • First observedsummarize-runtime-surfaces
    • First observedteach-mode
    • First observedtest-capabilities
    • First observedtool-plan
    • First observedtool-quality-audit
    • First observedtool-schema
    • First observedtrace-call-durations
    • First observedtrace-connection-function
    • First observedtrace-remote-traffic
    • First observedtrace-require-callers
    • First observedtype-text-box
    • First observedverify-path-exists
    • First observedvirtual-input
    • First observedvm-reset
    • First observedwatch-instance-property
    • First observedwatch-property-changes
    • First observedwatch-value
    • First observedworld-delta
    • First observedwrite-file
    • First observedwrite-path-value
    • First observedws-close
    • First observedws-connect
    • First observedws-list
    • First observedws-receive
    • First observedws-send

TDQS

B3.2/5.0

Scored across 291 tools

Disambiguation2/5

With 291 tools, many overlap heavily: list-strings/find-string-xrefs/find-constants-xref/find-functions-by-constant all search closure constants; remote-spy/monitor-remote/trace-remote-traffic/ensure-remote-spy/get-remote-spy-logs/configure-remote-spy all manage remote capture; hook-function/block-function/spoof-function-return/hook-and-log-function/count-function-calls/trace-call-durations all install function hooks; find-hidden-instances/find-detached-instances/get-nil-instances/find-hidden-remotes/find-hidden-scripts/find-running-scripts heavily overlap. Descriptions are detailed but the boundaries between these clusters are blurry, causing misselection.

Naming Consistency2/5

Conventions are mixed: many kebab-case verb_noun tools (find-string-xrefs, list-gc-tables, get-instance-properties), but also camelCase (search-instances, get-players, select-client), terse single-word tools (execute, script, run-luau, draw-create, ws-list, filter-gc), and others with irregular verbs. The pattern is not predictable across the set.

Tool Count1/5

291 tools is an extreme mismatch for any coherent server surface; many clusters duplicate functionality (multiple hidden-instance finders, multiple spy managers, multiple hooking variants), indicating over-expansion rather than well-scoped coverage. This is far beyond the 15-tool well-scoped range and well past the 25+ 'too many' threshold.

Completeness4/5

For the Roblox executor domain the surface is remarkably exhaustive: reflection/GC scanning, closure inspection, hooking, remote/signal monitoring, GUI driving, file I/O, crypto, drawing, packet, WebSocket, and agent-orchestration tools are all present. Minor lifecycle gaps (e.g. no first-class 'list all active hooks/spies' consolidated view) are workarounds, not dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to interact with a running Roblox game client, including executing Lua code, inspecting scripts, and spying on remotes, with a local dashboard for monitoring and control.
    79 npm
    MIT
  • -
    license
    Not graded
    quality
    Not graded
    maintenance
    Lets AI agents interact with a live Roblox game via an executor, supporting multi-agent sessions and tools for executing Lua, decompiling scripts, exploring game instances, and capturing screenshots.
    -