Skip to main content
Glama

Benchmark Rego query

rego_bench

Benchmark policy queries against your defined input and data paths with opa bench, then review iterations, ns/op, and allocation counts to pinpoint slow rules.

Instructions

Benchmark a Rego query against a policy + input with opa bench. Returns statistical timing data: iterations, ns/op, and allocation counts. Use this to spot slow rules.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
countNoNumber of times to repeat the benchmark (`--count N`). Defaults to OPA's built-in default of one. Above one, every repetition is returned in `runs`, `fastest` indexes the one the top-level figures come from, and `raw` is omitted since that document is in `runs`.
inputNoInline input document.
pathsNoPolicy / data paths to load. Each must be in an allowed root.
queryYesRego query to benchmark.
inputPathNoPath to a JSON input file.
v0CompatibleNoRead the policy as Rego v0 (`--v0-compatible`), the syntax OPA used before 1.0: rules without `if`, partial sets as `deny[msg] { ... }`. Needed for a policy that has not been migrated, which OPA 1.x otherwise refuses to load. Where the tool also takes a query, the query is read as v0 too, with the future keywords imported so `in`, `every` and `some x in` still work in it.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv0.8.0
    • changedInput schema / properties / v0Compatible / description
      Previous value: -"Read the policy as Rego v0 (`--v0-compatible`), the syntax OPA used before 1.0: rules without `if`, partial sets as `deny[msg] { ... }`. Needed for a policy that has not been migrated, which OPA 1.x otherwise refuses to load."New value: +"Read the policy as Rego v0 (`--v0-compatible`), the syntax OPA used before 1.0: rules without `if`, partial sets as `deny[msg] { ... }`. Needed for a policy that has not been migrated, which OPA 1.x otherwise refuses to load. Where the tool also takes a query, the query is read as v0 too, with the future keywords imported so `in`, `every` and `some x in` still work in it."
  2. Changed1 schema field changedv0.7.0
    • addedInput schema / properties / v0Compatible
      Added value: +{
      +  "description": "Read the policy as Rego v0 (`--v0-compatible`), the syntax OPA used before 1.0: rules without `if`, partial sets as `deny[msg] { ... }`. Needed for a policy that has not been migrated, which OPA 1.x otherwise refuses to load.",
      +  "type": "boolean"
      +}
  3. Changed1 schema field changedv0.6.0
    • changedInput schema / properties / count / description
      Previous value: -"Number of times to repeat the benchmark (`--count N`). Defaults to OPA's built-in default of one. Every repetition is returned in `runs`; the top-level figures come from the fastest of them."New value: +"Number of times to repeat the benchmark (`--count N`). Defaults to OPA's built-in default of one. Above one, every repetition is returned in `runs`, `fastest` indexes the one the top-level figures come from, and `raw` is omitted since that document is in `runs`."
  4. Changed1 schema field changedv0.5.0
    • changedInput schema / properties / count / description
      Previous value: -"Number of benchmark iterations. Defaults to OPA's built-in default."New value: +"Number of times to repeat the benchmark (`--count N`). Defaults to OPA's built-in default of one. Every repetition is returned in `runs`; the top-level figures come from the fastest of them."
  5. Addedv0.1.13
  6. Removedv0.1.5
  7. Addedv0.1.2
  8. Removedv0.1.1
  9. First observedv0.1.0

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and openWorldHint=true, so the description carries most of the behavioral load. It discloses the underlying mechanism (`opa bench`) and the shape of the result (iterations, ns/op, allocation counts), which is genuinely useful for an agent that has no output schema. It does not mention execution cost, timeouts, or the fact that benchmarking repeatedly executes the policy.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences that are front-loaded with the action, then the return shape, then the reason to use it. No filler and nothing repeated from the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description must say something about returns, and it does summarize the statistical fields. The count>1 behavior (runs/fastest, omitted raw) is only explained in the schema's count parameter, leaving the description slightly thin for an agent sizing up results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% across all six parameters, so the schema already documents query, input, paths, count, inputPath, and v0Compatible in detail. The description adds no parameter-level meaning beyond what the schema supplies, which is the baseline-3 case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (benchmark) and resource (a Rego query against a policy + input), names the underlying command (`opa bench`), and reports what comes back (iterations, ns/op, allocation counts). It does not explicitly name the nearest sibling (rego_eval_with_profile), so an agent must infer the distinction itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Use this to spot slow rules" gives a motivating use case, which implies the performance-investigation context. There is no explicit when-not guidance and no mention of alternatives such as rego_eval_with_profile, so the agent must decide between them on its own.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.