Skip to main content
Glama

Run Rego tests

rego_test

Execute Rego unit tests with opa test to get pass, fail, skip, and error counts plus per-test records. Filter runs by regex, enforce coverage thresholds, and debug failing cases.

Instructions

Run Rego unit tests with opa test. Returns aggregate pass/fail/skip/error counts plus per-test records. errored counts tests OPA could not evaluate (a rule conflict, a raising built-in); such a test is neither a pass nor a failure, and a suite with any is not passing. Tests live in *_test.rego files; rule names beginning with test_ are picked up automatically. Use runPattern to filter by name regex; when no tests match, the error hint includes the pattern you supplied. Use threshold to gate on minimum coverage (returns COVERAGE_BELOW_THRESHOLD on failure). Use varValues: true with verbose: true to include local variable bindings in the trace -- essential for debugging table-driven tests written with every tc in cases { ... } to identify which case caused a failure. When tests use the test_x[case] parameterized form, OPA reports the rule as a single test whatever the number of cases; parameterizedGroups maps the rule name to a record per case and caseCounts totals them, so a failing rule says which case failed. Use ignorePatterns to exclude generated or fixture files. Use bundle: true when testing bundle-structured policy directories. Use timeout to raise the per-test limit beyond OPA's default 5s. Note: enabling coverage or threshold switches OPA to coverage-report output mode -- per-test counts are unavailable but coverage and coveragePct fields are populated.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
countNoNumber of times to repeat the suite (`--count N`). Default is 1. Useful for catching flaky tests. OPA stops at the first repetition that fails, so `repetitions` in the output reports how many actually ran, and each test is listed once carrying its worst outcome across them.
pathsYesTest directories or files. `opa test` looks for `*_test.rego` siblings of source files.
bundleNoLoad paths as OPA bundle roots (`--bundle`). Required when testing policies structured as bundles with a `manifest.json` at the root. Not needed for plain policy directories.
explainNoAdd a query-explanation trace to test records (`--explain`). `fails` traces only failing tests, `full` traces everything, `notes` surfaces `trace()` notes, `debug` is most verbose. Populates each record's `trace` field; pair with `verbose: true` for the human-readable trace output too.
timeoutNoPer-test timeout as a Go duration string, e.g. `"30s"` or `"2m"` (`--timeout`). OPA's default is 5s. Increase for tests that load large policy sets or call slow built-ins.
verboseNoEmit per-test pass/fail details.
coverageNoInclude per-line coverage data. Switches output to coverage-report mode: test record counts are not available, but `coverage` and `coveragePct` fields are populated.
thresholdNoMinimum coverage percentage required (0–100). Returns COVERAGE_BELOW_THRESHOLD when actual coverage falls below this value. Implicitly enables coverage-report output mode.
varValuesNoInclude local variable bindings in trace output (`--var-values`). When a table-driven test using `every tc in cases { ... }` fails, the trace shows which `tc` triggered the failure. Has no effect unless `verbose: true` is also set (OPA only emits trace entries in verbose mode).
runPatternNoRun only tests whose names match this regular expression (passed as `--run`).
v0CompatibleNoRead the policy as Rego v0 (`--v0-compatible`), the syntax OPA used before 1.0: rules without `if`, partial sets as `deny[msg] { ... }`. Needed for a policy that has not been migrated, which OPA 1.x otherwise refuses to load. Where the tool also takes a query, the query is read as v0 too, with the future keywords imported so `in`, `every` and `some x in` still work in it.
v1CompatibleNoOpt in to OPA v1.0-compatible behaviors (`--v1-compatible`).
ignorePatternsNoGlob patterns for files to exclude from the test run (`--ignore <pattern>`). Pass one pattern per array element. Useful for excluding generated or fixture files that contain no tests (e.g. `["*_generated.rego", "fixtures/**"]`).

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv0.8.0
    • changedInput schema / properties / v0Compatible / description
      Previous value: -"Read the policy as Rego v0 (`--v0-compatible`), the syntax OPA used before 1.0: rules without `if`, partial sets as `deny[msg] { ... }`. Needed for a policy that has not been migrated, which OPA 1.x otherwise refuses to load."New value: +"Read the policy as Rego v0 (`--v0-compatible`), the syntax OPA used before 1.0: rules without `if`, partial sets as `deny[msg] { ... }`. Needed for a policy that has not been migrated, which OPA 1.x otherwise refuses to load. Where the tool also takes a query, the query is read as v0 too, with the future keywords imported so `in`, `every` and `some x in` still work in it."
  2. Changed1 schema field changedv0.7.0
    • addedInput schema / properties / v0Compatible
      Added value: +{
      +  "description": "Read the policy as Rego v0 (`--v0-compatible`), the syntax OPA used before 1.0: rules without `if`, partial sets as `deny[msg] { ... }`. Needed for a policy that has not been migrated, which OPA 1.x otherwise refuses to load.",
      +  "type": "boolean"
      +}
  3. Changed1 schema field changedv0.5.0
    • changedInput schema / properties / count / description
      Previous value: -"Number of times to repeat each test (`--count N`). Default is 1. Useful for measuring repeatability or catching flaky tests under load."New value: +"Number of times to repeat the suite (`--count N`). Default is 1. Useful for catching flaky tests. OPA stops at the first repetition that fails, so `repetitions` in the output reports how many actually ran, and each test is listed once carrying its worst outcome across them."
  4. Changed2 schema fields changedv0.1.20
    • addedInput schema / properties / explain
      Added value: +{
      +  "description": "Add a query-explanation trace to test records (`--explain`). `fails` traces only failing tests, `full` traces everything, `notes` surfaces `trace()` notes, `debug` is most verbose. Populates each record's `trace` field; pair with `verbose: true` for the human-readable trace output too.",
      +  "enum": [
      +    "fails",
      +    "full",
      +    "notes",
      +    "debug"
      +  ],
      +  "type": "string"
      +}
    • addedInput schema / properties / v1Compatible
      Added value: +{
      +  "description": "Opt in to OPA v1.0-compatible behaviors (`--v1-compatible`).",
      +  "type": "boolean"
      +}
  5. Changed4 schema fields changedv0.1.17
    • addedInput schema / properties / bundle
      Added value: +{
      +  "description": "Load paths as OPA bundle roots (`--bundle`). Required when testing policies structured as bundles with a `manifest.json` at the root. Not needed for plain policy directories.",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / count
      Added value: +{
      +  "description": "Number of times to repeat each test (`--count N`). Default is 1. Useful for measuring repeatability or catching flaky tests under load.",
      +  "minimum": 1,
      +  "type": "integer"
      +}
    • addedInput schema / properties / ignorePatterns
      Added value: +{
      +  "description": "Glob patterns for files to exclude from the test run (`--ignore <pattern>`). Pass one pattern per array element. Useful for excluding generated or fixture files that contain no tests (e.g. `[\"*_generated.rego\", \"fixtures/**\"]`).",
      +  "items": {
      +    "type": "string"
      +  },
      +  "type": "array"
      +}
    • addedInput schema / properties / timeout
      Added value: +{
      +  "description": "Per-test timeout as a Go duration string, e.g. `\"30s\"` or `\"2m\"` (`--timeout`). OPA's default is 5s. Increase for tests that load large policy sets or call slow built-ins.",
      +  "type": "string"
      +}
  6. Addedv0.1.13
  7. Removedv0.1.5
  8. Addedv0.1.2
  9. Removedv0.1.1
  10. First observedv0.1.0

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only supply readOnlyHint=false and openWorldHint=true, so the description carries most of the burden and does so well: it explains `errored` semantics (neither pass nor fail, suite not passing), the coverage/threshold output-mode switch and COVERAGE_BELOW_THRESHOLD error, the error-hint behavior echoing runPattern, and that varValues is inert without verbose. It stops short of warning that test execution runs arbitrary policy code with side-effect-capable built-ins.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and return shape, then a dense run of per-flag guidance; the length is justified by 13 parameters and no output schema. It is one unbroken block with no grouping, and a few points (verbose/varValues trace behavior) are restated verbatim from the schema, costing some efficiency.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and minimal annotations, the description does the necessary work of describing return fields (aggregate counts, per-test records, parameterizedGroups, caseCounts, coveragePct) and failure modes. It is nearly complete for a tool of this complexity; only the absence of sibling routing and execution-safety notes leaves a small gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with unusually rich per-parameter descriptions, so the baseline is 3. The description goes slightly beyond by consolidating cross-parameter interactions (varValues requires verbose; coverage/threshold switch output mode; count's worst-outcome aggregation) and by tying runPattern to the error hint, though much of this restates what the schema already says.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ("Run Rego unit tests with `opa test`") and immediately defines the scope of output (aggregate counts plus per-test records). It further distinguishes itself from generic evaluation siblings by describing test-discovery semantics (`*_test.rego`, `test_` prefix) that an agent can use to tell it apart from rego_eval or conftest_test.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use guidance for individual options: runPattern for filtering, threshold for coverage gating, varValues+verbose for table-driven debugging, ignorePatterns for generated files, bundle for bundle-structured dirs, timeout for slow suites. It does not, however, name an alternative tool (e.g. rego_test_multiroot for multi-root runs, or conftest_test for the conftest flow) to route between siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.