Skip to main content
Glama

Aegis DQ

CI PyPI Downloads Python License GitHub Marketplace Glama Open in Colab

The open-source agentic data quality framework. Point it at your policy docs and warehouse — it generates rules, validates your data, diagnoses every failure with LLM root-cause analysis, and proposes SQL fixes. Run from the CLI, Airflow, GitHub Actions, or conversationally via Hermes.

Real-world result: 12 AML policy docs → 55 rules generated → 11 BSA/OFAC violations detected → all diagnosed → $0.01 total LLM cost.

  • 31 rule types — completeness, uniqueness, validity, referential integrity, statistical, ML anomaly detection

  • 6 warehouse adapters — DuckDB, Postgres/Redshift, BigQuery, Databricks, AWS Athena, Snowflake

  • Pluggable LLMs — Anthropic Claude, OpenAI, Ollama (local), AWS Bedrock

  • Agentic pipeline — plan → parallel validation → LLM diagnose → RCA → SQL remediate → report

  • Hermes + MCP — run full pipelines conversationally; listed on Glama.ai


GitHub Actions — Quick Start

Add a data quality gate to any workflow in under 2 minutes:

# .github/workflows/data-quality.yml
name: Data Quality

on: [push, pull_request]

jobs:
  data-quality:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4

      - name: Validate data quality
        uses: aegis-dq/aegis-dq@v0.7.0
        with:
          rules-file: rules.yaml
          db: data/warehouse.duckdb
          anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}

The step fails the job automatically when any rules fail, blocking broken data from reaching production. Set fail-on-failure: 'false' to report without blocking.

Offline mode (no API key required):

      - name: Validate data quality (offline)
        uses: aegis-dq/aegis-dq@v0.7.0
        with:
          rules-file: rules.yaml
          db: data/warehouse.duckdb
          no-llm: 'true'

Action inputs

Input

Default

Description

rules-file

rules.yaml

Path to rules YAML

db

:memory:

DuckDB file path

warehouse

duckdb

duckdb · postgres · redshift

pg-dsn

—

PostgreSQL / Redshift connection DSN

no-llm

false

Skip LLM — free offline validation

llm

anthropic

anthropic · openai · ollama

llm-model

(provider default)

Override the default model

fail-on-failure

true

Fail the step when rules fail

version

(latest)

Pin a specific aegis-dq version

anthropic-api-key

—

Required when llm: anthropic

openai-api-key

—

Required when llm: openai

Action outputs

Output

Description

rules-checked

Total rules evaluated

passed

Rules that passed

failed

Rules that failed

pass-rate

Pass rate as a decimal (e.g. "91.67")

report-json

Absolute path to the full JSON report

Using outputs in downstream steps:

      - name: Validate data quality
        id: dq
        uses: aegis-dq/aegis-dq@v0.7.0
        with:
          rules-file: rules.yaml

      - name: Post summary
        run: echo "Pass rate: ${{ steps.dq.outputs.pass-rate }}%"

Related MCP server: dbt-doctor

Demo

Aegis DQ Demo

╭──────────────────────────────────────────────────────╮
│ Aegis DQ  —  RetailCo E-commerce Demo                │
│ LLM: amazon.nova-pro-v1:0 via AWS Bedrock            │
╰──────────────────────────────────────────────────────╯

✓ Pipeline complete in 7.1s · 12 rules · $0.0056 LLM cost

╭──────────────── Validation Summary ─────────────────╮
│  Rules checked  │  12                               │
│  Passed         │  1   │  Failed  │  11             │
│  Pass rate      │  8%  │  Cost    │  $0.005576      │
╰─────────────────────────────────────────────────────╯

LLM Diagnoses
  orders_customer_fk  →  Order placed with customer_id=99 that does not exist.
                         Likely cause: customer deleted or test record not cleaned up.

  products_sku_unique →  Duplicate SKU-001 — two products share the same identifier.
                         Likely cause: duplicate import from supplier feed.

Remediation SQL (LLM-generated)
  orders_status_valid          UPDATE orders SET status = 'SHIPPED' WHERE status = 'DISPATCHED';
  products_price_positive      UPDATE products SET price = ABS(price) WHERE price < 0;
  products_stock_non_negative  UPDATE products SET stock_quantity = 0 WHERE stock_quantity < 0;

Hermes integration — conversational data quality

Aegis DQ integrates with Hermes via MCP. Point Hermes at your context and your warehouse — it handles the rest.

You (Hermes chat)
     │
     ▼
Hermes — memory, scheduling, multi-channel delivery
     │  MCP
     ▼
Aegis DQ — rules engine, LLM diagnosis, audit trail
     │
     ▼
Your warehouse + Your LLM

Setup (2 steps):

pip install aegis-dq

Add to ~/.hermes/config.yaml:

mcp_servers:
  aegis:
    command: aegis
    args: [mcp]
    env:
      ANTHROPIC_API_KEY: "${ANTHROPIC_API_KEY}"

Define a pipeline manifest once:

# pipeline.yaml
name: orders-dq
rules: ./rules.yaml
database: ./warehouse.duckdb
kb:
  - ./policy.md     # business rules, SLAs, compliance docs
  - ./schema.md
goal: |
  Run all rules. For every failure explain the business impact,
  likely root cause, and a concrete remediation step.

Then just ask:

Load the pipeline at pipeline.yaml and run it.

Hermes calls load_pipeline → run_validation → returns a structured report. No flags, no re-explaining context on every run.

Full setup guide: aegis-dq.dev/integrations/hermes · MCP listing: glama.ai/mcp/servers/aegis-dq/aegis-dq


Why Aegis?

Aegis DQ

Great Expectations / Soda

Monte Carlo / Anomalo

Open source

✅ Apache 2.0

✅

❌ Commercial

Agentic LLM diagnosis + RCA

✅

❌

✅ Proprietary

SQL auto-fix proposals

✅

❌

❌

Audit trail (per-decision log)

✅

Partial

✅ Proprietary

Pluggable LLM (Anthropic, OpenAI, Bedrock, Ollama)

✅

❌

❌

dbt integration

✅

✅

Partial

Portable open rule standard

✅

Partial

❌

ML anomaly detection

✅ built-in

❌

✅ Proprietary


Install

pip install aegis-dq

Extra

What it adds

aegis-dq[bigquery]

BigQuery adapter

aegis-dq[databricks]

Databricks adapter

aegis-dq[athena]

AWS Athena adapter

aegis-dq[postgres]

PostgreSQL / Redshift adapter

aegis-dq[snowflake]

Snowflake adapter

aegis-dq[rest]

REST API server (FastAPI + uvicorn)

aegis-dq[openai]

OpenAI LLM provider

aegis-dq[airflow]

Airflow AegisOperator

aegis-dq[mcp]

MCP server for Hermes, Claude Desktop, and any MCP-compatible agent

aegis-dq[ml]

scikit-learn anomaly detection


5-minute quickstart

Step 1 — Install

pip install aegis-dq

Step 2 — Seed a demo database

import duckdb

con = duckdb.connect("demo.db")
con.execute("""
    CREATE TABLE orders AS
    SELECT i AS order_id, 'placed' AS status, i * 9.99 AS revenue
    FROM range(1, 10001) t(i)
""")
# introduce some bad data
con.execute("UPDATE orders SET order_id = NULL WHERE order_id % 200 = 0")
con.execute("UPDATE orders SET revenue = -5.00 WHERE order_id % 500 = 0")
con.close()

Step 3 — Generate rules from your schema (no hand-writing)

export ANTHROPIC_API_KEY=sk-ant-...

# Generate rules from table schema alone
aegis generate orders --db demo.db --output rules.yaml

# Or point it at a policy doc and get business validation rules too
aegis generate orders --db demo.db --kb docs/orders_policy.md --output rules.yaml

The LLM introspects your schema and generates not_null, accepted_values, between, and custom_sql rules automatically. Generated rules are stamped status: draft — review and promote to active.

Step 4 — Run

aegis run rules.yaml --db demo.db

Run without an API key (pass/fail only, no LLM diagnosis):

aegis run rules.yaml --db demo.db --no-llm

Pipeline

Every aegis run passes your data through a LangGraph pipeline:

rules (Python / YAML)
    │
    ▼
  plan ──► parallel_table ──► reconcile ──► remediate ──► report
                 │
         ┌──────────────────┐
         │  per table:      │
         │  execute         │
         │  classify        │
         │  diagnose        │  ← concurrent across all tables
         │  rca             │
         └──────────────────┘
  • plan — parse and validate rules, build an execution graph

  • parallel_table — concurrently fans out per table: execute all rules, classify failures by severity, diagnose with LLM, and trace root causes

  • reconcile — compare results against expected thresholds

  • remediate — LLM proposes a targeted SQL fix for each diagnosed failure

  • report — structured JSON + optional Slack notification


Rule types (31 total)

Category

Types

Completeness

not_null not_empty_string null_percentage_below

Uniqueness

unique composite_unique duplicate_percentage_below

Validity

sql_expression between min_value_check max_value_check regex_match accepted_values not_accepted_values no_future_dates column_exists

Referential

foreign_key conditional_not_null

Statistical

mean_between stddev_below column_sum_between

Timeliness

freshness date_order

Volume

row_count row_count_between custom_sql

Cross-table

reconcile_row_count reconcile_column_sum reconcile_key_match

ML / Anomaly

zscore_outlier isolation_forest learned_threshold

Example rule:

rules:
  - apiVersion: aegis.dev/v1
    kind: DataQualityRule
    metadata:
      id: orders_revenue_non_negative
      severity: critical
      owner: revenue-team
      tags: [revenue, validity]
    scope:
      warehouse: duckdb
      table: orders
    logic:
      type: sql_expression
      expression: "revenue >= 0"

Generate rules with the LLM

Instead of writing rules by hand, let Aegis introspect your table schema and generate a draft rules file:

# Schema-aware structural rules (not_null, between, unique, accepted_values...)
aegis generate orders --db warehouse.duckdb --output orders_rules.yaml

Add a --kb document — any plain text or markdown file describing your business logic — and the LLM generates business validation rules alongside structural ones:

aegis generate orders \
  --db warehouse.duckdb \
  --kb docs/orders_policy.md \
  --output orders_rules.yaml

What goes in a KB file? Anything your team knows about the data:

# orders_policy.md
- status must be one of: placed, confirmed, shipped, delivered, cancelled
- amount must be greater than 0; refunds are handled in a separate table
- customer_id must reference a valid customer (no test accounts: id > 1000)
- order_date must not be in the future
- discount_pct must be between 0 and 0.5 (max 50% discount)

The LLM turns these into accepted_values, sql_expression, between, and foreign_key rules automatically. Generated rules are stamped status: draft — review, promote to active, and commit.

All aegis generate options:

Flag

Default

Description

--db

—

DuckDB file for schema introspection

--kb

—

Business-context file (text/markdown)

--output

rules.yaml

Output YAML file

--max-rules

20

Cap on number of rules generated

--no-verify

false

Skip SQL verification of generated rules

--save-versions

false

Persist rules to version store

--provider

anthropic

LLM provider

--model

(default)

Override model


Warehouse adapters

Adapter

Install

Status

DuckDB

built-in

✅ GA

BigQuery

aegis-dq[bigquery]

✅ GA

Databricks

aegis-dq[databricks]

✅ GA

AWS Athena

aegis-dq[athena]

✅ GA

Postgres / Redshift

aegis-dq[postgres]

✅ GA

Snowflake

aegis-dq[snowflake]

✅ GA


LLM providers

Provider

Install

Default model

Anthropic (Claude)

built-in

claude-haiku-4-5

OpenAI

aegis-dq[openai]

gpt-4o-mini

Ollama (local)

aegis-dq[ollama]

llama3.2

AWS Bedrock

pip install boto3

amazon.nova-pro-v1:0

Switch providers at the CLI:

aegis run rules.yaml --llm openai --llm-model gpt-4o
aegis run rules.yaml --llm ollama --llm-model llama3.2
aegis run rules.yaml --llm bedrock --llm-model amazon.nova-pro-v1:0

Integrations

Integration

What it does

GitHub Action

CI/CD gate — fails the job when rules fail

aegis-dq[rest]

REST API server — aegis serve

aegis-dq[airflow]

AegisOperator — drop-in Airflow task

aegis-dq[mcp]

MCP server for Hermes, Claude Desktop, Cursor, and any MCP-compatible agent

aegis dbt generate

Convert dbt manifest.json to Aegis rules


CLI reference

Command

Description

aegis init

Generate a starter rules.yaml

aegis validate <config>

Check YAML syntax + schema (no warehouse needed)

aegis generate <table>

LLM-generate rules from table schema

aegis run <config>

Run validation, diagnose failures, produce a report

aegis rules list

Browse built-in rule templates

aegis audit trajectory <run-id>

Inspect the LLM decision trail for a past run

aegis audit search <query>

Full-text search across audit logs

aegis dbt generate <manifest>

Convert a dbt manifest to Aegis rules

aegis mcp

Start the MCP server for Hermes, Claude Desktop, or any MCP client

aegis run flags:

Flag

Default

Description

--db

:memory:

DuckDB file path

--llm

anthropic

LLM provider

--llm-model

(provider default)

Override model name

--no-llm

false

Skip LLM diagnosis entirely

--output-json

(none)

Write full JSON report to file

--notify

(none)

Slack webhook URL

--notify-on

failures

When to notify: all · failures · critical


Roadmap

Phase

Version

Items

Status

Foundation

v0.1

Core agent, DuckDB, CLI, audit trail

✅ Done

Differentiate

v0.5

BigQuery, Databricks, Athena, Airflow, Ollama, RCA, ShareGPT export, FTS5 search, dbt, MCP

✅ Done

Quality

v0.7

SQL verification pipeline, rule versioning, aegis generate (LLM + KB), GitHub Action, ML anomaly detection

✅ Done

Mature

v1.0

Postgres, REST API, parallel subagents, VS Code extension, eval suite, banking/healthcare packs

🚧 In progress

Full issue tracker: github.com/aegis-dq/aegis-dq/issues


Contributing

Contributions are welcome. See CONTRIBUTING.md to get started.

Good first issues: label:good first issue

License

Apache 2.0

Available Tools

6 tools
get_run_reportC

Get the ShareGPT-formatted trajectory for a run, including metadata.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided; description only states it retrieves data without mentioning side effects, permissions, or error conditions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence with 12 words, no fluff, but excessively brief for the context needed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Missing behavioral info and usage guidance; though output schema exists, the description is insufficient for a tool with a close sibling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter run_id has no description in schema (0% coverage) and the description adds no clarification beyond its name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the tool gets a trajectory in ShareGPT format with metadata, but doesn't differentiate from sibling 'get_trajectory'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool over alternatives like 'get_trajectory', or any prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_trajectoryB

Get the full decision trajectory for a completed run in ShareGPT format.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided; description adds output format context but fails to disclose read-only nature, error handling, or any behavioral constraints beyond the basic action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single, front-loaded sentence with no extraneous words. Efficient and to the point.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Adequate for a simple tool with one parameter and an output schema, but lacks details on error conditions, behavior for invalid inputs, and a clearer definition of 'ShareGPT format'.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not explain the run_id parameter or how to obtain it, leaving the agent without guidance on the single required input.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the action (Get), resource (decision trajectory), context (for a completed run), and output format (ShareGPT format), effectively distinguishing it from siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies usage for completed runs but provides no explicit guidance on when to use versus alternatives like get_run_report, nor exclusionary scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_runsB

List recent Aegis DQ run IDs, newest first.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry behavioral traits. It discloses the ordering (newest first) but does not mention if it is read-only or any edge cases. Adequate for a simple list operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single 8-word sentence that directly states the action and ordering. No wasted information; concise and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the basic behavior but omits explanation of the limit parameter. Given that an output schema exists, the return values are likely defined, but the parameter semantics are missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter 'limit' has no schema description (0% coverage) and is not mentioned in the description. The description adds no meaning beyond the schema, so the agent must guess its purpose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description explicitly states it lists recent Aegis DQ run IDs with newest first ordering. This clearly defines the action and resource, and distinguishes it from sibling tools like get_run_report which likely provide details for a specific run.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as search_decisions or get_run_report. The agent must infer usage from context alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

load_pipelineA

Load a pipeline manifest and return its configuration + goal as context.

Use this before run_validation to understand what a named pipeline does. After calling this, call run_validation with the rules_path and connection_params from the returned manifest.

Args: manifest_path: Path to a pipeline.yaml manifest file.

Returns: JSON with the pipeline config and a ready-to-use run_validation call.

ParametersJSON Schema
NameRequiredDescriptionDefault
manifest_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description carries full burden. It explains the return value includes a 'ready-to-use run_validation call' and describes the output structure. Could include more about side effects or authorization, but sufficient for the tool's straightforward purpose.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is 5 sentences, front-loaded with purpose. The 'Args' and 'Returns' sections are somewhat redundant with the schema but still helpful. Could be slightly more concise but not verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description provides sufficient context: it explains the return structure (config, goal, fields for run_validation) and the tool's role in a workflow. Complete for a one-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so description must explain the parameter. It describes 'manifest_path' as 'Path to a pipeline.yaml manifest file', adding clear meaning to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Load a pipeline manifest and return its configuration + goal as context', specifying the verb and resource. It distinguishes itself from sibling 'run_validation' by positioning as a prerequisite.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Use this before run_validation' and instructs to call run_validation with fields from the returned manifest. Provides clear when-to-use and next steps.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_validationA

Run Aegis DQ validation against a rules YAML file.

Args: rules_path: Path to the rules YAML file. warehouse: Warehouse type — one of: duckdb, bigquery, athena, databricks, postgres. Defaults to "duckdb" (in-memory). Set env vars for connection defaults (e.g. BQ_PROJECT + BQ_DATASET for BigQuery, POSTGRES_DSN for Postgres). connection_params: JSON object with warehouse connection kwargs. Overrides env var defaults. Examples: duckdb: {"path": "/data/prod.duckdb"} bigquery: {"project": "my-proj", "dataset": "analytics"} athena: {"s3_staging_dir": "s3://bucket/athena/", "region_name": "us-east-1"} databricks: {"server_hostname": "abc.azuredatabricks.net", "http_path": "/sql/1.0/warehouses/abc", "access_token": "dapi..."} postgres: {"dsn": "postgresql://user:pass@host:5432/db"} no_llm: If True, skip LLM diagnosis and run offline.

Returns: JSON-encoded validation report.

ParametersJSON Schema
NameRequiredDescriptionDefault
rules_pathYes
warehouseNoduckdb
connection_paramsNo{}
no_llmNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden for behavioral disclosure. It mentions returning a JSON report and the ability to skip LLM diagnosis, but does not clarify if the operation is read-only, whether auth is needed, or if any side effects occur (e.g., data modification).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with an initial purpose statement followed by parameter details and examples. It is somewhat lengthy but justified by the need to explain multiple warehouse types. Could be slightly more concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema, the description does not need to detail return values. It covers all parameters, explains behavior (no_llm), and provides connection guidance. No gaps are apparent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, yet the description provides detailed, actionable explanations for all four parameters, including defaults, environment variable hints, and concrete JSON examples for connection_params. This adds significant value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's action ('Run Aegis DQ validation') and target resource ('rules YAML file'). It is specific and distinct from sibling tools like 'get_run_report' or 'load_pipeline'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides some usage context (defaults, environment variables, no_llm flag) but does not explicitly explain when to use this tool versus alternatives (e.g., load_pipeline vs run_validation). The guidance is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_decisionsA

Full-text search over the audit decision trail.

Args: query: Search terms (e.g. "null ETL bug", "root cause orders") run_id: Optional — restrict to a specific run limit: Maximum number of results to return

Returns: JSON array of matching decision records.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes
run_idNo
limitNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description must carry the burden. It states that full-text search is performed and returns matching records, but lacks details on behavior such as pagination, sorting, or whether the search is fuzzy or exact.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the purpose and uses a structured docstring format. It is concise but slightly verbose with the 'Args:' and 'Returns:' sections; could be streamlined.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has a low complexity (3 params, 1 required) and an output schema is known to exist. The description covers inputs and return type adequately, though it does not explain output structure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description adds meaningful semantics: explains query with examples ('null ETL bug'), run_id as optional restriction, and limit as max results. This compensates for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Full-text search over the audit decision trail', specifying a verb (search), resource (audit decision trail), and scope. This distinguishes it from siblings like get_run_report or list_runs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives. Usage is implied by the description (searching decision records), but no exclusions or context for selecting among sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 6 tool updatesv0.7.0
    • First observedget_run_report
    • First observedget_trajectory
    • First observedlist_runs
    • First observedload_pipeline
    • First observedrun_validation
    • First observedsearch_decisions

TDQS

B3.4/5.0

Scored across 6 tools

Disambiguation2/5

The distinction between get_run_report and get_trajectory is unclear; both return ShareGPT-formatted trajectories with only a minor difference in metadata inclusion. This overlap causes ambiguity for an agent choosing between them.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern with lowercase underscores (e.g., get_run_report, list_runs, run_validation), making them predictable and easy to navigate.

Tool Count5/5

With six tools, the server is well-scoped for a data quality validation workflow. Each tool serves a distinct role without being overwhelming or insufficient.

Completeness4/5

The core workflow (load pipeline, run validation, list runs, retrieve trajectories, search decisions) is covered. Missing features like run deletion or pipeline management are minor gaps that do not hinder typical use.

Maintenance

ActivityInactive
ResponsivenessWithin a week

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    AI-driven MCP server that audits, profiles, detects schema drift, and auto-generates documentation for dbt projects, enabling natural language interaction with your dbt project's health.
    134
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Zero-config data quality monitoring as MCP tools. Profiles a warehouse (Postgres, BigQuery, Snowflake, MySQL, DuckDB), detects anomalies, and gates CI — read-only with the connection resolved server-side, never via the model.
    6
    305 PyPI
    10
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    A production-ready MCP server that provides comprehensive dbt project quality assessment for any GitHub repository, enabling AI agents to analyze dbt models, check metadata coverage, and map data lineage.
    9
    MIT