aegis-dq
OfficialThis server provides a conversational and programmatic interface for data quality validation pipelines, enabling you to run validations, inspect results, and audit decision trails.
run_validation: Execute data quality checks using a rules YAML file against various warehouses (DuckDB, BigQuery, Athena, Databricks, Postgres), with optional LLM-powered diagnosis or offline mode.load_pipeline: Load apipeline.yamlmanifest to inspect its configuration, rules path, connection parameters, and goal before running validation.list_runs: Retrieve recent validation run IDs, sorted newest first.get_run_report: Retrieve the full report and metadata for a specific completed run.get_trajectory: Fetch the complete LLM decision trajectory (ShareGPT format) for a run, allowing inspection of every reasoning step.search_decisions: Full-text search over the audit decision trail (e.g., root cause analyses), optionally scoped to a specific run.
Supports data quality checks on Databricks tables, including ML anomaly detection and auto-generated SQL remediation.
Integrates with dbt pipelines to run data quality checks as part of transformation workflows.
Allows running data quality validation rules against a DuckDB database, including completeness, uniqueness, and statistical checks.
Leverages local Ollama LLM for offline diagnosis and remediation, avoiding external API calls.
Uses OpenAI models to diagnose data quality failures and propose targeted SQL remediation steps.
Enables data quality validation on PostgreSQL or Redshift databases with support for all rule types and LLM diagnosis.
Delivers data quality reports and alerts directly to Slack channels for team visibility.
Provides data quality validation for Snowflake data warehouses, leveraging the same rule engine and LLM capabilities.
Aegis DQ
The open-source agentic data quality framework. Point it at your policy docs and warehouse — it generates rules, validates your data, diagnoses every failure with LLM root-cause analysis, and proposes SQL fixes. Run from the CLI, Airflow, GitHub Actions, or conversationally via Hermes.
Real-world result: 12 AML policy docs → 55 rules generated → 11 BSA/OFAC violations detected → all diagnosed → $0.01 total LLM cost.
31 rule types — completeness, uniqueness, validity, referential integrity, statistical, ML anomaly detection
6 warehouse adapters — DuckDB, Postgres/Redshift, BigQuery, Databricks, AWS Athena, Snowflake
Pluggable LLMs — Anthropic Claude, OpenAI, Ollama (local), AWS Bedrock
Agentic pipeline — plan → parallel validation → LLM diagnose → RCA → SQL remediate → report
Hermes + MCP — run full pipelines conversationally; listed on Glama.ai
GitHub Actions — Quick Start
Add a data quality gate to any workflow in under 2 minutes:
# .github/workflows/data-quality.yml
name: Data Quality
on: [push, pull_request]
jobs:
data-quality:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Validate data quality
uses: aegis-dq/aegis-dq@v0.7.0
with:
rules-file: rules.yaml
db: data/warehouse.duckdb
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}The step fails the job automatically when any rules fail, blocking broken data from reaching production. Set fail-on-failure: 'false' to report without blocking.
Offline mode (no API key required):
- name: Validate data quality (offline)
uses: aegis-dq/aegis-dq@v0.7.0
with:
rules-file: rules.yaml
db: data/warehouse.duckdb
no-llm: 'true'Action inputs
Input | Default | Description |
|
| Path to rules YAML |
|
| DuckDB file path |
|
|
|
| — | PostgreSQL / Redshift connection DSN |
|
| Skip LLM — free offline validation |
|
|
|
| (provider default) | Override the default model |
|
| Fail the step when rules fail |
| (latest) | Pin a specific |
| — | Required when |
| — | Required when |
Action outputs
Output | Description |
| Total rules evaluated |
| Rules that passed |
| Rules that failed |
| Pass rate as a decimal (e.g. |
| Absolute path to the full JSON report |
Using outputs in downstream steps:
- name: Validate data quality
id: dq
uses: aegis-dq/aegis-dq@v0.7.0
with:
rules-file: rules.yaml
- name: Post summary
run: echo "Pass rate: ${{ steps.dq.outputs.pass-rate }}%"Related MCP server: dbt-doctor
Demo

╭──────────────────────────────────────────────────────╮
│ Aegis DQ — RetailCo E-commerce Demo │
│ LLM: amazon.nova-pro-v1:0 via AWS Bedrock │
╰──────────────────────────────────────────────────────╯
✓ Pipeline complete in 7.1s · 12 rules · $0.0056 LLM cost
╭──────────────── Validation Summary ─────────────────╮
│ Rules checked │ 12 │
│ Passed │ 1 │ Failed │ 11 │
│ Pass rate │ 8% │ Cost │ $0.005576 │
╰─────────────────────────────────────────────────────╯
LLM Diagnoses
orders_customer_fk → Order placed with customer_id=99 that does not exist.
Likely cause: customer deleted or test record not cleaned up.
products_sku_unique → Duplicate SKU-001 — two products share the same identifier.
Likely cause: duplicate import from supplier feed.
Remediation SQL (LLM-generated)
orders_status_valid UPDATE orders SET status = 'SHIPPED' WHERE status = 'DISPATCHED';
products_price_positive UPDATE products SET price = ABS(price) WHERE price < 0;
products_stock_non_negative UPDATE products SET stock_quantity = 0 WHERE stock_quantity < 0;Hermes integration — conversational data quality
Aegis DQ integrates with Hermes via MCP. Point Hermes at your context and your warehouse — it handles the rest.
You (Hermes chat)
│
▼
Hermes — memory, scheduling, multi-channel delivery
│ MCP
▼
Aegis DQ — rules engine, LLM diagnosis, audit trail
│
▼
Your warehouse + Your LLMSetup (2 steps):
pip install aegis-dqAdd to ~/.hermes/config.yaml:
mcp_servers:
aegis:
command: aegis
args: [mcp]
env:
ANTHROPIC_API_KEY: "${ANTHROPIC_API_KEY}"Define a pipeline manifest once:
# pipeline.yaml
name: orders-dq
rules: ./rules.yaml
database: ./warehouse.duckdb
kb:
- ./policy.md # business rules, SLAs, compliance docs
- ./schema.md
goal: |
Run all rules. For every failure explain the business impact,
likely root cause, and a concrete remediation step.Then just ask:
Load the pipeline at pipeline.yaml and run it.Hermes calls load_pipeline → run_validation → returns a structured report. No flags, no re-explaining context on every run.
Full setup guide: aegis-dq.dev/integrations/hermes · MCP listing: glama.ai/mcp/servers/aegis-dq/aegis-dq
Why Aegis?
Aegis DQ | Great Expectations / Soda | Monte Carlo / Anomalo | |
Open source | ✅ Apache 2.0 | ✅ | ❌ Commercial |
Agentic LLM diagnosis + RCA | ✅ | ❌ | ✅ Proprietary |
SQL auto-fix proposals | ✅ | ❌ | ❌ |
Audit trail (per-decision log) | ✅ | Partial | ✅ Proprietary |
Pluggable LLM (Anthropic, OpenAI, Bedrock, Ollama) | ✅ | ❌ | ❌ |
dbt integration | ✅ | ✅ | Partial |
Portable open rule standard | ✅ | Partial | ❌ |
ML anomaly detection | ✅ built-in | ❌ | ✅ Proprietary |
Install
pip install aegis-dqExtra | What it adds |
| BigQuery adapter |
| Databricks adapter |
| AWS Athena adapter |
| PostgreSQL / Redshift adapter |
| Snowflake adapter |
| REST API server (FastAPI + uvicorn) |
| OpenAI LLM provider |
| Airflow |
| MCP server for Hermes, Claude Desktop, and any MCP-compatible agent |
| scikit-learn anomaly detection |
5-minute quickstart
Step 1 — Install
pip install aegis-dqStep 2 — Seed a demo database
import duckdb
con = duckdb.connect("demo.db")
con.execute("""
CREATE TABLE orders AS
SELECT i AS order_id, 'placed' AS status, i * 9.99 AS revenue
FROM range(1, 10001) t(i)
""")
# introduce some bad data
con.execute("UPDATE orders SET order_id = NULL WHERE order_id % 200 = 0")
con.execute("UPDATE orders SET revenue = -5.00 WHERE order_id % 500 = 0")
con.close()Step 3 — Generate rules from your schema (no hand-writing)
export ANTHROPIC_API_KEY=sk-ant-...
# Generate rules from table schema alone
aegis generate orders --db demo.db --output rules.yaml
# Or point it at a policy doc and get business validation rules too
aegis generate orders --db demo.db --kb docs/orders_policy.md --output rules.yamlThe LLM introspects your schema and generates not_null, accepted_values, between, and custom_sql rules automatically. Generated rules are stamped status: draft — review and promote to active.
Step 4 — Run
aegis run rules.yaml --db demo.dbRun without an API key (pass/fail only, no LLM diagnosis):
aegis run rules.yaml --db demo.db --no-llmPipeline
Every aegis run passes your data through a LangGraph pipeline:
rules (Python / YAML)
│
▼
plan ──► parallel_table ──► reconcile ──► remediate ──► report
│
┌──────────────────┐
│ per table: │
│ execute │
│ classify │
│ diagnose │ ← concurrent across all tables
│ rca │
└──────────────────┘plan — parse and validate rules, build an execution graph
parallel_table — concurrently fans out per table: execute all rules, classify failures by severity, diagnose with LLM, and trace root causes
reconcile — compare results against expected thresholds
remediate — LLM proposes a targeted SQL fix for each diagnosed failure
report — structured JSON + optional Slack notification
Rule types (31 total)
Category | Types |
Completeness |
|
Uniqueness |
|
Validity |
|
Referential |
|
Statistical |
|
Timeliness |
|
Volume |
|
Cross-table |
|
ML / Anomaly |
|
Example rule:
rules:
- apiVersion: aegis.dev/v1
kind: DataQualityRule
metadata:
id: orders_revenue_non_negative
severity: critical
owner: revenue-team
tags: [revenue, validity]
scope:
warehouse: duckdb
table: orders
logic:
type: sql_expression
expression: "revenue >= 0"Generate rules with the LLM
Instead of writing rules by hand, let Aegis introspect your table schema and generate a draft rules file:
# Schema-aware structural rules (not_null, between, unique, accepted_values...)
aegis generate orders --db warehouse.duckdb --output orders_rules.yamlAdd a --kb document — any plain text or markdown file describing your business logic — and the LLM generates business validation rules alongside structural ones:
aegis generate orders \
--db warehouse.duckdb \
--kb docs/orders_policy.md \
--output orders_rules.yamlWhat goes in a KB file? Anything your team knows about the data:
# orders_policy.md
- status must be one of: placed, confirmed, shipped, delivered, cancelled
- amount must be greater than 0; refunds are handled in a separate table
- customer_id must reference a valid customer (no test accounts: id > 1000)
- order_date must not be in the future
- discount_pct must be between 0 and 0.5 (max 50% discount)The LLM turns these into accepted_values, sql_expression, between, and foreign_key rules automatically. Generated rules are stamped status: draft — review, promote to active, and commit.
All aegis generate options:
Flag | Default | Description |
| — | DuckDB file for schema introspection |
| — | Business-context file (text/markdown) |
|
| Output YAML file |
|
| Cap on number of rules generated |
|
| Skip SQL verification of generated rules |
|
| Persist rules to version store |
|
| LLM provider |
| (default) | Override model |
Warehouse adapters
Adapter | Install | Status |
DuckDB | built-in | ✅ GA |
BigQuery |
| ✅ GA |
Databricks |
| ✅ GA |
AWS Athena |
| ✅ GA |
Postgres / Redshift |
| ✅ GA |
Snowflake |
| ✅ GA |
LLM providers
Provider | Install | Default model |
Anthropic (Claude) | built-in |
|
OpenAI |
|
|
Ollama (local) |
|
|
AWS Bedrock |
|
|
Switch providers at the CLI:
aegis run rules.yaml --llm openai --llm-model gpt-4o
aegis run rules.yaml --llm ollama --llm-model llama3.2
aegis run rules.yaml --llm bedrock --llm-model amazon.nova-pro-v1:0Integrations
Integration | What it does |
GitHub Action | CI/CD gate — fails the job when rules fail |
| REST API server — |
|
|
| MCP server for Hermes, Claude Desktop, Cursor, and any MCP-compatible agent |
| Convert dbt |
CLI reference
Command | Description |
| Generate a starter |
| Check YAML syntax + schema (no warehouse needed) |
| LLM-generate rules from table schema |
| Run validation, diagnose failures, produce a report |
| Browse built-in rule templates |
| Inspect the LLM decision trail for a past run |
| Full-text search across audit logs |
| Convert a dbt manifest to Aegis rules |
| Start the MCP server for Hermes, Claude Desktop, or any MCP client |
aegis run flags:
Flag | Default | Description |
|
| DuckDB file path |
|
| LLM provider |
| (provider default) | Override model name |
|
| Skip LLM diagnosis entirely |
| (none) | Write full JSON report to file |
| (none) | Slack webhook URL |
|
| When to notify: |
Roadmap
Phase | Version | Items | Status |
Foundation | v0.1 | Core agent, DuckDB, CLI, audit trail | ✅ Done |
Differentiate | v0.5 | BigQuery, Databricks, Athena, Airflow, Ollama, RCA, ShareGPT export, FTS5 search, dbt, MCP | ✅ Done |
Quality | v0.7 | SQL verification pipeline, rule versioning, | ✅ Done |
Mature | v1.0 | Postgres, REST API, parallel subagents, VS Code extension, eval suite, banking/healthcare packs | 🚧 In progress |
Full issue tracker: github.com/aegis-dq/aegis-dq/issues
Contributing
Contributions are welcome. See CONTRIBUTING.md to get started.
Good first issues: label:good first issue
License
Available Tools
6 toolsget_run_reportC
Get the ShareGPT-formatted trajectory for a run, including metadata.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided; description only states it retrieves data without mentioning side effects, permissions, or error conditions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence with 12 words, no fluff, but excessively brief for the context needed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Missing behavioral info and usage guidance; though output schema exists, the description is insufficient for a tool with a close sibling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter run_id has no description in schema (0% coverage) and the description adds no clarification beyond its name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the tool gets a trajectory in ShareGPT format with metadata, but doesn't differentiate from sibling 'get_trajectory'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool over alternatives like 'get_trajectory', or any prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_trajectoryB
Get the full decision trajectory for a completed run in ShareGPT format.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided; description adds output format context but fails to disclose read-only nature, error handling, or any behavioral constraints beyond the basic action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single, front-loaded sentence with no extraneous words. Efficient and to the point.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequate for a simple tool with one parameter and an output schema, but lacks details on error conditions, behavior for invalid inputs, and a clearer definition of 'ShareGPT format'.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain the run_id parameter or how to obtain it, leaving the agent without guidance on the single required input.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the action (Get), resource (decision trajectory), context (for a completed run), and output format (ShareGPT format), effectively distinguishing it from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies usage for completed runs but provides no explicit guidance on when to use versus alternatives like get_run_report, nor exclusionary scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_runsB
List recent Aegis DQ run IDs, newest first.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry behavioral traits. It discloses the ordering (newest first) but does not mention if it is read-only or any edge cases. Adequate for a simple list operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single 8-word sentence that directly states the action and ordering. No wasted information; concise and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the basic behavior but omits explanation of the limit parameter. Given that an output schema exists, the return values are likely defined, but the parameter semantics are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter 'limit' has no schema description (0% coverage) and is not mentioned in the description. The description adds no meaning beyond the schema, so the agent must guess its purpose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description explicitly states it lists recent Aegis DQ run IDs with newest first ordering. This clearly defines the action and resource, and distinguishes it from sibling tools like get_run_report which likely provide details for a specific run.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives such as search_decisions or get_run_report. The agent must infer usage from context alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
load_pipelineA
Load a pipeline manifest and return its configuration + goal as context.
Use this before run_validation to understand what a named pipeline does. After calling this, call run_validation with the rules_path and connection_params from the returned manifest.
Args: manifest_path: Path to a pipeline.yaml manifest file.
Returns: JSON with the pipeline config and a ready-to-use run_validation call.
| Name | Required | Description | Default |
|---|---|---|---|
| manifest_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. It explains the return value includes a 'ready-to-use run_validation call' and describes the output structure. Could include more about side effects or authorization, but sufficient for the tool's straightforward purpose.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is 5 sentences, front-loaded with purpose. The 'Args' and 'Returns' sections are somewhat redundant with the schema but still helpful. Could be slightly more concise but not verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description provides sufficient context: it explains the return structure (config, goal, fields for run_validation) and the tool's role in a workflow. Complete for a one-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so description must explain the parameter. It describes 'manifest_path' as 'Path to a pipeline.yaml manifest file', adding clear meaning to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Load a pipeline manifest and return its configuration + goal as context', specifying the verb and resource. It distinguishes itself from sibling 'run_validation' by positioning as a prerequisite.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Use this before run_validation' and instructs to call run_validation with fields from the returned manifest. Provides clear when-to-use and next steps.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_validationA
Run Aegis DQ validation against a rules YAML file.
Args: rules_path: Path to the rules YAML file. warehouse: Warehouse type — one of: duckdb, bigquery, athena, databricks, postgres. Defaults to "duckdb" (in-memory). Set env vars for connection defaults (e.g. BQ_PROJECT + BQ_DATASET for BigQuery, POSTGRES_DSN for Postgres). connection_params: JSON object with warehouse connection kwargs. Overrides env var defaults. Examples: duckdb: {"path": "/data/prod.duckdb"} bigquery: {"project": "my-proj", "dataset": "analytics"} athena: {"s3_staging_dir": "s3://bucket/athena/", "region_name": "us-east-1"} databricks: {"server_hostname": "abc.azuredatabricks.net", "http_path": "/sql/1.0/warehouses/abc", "access_token": "dapi..."} postgres: {"dsn": "postgresql://user:pass@host:5432/db"} no_llm: If True, skip LLM diagnosis and run offline.
Returns: JSON-encoded validation report.
| Name | Required | Description | Default |
|---|---|---|---|
| rules_path | Yes | ||
| warehouse | No | duckdb | |
| connection_params | No | {} | |
| no_llm | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden for behavioral disclosure. It mentions returning a JSON report and the ability to skip LLM diagnosis, but does not clarify if the operation is read-only, whether auth is needed, or if any side effects occur (e.g., data modification).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with an initial purpose statement followed by parameter details and examples. It is somewhat lengthy but justified by the need to explain multiple warehouse types. Could be slightly more concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the presence of an output schema, the description does not need to detail return values. It covers all parameters, explains behavior (no_llm), and provides connection guidance. No gaps are apparent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, yet the description provides detailed, actionable explanations for all four parameters, including defaults, environment variable hints, and concrete JSON examples for connection_params. This adds significant value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's action ('Run Aegis DQ validation') and target resource ('rules YAML file'). It is specific and distinct from sibling tools like 'get_run_report' or 'load_pipeline'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides some usage context (defaults, environment variables, no_llm flag) but does not explicitly explain when to use this tool versus alternatives (e.g., load_pipeline vs run_validation). The guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_decisionsA
Full-text search over the audit decision trail.
Args: query: Search terms (e.g. "null ETL bug", "root cause orders") run_id: Optional — restrict to a specific run limit: Maximum number of results to return
Returns: JSON array of matching decision records.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| run_id | No | ||
| limit | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description must carry the burden. It states that full-text search is performed and returns matching records, but lacks details on behavior such as pagination, sorting, or whether the search is fuzzy or exact.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the purpose and uses a structured docstring format. It is concise but slightly verbose with the 'Args:' and 'Returns:' sections; could be streamlined.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has a low complexity (3 params, 1 required) and an output schema is known to exist. The description covers inputs and return type adequately, though it does not explain output structure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description adds meaningful semantics: explains query with examples ('null ETL bug'), run_id as optional restriction, and limit as max results. This compensates for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Full-text search over the audit decision trail', specifying a verb (search), resource (audit decision trail), and scope. This distinguishes it from siblings like get_run_report or list_runs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives. Usage is implied by the description (searching decision records), but no exclusions or context for selecting among sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.7.0- First observed
get_run_report - First observed
get_trajectory - First observed
list_runs - First observed
load_pipeline - First observed
run_validation - First observed
search_decisions
TDQS
Scored across 6 tools
The distinction between get_run_report and get_trajectory is unclear; both return ShareGPT-formatted trajectories with only a minor difference in metadata inclusion. This overlap causes ambiguity for an agent choosing between them.
All tool names follow a consistent verb_noun pattern with lowercase underscores (e.g., get_run_report, list_runs, run_validation), making them predictable and easy to navigate.
With six tools, the server is well-scoped for a data quality validation workflow. Each tool serves a distinct role without being overwhelming or insufficient.
The core workflow (load pipeline, run validation, list runs, retrieve trajectories, search decisions) is covered. Missing features like run deletion or pipeline management are minor gaps that do not hinder typical use.
Maintenance
Related MCP Connectors
Query your warehouse or a CSV with Claude/ChatGPT over MCP, governed by table-level ACL + audit.
The grounded data layer for any LLM: governed SQL, metrics, lineage and catalog over your data.
Deterministic validation for AI-generated artifacts: JSON Schema, OpenAPI response, SQL syntax.
Deterministic safety, correctness & cost gate that vets Postgres SQL before your AI agent runs it.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceData observability for AI agents. Query alerts, monitor freshness, investigate schema drift, and trace lineage across your data warehouse via 53 MCP tools.2MIT
- AlicenseNot gradedqualityCmaintenanceAI-driven MCP server that audits, profiles, detects schema drift, and auto-generates documentation for dbt projects, enabling natural language interaction with your dbt project's health.134MIT
- AlicenseAqualityAmaintenanceZero-config data quality monitoring as MCP tools. Profiles a warehouse (Postgres, BigQuery, Snowflake, MySQL, DuckDB), detects anomalies, and gates CI — read-only with the connection resolved server-side, never via the model.6305 PyPI10MIT
- AlicenseNot gradedqualityDmaintenanceA production-ready MCP server that provides comprehensive dbt project quality assessment for any GitHub repository, enabling AI agents to analyze dbt models, check metadata coverage, and map data lineage.9MIT