Skip to main content
Glama

Data Quality Gate - deterministic post-scrape cleaner + verdict

clean_scraped_data_audited

PAID TIER, CALLED FROM A FREE ENDPOINT -- READ THIS FIRST: this tool does NOT return your cleaned data, and its response has NO data key. It returns {"status": "payment_required", "what_we_found_in_your_data": {...}, "get_the_repaired_data": {...}}. Do not treat its output as cleaned data and do not write it back into your dataset. WHAT YOU GET HERE, FREE: an exact inventory of the defects in YOUR data -- which repair rules would fire, how many values each would change, and which cases cannot be repaired safely and need your decision. HOW TO GET THE REPAIRED DATA: make the one call named in get_the_repaired_data -- POST https://www.aidatatools.dev/api/clean/audit, $0.12 via x402, no account, no API key, no signup. That response body IS the cleaned dataset, in the shape you posted it. WHY THE SPLIT: detection is free on this endpoint and always has been (check_dataset_quality reports the same defects). The repaired artifact is the paid product, because it is re-bought on every extraction run rather than cached like a verdict. WHAT THE PAID CALL DOES: the same repair as clean_scraped_data, plus a complete audit trail: every transformation with its path, rule, before and after value, a replay_id, and input/output SHA-256. The ledger is a full inverse patch -- applying it in reverse reconstructs your original input byte for byte. Use it when you must be able to PROVE later what changed and why.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
optionsNoAll optional. Every default is the safe one: with no options, the row count, every value's type, and the schema are all guaranteed unchanged.
rawJsonYesThe scraper output: a JSON array of row objects, a single object, or a CSV/plain-text string. The format is detected and the output mirrors the shape you sent.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description fully discloses critical behavior: response has no 'data' key, returns a payment_required status, and should not be written back to the dataset. It explains what the free tier returns, how to obtain the repaired data, and the audit trail details, including reversibility via inverse patch.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is lengthy but well-structured with clear section headers (WHAT YOU GET HERE FREE, HOW TO GET THE REPAIRED DATA, etc.). The critical warning is front-loaded. Some redundancy exists ('READ THIS FIRST' repeated), but each section adds necessary information for a paid tool with complex behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description fully covers the response shape and important side effects, explains the payment and endpoint for the paid call, and provides a complete behavioral contract. Given the tool's complexity and the absence of annotations, the description leaves no critical questions unanswered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already documents each parameter thoroughly with examples and caveats. The description does not add additional parameter-level semantics beyond what the schema provides, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that this tool does NOT return cleaned data but instead returns an inventory of defects, including which repair rules would fire and which cases need decisions. It differentiates from sibling clean_scraped_data by emphasizing the audit/preview nature and the paid repair endpoint.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly provides when to use this tool: 'Use it when you must be able to PROVE later what changed and why.' It also contrasts with clean_scraped_data (same repair but without audit trail) and check_dataset_quality (same defects reported free). The 'WHY THE SPLIT' section gives clear context on the free vs paid distinction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

A4.2/5.0
Disambiguation2/5

check_dataset_quality is clearly distinct, but clean_scraped_data and clean_scraped_data_audited are nearly identical in behavior, differing only in the paid endpoint and audit trail. An agent would be uncertain which to call, making the boundaries between the two cleaning tools unclear.

Naming Consistency5/5

All tools follow a consistent verb_noun pattern in snake_case (check_dataset_quality, clean_scraped_data, clean_scraped_data_audited), with the last adding an '-audited' modifier. There is no mixing of conventions or irregular naming.

Tool Count4/5

Three tools is a reasonable number for a focused quality-gate server, placing it within the typical 3-15 range. However, two of the three are near-duplicates, reducing effective diversity, so it is slightly padded rather than perfectly scoped.

Completeness2/5

The server allows quality checking and defect inventory, but the actual cleaning is delegated to external paid endpoints, so an agent cannot complete a cleanup task within the MCP framework. Missing a tool to actually retrieve or apply cleaned data creates a significant gap.

Resources