Data Quality Gate - deterministic post-scrape cleaner + verdict
Server Details
Post-scrape data cleaner, no LLM: repairs mojibake, HTML, invisible chars. Plus a verdict.
/.well-known/glama.json file. Claimed server authors can inspect health checks, view analytics, and manage their connector listing.- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP
- URL
Glama MCP Gateway
Connect through Glama MCP Gateway for full control over tool access and complete visibility into every call.
Full call logging
Every tool call is logged with complete inputs and outputs, so you can debug issues and audit what your agents are doing.
Tool access control
Enable or disable individual tools per connector, so you decide what your agents can and cannot do.
Managed credentials
Glama handles OAuth flows, token storage, and automatic rotation, so credentials never expire on your clients.
Usage analytics
See which tools your agents call, how often, and when, so you can understand usage patterns and catch anomalies.
Tool Definition Quality
Average 4.6/5 across 3 of 3 tools scored.
check_dataset_quality is clearly distinct, but clean_scraped_data and clean_scraped_data_audited are nearly identical in behavior, differing only in the paid endpoint and audit trail. An agent would be uncertain which to call, making the boundaries between the two cleaning tools unclear.
All tools follow a consistent verb_noun pattern in snake_case (check_dataset_quality, clean_scraped_data, clean_scraped_data_audited), with the last adding an '-audited' modifier. There is no mixing of conventions or irregular naming.
Three tools is a reasonable number for a focused quality-gate server, placing it within the typical 3-15 range. However, two of the three are near-duplicates, reducing effective diversity, so it is slightly padded rather than perfectly scoped.
The server allows quality checking and defect inventory, but the actual cleaning is delegated to external paid endpoints, so an agent cannot complete a cleanup task within the MCP framework. Missing a tool to actually retrieve or apply cleaned data creates a significant gap.
Available Tools
3 toolscheck_dataset_qualityAInspect
Call this before using any dataset. Returns a deterministic quality verdict (RELIABLE / USABLE_WITH_CLEANING / UNRELIABLE) with exact facts: completeness, nulls, type consistency, impossible values, duplicates, outliers, and (on financial/trading data) cross-source price divergence. 100% deterministic, no LLM. Free -- this MCP endpoint runs the engine directly; POST /api (plain REST, same engine) is x402-gated at $0.01/call instead. Input: rawJson (a JSON array of row objects, or a single object); datasetId is accepted but not resolvable on this deployment -- pass rawJson instead.
| Name | Required | Description | Default |
|---|---|---|---|
| rawJson | No | The dataset: a JSON array of row objects, or a single object. | |
| datasetId | No | An Apify dataset id. Not resolvable on this deployment; pass rawJson instead. |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses important behavioral traits: 100% determinism, no LLM, freeness of this endpoint, and the limitation that datasetId is not resolvable. It also lists the exact output facts and conditional cross-source price divergence, providing a detailed behavioral model.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with a clear call-to-action and efficiently lists output facts. The cost comparison with the POST /api endpoint adds useful context but also extra length. Overall, it is well-organized but could be slightly tightened without losing essential information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the tool's purpose, input format, output details, determinism, and the datasetId limitation. However, it does not explicitly state whether rawJson is required (the schema lists 0 required parameters), nor does it describe error behavior or what happens when both parameters are absent. Given no output schema, these are minor gaps in an otherwise complete description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The tool description largely repeats the schema's parameter details (rawJson as JSON array or object; datasetId not resolvable) without adding new syntax, validation, or usage nuances. The only addition is a slight emphasis on rawJson, but it does not go beyond what the schema already states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool checks dataset quality and returns a deterministic verdict (RELIABLE / USABLE_WITH_CLEANING / UNRELIABLE) with specific quality facts. It distinguishes itself from the alternative paid POST /api endpoint, making the purpose and scope unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs to call before using any dataset, and provides clear alternatives: the paid POST /api endpoint and the recommendation to pass rawJson instead of datasetId. This gives strong contextual guidance for when and how to use the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clean_scraped_dataAInspect
PAID TIER, CALLED FROM A FREE ENDPOINT -- READ THIS FIRST: this tool does NOT return your cleaned data, and its response has NO data key. It returns {"status": "payment_required", "what_we_found_in_your_data": {...}, "get_the_repaired_data": {...}}. Do not treat its output as cleaned data and do not write it back into your dataset.
WHAT YOU GET HERE, FREE: an exact inventory of the defects in YOUR data -- which repair rules would fire, how many values each would change, and which cases cannot be repaired safely and need your decision.
HOW TO GET THE REPAIRED DATA: make the one call named in get_the_repaired_data -- POST https://www.aidatatools.dev/api/clean, $0.04 via x402, no account, no API key, no signup. That response body IS the cleaned dataset, in the shape you posted it.
WHY THE SPLIT: detection is free on this endpoint and always has been (check_dataset_quality reports the same defects). The repaired artifact is the paid product, because it is re-bought on every extraction run rather than cached like a verdict.
WHAT THE PAID CALL DOES: removes leftover HTML tags and entities, decodes mojibake ('Café' -> 'Café'), strips invisible characters (zero-width, BOM, soft hyphen), normalises non-breaking spaces and trims values -- across nested objects and arrays too. 100% deterministic, no LLM: the same input always yields byte-identical output, and cleaning twice equals cleaning once. It repairs how data was ENCODED, never what it SAYS: masked placeholders ('N/A', 'None'), near-duplicate rows and failed extractions ('access denied', 'captcha', which mean that record must be re-scraped) are reported with a proposal, never silently deleted or rewritten. The full boundary -- 7 rules applied automatically, 5 needing an explicit opt-in, 8 only ever reported -- is at GET https://www.aidatatools.dev/api/clean.
| Name | Required | Description | Default |
|---|---|---|---|
| options | No | All optional. Every default is the safe one: with no options, the row count, every value's type, and the schema are all guaranteed unchanged. | |
| rawJson | Yes | The scraper output: a JSON array of row objects, a single object, or a CSV/plain-text string. The format is detected and the output mirrors the shape you sent. |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and excels: it discloses the response shape (no data key), the deterministic/no-LLM nature, the explicit list of repairs (HTML tags, mojibake, invisible chars), and the boundary of what is repaired vs. reported rather than silently changed. It also reveals cost/authentication requirements for the paid endpoint, which is valuable behavioral context beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Though long, the description is densely informative and structured via headers and a prominent warning. Every section ('WHAT YOU GET HERE', 'HOW TO GET THE REPAIRED DATA', 'PROVENANCE', 'PAID CALL BEHAVIOR') earns its place, and the front-loaded warning prevents misuse. The structure aids scannability despite the length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema, but the description fully compensates: it explains the exact output structure, the fields returned, the calling flow for the paid service, the deterministic nature, and the repair/report boundary. It also preempts common mistakes (e.g., writing output back into the dataset) and provides a URL for the full rules. This is a complete picture for an agent to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters (rawJson and options with detailed defaults/enums). The main description does not add parameter-level semantics beyond what the schema provides; it focuses on tool behavior. This meets the baseline 3 for high schema coverage, but does not exceed it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states the tool does NOT return cleaned data, instead returning an inventory of defects and instructions for the paid repair endpoint. The verb 'returns' plus the precise resource ('exact inventory of the defects in YOUR data') makes the purpose unambiguous, and the contrast with the paid call distinguishes it from sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly explains when to use this free endpoint (detection, no account needed) versus the paid call for repaired data, and explicitly names check_dataset_quality as reporting the same defects. It also provides a concrete alternative path ('make the one call named in get_the_repaired_data') and a rationale for the split.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clean_scraped_data_auditedAInspect
PAID TIER, CALLED FROM A FREE ENDPOINT -- READ THIS FIRST: this tool does NOT return your cleaned data, and its response has NO data key. It returns {"status": "payment_required", "what_we_found_in_your_data": {...}, "get_the_repaired_data": {...}}. Do not treat its output as cleaned data and do not write it back into your dataset.
WHAT YOU GET HERE, FREE: an exact inventory of the defects in YOUR data -- which repair rules would fire, how many values each would change, and which cases cannot be repaired safely and need your decision.
HOW TO GET THE REPAIRED DATA: make the one call named in get_the_repaired_data -- POST https://www.aidatatools.dev/api/clean/audit, $0.12 via x402, no account, no API key, no signup. That response body IS the cleaned dataset, in the shape you posted it.
WHY THE SPLIT: detection is free on this endpoint and always has been (check_dataset_quality reports the same defects). The repaired artifact is the paid product, because it is re-bought on every extraction run rather than cached like a verdict.
WHAT THE PAID CALL DOES: the same repair as clean_scraped_data, plus a complete audit trail: every transformation with its path, rule, before and after value, a replay_id, and input/output SHA-256. The ledger is a full inverse patch -- applying it in reverse reconstructs your original input byte for byte. Use it when you must be able to PROVE later what changed and why.
| Name | Required | Description | Default |
|---|---|---|---|
| options | No | All optional. Every default is the safe one: with no options, the row count, every value's type, and the schema are all guaranteed unchanged. | |
| rawJson | Yes | The scraper output: a JSON array of row objects, a single object, or a CSV/plain-text string. The format is detected and the output mirrors the shape you sent. |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully discloses critical behavior: response has no 'data' key, returns a payment_required status, and should not be written back to the dataset. It explains what the free tier returns, how to obtain the repaired data, and the audit trail details, including reversibility via inverse patch.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is lengthy but well-structured with clear section headers (WHAT YOU GET HERE FREE, HOW TO GET THE REPAIRED DATA, etc.). The critical warning is front-loaded. Some redundancy exists ('READ THIS FIRST' repeated), but each section adds necessary information for a paid tool with complex behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite no output schema, the description fully covers the response shape and important side effects, explains the payment and endpoint for the paid call, and provides a complete behavioral contract. Given the tool's complexity and the absence of annotations, the description leaves no critical questions unanswered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema already documents each parameter thoroughly with examples and caveats. The description does not add additional parameter-level semantics beyond what the schema provides, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states that this tool does NOT return cleaned data but instead returns an inventory of defects, including which repair rules would fire and which cases need decisions. It differentiates from sibling clean_scraped_data by emphasizing the audit/preview nature and the paid repair endpoint.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly provides when to use this tool: 'Use it when you must be able to PROVE later what changed and why.' It also contrasts with clean_scraped_data (same repair but without audit trail) and check_dataset_quality (same defects reported free). The 'WHY THE SPLIT' section gives clear context on the free vs paid distinction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Claim this connector by publishing a /.well-known/glama.json file on your server's domain with the following structure:
{
"$schema": "https://glama.ai/mcp/schemas/connector.json",
"maintainers": [{ "email": "your-email@example.com" }]
}The email address must match the email associated with your Glama account. Once published, Glama will automatically detect and verify the file within a few minutes.
Control your server's listing on Glama, including description and metadata
Access analytics and receive server usage reports
Get monitoring and health status updates for your server
Feature your server to boost visibility and reach more users
For users:
Full audit trail – every tool call is logged with inputs and outputs for compliance and debugging
Granular tool control – enable or disable individual tools per connector to limit what your AI agents can do
Centralized credential management – store and rotate API keys and OAuth tokens in one place
Change alerts – get notified when a connector changes its schema, adds or removes tools, or updates tool definitions, so nothing breaks silently
For server owners:
Proven adoption – public usage metrics on your listing show real-world traction and build trust with prospective users
Tool-level analytics – see which tools are being used most, helping you prioritize development and documentation
Direct user feedback – users can report issues and suggest improvements through the listing, giving you a channel you would not have otherwise
The connector status is unhealthy when Glama is unable to successfully connect to the server. This can happen for several reasons:
The server is experiencing an outage
The URL of the server is wrong
Credentials required to access the server are missing or invalid
If you are the owner of this MCP connector and would like to make modifications to the listing, including providing test credentials for accessing the server, please contact support@glama.ai.
Discussions
No comments yet. Be the first to start the discussion!
Related MCP Servers
- AlicenseAqualityBmaintenanceCleans HTML from URLs or raw HTML into LLM-ready text, reducing token burn for AI agents.376MIT
- AlicenseAqualityAmaintenanceWeb scraping, crawling, and structured data extraction for AI agents. 5 tools: scrape (clean markdown from any URL), crawl (entire sites), map (discover URLs), extract (structured JSON), and search. 833ms avg latency, single binary, self-hostable.8852AGPL 3.0
- AlicenseBqualityFmaintenanceWeb scraper and token compressor that converts HTML to clean markdown with 70-80% fewer tokens. Single-page compression and multi-page BFS crawling with auto-fallback fetch modes.2241MIT
- AlicenseAqualityBmaintenanceDeterministic fluff detector for AI-generated prose. No model, no API key.3MIT