Skip to main content
Glama

data.clean

Read-onlyIdempotent

Clean CSV, TSV, JSON, JSON Lines, or YAML records with collision-safe header normalization, deterministic integrity hashes, and exact transformation counts.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
formatYesSource format; use jsonl for JSON Lines or NDJSON
contentYes
trim_stringsNo
empty_to_nullNo
normalize_headersNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
dataYesStructured Dataset cleaning result
metaYes
serviceYes
versionYes
request_idYesUnique request identifier

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish read-only, idempotent, non-destructive behavior. The description adds specific behavioral details beyond the annotations, such as collision-safe header normalization, deterministic integrity hashes, and exact transformation counts, giving agents a clearer picture of what the operation does and returns. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that front-loads the core purpose and lists supported formats and key features without redundancy. It is dense but not bloated; a slight structural breakdown could improve scannability, but it remains concise and effective.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema exists, so the description does not need to explain return values. It covers supported formats and the main transformation behaviors, while the schema handles limits like maxLength and the enum for format. The main gap is insufficient explanation of parameter semantics, but this is partly mitigated by the output schema and self-explanatory parameter names.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20% (only 'format' is described). The description mentions 'collision-safe header normalization,' which partially explains normalize_headers, but trim_strings and empty_to_null are left entirely to their names. There is no clarification of how empty strings become null or how trimming behavior works, so parameter semantics remain under-specified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'clean' with a clear resource (CSV/TSV/JSON/JSONL/YAML records) and elaborates on what cleaning entails: collision-safe header normalization, deterministic integrity hashes, and exact transformation counts. This distinguishes it from sibling tools like data.convert and data.deduplicate, which have different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use or alternative guidance is provided. The intent is implied by the purpose, suggesting use when records need standardization or transformation metrics, but it does not mention when not to use it or how it compares to closely related tools such as data.deduplicate or data.json-repair.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

A3.8/5.0
Disambiguation4/5

Tools are grouped into clear domain prefixes (crypto, data, developer, document, research, web) and each tool name describes a specific function; however, a few umbrella tools like web.full-audit and data.contract overlap with their more targeted counterparts, creating minor ambiguity.

Naming Consistency5/5

All tool names follow a consistent pattern: a domain prefix, a dot, and a hyphenated lowercase compound name (e.g., crypto.base-block-inspect, web.seo-audit). This makes naming predictable and easy to scan.

Tool Count1/5

At 63 tools, the surface area is very large and exceeds the 50+ threshold for extreme mismatch. While the tools are organized into six domains, the sheer number makes it difficult for an agent to select efficiently, and some tools are bundled combinations of others.

Completeness5/5

Each domain offers a thorough set of operations: crypto covers address, account, block, contract, events, gas, and transaction inspection; data covers cleaning, conversion, schema, and validation; developer covers code review, dependency/license audits, and test generation; research covers SEC, OFAC, GLEIF, and USAspending; web covers extraction, SEO, security, and performance. No obvious dead ends exist for the read-only/inspection purpose.

Resources