Skip to main content
Glama

Data quality scorecard

data_quality_scorecard
Read-onlyIdempotent

How clean is the data a buyer would receive in a state — numbers, not adjectives.

The scorecard grades the records a buyer would receive on mechanical
conformance across four dimensions — format (state/phone/email/zip),
completeness (a name for who filed it, an address present), consistency (names in
CRM-ready Title Case, not ALL-CAPS; the address's own state agrees with its
ZIP), and standardization (how much of the
state's raw status / entity-type vocabulary is mapped into the canonical
cross-state values that `status` / `entity_type` filters match on — an
unmapped row is one a canonical filter silently misses). It returns an
overall 0–100 score, the
per-dimension breakdown, and per-check pass rates with sample offenders you
can click through. Use it to answer "how clean is the data we're selling in
{state}?" and to track data-quality work the way classification is tracked.

The payload also carries a fifth, record-centric **coverage** dimension
(the `coverage` block + `dimensions.coverage`): per pipeline stage, how
many records that brain has NEVER stamped (`gap`), plus a stale-version
count where the brain persists one. Free-chain checks are scored; paid
stages (skip trace / gap-fill / validation / LLM passes) are reported but
unscored — enrichment is spent per order, so an un-enriched set of records
is posture, not a defect. Coverage deliberately does not move the headline
`score`. Each check includes `browse_filters` (a `missing_stage` filter):
the exact set of records works on `browse_leads` and scopes a surgical repair run
on the pipeline trigger. Stages whose brains leave no per-record mark are
listed under `coverage.unmeasured` with reasons rather than pretended into
numbers.

Args:
    state: Two-letter state code (e.g. `FL`, `CO`). Omit to get every state,
        worst score first.
    sample_limit: Max sample offenders to return per check (0–50, default 8).

Returns:
    A scorecard dict for one state, or `{"states": [...]}` for all states.
    Either shape carries a `_meta` provenance block (schema_version,
    freshness, source, score_versions, access_level). Served from a cache
    refreshed in the background: `computed_at` / `age_seconds` date it,
    `stale: true` means the refresh is behind (report the numbers with their
    age), `status: "computing"` means none exists yet. Never call it again for a fresher one.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
stateNo
sample_limitNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent safety, and the description goes well beyond them: cache-backed freshness semantics ('stale: true', 'status: computing'), the rule that paid stages are reported but unscored and do not move the headline score, unmeasured stages listed with reasons, and an explicit 'never call it again for a fresher one' instruction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core question and the dimension list, and the layout is scannable despite the length. It is verbose with bolding and long parentheticals, but nearly every clause conveys non-redundant semantics for a complex tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Complete for a two-parameter read tool: covers input semantics, the caching/provenance behavior, the scoring rubric, and the fact that unscored coverage will not move the number. Nothing an agent needs in order to call and interpret it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must carry the burden — and it does: `state` is specified as a two-letter code with omit-to-return-all (worst first), and `sample_limit` as a 0–50 range defaulting to 8. Both parameters are fully documented in prose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('grades the records a buyer would receive') and enumerates the four dimensions it scores plus the fifth coverage dimension. An agent can distinguish it from browse_leads/quote_list siblings without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly frames the use case ('how clean is the data we're selling in {state}?') and ties it to tracking data-quality work, and it names where browse_filters/the pipeline trigger come into play. It stops short of stating when NOT to use it versus a sibling, so it is clear context rather than full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources