Skip to main content
Glama
ahines99

portco-mcp

by ahines99

profile_schema

Profile a data source read-only to get row counts, null rates, distinct counts, patterns, and PII classes from aggregates only, then pause the run for mapping review.

Instructions

Profile a source read-only: row counts, null rates, distinct counts, patterns, PII classes.

    Aggregates only (`sample_rows=0` by construction); values never leave the source adapter.
    Starts a run that pauses after profiling; continue it with propose_canonical_mapping.
    

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
schemasNo
connection_idYes

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
run_idYes
statusYes
tablesYes
findingsYes
profile_idYes
resource_uriYes
pii_column_countYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds real behavioral context beyond the annotations: aggregates only, sample_rows=0 by construction, and values never leaving the source adapter — important data-governance behavior. It also discloses the run-state side effect (pauses after profiling). The only friction is the word "read-only", which sits in mild tension with readOnlyHint=false; the description resolves this by explicitly stating it starts a run, but the phrasing could momentarily mislead.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight lines, front-loaded with the purpose, then privacy behavior, then workflow — order is sensible and there is no filler. Minor deduction because `sample_rows=0` references a parameter that does not exist in the input schema, which adds a small decoding cost for the agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. The description covers purpose, data-handling constraints, and the required next step for a run-based tool, which is most of what an agent needs. The remaining gap is the undocumented `schemas` parameter, which is the one thing a caller needs in order to control scope.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and neither of the two parameters (connection_id, schemas) is explained in the description. In particular the optional `schemas` scoping parameter — which lets the caller narrow profiling to specific schemas — is never mentioned, so an agent could miss that it can limit scope.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Profile a source") and enumerates the exact outputs computed — row counts, null rates, distinct counts, patterns, PII classes. It also positions itself in the pipeline by naming propose_canonical_mapping as the follow-up step, which distinguishes it from siblings like generate_dbt_artifacts or run_sandbox_tests.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly says the tool "starts a run that pauses after profiling" and that the agent must "continue it with propose_canonical_mapping", which is actionable sequencing guidance. However, it never states when NOT to use it or what prerequisites exist (e.g., whether an onboarding run must be started first), so it stops short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.