Skip to main content
Glama
ahines99

portco-mcp

by ahines99

Portfolio Company Data Onboarding Agent

CI

Project page and recorded scripted replay · Case study · Sample reports · Release status

Synthetic-data portfolio prototype by Alex Hines, developed with substantial AI assistance. The recorded fixture demo uses automated reviewer decisions. Six real model sessions and a captioned synthetic-voice tour are available. Alex approved finalization; final acceptance records delegated workflow completion, accepted annotations and known scope limits. No required owner actions remain.

Onboards an unfamiliar portfolio company's data into a canonical private-equity data model:

  1. Profiles the source through aggregates and explicitly approved category domains.

  2. Infers entities, keys and joins, and maps columns to the ontology.

  3. Routes anything uncertain to a human.

  4. Generates a tested dbt project and semantic layer.

  5. Executes every generated metric in a disposable sandbox and reconciles monetary amounts to the cent.

  6. Publishes only what a human reviewer certified.

It is exposed to agents (for example Claude Code) through a typed MCP server and a set of Agent Skills. The agent does the discovery; deterministic code does every calculation; humans hold every approval.

uv sync --all-extras && uv run poe demo
== 1. Success path: fixture A (synthetic B2B SaaS company)
   agent profiled, inferred entities/joins and proposed mappings -> paused at 'mapping_review' with 22 review items
   trap caught -> billing.invoice_lines.amount -> invoice_line.amount (suggest cents_to_major) [UNIT_MISMATCH]
   trap caught -> crm.opportunities.rev -> opportunity.amount [LOW_CONFIDENCE, SEMANTIC_TRAP]
   agent tried to approve its own proposals -> Forbidden
   reviewer 'alice' recorded approval a132e212 (22 decisions)
   generated dbt + semantic layer; sandbox build exit 0; 26/26 reconciliation checks passed (money exact to the cent)
   certified (6232731a) and published v0001: active_customers, arr, billings, cogs, ebitda, gross_margin_pct, mrr, opex, revenue_recognized
   audit log: 46 events, hash chain intact

== 2. Controlled failure path: malformed invoice lines + an injected source timeout
   injected timeout during profiling -> retried (attempts: 2), run continued to 'mapping_review'
   sandbox tests failed: non_negative_stg_billing__invoice_lines_amount, non_negative_stg_billing__invoice_lines_quantity
   run stopped at 'test_failures' - nothing is published unless a reviewer waives or the mapping changes

== 3. Prompt injection: instructions planted in a source table comment
   1 comment(s) flagged as untrusted; they are quoted as data, never followed
   mappings identical to the clean run: yes (68 columns)
   planted text in any artifact or finding: NONE

Why this is not just a chatbot

A chatbot would…

This system…

Keep state in the conversation

Persists every run, step attempt, output, evidence record, finding, approval and audit event in a database. Runs pause, resume, retry and rerun idempotently under an execution lease, and the same inputs give identical content hashes.

Read your data to "understand" it

The profiling adapter returns aggregates (counts, ratios, pattern-match counts computed in SQL). Category labels require an operator-approved domain for the exact column; undeclared or unexpected values are withheld (ADR-0003). A PII guard scans MCP responses, and the fixture tests check planted canaries across output surfaces.

Do arithmetic in the prompt

Computes every number in code. Generated dbt marts are reconciled against independent Python reference calculators at zero tolerance, and against the fixture generator's own bookkeeping in golden tests.

Guess when unsure

Routes low-confidence, metric-bearing, PII, conflicting, unit-mismatched and "looks like revenue but isn't" mappings to review. Metrics without evidence are NEEDS_EVIDENCE, never estimated.

Let you type "approved"

Accepts approvals only from reviewer principals, never from the agent, and never from whoever started the run. Each approval is bound to a hash of exactly what was reviewed, and any later change revokes it. Publishing without a valid certification fails closed.

Follow instructions it reads

Treats source text as untrusted data. Instruction-like comments and values are flagged and withheld, and cannot change workflow state (eval case G15).

Be judged by vibes

Has CI gates for 37 golden evaluation cases across seven dimensions, plus 459 non-Postgres tests and four PostgreSQL tests, including security and failure injection. See the dated release evidence for exact scope.

Related MCP server: RecoSearch

Quickstart

Requires Python 3.12 and uv. No Docker and no API keys are needed for the demo, the tests or the evals.

If uv is installed as a Python module but is not on PATH, replace uv below with python -m uv.

uv sync --all-extras          # environment from uv.lock
uv run poe fixtures           # generate synthetic fixture databases into var/fixtures
uv run poe demo               # happy path, controlled failure path, prompt injection (~25 s)
uv run poe test               # full test suite (~3 min; dbt runs in sandboxes)
uv run poe eval               # 37 golden cases; report in evals/reports/latest.md (--snapshot for a dated copy)
uv run poe cov-metrics        # metric reference calculators at 100% branch coverage
uv run poe lint && uv run poe typecheck

Drive it by hand (CLI)

uv run portco run --fixture portco_a                      # agent: runs until the first gate
uv run portco review <run_id> --export review.yaml        # human: see what needs deciding
uv run portco review <run_id> --import review.yaml --reviewer alice --default approve
uv run portco resume <run_id>                             # agent: generate, sandbox-test, stop for certification
uv run portco review <run_id> --export cert.yaml && uv run portco review <run_id> --import cert.yaml --reviewer alice --default approve
uv run portco resume <run_id>                             # publish
uv run portco report <run_id> --out report.md             # run report / certification packet
uv run portco audit <run_id> --verify                     # hash-chain check

Use it from Claude Code (MCP + Skills)

.mcp.json registers the stdio server (uv run portco-mcp) as an agent principal. Copy skills/* into .claude/skills/. See docs/agent_walkthrough.md. Over HTTP, run uvicorn src.mcp_server:app with PORTCO_HTTP_TOKENS set (the app refuses to start without it); each bearer token carries a role and an explicit tenant scope (token=principal:role:company1|company2, * for all). GET /healthz is the unauthenticated liveness probe.

MCP surface

Kind

Name

Notes

Tools

start_onboarding_run, get_run_status, resume_run, list_pending_reviews, profile_schema, propose_canonical_mapping, generate_dbt_artifacts, run_sandbox_tests, publish_run, healthcheck

Agent-callable; typed inputs and outputs; typed error contract

Tools (human)

submit_mapping_review, certify_run

Reviewer principals only

Resources

project://policies, ontology://pe/v1, ontology://pe/v1/metrics/{metric}, run://{run_id}/{summary,profile,entities,joins,mapping,resolved-mapping,test-report,certification-packet,findings,audit,metrics}, run://{run_id}/artifacts/{+path}, evidence://{evidence_id}, finding://{finding_id}/lineage

Tenant-scoped; path traversal rejected

Prompts

onboarding_kickoff, review_run, explain_mapping

Reference resources by URI; never embed data

How it works

connection_validation → schema_profiling → entity_inference → join_inference → canonical_mapping
   → [Gate A: mapping_review] → artifact_generation → automated_tests (sandbox) → [test_failures waiver?]
   → [Gate B: human_certification] → publish
  • Ontology as data: ontology/pe_canonical_v1.yaml defines 15 entities, synonyms, relationships and 18 metrics with PE pitfalls. Scoring weights live in ontology/scoring.yaml.

  • Fixtures with answer keys: fixture A (SaaS with planted traps), fixture B (SAP-style naming), and 10 adversarial or fault variants, each with committed ground truth.

  • Generated dbt: staging (rename, cast, reviewed transforms, PII hashed or dropped, filtered rows flagged), intermediate (cross-entity scoping), marts, MetricFlow semantic models and metrics.

Details: docs/architecture.md · decisions: docs/adr/ · contracts: docs/data_contracts.md · threats: docs/threat_model.md · roadmap: docs/ROADMAP.md · acceptance audit: docs/acceptance.md · portfolio finalization: docs/PORTFOLIO-ROADMAP.md · September audit and fixes: docs/REMEDIATION-2026-09-27.md

Setup, tokens, upgrades and recovery: operations.

Results (at time of writing)

Measure

Result

Golden eval cases

37/37 passing; every dimension at 100%

Mapping top-1 accuracy

1.00 on fixture A (68 columns); fixture B (SAP naming) 34/35 exact, with the extra one flagged for review

Joins

8/8 on fixture A and 4/4 on fixture B, no false positives; orphan rates exact

Metric reconciliation

All nine generated metric definitions executed and reconciled; monetary amounts exact to the cent; 26/26 demo checks including structural and mart checks

PII

10/10 fixture A PII columns classified; zero canary leaks across outputs, artifacts, audit, evidence and published bundles

Limitations and v0.2

  • Sources are DuckDB fixtures. Snowflake, Airbyte and catalog publishing are designed as further capability modules behind the same adapter and policy interfaces.

  • Publish target is a versioned local directory; warehouse deployment needs environment-scoped approvals.

  • The optional LLM mapping judge (PORTCO_LLM_ENABLED=true) is off by default. It can only reorder deterministic candidates or abstain, and it stays opt-in until a recorded live comparison beats the deterministic baseline (ADR-0010).

  • HTTP auth uses static dev bearer tokens; production needs an OAuth/JWT verifier.

  • Postgres-specific tests are separate from the default local suite. Four tests passed against an isolated PostgreSQL 14.24 instance locally and PostgreSQL 16 in hosted CI on 2026-09-27.

  • Hosted CI passed minimal wheel/sdist checks on Linux and Windows, and actual Docker tests for migrations, auth, certified publication, file hashes, audit integrity, MCP restart and full stack recreation with both named volumes preserved. Local Docker is unavailable; the CI run is the proof.

  • An additional unfamiliar-schema probe is reported in full, including missed predictions, in heldout-probe.json. Fixture accuracy is not real-world accuracy.

  • MetricFlow 0.15.0 configuration validation failed in the isolated compatibility experiment; the supported local semantic executor is not a claim of full MetricFlow runtime compatibility.

  • A recorded Claude Code session and the live with-vs-without-Skill comparison are still to do (they need a model session); the static Skill checks run in the test suite. The three-prompt comparison harness records and scores real transcripts once collected; it does not substitute synthetic results for live evidence.

Project layout

src/domain/        contracts, ontology loader, metric reference calculators, policies, PII guard
src/adapters/      DuckDB read-only adapter + SQL guard, repositories, artifact store, fault injection
src/services/      one module per workflow step, approvals, optional LLM judge
src/workflows/     persistent engine, step wiring, facade used by CLI and MCP
src/capabilities/  MCP tools, resources, prompts, PII-guard middleware, auth
skills/            Agent Skills (SKILL.md + references/)
ontology/          canonical PE ontology, abbreviations, scoring weights
templates/dbt/     dbt project templates and generic tests
fixtures/          committed ground truth for generated fixtures
evals/             golden cases, drivers, checks, reports
tests/             unit, integration, golden, security, MCP, snapshot tests

Available Tools

12 tools
certify_runA

Certify the generated bundle and its metrics. Reviewer principals only.

    `subject_hash` is the `subject_hash` of the pending certification items: it binds the approval to the
    exact packet under review (bundle, tests, waivers, open findings). `metric_decisions` maps
    metric -> approve|reject; unknown metric names are rejected, unlisted metrics are approved. Rejected
    metrics (and metrics derived from them) are not published.
    
ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYes
commentNo
subject_hashYes
bundle_decisionNoapprove
metric_decisionsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
gateYes
run_idYes
reviewerYes
decisionsYes
approval_idYes
next_actionYes
subject_hashYes

TDQS

A3.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses the default-approve semantics for unlisted metrics, the rejection of unknown metric names, and the consequence that rejected metrics and their derivatives are not published. It also flags the authorization requirement. These are exactly the behavioral facts an agent needs before invoking a non-idempotent mutation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads action and authorization in the first two sentences, then details the two parameters that carry the real semantics. Dense but every clause adds information; no filler or restatement of the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with an output schema (so return values need no explanation), the description covers the critical decision semantics and side effects. The remaining gap is the unexplained run_id, comment, and bundle_decision parameters, which leaves the argument surface partly opaque.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does well for two of five parameters: subject_hash (binds approval to the exact packet under review) and metric_decisions (map of metric -> approve|reject, with rejection cascade). run_id, comment, and bundle_decision (default 'approve') are never explained, so coverage remains partial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('certify the generated bundle and its metrics') and adds an authorization scope ('Reviewer principals only'). An agent can distinguish it from most siblings, though it never explicitly contrasts with publish_run, which is the nearest neighbor semantically.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'reviewer principals only' clause is a real eligibility constraint, and the metric_decisions paragraph implies usage context (acting on pending certification items). However, there is no explicit when-to-use/when-not guidance and no routing to publish_run or list_pending_reviews, leaving the workflow position to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_dbt_artifactsA

Generate the dbt project and semantic layer from the reviewed mapping.

    Requires `approval_id` from submit_mapping_review; the server verifies it exists, is not revoked
    or expired, belongs to this run, and is bound to the current mapping content (no boolean flags).
    
ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYes
approval_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
filesYes
modelsYes
run_idYes
statusYes
resource_uriYes
manifest_hashYes
not_generatedYes
generated_metricsYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare the generic write profile (readOnlyHint=false, idempotentHint=false); the description goes further by disclosing the server-side gate: the approval must exist, be unrevoked, unexpired, scoped to this run, and content-bound. It does not say what files/artifacts get written, where, or whether existing output is overwritten, which is the main remaining behavioral gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded: the action and its output are stated first, then the precondition. The second sentence is dense but every clause is load-bearing; the parenthetical '(no boolean flags)' is slightly cryptic but adds an anti-pattern warning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A write tool with a real prerequisite is documented adequately, and an output schema exists so return values need not be explained. The only missing pieces are artifact destination/overwrite behavior, which are minor given annotations and output schema coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the two params have only titles, so the description must carry the load. It substantially defines approval_id (provenance and the four validation checks it must pass) but says nothing about run_id beyond the implicit; the coverage of one of two parameters is rich, so this lands above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Generate) and two concrete artifacts (dbt project and semantic layer) plus the input they derive from (the reviewed mapping). An agent can distinguish this from propose_canonical_mapping, submit_mapping_review, or publish_run without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear precondition and names the upstream sibling that produces it: approval_id comes from submit_mapping_review. It does not state when NOT to use the tool (e.g., versus publish_run or run_sandbox_tests as a next step), so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_run_statusC
Read-onlyIdempotent

Current status, step history, gate and pending review items of a run.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
gateYes
errorNo
stepsYes
run_idYes
statusYes
resourcesYes
company_idYes
next_actionYes
current_stepYes
connection_idYes
pending_itemsYes
pending_review_countYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is fully covered. The description adds no behavioral context beyond that — nothing about how fresh the status is, whether it should be polled, or any access constraints — and the returned fields it names are already captured by the output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short fragment with zero filler, and the most important element (current status) is front-loaded. It is efficient, though it reads as a noun phrase rather than a complete statement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema describing the return shape and annotations covering safety, the description does not need to explain returns or side effects. However, it omits any usage routing and any explanation of the one required identifier, leaving gaps an agent must guess at.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one required parameter, run_id, with 0% schema description coverage, and the description does not explain what a run_id is, where it comes from, or how to obtain it. Referencing "a run" is the only implicit connection, which is a weak substitute for real parameter guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource (a run) and enumerates the exact contents returned: current status, step history, gate, and pending review items. It clearly reads as a retrieval tool, though it never explicitly differentiates itself from the overlapping sibling list_pending_reviews.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as list_pending_reviews for review items or resume_run/certify_run for run lifecycle actions. The agent must infer all routing from the name and the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

healthcheckA
Read-onlyIdempotent

Return service health for diagnostics.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
versionYes
databaseYes
schema_versionYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered. The description adds little behavioral detail beyond the return purpose and does not mention auth, rate limits, or return format. It does not contradict annotations, but it mostly restates the obvious.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with no filler, and the core action is front-loaded. It is appropriately sized for a zero-parameter health check.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only diagnostic with no inputs and an output schema, the description is largely sufficient. Annotations cover behavior and the output schema can describe return values, though adding auth or usage prerequisites would make it complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema baseline is 4. There are no parameter semantics for the description to clarify, and schema coverage is already 100%.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Return service health.' It is clear what the tool does and is implicitly distinct from the workflow-oriented siblings. It does not explicitly name an alternative, but no sibling overlaps with its diagnostic purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for diagnostics' gives an implied usage context, suggesting it should be called when checking service health. However, it does not state when-not to use it, prerequisites, or alternatives, so guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_pending_reviewsC
Read-onlyIdempotent

Review items a human must decide before the run can continue.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered by structured data. The description does add domain context — these are approval/decision items that block run progression — which is genuinely useful beyond the annotations, but it discloses nothing about ordering, counts, or what happens after a decision.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, which is structurally fine. However, brevity here comes at the cost of under-specification rather than economy — there is no wasted wording, but not enough wording either.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the annotations cover the safety profile. Even so, for a run-scoped read tool the description omits the run_id scoping and any usage cue, leaving it minimally adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the only parameter is required (run_id), documented solely by its title. The description never mentions run_id or explains that the listing is scoped to one run, so it fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description conveys what the items are — human decisions that gate the run — but never states the action the tool performs (list/return pending review items). It is inferable from the name list_pending_reviews, yet the sentence alone could be read as describing a review workflow rather than a retrieval call. It also never ties the items to a specific run, so differentiation from siblings like get_run_status is only partial.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as get_run_status or submit_mapping_review. The phrase 'before the run can continue' hints at the blocking context but stops short of saying when an agent should call this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

profile_schemaA

Profile a source read-only: row counts, null rates, distinct counts, patterns, PII classes.

    Aggregates only (`sample_rows=0` by construction); values never leave the source adapter.
    Starts a run that pauses after profiling; continue it with propose_canonical_mapping.
    
ParametersJSON Schema
NameRequiredDescriptionDefault
schemasNo
connection_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
run_idYes
statusYes
tablesYes
findingsYes
profile_idYes
resource_uriYes
pii_column_countYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds real behavioral context beyond the annotations: aggregates only, sample_rows=0 by construction, and values never leaving the source adapter — important data-governance behavior. It also discloses the run-state side effect (pauses after profiling). The only friction is the word "read-only", which sits in mild tension with readOnlyHint=false; the description resolves this by explicitly stating it starts a run, but the phrasing could momentarily mislead.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight lines, front-loaded with the purpose, then privacy behavior, then workflow — order is sensible and there is no filler. Minor deduction because `sample_rows=0` references a parameter that does not exist in the input schema, which adds a small decoding cost for the agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. The description covers purpose, data-handling constraints, and the required next step for a run-based tool, which is most of what an agent needs. The remaining gap is the undocumented `schemas` parameter, which is the one thing a caller needs in order to control scope.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and neither of the two parameters (connection_id, schemas) is explained in the description. In particular the optional `schemas` scoping parameter — which lets the caller narrow profiling to specific schemas — is never mentioned, so an agent could miss that it can limit scope.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Profile a source") and enumerates the exact outputs computed — row counts, null rates, distinct counts, patterns, PII classes. It also positions itself in the pipeline by naming propose_canonical_mapping as the follow-up step, which distinguishes it from siblings like generate_dbt_artifacts or run_sandbox_tests.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly says the tool "starts a run that pauses after profiling" and that the agent must "continue it with propose_canonical_mapping", which is actionable sequencing guidance. However, it never states when NOT to use it or what prerequisites exist (e.g., whether an onboarding run must be started first), so it stops short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

propose_canonical_mappingA

Infer entities and joins and propose canonical mappings for a profiled run.

Deterministic scoring; anything uncertain, PII-bearing, metric-bearing or conflicting is routed to human review rather than guessed. Stops at the mapping-review gate.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
joinsYes
run_idYes
statusYes
proposalsYes
row_filtersYes
mapping_hashYes
resource_uriYes
review_itemsYes
auto_acceptedYes
requires_reviewYes
unmapped_requiredYes
metrics_needing_evidenceYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations establish this is a non-read-only, non-idempotent, non-destructive write, so the safety bar is partly met. The description adds real value beyond that: deterministic scoring, routing of uncertain/PII-bearing/metric-bearing/conflicting cases to human review instead of guessing, and halting at the mapping-review gate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, followed by two focused clauses on scoring behavior and the stopping point. Slightly padded by line-break formatting, but every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description covers the action, its determinism policy, and its terminal gate. The remaining gap is the run_id contract and any prerequisite that the run be profiled first, which is only hinted at.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single run_id parameter has no schema description, so the description must carry the burden. 'For a profiled run' adds a useful constraint (the run must already be profiled), but format, source, or expected value of run_id is never explained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States specific verbs (infer, propose) and resources (entities, joins, canonical mappings) tied to a profiled run. The 'stops at the mapping-review gate' clause distinguishes it from the downstream submit_mapping_review, though that routing is implied rather than named as a sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: the tool operates on a profiled run and hands off at the mapping-review gate, so the agent can infer it precedes submit_mapping_review. However, there is no explicit when-to-use, when-not, or named alternative, leaving pipeline positioning to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

publish_runA
DestructiveIdempotent

Publish certified artifacts. Fails closed without a valid certification bound to the bundle.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYes
certification_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
pathYes
reusedYes
run_idYes
statusYes
versionYes
manifest_hashYes
excluded_metricsYes
published_metricsYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and the title notes the certification requirement, so the bar is lower. The description still adds a genuinely new behavioral trait beyond the structured fields: fail-closed behavior on missing/invalid certification. It does not elaborate on what 'destructive' means here (e.g., what prior published state is overwritten), which caps it below 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, purpose first, precondition second, with zero filler. Every clause carries information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be described, and annotations cover safety/idempotency. The description covers purpose, precondition, and failure mode adequately for a two-parameter tool, though it leaves the destructive effect of publishing unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for two required parameters, so the description must compensate. It clarifies that certification_id must reference a valid certification bound to the bundle, which adds real semantics for that parameter, but run_id is never explained and no format/typing detail is given.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (publish certified artifacts) and adds a scope qualifier (certified only). It is discernible from the sibling certify_run by framing this as the downstream publish step, but it never names or contrasts that sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The precondition 'without a valid certification bound to the bundle' implies the tool cannot be used until certify_run has produced a certification, which is useful gating context. However, it never explicitly says when to use this versus certify_run, generate_dbt_artifacts, or the other publishing-adjacent siblings, leaving the routing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resume_runA

Continue a paused or failed run. Gates re-check approvals; completed steps are reused, not recomputed.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
gateYes
errorNo
stepsYes
run_idYes
statusYes
resourcesYes
company_idYes
next_actionYes
current_stepYes
connection_idYes
pending_itemsYes
pending_review_countYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (not read-only, not idempotent, non-destructive), and the description adds real behavior beyond them: gates re-check approvals and completed steps are reused rather than recomputed. It still omits auth/permission requirements and what happens to already-failed steps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the action and followed by the key behavioral guarantee. Every clause earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description conveys the resume semantics well for a simple one-parameter tool. The only material gap is that the run_id parameter is left entirely undocumented.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one parameter with 0% schema description coverage, and the description says nothing about run_id at all. Although the name is fairly self-explanatory, the low-coverage rule requires the description to compensate and it does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (continue/resume) and resource (a run) plus the qualifying state (paused or failed). It implicitly separates itself from start_onboarding_run and get_run_status, but never names a sibling to make the distinction explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear trigger condition: use it for a run that is paused or failed, which contrasts with starting a new run. No explicit when-not guidance or named alternatives are provided, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_sandbox_testsA

Build the generated project with dbt in a disposable sandbox and reconcile metrics exactly.

On success the run pauses at the certification gate; on failure at the test-failures gate.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
cachedYes
passedYes
run_idYes
statusYes
data_testsYes
resource_uriYes
dbt_exit_codeYes
waived_checksYes
failing_checksYes
reconciliationYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the safety profile (not read-only, not destructive, not idempotent, closed-world), so the bar is lower. The description adds real value beyond them: it runs in a disposable sandbox and pauses at a specific gate depending on outcome, which is meaningful state-transition context for a non-idempotent tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly worded sentences with zero filler. The core action is front-loaded and the outcome-gate behavior follows as supporting detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and annotations cover safety. The description conveys the sandbox behavior and gate outcomes, but omits prerequisites (e.g. whether dbt artifacts must be generated first), leaving a small gap for a stateful pipeline step.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description says nothing about run_id. However, it is a single, self-explanatory identifier, so the practical risk of misuse is low. The description adds no parameter meaning, which caps this below the baseline-4 case for zero-parameter tools.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource combination: 'Build the generated project with dbt in a disposable sandbox and reconcile metrics.' This clearly separates it from siblings like generate_dbt_artifacts and certify_run. It stops short of explicitly naming which sibling it is not, so a 5 isn't warranted.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the mention of the certification gate (success) and test-failures gate (failure), which places the tool in a pipeline, but it never states when to call it versus alternatives or what prerequisites must exist first. No explicit when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_onboarding_runA

Start read-only onboarding of a registered source (e.g. fixture:portco_a).

    Runs connection validation, profiling, entity and join inference and canonical mapping,
    then stops at the first human review gate. Returns the run status and pending review items.
    
ParametersJSON Schema
NameRequiredDescriptionDefault
company_idNo
connection_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
gateYes
errorNo
stepsYes
run_idYes
statusYes
resourcesYes
company_idYes
next_actionYes
current_stepYes
connection_idYes
pending_itemsYes
pending_review_countYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the safety profile (not read-only, not destructive, not idempotent, closed world), so the description's added value is the workflow disclosure: which phases run, where execution halts, and what is returned. The phrase 'read-only onboarding' sits in tension with readOnlyHint=false, though it reads as describing source-side non-mutation (consistent with destructiveHint=false) rather than the tool's own side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short lines, front-loaded with the action and followed by execution scope and return behavior; no filler. Slightly marred by the ambiguous 'read-only' qualifier, which is the one phrase that could mislead.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a stateful pipeline-start tool in an eleven-tool suite, the description covers execution and terminal state well, and an output schema exists so return details needn't be repeated. However it omits prerequisites, permission requirements, and how it relates to resume_run / list_pending_reviews, which an agent needs to sequence calls correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for both parameters, so the description carries the full burden, yet only 'connection_id' is implied via the 'registered source (e.g. fixture:portco_a)' example. The optional 'company_id' parameter is never mentioned or explained, leaving its scoping role undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Start ... onboarding of a registered source') and enumerates the pipeline phases it executes (validation, profiling, entity/join inference, canonical mapping). It also distinguishes itself from the sibling 'resume_run' by framing this as the initial start that halts at the first human review gate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The stopping condition ('stops at the first human review gate') implies the natural follow-up (list_pending_reviews / resume_run), but no explicit when-to-use or when-not-to-use guidance is given and no alternative is named. An agent must infer the relationship to the ten sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

submit_mapping_reviewA

Record a human reviewer's decisions for the run's current gate. Reviewer principals only.

    Decisions: approve | reject | approve_with_override (override keys: canonical_field, transform,
    pii_handling). The approval is bound to the content hash of the items under review.
    
ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYes
commentNo
decisionsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
gateYes
run_idYes
reviewerYes
decisionsYes
approval_idYes
next_actionYes
subject_hashYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (non-read-only, non-idempotent, non-destructive), and the description adds genuinely new context: reviewer-principal authorization and the fact that approval is bound to the content hash of items under review (implying stale approvals are rejected). It stops short of describing retry/idempotency behavior or what happens to prior decisions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Short and front-loaded: purpose sentence first, then eligibility and decision vocabulary. The embedded line breaks and indentation are slightly noisy but no sentence is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists, so return values need no explanation, and annotations carry the mutation safety profile. For a 3-param mutation tool the description covers the essential decision vocabulary and the content-hash binding, leaving only minor gaps around run state prerequisites and array completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does document the decision values (approve | reject | approve_with_override) and the override keys, which mirrors the sparse schema description. But run_id and comment semantics, and whether the decisions array must cover every pending item, remain undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Record a human reviewer's decisions for the run's current gate', with a clear scope (current gate of a run). It is distinguishable from list_pending_reviews and certify_run by its reviewer-decision focus, though it does not name a sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Reviewer principals only' gives an eligibility constraint, and the enumerated decisions imply what the call is for. However, there is no explicit when-to-use guidance versus adjacent tools such as certify_run or resume_run, nor an indication of what state the run must be in before submission.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 12 tool updatesv0.1.0
    • First observedcertify_run
    • First observedgenerate_dbt_artifacts
    • First observedget_run_status
    • First observedhealthcheck
    • First observedlist_pending_reviews
    • First observedprofile_schema
    • First observedpropose_canonical_mapping
    • First observedpublish_run
    • First observedresume_run
    • First observedrun_sandbox_tests
    • First observedstart_onboarding_run
    • First observedsubmit_mapping_review

TDQS

A3.6/5.0

Scored across 12 tools

Disambiguation4/5

Most tools map to distinct pipeline stages (profile, propose, review, generate, test, certify, publish), so an agent can largely tell them apart. However, start_onboarding_run and profile_schema both 'start a run' with profiling, and get_run_status overlaps list_pending_reviews on surfacing pending review items, creating minor boundary ambiguity.

Naming Consistency4/5

Tools follow a clear snake_case verb_noun pattern (start_onboarding_run, get_run_status, resume_run, submit_mapping_review, certify_run, publish_run, etc.). The lone exception is healthcheck, a bare noun that breaks the otherwise consistent convention.

Tool Count5/5

12 tools is well-scoped for a governed data-onboarding pipeline with multiple human gates. Each tool corresponds to a meaningful lifecycle stage or control operation, with no obvious redundancy bloating the surface.

Completeness4/5

The surface covers the full happy-path lifecycle: start, status, resume, profile, propose mapping, submit review, generate artifacts, sandbox test, certify, and publish, plus review listing and healthcheck. Missing abort/cancel and any run-listing or cleanup operations are minor gaps agents can work around.

Maintenance

ActivityNo data
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    B
    maintenance
    A read-only MCP server that exposes dbt project artifacts and data quality result tables (BigQuery/Postgres) to LLM clients, enabling deep introspection, run-history analysis, source freshness, test coverage, and lineage walks.
    27
    36 npm
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    A deterministic MCP server that governs read-only queries across multiple data sources, returning answers with full provenance (every row cited) or a typed refusal, ensuring LLM answers are traceable and contract-enforced.
    1
    Apache 2.0
  • A
    license
    Not graded
    quality
    C
    maintenance
    Local-first code intelligence and safety layer for AI coding agents. MCP server exposes dependency graph, impact analysis, and AST-compressed repo context, backed by typed local memory, patch-scope safety gates, and git-independent transaction rollback.
    1
    MIT
  • A
    license
    B
    quality
    B
    maintenance
    A local, evidence-driven MCP runtime and control plane for open-source maintainers that provides workspace-bounded tools including controlled file operations, command execution, validation primitives, durable execution records, and human review workflows via stdio and Streamable HTTP transports.
    33
    MIT