Skip to main content
Glama
kewinall

Data Platform MCP Server

by kewinall

Data Platform MCP Server

目前版本:v0.5.0

互動式架構與專案總覽
GitHub Pages · Repository HTML

這是一套 production-oriented Tool / Integration Layer,將 PostgreSQL、Vertica、Airflow、logs、metadata、lineage 與 normalized ETL metadata,透過受治理的 Model Context Protocol(MCP)tools 提供給 AI Agent、IDE 與 MCP Host。

預設使用 synthetic demo data,不需要公司 / 客戶資料、真實 credential 或 internal URL。

專案定位

本 Repository 負責 Portfolio 中的 Canonical MCP Access Layer:

  • MCP protocol transport 與 tool contracts

  • PostgreSQL / Vertica catalog access

  • read-only SQL guardrails

  • Airflow 與 log integrations

  • metadata 與 lineage access

  • OIDC/JWT authentication

  • RBAC 與 per-tool scopes

  • tenant-aware source isolation

  • structured audit events

  • OpenTelemetry observability

  • 消費 enterprise-etl-platform 產生的 normalized ETL metadata

本專案刻意不實作 autonomous agent orchestration、RAG chat、model routing,或 ETL parsing / migration truth。

Related MCP server: DataLakeHouseMCP

架構

ChatGPT / Claude / Codex / MCP Host
                 |
                 | MCP Streamable HTTP
                 v
        +----------------------+
        | MCP SDK Bearer Gate  |
        | Static / OIDC JWT    |
        +----------+-----------+
                   |
             Role / Scope
                   |
             Tenant Policy
                   |
             Audit + OTel
                   |
        +----------v-----------+
        | DataPlatformService  |
        +-----+-----------+----+
              |           |
       Catalog layer   DataOps layer
              |           |
       +------+-----+     +----------------+
       |            |     |                |
 PostgreSQL      Vertica Airflow       Logs/Runbooks

核心能力

  • MCP SDK v2 / Streamable HTTP

  • PostgreSQL + Vertica catalog adapters

  • multi-source catalog abstraction

  • metadata 與 table / projection inspection

  • catalog-backed 與 SQL AST lineage

  • SQLGlot-based read-only SQL policy

  • database read-only session enforcement

  • Airflow 3 /api/v2 integration

  • OpenSearch / Loki log search

  • OIDC/JWT + JWKS validation

  • roles:reader、analyst、operator、admin

  • per-tool scopes

  • tenant-aware source isolation

  • structured audit JSONL

  • Kubernetes Helm deployment

  • NetworkPolicy / PDB / optional HPA

  • External Secrets examples

  • OpenTelemetry traces + metrics

  • air-gapped bundle support

ETL Metadata Contract

enterprise-etl-platform 是 ETL metadata 與 lineage truth 的 producer;本 Server 只消費已發布的 normalized artifacts。

Enterprise ETL Platform
  parser / migration / lineage truth
              |
              | normalized artifact
              v
Data Platform MCP Server
  governed read-only access
              |
              v
AI Agent / IDE / MCP Host

MCP layer 會保留 producer 的 structural、inferred-deterministic 與 capability boundary 等分類,不會自行補出缺失的 lineage。

關鍵工程決策

決策

原因 / 效益

Trade-off

MCP Tool Contract 作為 integration boundary

Client 使用受治理 capability,而不是直接取得 raw backend credentials / APIs

增加 protocol / schema compatibility 維護成本

Adapter Pattern 隔離 platform backends

Tool contract 穩定,backend 可獨立替換

Backend-specific behavior 仍需在 adapter 層維護

Read-only by design + AST policy + DB session

安全控制位於 prompt 之下,形成 defense in depth

保守 policy 可能拒絕部分複雜但合法的 SQL

RBAC 決定 what,tenant policy 決定 which source

operation authorization 與 data-source isolation 分離

Source-level tenant isolation 不等同 row-level RLS

OIDC/JWT 在 resource boundary 驗證

集中處理 identity 與 role mapping

IdP / JWKS availability 成為 auth dependency

Producer-owned ETL metadata contract

避免 parser / lineage truth 被複製成第二套

需要維護 schema / version compatibility

失敗語意與復原原則

  • malformed 或 mutation-capable SQL 在 database execution 前直接拒絕

  • cross-tenant source access 直接 deny,不 fallback 到 unrestricted access

  • OIDC / JWKS 驗證失敗時 fail closed

  • backend outage 回傳 bounded tool failure,不升高權限

  • invalid ETL metadata artifact fail closed

  • lineage 缺失時維持缺失,不由 MCP Server fabricate edge

可驗證 Evidence

Claim

Repository Evidence

OIDC / RBAC / audit regression

tests/test_auth_audit.py, tests/test_security.py, src/data_platform_mcp/security.py

MCP protocol contract

tests/test_mcp_protocol.py, tests/test_service.py

PostgreSQL / Vertica boundary

tests/test_vertica.py, tests/test_integrations.py

Observability / audit correlation

tests/test_observability.py, src/data_platform_mcp/observability.py, src/data_platform_mcp/audit.py

Kubernetes baseline

deploy/helm/data-platform-mcp-server/, tests/helm-values.yaml, .github/workflows/ci.yml

Security gate

.github/workflows/security.yml

ETL metadata contract

src/data_platform_mcp/adapters/etl_metadata.py, tests/test_etl_metadata.py

ETL MCP integration

tests/test_mcp_protocol.py, docs/etl-metadata-integration.md

快速開始

git clone https://github.com/kewinall/data-platform-mcp-server.git
cd data-platform-mcp-server
python -m venv .venv
source .venv/bin/activate
pip install -e '.[dev]'
cp .env.example .env

stdio:

DPMCP_TRANSPORT=stdio data-platform-mcp

HTTP:

DPMCP_TRANSPORT=streamable-http data-platform-mcp

Endpoint:http://127.0.0.1:8000/mcp

主要 Tools

Tool

Scope

用途

health

platform:read

health / version

whoami

platform:read

identity、role、tenant、scopes

list_data_sources

catalog:read

tenant-visible sources

describe_table

catalog:read

column metadata

get_table_lineage

lineage:read

catalog lineage

analyze_sql_lineage

lineage:read

SQL AST lineage

explain_sql

sql:explain

guarded EXPLAIN

list_dags

operations:read

Airflow DAG discovery

search_etl_logs

logs:read

ETL log search

list_etl_pipelines

catalog:read

normalized ETL pipeline discovery

get_etl_table_lineage

lineage:read

producer lineage + capability boundary

工程文件

docs/ 內包含 OIDC、MCP protocol、ETL metadata integration、deployment、observability、security 與 operational guidance。

Portfolio 責任邊界

  • Data Platform MCP Server:standardized governed tool / integration access

  • Enterprise ETL Platform:ETL metadata producer、migration truth、lineage truth

  • Agentic DataOps Copilot:operational reasoning client

  • Enterprise RAG Platform:enterprise knowledge 與 retrieval

  • Multi-LLM AI Gateway:model control plane

授權

MIT

Available Tools

21 tools
analyze_sql_lineageA

Parse read-only SQL and extract referenced input tables and CTEs.

ParametersJSON Schema
NameRequiredDescriptionDefault
sqlYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description must disclose behavioral traits itself. The verb 'parse' clearly indicates a non-mutating, analytical operation, and 'read-only SQL' suggests the input is not executed. This conveys safety and purpose beyond the schema, though it does not detail error handling or dialect support, which is acceptable for a simple tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. It states the purpose and the input constraint in a compact way, earning every word.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has a single parameter and an output schema (not shown but noted in context), so the description need not explain return values. It covers the core functionality and the read-only constraint, leaving only minor gaps like error behavior or supported SQL dialects, which are not critical for a parsing tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds a meaningful constraint by specifying 'read-only SQL', which informs the agent about acceptable input. However, it does not provide examples, dialect limitations, or format details, so it only partially compensates for the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action (parse) on a resource (read-only SQL) with a defined outcome (extract referenced input tables and CTEs). This is sufficiently distinct from siblings like get_table_lineage (which likely works on a catalog) and explain_sql (which explains query plans).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a clear context: use when you have a SQL string and need its input tables/CTEs. It implies a static analysis use case without explicitly naming alternatives or exclusions, but the 'read-only SQL' qualifier gives enough context for an agent to select it over catalog-based lineage tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

describe_tableC

Describe table columns and data types.

ParametersJSON Schema
NameRequiredDescriptionDefault
tableYes
schemaYes
sourceYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries the full burden. 'Describe' implies a read-only metadata operation, but the description does not disclose error behavior, permissions, or any other operational traits. It only restates the basic purpose.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler or repetition. Every word contributes to stating the tool's purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has three required parameters with no descriptions, no annotations, and no usage context. The description is too thin to fully support correct invocation, especially for understanding how to specify a source. The output schema helps, but does not compensate for missing parameter and routing guidance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not explain what 'source' means or how it relates to schema and table. The parameter names are self-explanatory to some degree, but the description adds almost no semantic value beyond the schema fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the operation ('Describe') and the resource ('table columns and data types'). It is easy to understand what the tool does, though it does not explicitly differentiate itself from siblings like list_tables or table_statistics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives such as list_tables, table_statistics, or list_schemas. It does not mention prerequisites, intended context, or when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

explain_sqlB

Explain tenant-authorized read-only SQL after SQLGlot AST validation.

ParametersJSON Schema
NameRequiredDescriptionDefault
sqlYes
sourceYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the load and does disclose important behavioral traits: the operation is read-only and tenant-authorized, and SQLGlot AST validation happens first. It does not mention error behavior or whether the query is executed/parsed, but the core safety and pipeline behavior are stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is tight and front-loads the main action, with each qualifier adding information. It is concise, though brief to the point of omitting parameter and usage detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema exists, so return format is covered, and the description covers read-only/auth. But with undefined 'source' and no usage differentiation from sibling tools, an agent cannot reliably select and invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Both parameters are bare strings with zero schema descriptions. 'sql' is self-explanatory, but 'source' is ambiguous (data source ID, dialect, or something else) and the description does not clarify it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a clear verb ('Explain') with a specific resource ('SQL') and adds meaningful qualifiers: tenant-authorized, read-only, and post-validation. It does not specify what kind of explanation is produced or how it differs from analyze_sql_lineage, so it is not fully differentiating.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides implied use conditions: the SQL must be tenant-authorized, read-only, and processed through SQLGlot AST validation. There is no explicit statement of when to prefer this tool over analyze_sql_lineage or other siblings, and no exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_dag_statusB

Return the latest known state of a DAG.

ParametersJSON Schema
NameRequiredDescriptionDefault
dag_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It does communicate that the result is the 'latest known' state, hinting at possible staleness, but it does not state whether the tool errors on unknown DAG IDs, whether it is read-only, or any rate/freshness behavior. The one-line description is insufficient for transparent behavioral expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no redundant wording. It states the core action and object efficiently, earning its place without filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple (one parameter) and has an output schema, so return values need not be described. However, the description lacks any usage context, alternatives, or error/edge-case behavior, and the sole parameter is undocumented. It is minimally viable but has clear gaps for an agent deciding when to call it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage for 'dag_id', and the description does not explain the parameter beyond implying it identifies a DAG. However, the parameter name 'dag_id' is quite self-explanatory, and there is only one required parameter, so the meaning is reasonably recoverable though not explicitly documented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb ('Return'), a resource ('a DAG'), and what is returned ('the latest known state'). It distinguishes itself from sibling tools like 'list_dags' by focusing on a single DAG's status rather than listing DAGs, though it does not explicitly name the distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as 'list_dags', 'health', or 'search_etl_logs'. The phrase 'Return the latest known state' weakly implies querying a specific DAG's status, but there is no explicit context, prerequisite, or exclusion to steer an agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_etl_pipelineC

Return the producer-owned normalized ETL metadata document.

ParametersJSON Schema
NameRequiredDescriptionDefault
pipeline_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden. The verb 'Return' weakly implies a read‑only operation, but the description does not explicitly declare read-only status, failed-handling, permission requirements, rate‑limits, or any other behavioral trait. This is a bare minimum and offers no real transparency for an unannotated tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that wastes no words; it clearly states the main action and object. It is front-loaded and easy to parse. It could have included more useful detail, but that would not negatively affect conciseness; therefore it is an effective but minimal structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema, so the return shape is covered. However, the description does not establish why this tool is the correct one among the many list_* and get_* inputs in the sibling set. It also does not define the scope of the retrieved document (e.g., whether it includes steps, dependencies, lineage) making the tool's capabilities ambiguous. The context is therefore not complete for reliable tool selection and invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage, and the description adds no meaning to pipeline_id. It does not explain what a pipeline_id is, whose identifier it is, the format, or how to obtain it. With such low schema coverage, the description must compensate completely, but it fails to tell the agent how to populate the only required parameter correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the action (Return) and the resource ('the producer-owned normalized ETL metadata document'), giving a specific enough purpose to distinguish it from immediate siblings like get_etl_pipeline_steps or get_etl_table_lineage. However, 'producer-owned' is domain-specific and not explicitly put in contrast with other pipeline tools, so it is clear but not fully self-contained.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to choose this tool over list_etl_pipelines, get_etl_pipeline_steps, get_etl_pipeline_dependencies, or other related tools. The description alone does not say which scenario calls for the full metadata document rather than a subset or a search result, so the agent is left with incomplete routing information.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_etl_pipeline_dependenciesC

Return producer-classified ETL step/workflow dependencies.

ParametersJSON Schema
NameRequiredDescriptionDefault
pipeline_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the behavioral burden. It only says 'Return' and gives no detail on read-only safety, failure modes, required permissions, or how dependencies are classified; the vague 'producer-classified' qualifier adds minimal behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence with no filler; the action verb is front-loaded. It is efficient but arguably too terse to carry the explanatory weight needed elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read tool with an output schema, the description is minimally usable: the agent can infer to pass pipeline_id. But it lacks usage guidance and any explanation of 'producer-classified', so it is not complete enough to confidently distinguish from closely related pipeline tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description never explains pipeline_id or how it relates to the returned dependencies. The description adds no parameter meaning beyond the bare parameter name and type.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb and resource: return dependencies of an ETL pipeline, qualified as producer-classified step/workflow dependencies. It does not explicitly contrast with sibling tools like get_etl_pipeline_steps or get_etl_table_lineage, so it misses the differentiation needed for a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance about when to use this tool versus siblings such as get_etl_pipeline_steps or get_etl_table_lineage. There are no prerequisites, exclusions, or alternative-tool routing in the description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_etl_pipeline_stepsB

Return deterministic ETL step metadata for one pipeline.

ParametersJSON Schema
NameRequiredDescriptionDefault
pipeline_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden. It does disclose a useful behavioral trait ('deterministic') and implies a read-only metadata operation, but it omits error behavior, permissions, and any side-effect guarantees.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. The key scoping word ('one pipeline') appears at the end but the sentence remains efficient and informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema available and a single parameter, the description is mostly adequate, but it lacks guidance on pipeline identification and error cases)Skip and does not clarify the relationship to sibling ETL pipeline tools. Adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description adds almost nothing beyond the schema's title. It confirms the operation targets 'one pipeline,' but does not explain pipeline_id semantics, source, or expected format.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Return') plus a precise object ('deterministic ETL step metadata') and scope ('for one pipeline'). This clearly distinguishes it from sibling tools like get_etl_pipeline and get_etl_pipeline_dependencies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool versus alternatives, how to obtain a valid pipeline_id, or what conditions make this the right choice. The description states what it does but not when to call it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_etl_table_lineageC

Return producer-owned ETL lineage edges and capability boundaries.

ParametersJSON Schema
NameRequiredDescriptionDefault
tableYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description must carry the behavioral transparency burden. It indicates the return content (producer-owned lineage edges, capability boundaries) but does not disclose side effects, authorization needs, what 'producer-owned' means, or what 'capability boundaries' entail. The output schema exists, but behavior is under-described.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no redundancy. Every phrase ('producer-owned', 'ETL lineage edges', 'capability boundaries') contributes information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations and one un-documented parameter, the description is too thin. An agent needs to know what 'producer-owned' means, why this differs from get_table_lineage, and what 'capability boundaries' refers to. The output schema covers return values, but usage context is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description does not explain the lone table parameter. It only reuses the word 'table' from the schema property name, adding no meaning about expected format, ownership scope, or relationship to ETL pipelines.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('get') and resource ('ETL lineage edges and capability boundaries'), and includes the scoping qualifier 'producer-owned' which distinguishes it from the sibling get_table_lineage. However, 'capability boundaries' is jargon that is not explained, and the short phrase doesn't fully disambiguate from closely related lineage tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool instead of alternatives like get_table_lineage, get_etl_pipeline, or get_etl_pipeline_dependencies. The description provides no exclusions, prerequisites, or recommended contexts.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_table_lineageC

Return catalog-backed upstream lineage within the tenant source boundary.

ParametersJSON Schema
NameRequiredDescriptionDefault
tableYes
schemaYes
sourceYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose meaningful behavioral traits: the lineage is catalog-backed and stays within the tenant source boundary, which hint at safe, read-only behavior. However, it doesn't mention permissions, edge cases, or what happens when no lineage exists, leaving some behavioral ambiguity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The entire description is a single focused sentence. It leads with the action and object, then adds a scope qualifier. There is no filler, redundancy, or unnecessary detail – exactly the right size for a straightforward endpoint.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having an output schema, the description lacks essential usage context: no parameter semantics, no alternative routing, and no annotation safety signals. An agent cannot confidently infer how to construct a valid request or how this tool differs from related lineage tools. It feels incomplete for a tool with three required parameters and no schema documentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description contains no parameter explanations. Parameters 'source', 'schema', and 'table' are only bare names; 'source' especially is ambiguous and could refer to a data source or lineage source. The description fails to clarify how these parameters should be populated or how they relate to the lineage lookup.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb with a precise object: 'Return catalog-backed upstream lineage.' It clearly communicates a read operation that yields upstream lineage data from the catalog, and 'within the tenant source boundary' adds a meaningful scope. While it doesn't explicitly name sibling tools, 'catalog-backed' helps differentiate it from SQL-based lineage tools like analyze_sql_lineage.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no direct guidance about when to use this tool versus alternatives. It implies lineage is read from the catalog rather than computed from SQL, but it never states exclusions or recommends another tool for SQL-based lineage. The tenant boundary qualifier is a constraint, not a usage rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_table_metadataC

Return governed metadata for a tenant-authorized table or view.

ParametersJSON Schema
NameRequiredDescriptionDefault
tableYes
schemaYes
sourceYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are absent, so the description must disclose behavior. It only says 'Return', implying read-only, but doesn't mention authentication, error handling, or whether it accesses only tenant-authorized data. Lacks detail on side effects or requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with no fluff, front-loading the action and resource. It's appropriately concise but sacrifices content for brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having an output schema, the description is incomplete for a tool with three undocumented parameters. An agent cannot determine valid inputs or usage context. The description alone is insufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description provides no parameter details. Source, schema, and table are left undefined, forcing the agent to guess format or meaning. The description adds zero value for parameter understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the action (Return) and resource (governed metadata for a table or view). The term 'governed metadata' differentiates it from siblings like describe_table (schema) and table_statistics (stats), making purpose distinct.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus describe_table, table_statistics, or get_table_lineage. The description implies use when governed metadata is needed but offers no explicit alternatives or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

healthA

Return MCP server health and version.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It indicates a read-only status check, which is useful behavioral information, but it does not disclose error behavior, latency, or availability semantics. For a simple health tool this is acceptable but minimal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, concise sentence that states exactly what the tool returns. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With zero parameters and an output schema available, the description is complete for an agent to understand and invoke this tool correctly. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there is nothing to document beyond the schema. The description correctly omits parameter details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Return') and identifies the exact resource ('MCP server health and version'). It is unambiguous and clearly distinguishable from sibling tools, which are all data and schema related.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There are no alternative health-related siblings, so exclusion guidance is not needed. The description makes the usage context clear: call this tool to check server health and version, likely before other operations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_dagsA

List ETL orchestration DAGs known to the server.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral burden. It implies a safe, read-only operation via the word 'List' and scopes results to DAGs 'known to the server.' However, it does not mention potential pagination, authentication requirements, or any edge-case behavior; for a simple no-parameter list tool, this level is adequate but not thorough.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-formed sentence with no unnecessary words. It front-loads the action ('List') and immediately specifies the object and scope.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has no parameters, an output schema exists, and the operation is a simple list, the description is largely complete. It could be slightly more complete by mentioning whether results are paginated or limited, but that may be covered by the output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there is nothing for the description to add beyond the schema. Per the rubric, 0 parameters baselines at 4, and the description says nothing contradictory or redundant about parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') and a specific resource ('ETL orchestration DAGs known to the server'), making the tool's purpose clear. It also distinguishes itself from sibling tools like get_dag_status, which focuses on status rather than enumeration.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not explicitly state when to use this tool versus alternatives or when not to use it. The phrase 'known to the server' implies a discovery/enumeration use case, and siblings like get_dag_status and search_etl_logs signal obvious alternatives, but no explicit guidance is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_data_sourcesA

List data sources visible to the current tenant.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full burden of behavioral disclosure. The description only states the action and scope, without indicating whether the operation is read-only, any performance implications, or side effects. It does not disclose the return format or any potential pagination, leaving the agent to guess at behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence with no redundancy. It front-loads the action ('List') and includes a useful scope qualifier, making it efficient and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has no parameters, an output schema exists, and the description clearly states the scope, it is reasonably complete for a simple list operation. However, it omits any caveats like pagination or sorting, but these may not be necessary for such a basic tool. The presence of the output schema compensates for lack of return-value details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema already covers all (100% coverage). According to guidelines, the baseline for 0 parameters is 4. The description adds no parameter-related information, but none is needed since there are no parameters to clarify.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action ('List') and resource ('data sources') with a scope qualifier ('visible to the current tenant'), making it distinct from sibling tools like list_schemas or list_tables. It precisely communicates what the tool does without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus alternatives. The description does not mention any prerequisites, conditions, or compare it to other list tools, leaving the agent to infer usage solely from the resource name. No 'when to use' or 'when not to use' is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_etl_pipelinesA

List normalized ETL pipeline IDs published by the ETL metadata producer.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. 'List ... IDs' clearly indicates a read-only operation, and 'published by the ETL metadata producer' scopes the source. However, it does not disclose pagination, ordering, or whether the list is exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, well-structured sentence that leads with the action and object, then adds the qualifiers 'normalized' and 'published by the ETL metadata producer.' No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter list tool with an output schema, the description covers the essential purpose and source of results. It does not explicitly distinguish from siblings or mention return-shape caveats, but the simple read-only nature makes this largely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters and 100% schema coverage, so there is no parameter behavior to document. The description correctly implies an unqualified list operation, making the empty schema unsurprising.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') and identifies the exact resource returned ('normalized ETL pipeline IDs published by the ETL metadata producer'). This clearly distinguishes it from siblings like get_etl_pipeline, list_dags, and search_etl_metadata.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance is given about when to use this tool versus alternatives such as get_etl_pipeline, list_dags, or search_etl_metadata. The intended use is implied by the name and list semantics, but not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_schemasB

List schemas in a configured and tenant-authorized data source.

ParametersJSON Schema
NameRequiredDescriptionDefault
sourceYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full behavioral disclosure. It mentions 'configured and tenant-authorized' but does not state whether the operation is read-only, what errors might occur, or any side effects. The behavior is minimally described.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence with no superfluous words. The core action and resource are front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

While an output schema exists and covers return values, the description does not adequately explain the input parameter or usage context. For a simple one-parameter tool, it lacks essential detail on how to specify 'source' and any prerequisites beyond a vague mention of configuration and authorization.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has no description for the 'source' parameter (0% coverage). The description hints at 'data source' but does not explicitly define what 'source' accepts, its format, or examples. It fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (list) and resource (schemas) and adds context ('configured and tenant-authorized data source'). It is specific enough to distinguish from siblings like list_tables and list_data_sources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives like list_tables or list_data_sources. The description only states the action, leaving usage entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_tablesB

List tables or views in a tenant-authorized schema.

ParametersJSON Schema
NameRequiredDescriptionDefault
schemaYes
sourceYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses that the listing is scoped to a 'tenant-authorized schema,' which is a useful behavioral constraint. However, it does not mention edge cases such as empty results, errors for non-existent schemas, or any limits on the number of returned items. Since this is a read-only listing operation, the absence of side-effect warnings is acceptable, but the description could be more explicit about read-only behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. It conveys the core action and scope efficiently. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

While an output schema exists, reducing the need to describe return values, the description is incomplete for a tool with two required parameters that are unexplained. It also does not clarify how 'schema' and 'source' relate to each other or to the tenant context. An agent would need external knowledge or experimentation to call this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description provides no explanation for either parameter. 'schema' is only mentioned in the scope phrase, but not defined as a parameter. 'source' is completely unexplained. The description fails to compensate for the schema's lack of parameter details, leaving the agent to guess what values are expected.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('List') and resource ('tables or views') within a defined scope ('tenant-authorized schema'). It clearly distinguishes from siblings like list_schemas (lists schemas) and list_data_sources, and even mentions that views are included, which is not obvious from the name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives. The description does not mention alternatives like get_table_metadata, describe_table, or list_schemas, nor does it state conditions that would favor one tool over another. An agent would have to infer usage from the name and sibling context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_etl_logsC

Search the configured read-only ETL log backend.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral disclosure burden. It does mention 'read-only', which is a useful safety signal, but it omits query behavior, error handling, rate limits, and what 'configured' backend means. This is minimal disclosure for a tool with no annotation support.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single tight sentence that front-loads the verb and resource with no wasted words. It is not a tautology and is easily parsed, though it is sparse in content—an issue better captured by other dimensions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter tool with no annotations and a sibling search tool, the description is barely more than a label. It leaves out query format, limit semantics, relationship to search_runbooks, and any details about what the output schema provides. An agent could identify the tool's purpose but would be guessing on invocation details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but neither 'query' nor 'limit' is mentioned. The word 'Search' weakly implies query is a search string, but there is no explanation of query syntax or limit's default behavior. The description completely fails to make up for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Search') and a clear resource ('configured read-only ETL log backend'), making the tool's function clear. It does not explicitly differentiate from sibling search_runbooks, but the resource distinction is strong enough that an agent can infer the difference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided about when to use this tool versus alternatives. The sibling search_runbooks also performs a search, but the description neither contrasts them nor states conditions for choosing one. Only a weak implied context (when you need ETL logs) exists.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_etl_metadataC

Search ETL pipeline, step, table, and dependency names.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are absent, so the description carries full responsibility for disclosing behavior. It only says 'search' without detailing case sensitivity, partial matching, result limits, pagination, or the response structure. There is no mention of auth requirements, rate limits, or any side effects. The description is too sparse to inform an agent about operational behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no redundant words, which is commendable for conciseness. However, it is so brief that it sacrifices substance. Still, for a one-line definition, it is well-structured and front-loaded with the core action and target.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (2 parameters, no nested objects) and the existence of an output schema, the description is still incomplete. It does not explain what the search returns, how results are ordered, or any limitations. For a search tool with many siblings, an agent would need more context to use it correctly, but the description provides minimal information.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does not explain that 'query' is the search string or that 'limit' caps the number of results. While these are somewhat intuitive, the description adds no additional meaning beyond the parameter names themselves, leaving interpretation to the agent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (search) and the target resources (ETL pipeline, step, table, and dependency names). This distinguishes it from sibling tools that list or get specific entities, and from search_etl_logs and search_runbooks which target different domains. The verb and object are explicit and specific.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. With many sibling tools like get_etl_pipeline, list_tables, and search_etl_logs, an agent would benefit from knowing that this tool is for name-based search across metadata entities, but the description does not mention any conditions, exclusions, or alternative selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_runbooksC

Search operational runbooks and troubleshooting knowledge.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure, but it only says 'Search'. It does not mention whether results are paginated, sorted, limited, or how the limit parameter affects behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence with no fluff, and it is front-loaded with the action verb. However, the conciseness comes at the cost of missing usage and parameter guidance, making it adequate but not exemplary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple and an output schema exists, so return values need not be described. Still, the description lacks behavioral details, parameter semantics, and differentiation from sibling tools, leaving noticeable gaps for an agent deciding whether to invoke it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it adds no explanation for the 'query' or 'limit' parameters. The noun 'runbooks' weakly implies what 'query' searches, but the optional 'limit' behavior is completely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb ('Search') and a specific resource ('operational runbooks and troubleshooting knowledge'), which tells an agent what this tool does. It does not explicitly distinguish it from search_etl_logs, but the resource is distinct enough that the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives like search_etl_logs, nor any exclusions or prerequisites. The only hint is the searchable resource itself, which is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

table_statisticsA

Return lightweight table statistics without scanning business rows.

ParametersJSON Schema
NameRequiredDescriptionDefault
tableYes
schemaYes
sourceYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral disclosure. It reveals that the tool does not scan business rows, implying efficiency and perhaps approximate or metadata-based statistics. However, it does not mention whether the statistics are up-to-date, if permissions are required, or what happens when statistics are unavailable. The description offers one useful behavioral trait but lacks depth.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that immediately states the purpose and the key differentiator. There is no fluff, and the most important information is front-loaded. Every word contributes to the tool's understanding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (three obvious parameters) and the presence of an output schema that likely details the returned statistics, the description is adequate but minimal. It does not mention potential caveats like staleness or permissions, but for a simple statistics tool, the coverage is sufficient, though not comprehensive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description provides no additional meaning for the three parameters (source, schema, table). While the names are self-explanatory, the description does not clarify any constraints, defaults, or format expectations, leaving the agent to infer everything from the parameter names alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'return' and the resource 'table statistics', and the qualifier 'without scanning business rows' differentiates it from heavier operations like full scans. This effectively distinguishes it from sibling tools like get_table_metadata or describe_table, which likely involve more detailed metadata or schema inspection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies a lightweight, non-scanning use case, but it does not explicitly state when to prefer this tool over alternatives like get_table_metadata or describe_table, nor does it mention any exclusions. The 'lightweight' and 'without scanning' hints provide context but fall short of explicit guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

whoamiA

Return the authenticated principal, tenant, and effective RBAC scopes.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It states the tool returns identity information, which implies a read-only operation, but it does not explicitly say it has no side effects, require authentication (though 'authenticated' is mentioned), or disclose error behavior. It provides minimal but sufficient transparency for a simple query tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, focused sentence that front-loads the verb and clearly states the returned data. It has no filler or redundant wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is extremely simple (no parameters) and has an output schema, so the description does not need to explain return values. It fully states what the tool does, making it complete for an agent to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there is nothing to explain. The description need not add parameter details; the baseline of 4 applies because there are no parameters for the schema to cover.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Return') and a specific resource ('authenticated principal, tenant, and effective RBAC scopes'). It clearly distinguishes from all sibling tools, which focus on data assets, DAGs, and logs, none of which relate to identity or access.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not explicitly state when to use this tool versus alternatives, but its purpose is so distinct that usage is implied. There is no mention of context such as 'call this at session start' or 'use this before making other calls to verify permissions,' leaving the agent to infer the appropriate timing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 6 tool updatesv0.5.0
    • Addedget_etl_pipeline
    • Addedget_etl_pipeline_dependencies
    • Addedget_etl_pipeline_steps
    • Addedget_etl_table_lineage
    • Addedlist_etl_pipelines
    • Addedsearch_etl_metadata
  2. 4 tool updatesv0.4.0
    • Addedanalyze_sql_lineage
    • Addedget_table_lineage
    • Addedget_table_metadata
    • Addedwhoami
  3. 11 tool updatesv0.2.0
    • First observeddescribe_table
    • First observedexplain_sql
    • First observedget_dag_status
    • First observedhealth
    • First observedlist_dags
    • First observedlist_data_sources
    • First observedlist_schemas
    • First observedlist_tables
    • First observedsearch_etl_logs
    • First observedsearch_runbooks
    • First observedtable_statistics

TDQS

B3.3/5.0

Scored across 21 tools

Disambiguation4/5

Tools generally target distinct resources or aspects: data source/schema/table discovery, ETL pipeline metadata, lineage, SQL analysis, and observability. The closest overlap is among get_etl_table_lineage, get_table_lineage, and analyze_sql_lineage, but their producer/catalog/SQL scopes are described enough to avoid frequent misselection.

Naming Consistency4/5

Most tools follow a consistent snake_case verb_noun pattern (list_*, get_*, search_*, describe_*). Minor deviations are health and whoami, which are bare commands, and table_statistics, which lacks a get_ prefix.

Tool Count4/5

21 tools is slightly above the ideal 3-15 range, but the server spans several coherent subdomains: metadata discovery, ETL, lineage, SQL analysis, and operations. Each tool maps to a distinct operation, so the count does not feel padded.

Completeness4/5

For a governed, read-only metadata server the surface covers listing/discovery, table details, ETL lifecycle metadata, lineage, DAG status, logs, and runbooks. Minor gaps such as no dedicated getters for individual data sources/schemas and no downstream lineage are workaround-able but would improve full coverage.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables AI agents to query live schema, lineage, and query-context across data warehouses, dbt projects, orchestration systems, and BI tools via MCP tools.
    Apache 2.0
  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables AI-powered MCP clients to interact with data lakehouse components including Kafka, Flink, and Trino/Iceberg for managing topics, jobs, catalogs, and executing queries.
    2
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    MCP server providing read-only Snowflake metadata tools (schemas, tables, queries, lineage) for agentic data pipeline generation, enabling natural-language-to-pipeline workflows with dbt, Airflow, and Great Expectations.
    1
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables enterprise AI agents to query governed data lineage, PII-aware schema documentation, and semantic metadata from SQL logs via MCP, with role-based access and vector search.
    MIT