Skip to main content
Glama

Onboard a new data source, step 2: run analysis after uploading

finish_data_source_onboarding
Idempotent

Call after uploading the file(s) returned by onboard_data_source — kicks off AI analysis and waits until the spec reaches "ready" or "failed". If it returns before that (timedOut: true), do NOT call this tool again just to keep checking — that re-attempts starting analysis. Poll with get_status (specId) instead until it reaches a terminal status.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
specIdYesspecId returned by onboard_data_source.
specNameYesName of the data spec being onboarded.
workspaceIdNoWorkspace to act on. Defaults to your only workspace if you have exactly one.
loadSampleDataNoWhether to load the sample file and trigger the data-load job once analysis finishes (default true).

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
errorNo
specIdNo
statusNo"processing" | "ready" | "failed"
messageYes
progressNo
timedOutNo
lastJobIdNo
workspaceIdNo
errorDetailsNo
statusMessageNo
analysisEndTimeNo
analysisStartTimeNo
analysisDurationMsNo
hasTransformationConfigNo

TDQS

A3.9/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description directly contradicts the annotation idempotentHint: true. It warns 'do NOT call this tool again just to keep checking — that re-attempts starting analysis,' implying repeated calls have a side effect and do not return a cached/terminal result. This undermines the idempotency annotation and is a serious consistency problem.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three focused sentences and immediately gives the required trigger, the behavior, and the critical warning. Every sentence earns its place and the most important actionable information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with a rich input schema and output schema, the description covers the necessary operational context: after which step to call it, what it waits for, what happens on timeout, why re-calling is wrong, and what to use instead. No essential behavior is left unexplained for the agent to invoke it safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and each parameter already has a meaningful description (e.g., 'specId returned by onboard_data_source', 'loadSampleData ... default true'). The tool description adds some usage context around the parameters but does not need to compensate for missing schema info.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's specific role: 'Call after uploading the file(s) returned by onboard_data_source — kicks off AI analysis and waits until the spec reaches "ready" or "failed"'. This distinguishes it from the sibling onboarding step and from status-polling tools like get_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to call it ('after uploading the file(s) returned by onboard_data_source') and provides an explicit alternative: 'Poll with get_status (specId) instead until it reaches a terminal status.' It also warns against repeated calls, making the usage boundary very clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

A4.2/5.0
Disambiguation4/5

Most tools map to distinct lifecycle phases and the descriptions explicitly separate overlapping-sounding concepts, such as list_data versus submit_query and the generic call_dpf_api from dedicated tools. The three finish_* tools are similarly worded but each is clearly tied to a specific preceding operation, so confusion should be limited.

Naming Consistency4/5

The tool names are uniformly snake_case and mostly follow a readable verb_noun pattern like delete_data_spec, create_workspace, and run_data_job. It is not a perfect 5 because broader names like manage_connection and manage_trigger, the generic call_dpf_api, and list_my_workspaces with its pronoun make the naming pattern less predictable.

Tool Count4/5

At 16 tools, the set is just slightly above the ideal range, but the tools generally earn their place by representing distinct steps or workflow boundaries. The start/finish pairs create some apparent redundancy, but that is a natural consequence of the multi-step file-upload flow.

Completeness4/5

The toolset provides solid coverage of the core data-platform lifecycle: workspaces, data specs, jobs, connections, triggers, scheduled pulls, status polling, and SQL querying. Some additional DPF capabilities are only reachable through the generic call_dpf_api rather than dedicated tools, and billing mutations are explicitly left outside the MCP surface, so coverage is strong but not absolute.

Resources