Skip to main content
Glama

propose_dataset_ingestion

Generate review-ready idempotent SQL for ingesting a dataset with its manifest and file map, so a human can validate before applying.

Instructions

Propone el SQL de alta de un dataset real ya organizado con su dataset.yaml y su mapa files (decisión 46, cierra el ciclo que la decisión 44 dejó preparado: manifiesto real -> lector real -> SQL). dataset_dir es una ruta LOCAL en la máquina donde corre este servidor MCP (la misma que ya aloja la biblioteca de datos, sección 23) -- nunca una ruta remota ni relativa a este proceso.

Solo funciona si el format del manifiesto ya es una de las cuatro etiquetas de backend/ingestion/datasets/formats.py::SUPPORTED_FORMATS (nunca unsupported_pending_adapter, que necesita su propio adaptador nuevo primero, ver docs/protocolo-ingesta-ia.md) y si su mapa files cubre todos los roles que ese formato necesita.

NUNCA toca la base de datos real ni la conexión con ella: solo genera y devuelve el SQL de alta (idempotente, ON CONFLICT DO UPDATE, mismo estilo que todos los scripts/register_*.py) para que una persona lo revise antes de aplicarlo -- ninguna automatización de este protocolo sustituye esa revisión humana (decisión 17). study_id/study_name son opcionales pero, si se da uno, hace falta el otro: sin ellos, el atlas se da de alta sin enlazar ningún estudio (nunca uno inventado).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
study_idNo
study_doiNo
study_nameNo
study_yearNo
dataset_dirYes
study_authorsNo
study_journalNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
sqlYes
formatYes
dataset_idYes
dataset_dirYes
region_countYes
network_countYes
coordinate_countYes
membership_countYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It excels: it explicitly states the tool NEVER touches the real database or its connection, only generates idempotent SQL (ON CONFLICT DO UPDATE), requires human review before application (decision 17), and that dataset_dir must be local, never remote. This is rich, safety-relevant behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is verbose (three paragraphs) but front-loaded with the core purpose in the first sentence, and every sentence carries substantive information about constraints or behavior. It could be tightened, but there is no wasted or filler content; the density justifies the length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists (so return values are covered) and the tool is complex (7 params, 0% schema coverage), the description is largely complete: it covers the SQL generation mechanics, idempotency, human-review requirement, format preconditions, and study pairing. The main gap is the four undocumented study metadata parameters, which prevents a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does well for dataset_dir (local path semantics) and the study_id/study_name pairing rule, but says nothing about study_doi, study_year, study_authors, or study_journal. Four of seven parameters remain undocumented in both schema and description, so the compensation is only partial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Propone el SQL de alta de un dataset real' (proposes the registration SQL for a real dataset). It is clearly differentiated from the sibling tools, which are all analysis/rendering tools (render_brain, search_region, get_connectivity), making this the only ingestion/registration tool in the set.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description specifies clear preconditions for use: the manifest format must already be one of the four labels in SUPPORTED_FORMATS (never unsupported_pending_adapter), and the files map must cover all required roles. It also states the study_id/study_name pairing constraint and that dataset_dir must be a local path. It does not name explicit alternative tools, but the siblings are all unrelated analysis tools, so no true alternative exists.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.