hydra-etl
This MCP server lets an AI assistant author, inspect, validate and run Hydra ETL declarative pipelines — in plain language, with validation before anything is written or run.
Explore the building blocks: list the 18 transformation operations, list connectors accepted by sources/destinations, list workflow actions, and get the full JSON schema (parameters, types, defaults) of any single operation.
Discover the workspace: list existing jobs (folders with a
pipeline.yaml), list workflows, and read the raw YAML of an existing job or workflow before editing it.Learn from real examples: find the closest existing jobs to a request and return their manifests, to avoid reconstructing from memory.
Inspect data first: preview columns and first rows of a CSV, JSON or Parquet file — before writing a job (to get real column names) and after a run (to check results).
Write jobs safely: create the four manifests (sources, destinations, pipeline, optional transformations) as YAML, validated before writing — if validation fails, nothing is written and errors are returned.
Write workflows: author
workflow.yamlfiles with job or action steps,depends_onlists, parallel branches and retries — also written only after validation.Validate: run the official validator on a job for manifest structure and
pipeline.from/pipeline.toresolution; this is the ground truth for whether a job will run.Verify intent: check a written job against the original user request (12 deterministic rules — missing requested operation, numeric comparison without
cast, contradicted load mode, mismatched file extension, plaintext secret, unsupported technology).Explain failures: translate validation errors into plain language and propose a fix.
Execute: run a job or a workflow (returning logs / step-by-step trace) — but only when the user explicitly asks, since it writes real data.
Safety by design: only the write and run tools are non-read-only; reads and validations are idempotent, and destructive run tools are clearly flagged.
Allows creating Hydra ETL pipelines that read from or write to MariaDB databases.
Allows creating Hydra ETL pipelines that read from or write to MongoDB databases.
Allows creating Hydra ETL pipelines that read from or write to MySQL databases.
Allows creating Hydra ETL pipelines that read from or write to PostgreSQL databases.
Hydra ETL
Your data pipelines are described, not programmed.
Website · Docs & playground · MCP server · VS Code extension · Changelog
Hydra is an open-source declarative ETL engine. You write what a pipeline is in YAML;
Hydra validates it before touching any data, then runs it. The same manifests run from
the terminal, a REST API, or a visual editor in your browser, with one pip install and nothing else to deploy.
Try it without installing anything
Two sandboxes run Hydra in your browser. Nothing is installed on your machine, and nothing is left behind when you close the tab.
Guided labs — no account needed
Killercoda gives you a Linux terminal inside your browser, next to a short lesson that tells you exactly what to type. Each lab takes ten to fifteen minutes and checks your work as you go. There is nothing to sign up for.
Start with Meet hdrctl to see what a job is, then Parameters and secrets to run one pipeline against two databases.
A full environment, with the visual editor
GitHub Codespaces starts a real machine in the cloud and opens an editor in your browser. It needs a free GitHub account, and it runs within GitHub's free monthly allowance.
This one comes with a MySQL database already running beside it and a working
pipeline. Once it opens, type ./tour — a five-minute walkthrough that
shows you each command before running it. It is the only thing you have to type.
Unlike the labs, this environment includes Hydra Studio, the visual editor, so you can see the same pipeline as a diagram and run it from there.
Related MCP server: Datris MCP Server
Quick start
pip install "hydra-etl[server]"
hdrctl serve --open # Studio + API on http://localhost:5678Or stay in the terminal:
hdrctl init my_job # scaffold one of six templates
hdrctl validate my_job # strict validation, no data touched
hdrctl run my_job # executeNo Node, no build step, no database, no message broker. Python 3.9+ on Linux, macOS or Windows.
Why Hydra
Validate before you run.
hdrctl validatechecks sources, steps, types and destinations deterministically, before any data is read or written.Declarative, visual, zero deployment. YAML manifests, a browser editor that outputs the same YAML, and a single Python process. No cluster, no JVM, no container required.
Built for CI/CD. Manifests are versioned and reviewed like code;
hdrctlscaffolds, validates, tests and runs from a terminal or a CI job.Built-in scheduler. DAG workflows with dependencies, parallel branches, delayed retries, cron triggers, conditional guards and eleven actions (webhook, email, Bash, PowerShell, SSH, Python…).
One job, many environments.
{{ param: }},{{ env: }}and${SECRET:}keep the manifest identical across dev, staging and production.Two engines per operation. pandas or DuckDB, chosen step by step, with optional Rust acceleration (see Native acceleration).
AI-ready. An MCP server lets Claude, Cursor or VS Code write pipelines that Hydra validates.
How it compares
Hydra | Airflow | dbt | Airbyte | NiFi / Apache Hop | |
Pipelines defined in | YAML | Python | SQL + YAML | UI / config | UI flows |
Scope | Extract, transform, load | Orchestration | Transform in the warehouse | Extract & load | Extract, transform, load |
Visual editor | Included | Monitoring UI | Not in dbt Core | Included | Included |
To get started |
| Scheduler, webserver, metadata DB | A data warehouse | Docker / Kubernetes | JVM |
Hydra is not a replacement for all of these. It targets file-and-database pipelines that should stay readable and run without extra infrastructure.
What you get
Component | Description |
Engine | Declarative jobs: one source, N transformations, one destination |
CLI |
|
API | FastAPI, with interactive docs at |
Studio | Visual editor for jobs and workflows, served by the same process |
Workflows | Multi-job DAG with dependencies, actions, retries and runtime parameters |
MCP | Natural-language pipeline authoring, validated by Hydra |
Connectors (sources and destinations):
CSV | JSON | Parquet | MySQL / MariaDB | PostgreSQL | MongoDB | Web API |
Transformation engines: pandas and DuckDB.
Who it is for
Data engineers and analysts who need repeatable file-and-database pipelines without standing up infrastructure for them.
Teams where pipelines must stay readable by people who do not write Python: a YAML manifest reviewed in a pull request, not a script.
Developers embedding ETL in a product, who want a CLI and a REST API over the same engine.
A job in four files
A job is a folder. Four manifests describe it, and each one answers a single question.
sources.yaml: where the data comes from
version: "1.0"
sources:
src_input:
type: csv
extract:
table: ./input.csvtransformations.yaml: how it is reshaped
version: "1.0"
steps:
- cast:
mapping:
amount: float
- filter:
expr: "amount > 0"
- aggregate:
by: [name]
agg:
total: { func: sum, col: amount }A CSV carries no types, so cast comes before any numeric comparison.
destinations.yaml: where it goes
version: "1.0"
destinations:
dest_output:
type: csv
load:
table: ./output.csv
mode: replace # append | replace | upsertpipeline.yaml: which source feeds which destination
version: "1.0"
pipeline:
from: src_input
to: dest_outputhdrctl validate my_job && hdrctl run my_jobWorkflows: order several jobs
version: "1.0"
workflow:
name: daily_etl
trigger:
type: schedule
cron: "0 8 * * *"
steps:
- name: extract
type: job
job: ./jobs/extract
depends_on: []
- name: transform
type: job
job: ./jobs/transform
depends_on: ["extract"] # always a list, supports fan-in
- name: notify
type: action
action: webhook
params: { url: "{{ env:WEBHOOK_URL }}" }
depends_on: ["transform"]
on_failure: skiphdrctl workflow validate ./workflow.yaml
hdrctl workflow run ./workflow.yamlAn edge is a dependency, not a pipe: it decides when a job runs, never what data reaches it. Steps that share no dependency run in parallel.
Parameters
Values can be declared once and reused, or created while the workflow runs.
- filter:
expr: "region == '{{ param:region }}'"{{ param:NAME }} reads a parameter and {{ env:NAME }} an environment variable.
${SECRET:NAME} reads a secret: it is looked up in the injected secret store when one is configured, and
otherwise in the environment, under the name upper-cased with dots and dashes turned into underscores — so
${SECRET:db.password} reads DB_PASSWORD. Secrets therefore come from the process environment (CI
variables, a secret manager, your shell) and never from a file in the repository.
The set_param and assign_param actions create and change parameters mid-run, so two jobs can share a
placeholder and produce different results.
Install what you need
The base install is the engine and the CLI. Everything else is opt-in.
pip install hydra-etl # engine + CLI
pip install "hydra-etl[server]" # + API + Studio
pip install "hydra-etl[postgres]" # + PostgreSQL driver
pip install "hydra-etl[mssql]" # + Microsoft SQL Server driver
pip install "hydra-etl[all]" # everythingAvailable extras: server, native, duckdb, parquet, mysql, postgres, mssql, mongodb, http, mcp, all.
Or run it in a container — optional
Hydra needs no container: it stays one pip install and one Python process.
If your team standardises on containers, or you would rather keep Python off
the host, the image carries the engine, the CLI, the API and the Studio:
docker run --rm -p 5678:5678 -v hydra-workspace:/workspace \
ghcr.io/bejaouibechir/hydra:latestBuilding it yourself, per system: Linux · Windows · macOS.
Serving
hdrctl serve # Studio and API on port 5678
hdrctl serve --open # and open the browser
hdrctl serve --no-studio # API only, for a headless server
hdrctl serve --port 8080The server writes projects into the directory you launch it from. Interactive API docs are at /docs.
Your AI assistant, connected (MCP)
Hydra ships an MCP server. Point Claude Desktop, Cursor, VS Code or any MCP client at it, and ask for a pipeline in plain language.
pip install "hydra-etl[mcp]"
hydra-mcpYour assistant writes the manifests; Hydra validates them before anything is written, and nothing runs until you ask. Requires Python 3.10+. See docs/MCP.md for client configuration.
Graded by Glama, which builds the server in a sandbox, runs security checks and scores the tool definitions:
Native acceleration (optional)
Parts of the engine have a Rust implementation. It is optional and off by default: without it, Hydra behaves exactly the same.
pip install "hydra-etl[native]"
HYDRA_BACKEND=rust hdrctl run ./jobs/salesOr per operation, with a hydra.backends.yaml file next to the job:
default: python
overrides:
csv.read: rust # currently the only accelerated operationOn one million rows, CSV reading is about four times faster, with byte-for-byte identical results
(parity checked on 26,000 CSV files and one million floats against CPython repr()). Hydra falls back to
Python automatically whenever that guarantee cannot be kept. Benchmarks: bench/.
VS Code extension
Completion and validation for Hydra manifests inside the editor, without installing Hydra.
Download the .vsix from the latest release,
or build it yourself from vscode-extension/ (python build_vsix.py, no Node required).
Project status
Beta. The documented connectors and steps are implemented and covered by tests. The shape of the YAML DSL is settled; any breaking change will be announced in the release notes before 1.0. Use it on real work, and pin your version.
Contributing
Issues, ideas and pull requests are welcome. See CONTRIBUTING.md. If Hydra is useful to you, a ⭐ on GitHub helps others find it.
License
Hydra ETL is open source under the GNU Affero General Public License v3.0 or later (AGPL-3.0-or-later).
You can, for free
Use Hydra ETL for any purpose, including inside a company and for commercial work.
Run your own pipelines with it, internally or for clients.
Modify it and redistribute it.
Using Hydra ETL to move your data does not make your data, your YAML pipelines or your other applications subject to the AGPL.
The one obligation
If you modify Hydra ETL itself and make it available to others — including over a network, for example as a hosted service — you must publish your modified source code under the same license.
Commercial license
If that obligation does not fit your case — typically, you embed a modified Hydra ETL in a proprietary product or offer it as a service without publishing your changes — a commercial license is available.
Contact: admin@hydraetl.com — reply within 2 business days.
Support and services
Support with response time commitments
Help migrating existing scripts or tools to Hydra ETL
Available Tools
17 toolshydra_check_jobARead-onlyIdempotent
Check that a job does what the user actually asked for. This completes hydra_validate_job: that one says whether the YAML is correct, this one whether the job answers the request. Twelve deterministic rules: operation requested but missing, numeric comparison without a cast, load mode contradicted, inconsistent file extension, plaintext secret, unsupported technology. Call this AFTER writing a job, before presenting it.
| Name | Required | Description | Default |
|---|---|---|---|
| job_path | Yes | ||
| user_request | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, and the description does not contradict them. It adds useful behavioral context by revealing twelve deterministic rules and giving examples of what is checked (missing operations, cast issues, plaintext secrets). It doesn't over-explain outputs, which the output schema covers.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than strictly necessary, but each sentence earns its place: definition, sibling contrast, rule examples, and call timing. It is front-loaded with the core purpose and remains readable. Minor over-specificity in the rule list keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only checker with two string parameters and an output schema, the description covers the main things an agent needs: what the tool does, how it differs from the validation sibling, and when to call it. It doesn't describe the exact response shape, but the output schema can carry that. The missing parameter format guidance is the main completeness gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description needed to compensate for both parameters. It only implies that user_request is what the user asked for and job_path refers to the written job; it never states expected formats or how they should be supplied. This is a clear gap, since no parameter descriptions exist in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Check that a job does what the user actually asked for.' It explicitly contrasts with hydra_validate_job, making the tool's role distinct from a sibling. The reference to twelve deterministic rules further sharpens what 'check' means.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit sequencing: 'Call this AFTER writing a job, before presenting it.' It also names the alternative hydra_validate_job and explains the division of labor (YAML correctness vs. answering the request), so an agent can decide between the two. No misleading conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hydra_describe_operationARead-onlyIdempotent
Return the full JSON schema of one transformation operation: every parameter, its type and its default value. Call this when hydra_list_operations is not enough.
| Name | Required | Description | Default |
|---|---|---|---|
| operation | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this read-only, idempotent, and non-destructive, so the bar is lower. The description adds that the return value is the full schema with every parameter, type, and default, but it does not disclose other behavioral traits such as errors or availability.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no repetition: the first states the return contract and the second gives routing guidance. The core information is front-loaded and every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, read-only description tool with an output schema, the definition is nearly complete: it explains what is returned and when to use it over the sibling. The main gaps are the vague trigger condition and no explicit statement about where accepted operation names come from.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only a required string 'operation' with no description, and schema description coverage is 0%. The description contextually identifies the parameter as selecting a transformation operation, but it does not provide acceptable values, format, or examples, leaving the agent to infer them from sibling tools.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pair ('Return the full JSON schema of one transformation operation') and scopes the tool to a single operation. It also differentiates from the sibling hydra_list_operations by promising per-operation parameter details.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly names hydra_list_operations as the alternative and gives a conditional rule: use this when listing is not enough. However, 'not enough' is somewhat vague and does not specify exactly what missing detail should trigger this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hydra_explain_errorARead-onlyIdempotent
Explain a Hydra ETL validation error in plain language and propose the fix. Use this when hydra_validate_job or hydra_write_job returns a message you cannot interpret.
| Name | Required | Description | Default |
|---|---|---|---|
| error_message | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint, idempotentHint, and destructiveHint annotations already establish the safety profile. The description adds behavioral context by specifying that the tool translates raw errors into plain language and proposes a fix, which goes beyond the schema and annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences carry all essential information: what the tool does, the output form, and the exact usage context. No filler or repetition of schema fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, read-only analysis tool with an output schema, the description is fully sufficient. It names the triggering sibling tools and explains the parameter source, leaving no ambiguity about when or how to call it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates by tying error_message to the uninterpretable message returned by hydra_validate_job or hydra_write_job. This gives the single parameter practical meaning beyond its generic 'Error Message' label.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Explain'), names the exact resource (Hydra ETL validation error), and states the intended outcome ('propose the fix'). This clearly differentiates it from sibling tools like hydra_validate_job and hydra_write_job, which produce rather than explain errors.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit trigger condition: use when hydra_validate_job or hydra_write_job returns an uninterpretable message. This tells the agent both when to invoke the tool and which sibling outputs it is meant to clarify.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hydra_find_exampleARead-onlyIdempotent
Find, among real Hydra ETL jobs, the ones closest to the request, and return their manifests. Call this BEFORE writing an unusual job: an example that runs beats a reconstruction from memory.
| Name | Required | Description | Default |
|---|---|---|---|
| count | No | ||
| user_request | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds that the tool searches real, runnable Hydra jobs and returns manifests, but it does not explain similarity behavior, edge cases, or any constraints beyond the read-only annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler: the purpose is front-loaded, and the second sentence earns its place by stating when to call and why. Not a single wasted word.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema, rich annotations, and a single required parameter, the tool is largely self-explanatory. The only real gap is the undocumented count parameter, but its optional nature and default soften the impact.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the text needed to compensate for the undocumented user_request and count parameters. 'Closest to the request' hints at user_request, but count is never explained and neither parameter gets explicit format, meaning, or example semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Find'), a concrete resource ('real Hydra ETL jobs'), a selection criterion ('closest to the request'), and a clear result ('return their manifests'). It also frames the tool as a pre-write aid, which distinguishes it from siblings like hydra_write_job or hydra_list_jobs without needing to open their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence gives explicit timing guidance: call this BEFORE writing an unusual job, with a rationale ('an example that runs beats a reconstruction from memory'). It does not explicitly list alternatives or say when not to use it, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hydra_list_actionsARead-onlyIdempotent
List the actions usable in a workflow step of type 'action'. Careful: an unknown action is silently ignored at run time and the step is still counted as successful, so check the name before writing it.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes beyond the read-only and idempotent annotations by disclosing a critical runtime behavior: unknown actions are silently ignored and the step still counts as successful. This warning is essential for correct tool usage and is not conveyed by the structured annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences deliver the core purpose and a critical warning with no extraneous filler. The primary function is front-loaded, and the warning earns its place by preventing a subtle failure mode.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters, a read-only/idempotent annotation set, and an output schema present, the description provides everything an agent needs: a precise purpose and a vital operational caveat. Nothing important is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, so the baseline is 4. There is nothing for the description to add about parameter semantics, and it correctly does not invent any.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('List the actions usable in a workflow step of type action'), which clearly distinguishes this tool from siblings like hydra_list_operations or hydra_list_connectors. It names the exact context in which these actions apply, leaving no ambiguity about its purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool: before writing an action name, since unknown actions are silently ignored. It provides clear context that the agent should call this tool to verify valid action names prior to authoring a workflow step, though it does not explicitly name alternatives or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hydra_list_connectorsARead-onlyIdempotent
List the connector types accepted by the 'type' key of a source or a destination. Any other type fails at run time: S3, BigQuery and Snowflake are not supported.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds useful behavioral context by warning that unsupported types (S3, BigQuery, Snowflake) fail at run time, which tells the agent why consulting this list matters.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences lead with the primary purpose and follow with only the critical caveat. Every clause earns its place; no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only list tool with an output schema and full annotation coverage, the description covers purpose and important edge behavior. Nothing an agent needs in order to invoke this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so baseline is 4 by rubric. There is no parameter information to add; the description does not need to compensate for any schema gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('connector types accepted by the 'type' key of a source or a destination'), which clearly distinguishes it from sibling list tools such as hydra_list_operations and hydra_list_jobs. The additional unsupported types reinforce what the tool is about and avoid confusion with connector management.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear context is given: the tool is for discovering valid values of the 'type' key on sources/destinations. It doesn't explicitly name alternatives, but sibling tools differ by resource (operations, actions, jobs, workflows), so the usage context is reasonably unambiguous. The warning about unsupported types provides a practical usage caution.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hydra_list_jobsARead-onlyIdempotent
List the Hydra ETL jobs present in the workspace. A job is a folder holding a pipeline.yaml. Returns relative paths, usable as they are in the other tools.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds meaningful behavioral context beyond annotations by explaining what constitutes a job and what the output looks like ('relative paths, usable as they are in the other tools'). This is useful for an agent deciding how to chain results downstream.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences carry all the essential information: the action, the scope, the defining characteristic of a job, and the output format. No filler or redundant restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter listing tool with full annotation coverage and an output schema, the description tells the agent everything it needs: what is listed, how jobs are recognized, and how the returned paths can be used downstream. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there is no parameter semantics burden on the description. The baseline of 4 applies because no parameter documentation is needed at all.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'List the Hydra ETL jobs present in the workspace.' It also defines a job as 'a folder holding a pipeline.yaml,' which disambiguates it from sibling list tools like hydra_list_operations, hydra_list_connectors, and hydra_list_actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: this tool is for discovering Hydra ETL jobs in the workspace. It also explains that the returned relative paths are directly usable in other tools, which is practical guidance for when to use the result. It does not explicitly name alternatives or exclusion conditions, but the scope is clear enough from the definition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hydra_list_operationsARead-onlyIdempotent
List the 18 transformation operations of Hydra ETL with their required parameters. Call this BEFORE writing a transformations.yaml: any operation missing from this list does not exist and will be rejected.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the description need not repeat that. The description adds useful behavioral context by stating that operations missing from the list will be rejected, which is valuable. However, it does not disclose the return format or whether the list includes enums or default values.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded with the core purpose, then adds a crucial usage directive. Every sentence earns its place, and it is appropriately sized for the tool's simplicity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero parameters and an output schema, the description is quite complete: it tells the agent when to call it and the key fact that missing operations are rejected. It could potentially mention that hydra_describe_operation provides details, but that is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the description does not need to explain parameter semantics. The baseline for 0 parameters is 4, and the description appropriately focuses on the output rather than parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists the 18 transformation operations of Hydra ETL with their required parameters, using a specific verb and resource. It does not explicitly differentiate from sibling tools like hydra_describe_operation, but the distinction is implied by the focus on listing operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly instructs to call this BEFORE writing a transformations.yaml, providing clear usage context. It also implicitly excludes operations not in the list, but does not explicitly mention alternatives like hydra_describe_operation for detailed info on a specific operation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hydra_list_workflowsARead-onlyIdempotent
List the workflows in the workspace. A workflow orchestrates several jobs: dependencies, parallel execution, retries, cron triggering. Use this as soon as the request chains several jobs.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint false, so the safety profile is covered. The description adds conceptual context about workflows orchestrating jobs but does not disclose additional behavioral details such as pagination, rate limits, or authorization requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the first states the operation, the second explains the resource concept and provides a directly actionable usage trigger. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a simple parameterless listing tool with an output schema and read-only annotations. The description explains what a workflow is and when to use the tool, which is enough for an agent to select and invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline is 4. The description adds no parameter-specific detail, but none is needed; the empty schema is fully self-explanatory.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('List the workflows in the workspace') and clarifies the distinct role of a workflow versus individual jobs. This clearly differentiates it from sibling tools like hydra_list_jobs and hydra_list_operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit usage trigger: 'Use this as soon as the request chains several jobs.' However, it does not name sibling alternatives or state when not to use the tool, so it falls just short of full guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hydra_preview_dataARead-onlyIdempotent
Show the columns and the first rows of a data file (CSV, JSON, Parquet). Call this BEFORE writing a job, to learn the real column names instead of guessing them — and AFTER a run, to check the result.
| Name | Required | Description | Default |
|---|---|---|---|
| rows | No | ||
| file_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds that the tool returns only columns and first rows, which is some behavioral context, but it does not disclose limits, error cases, or behavior on unsupported files beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded: it states the core capability first, then adds high-value usage timing. Every sentence earns its place, and no redundant or filler content is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter preview tool with an output schema available, the description gives enough context to invoke it correctly: what it previews, which formats it supports, and when to call it. It could be more complete by explaining the row parameter and any file-access requirements, but the low complexity and existing annotations reduce the need.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only hints at parameter meaning: file_path is a data file of certain formats and rows relates to 'first rows.' It does not clarify the row-count semantics, default behavior, or path expectations, leaving the agent to rely on the schema's minimal titles.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('Show the columns and the first rows'), a clear resource ('data file'), and accepted formats (CSV, JSON, Parquet). It clearly distinguishes this from sibling tools by focusing on previewing data rather than listing, writing, or running jobs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says to call this 'BEFORE writing a job' to learn column names and 'AFTER a run' to check results, giving clear situational guidance. It does not explicitly name alternatives or state when not to use it, so it falls just short of a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hydra_read_jobARead-onlyIdempotent
Read the manifests of an existing job. Call this before modifying a job, so you start from its real content instead of rewriting it from memory.
| Name | Required | Description | Default |
|---|---|---|---|
| job_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds only modest behavioral context by implying the tool returns the actual state of the job, but does not disclose error behavior, authentication needs, or other side effects beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with zero filler. The core action is front-loaded, and the usage rationale is woven in efficiently without repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read tool with one parameter, provided output schema, and strong annotations, the description covers purpose and usage adequately. It omits failure behavior when the job does not exist, but that is minor given the other structured context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the single parameter job_path. It does not explain the expected format, allowed values, or how to obtain a valid path, leaving the agent to infer from the tool name and phrasing 'existing job'. This is a significant gap for a parameter that is required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Read the manifests of an existing job') with a clear verb and resource. It is immediately distinguishable from siblings like hydra_write_job, hydra_run_job, and hydra_list_jobs, especially with the explicit 'before modifying' context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Call this before modifying a job' and gives the rationale ('start from its real content instead of rewriting it from memory'), providing clear when-to-use guidance. It does not name alternatives or state when NOT to use it, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hydra_read_workflowARead-onlyIdempotent
Read an existing workflow.yaml file and return its raw YAML. Call this before modifying a workflow, so you start from its real content instead of rewriting it from memory.
| Name | Required | Description | Default |
|---|---|---|---|
| workflow_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds value by stating that the tool returns raw YAML and is intended as a pre-modification read step, which aligns with the annotations and gives operational context beyond the annotations alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two efficient sentences, with the core action and output front-loaded in the first sentence and the usage rationale in the second. No redundant information or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with a single parameter, rich annotations, and an output schema, the description covers the essential context: read before modifying, pass a path to an existing workflow.yaml, and expect raw YAML. Minor ambiguity remains about how workflow_path should be specified, but it is not a significant gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the undocumented workflow_path parameter. It identifies the parameter as an 'existing workflow.yaml file', which gives necessary context, but it does not clarify path conventions, file location, or valid formats. For a single required string parameter, this is sufficient but not thorough.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a specific action ('Read'), a specific resource ('existing workflow.yaml'), and the output ('raw YAML'). It clearly differentiates from siblings like hydra_write_workflow, hydra_list_workflows, and hydra_run_workflow by focusing on reading an existing file's content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly tells the agent when to call this tool: before modifying a workflow, so it starts from real content rather than memory. It gives clear usage context, though it does not provide when-not-to-use guidance or mention alternative tools explicitly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hydra_run_jobADestructive
Run a Hydra ETL job and return the execution log. Only call this tool if the user explicitly asked for the run: it writes real data to the destination.
| Name | Required | Description | Default |
|---|---|---|---|
| job_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as destructiveHint=true and readOnlyHint=false, and the description reinforces that it writes real data to the destination. It adds the 'explicit user request' caution, but does not disclose other behavioral traits like execution duration, blocking behavior, or partial failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The first sentence states the core action and output; the second delivers an important safety cue. The most critical behavioral warning is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter destructive tool with an output schema and annotations, the description covers the main risk and return behavior. However, the missing semantics of job_path leave an agent uncertain about what value to pass, and there is no pointer to validation or preview alternatives before performing a real data write.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the one parameter, job_path. It does not explain the expected path format, whether it is absolute/relative, or how to obtain valid job paths from sibling tools like hydra_list_jobs. The parameter name is somewhat self-explanatory, but the description adds no real semantic value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Run'), the resource ('a Hydra ETL job'), and the return value ('return the execution log'). This distinguishes it from sibling tools like hydra_read_job, hydra_validate_job, and hydra_write_job.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an explicit when-not-to-use condition: only call when the user explicitly asked for the run, because it writes real data. It does not name alternatives such as validate_job or preview_data, but the consent requirement is a strong usage signal.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hydra_run_workflowADestructive
Run a workflow and return the trace, step by step. Only call this tool if the user asked for the run: the jobs it contains write real data to their destinations.
| Name | Required | Description | Default |
|---|---|---|---|
| workflow_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false; the description adds meaning by explaining that the contained jobs write real data to destinations and that the output is a step-by-step trace. It aligns with the annotations and adds context, though it omits details like reversibility or prerequisites.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with the core action front-loaded and a necessary safety caveat immediately after. There is no filler or redundant repetition of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter destructive run tool with an output schema, the description covers the main behavior and side-effect warning. However, it lacks guidance on how to obtain or validate workflow_path, and whether validation should happen before running, which would round out the context for safe invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain the workflow_path parameter at all. The title 'Workflow Path' gives minimal meaning, but the agent gets no guidance on path format, origin, or how to obtain a valid value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb and resource: 'Run a workflow' and says it returns a step-by-step trace. It differentiates from write/validate siblings by emphasizing actual execution, though it does not explicitly distinguish itself from hydra_run_job.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says 'Only call this tool if the user asked for the run,' which is a clear when-not condition. It also warns that jobs write real data. It does not name an alternative such as validate or preview, but the exclusion is strong enough for an agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hydra_validate_jobARead-onlyIdempotent
Validate a job with the official Hydra ETL validator: the structure of the four manifests, and the resolution of pipeline.from to a declared source and of pipeline.to to a declared destination. This is the ground truth — if this tool refuses, the job will not run.
| Name | Required | Description | Default |
|---|---|---|---|
| job_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds valuable behavioral context beyond annotations: the specific validation semantics (manifests, pipeline source/destination resolution) and the consequence that refusal means the job will not run.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler. The first sentence front-loads the tool's purpose and scope; the second reinforces its authoritative role. Every clause adds information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with an output schema and read-only annotations, the description covers purpose, validation scope, and behavioral consequence. The only notable gap is not explicitly distinguishing this from hydra_check_job or explaining job_path format, but overall the agent has enough to call it appropriately.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single job_path parameter, and the description only indirectly refers to it as 'a job.' The mapping to job_path is obvious from the tool name and parameter title, but the description does not explain path format, expected file types, or how job_path should be resolved. It adds minimal meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Validate a job with the official Hydra ETL validator.' It details exactly what validation covers (four manifests, pipeline.from/pipeline.to resolution), and calls itself 'the ground truth,' which sets it apart from sibling tools like hydra_check_job.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage context by framing this as the authoritative validation step: 'if this tool refuses, the job will not run.' It does not explicitly name alternatives or state when not to use it, but the ground-truth framing gives an agent enough context to select it for definitive validation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hydra_write_jobA
Write the manifests of a job, AFTER validation. Each manifest is passed as YAML text. If validation fails, nothing is written and the errors are returned: fix them and call the tool again. Writing does not run the job.
| Name | Required | Description | Default |
|---|---|---|---|
| job_path | Yes | ||
| sources_yaml | Yes | ||
| pipeline_yaml | Yes | ||
| destinations_yaml | Yes | ||
| transformations_yaml | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=false, destructiveHint=false, and idempotentHint=false, so the description doesn't need to restate those. The description adds valuable behavioral context: the atomicity of the write (nothing is written if validation fails), the error-return behavior, and the fact that writing does not trigger execution. This goes beyond the annotations and helps the agent understand side effects and failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, each earning its place: the first states the core action and precondition, the second explains the failure mode and retry guidance, and the third clarifies a key boundary (no execution). It is front-loaded with the most important information and contains no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 5 parameters, no schema descriptions, and an output schema that presumably describes the return value. The description covers the key behavioral context: validation-before-write, atomic failure, and no execution. It does not explain what the output schema contains or what a successful write returns, but the output schema likely covers that. The main gap is the lack of parameter-level detail for job_path and transformations_yaml, but overall the description is sufficient for an agent to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the schema's lack of parameter documentation. The description explains that each manifest is passed as YAML text, which covers the general format of the YAML parameters (sources_yaml, destinations_yaml, pipeline_yaml, transformations_yaml). However, it does not explain the role of job_path or the optional transformations_yaml parameter, nor does it clarify the relationship between the four required manifests. The description adds some meaning but leaves gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('write'), a specific resource ('manifests of a job'), and a critical precondition ('AFTER validation'). It clearly distinguishes itself from siblings like hydra_validate_job and hydra_run_job by stating that writing does not run the job and that validation must happen first. This is a clear, non-tautological definition.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says to call this tool after validation, and explains the failure behavior: if validation fails, nothing is written and errors are returned, so the agent should fix them and call again. It also clarifies that writing does not run the job, which helps the agent choose between this and hydra_run_job. However, it does not explicitly name the alternative validation tool or state when to use hydra_validate_job instead, so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hydra_write_workflowA
Write a workflow, AFTER validation. A step is either type='job' with the path of a job folder, or type='action'. depends_on is ALWAYS a list: steps with no dependency in common run in parallel. If validation fails, nothing is written.
| Name | Required | Description | Default |
|---|---|---|---|
| workflow_path | Yes | ||
| workflow_yaml | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnlyHint=false, destructiveHint=false), the description discloses the all-or-nothing behavior: 'If validation fails, nothing is written.' It also explains the parallelism semantics of depends_on, which is helpful behavioral context not present in the structured metadata.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the key operation and validation requirement. Every sentence adds useful information: the step model, depends_on behavior, and failure behavior. There is no filler or repetition of the name beyond the opening verb phrase.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given only two string parameters, an output schema is present, and the description explains the important workflow YAML semantics and failure behavior, the definition is complete enough for an agent to invoke the tool correctly. No critical calling prerequisites are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does meaningfully, explaining that workflow_yaml contains steps with type='job' or type='action' and that '`depends_on` is ALWAYS a list.' It leaves the workflow_path parameter to its name, but the core YAML semantics are well covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific operation on a specific resource: 'Write a workflow' and adds the critical condition 'AFTER validation.' It also clarifies the content model (steps of type job or action), which distinguishes it from sibling workflow tools like hydra_read_workflow and hydra_run_workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly indicates when to use this tool: after validation, and 'If validation fails, nothing is written.' It does not explicitly name a validation alternative or list exclusions, but the sequencing and failure behavior provide solid usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
17 tool updates
- First observed
hydra_check_job - First observed
hydra_describe_operation - First observed
hydra_explain_error - First observed
hydra_find_example - First observed
hydra_list_actions - First observed
hydra_list_connectors - First observed
hydra_list_jobs - First observed
hydra_list_operations - First observed
hydra_list_workflows - First observed
hydra_preview_data - First observed
hydra_read_job - First observed
hydra_read_workflow - First observed
hydra_run_job - First observed
hydra_run_workflow - First observed
hydra_validate_job - First observed
hydra_write_job - First observed
hydra_write_workflow
TDQS
Scored across 17 tools
Each tool targets a distinct resource and action: metadata discovery, job lifecycle, workflow lifecycle, data preview, and error explanation. The only near-overlap, validate_job vs check_job, is explicitly differentiated in the descriptions.
All tools share the hydra_ prefix and follow a verb_noun pattern (list_, describe_, read_, validate_, write_, run_, explain_, preview_, check_, find_). Verbs and nouns are singular/plural only when logically appropriate.
17 tools is slightly above the ideal 3-15 range, but each tool earns its place across job authoring, workflow orchestration, validation, execution, and discovery. The count feels a little heavy but not bloated.
The surface covers the full job and workflow authoring lifecycle: discover, describe, read, validate, write, run, check, and preview. Minor gaps exist: there is no delete or cancel tool for jobs/workflows, and workflow validation is only mentioned as a precondition of write rather than exposed as a separate tool.
Maintenance
Related MCP Connectors
MCP server for progressive tool usage at any scale (see https://klavis.ai)
MCP server that lets AI assistants use all OneSchema features exposed via the public API.
MCP server for secureFlows: token-free URL builders and integration-linting tools for AI agents.
Hosted MCP server for AI-driven data ops. Create apps, manage schemas, and CRUD structured data.
Related MCP Servers
- AlicenseAqualityBmaintenanceA token-efficient, schema-aware MCP server that enables AI assistants to safely read, modify, query, and validate JSON, YAML, and TOML files with automatic schema detection and format conversion capabilities.826 PyPI10MIT
- AlicenseNot gradedqualityBmaintenanceMCP server with 32 tools for ETL ingestion, AI-generated data quality rules, AI transformations, vector search, and natural-language SQL. Works across Postgres, MongoDB, Kafka, S3/MinIO, HashiCorp Vault, and five vector stores (Qdrant, Weaviate, Milvus, Chroma, pgvector).12AGPL 3.0
- AlicenseAqualityDmaintenanceMCP server that validates LLM-generated tool-call arguments, lints tool definitions, and produces retry messages for AI assistants.338 npm1MIT
- AlicenseNot gradedqualityDmaintenanceAn MCP server that gives AI assistants the ability to connect to, query, profile, and monitor data sources — turning any LLM into an interactive data engineering copilot.MIT