Skip to main content
Glama

stage_dataset

Download datasets directly to compute nodes instead of the login node to avoid shared bandwidth. Supports direct URLs, Hugging Face, and git repositories.

Instructions

Telecharge un jeu de donnees depuis un noeud de calcul plutot que depuis le noeud de login, dont la bande passante est partagee. Accepte une URL directe, un depot Hugging Face ou un depot git.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
archNox64cpu
kindNoauto
sourceYes
confirmNo
minutesNo
time_limitNo
destinationYes

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.4.0

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and destructiveHint=false, covering the mutation/safety profile. The description adds useful context: the download occurs on a compute node rather than the login node, and supports URL/Hugging Face/git sources. However it says nothing about the confirm flag, timeout/minutes behavior, or prerequisites, which are exactly the behavioral traits an agent would need beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the key rationale (compute node vs shared-bandwidth login node) front-loaded and the accepted source types following. No filler, though the second sentence could be folded more economically.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter tool with 0% schema coverage, an output schema, and no parameter documentation, the description is not complete enough: the role of confirm, minutes, time_limit, and arch is unexplained, and there is no indication of what staging returns or how long it takes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 7 parameters, so the description must compensate, and it only partially does: it explains the source kinds (direct URL, Hugging Face repo, git repo), which maps to source/kind. It says nothing about destination, arch, confirm, minutes, or time_limit, leaving most parameters undocumented in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action (download/stage a dataset) and a specific resource, plus a distinctive rationale (run on a compute node instead of the bandwidth-shared login node) and the accepted source types. It is clear what the tool does, though it does not explicitly name the sibling it replaces (e.g. download_from_romeo/upload_to_romeo) to remove ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies the when-to-use case (you want staging to happen on a compute node because the login node's bandwidth is shared), but there is no explicit when-not-to-use guidance and no alternatives such as download_from_romeo or run_login_command are named. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.