Skip to main content
Glama

ROMEO MCP

Du dialogue avec votre IA au calcul scientifique sur ROMEO.

Un serveur MCP local pour préparer, soumettre, suivre et comprendre vos calculs Slurm.

Démarrer · Tous les guides · Tools · Documentation ROMEO

Configuration · Dépannage · Mises à jour · Contribuer

Vérifications Python 3.11 et plus Transport stdio Licence MIT

Source : Centre de Calcul Régional ROMEO / URCA, article du 21 novembre 2013.

Navigation rapide

Choisissez votre besoin pour ouvrir le guide ou la fiche correspondante. L’index de la documentation rassemble tous les parcours.

Je veux…

Guide

Accès direct

Installer et connecter le MCP

Prise en main

Accès SSH · Configuration · Clients IA

Découvrir les outils

Catalogue Tools

Fiches par catégorie · Profils d’outils

Préparer et suivre un calcul

Premier essai

Préparer · Soumettre · Suivre

Utiliser Python, MPI ou les GPU

Calcul parallèle

Python · OpenMPI · GPU

Gérer les fichiers et le stockage

Fichiers et transferts

Envoyer · Récupérer · Quotas

Ouvrir un notebook ou un service

Services interactifs

Préparer un service · Connexion · JupyterLab

Comprendre un échec ou mesurer un job

Dépannage

Diagnostic · Journaux · Efficacité

Conserver les preuves d’une expérience

Reproductibilité

Collecter une fiche · Exporter · Limites des mesures

Consulter une procédure ROMEO

Parcours ROMEO 2025

Sommaire officiel · Chercher · Lire une page

Mettre à jour ou contribuer

Mises à jour

Versions · Contribution · Tests · Sécurité

Related MCP server: tacc-mcp-bio

Pour qui ?

Étudiants, enseignants et chercheurs disposant d’un accès déjà autorisé à ROMEO : simulation numérique, chimie, bio-informatique, statistiques, calcul MPI ou apprentissage automatique.

Décrivez votre besoin à votre assistant IA. Le MCP lui fournit les outils pour consulter la documentation, construire un script Slurm, vérifier les ressources demandées et suivre le résultat. Il utilise votre connexion SSH et les droits de votre compte.

Projet communautaire indépendant. Ce dépôt n’est pas un service officiel de l’Université de Reims Champagne-Ardenne. L’accès au calculateur et ses règles restent ceux de ROMEO.

Votre besoin

Ce que fournit le MCP

Préparer un calcul

Script Slurm, choix de partition, contrôles CPU, RAM et architecture

Utiliser les GPU

Prise en compte de l’architecture ARM des nœuds GPU et des environnements Spack

Comprendre un échec

État du job, extraits ciblés des journaux, diagnostic et efficacité

Lancer plusieurs expériences

Tableaux de paramètres, étapes dépendantes et points de reprise

Trouver la bonne commande

Documentation embarquée, recherche locale et lecture par section

Reprendre une conversation

Registre local des jobs soumis par le MCP

Commencer avec peu d’outils

Profil essentiel, avec accès au catalogue complet à la demande

Conserver les preuves d’un calcul

Fiche JSON et Markdown : script filtré, code, environnement, ressources et empreintes

Comment ça fonctionne

Du client IA aux nœuds de calcul, via le MCP local, SSH et Slurm

  1. Votre client IA lance le serveur Python local et lui parle par stdio.

  2. Le MCP consulte son corpus local ou utilise le client OpenSSH de votre poste.

  3. Slurm attribue les ressources et exécute les calculs sur les nœuds appropriés.

  4. Le MCP restitue au client les états, résultats et extraits demandés.

Le MCP n’embarque aucun modèle IA. Les contenus renvoyés par ses outils peuvent entrer dans le contexte de votre assistant et être transmis à son fournisseur. Choisissez les fichiers et journaux que vous lui donnez en fonction des règles de votre équipe.

Prise en main

1. Préparer les accès

Il vous faut :

  • Python 3.11 ou plus, Git et un client OpenSSH (ssh -V).

  • Un compte ROMEO actif, une clé SSH enregistrée et un projet de calcul autorisé.

  • Un client acceptant les serveurs MCP locaux en stdio.

L’installation automatisée prévoit Codex, Claude Code et Claude Desktop. La disponibilité du MCP dépend aussi de la version et de la configuration de votre client.

Consultez la documentation officielle embarquée pour la création du compte et la connexion SSH.

Dans ~/.ssh/config (Windows : %USERPROFILE%\.ssh\config), ajoutez une entrée en remplaçant les deux valeurs VOTRE_… :

Host romeo1
    HostName romeo1.univ-reims.fr
    User VOTRE_IDENTIFIANT
    IdentityFile ~/.ssh/VOTRE_CLE_PRIVEE
    IdentitiesOnly yes

Testez dans un terminal :

ssh romeo1

Vérifiez l’empreinte de l’hôte selon les instructions ROMEO lors de la première connexion. Une clé protégée par une phrase secrète doit être disponible via votre agent SSH avant de lancer le client IA. Ne copiez jamais la clé privée dans ce dépôt. Fermez cette session avec exit après vérification.

2. Installer le serveur

Windows : PowerShell

git clone https://github.com/Gotman08/romeo-mcp.git
cd romeo-mcp
py -3 -m venv .venv
.\.venv\Scripts\python.exe -m pip install -e .
.\.venv\Scripts\python.exe -m romeo_mcp configure --account VOTRE_PROJET --profile essential
.\.venv\Scripts\python.exe -m romeo_mcp doctor

Linux / macOS : terminal

git clone https://github.com/Gotman08/romeo-mcp.git
cd romeo-mcp
python3 -m venv .venv
./.venv/bin/python -m pip install -e .
./.venv/bin/python -m romeo_mcp configure --account VOTRE_PROJET --profile essential
./.venv/bin/python -m romeo_mcp doctor

Remplacez VOTRE_PROJET par le code de votre projet, trouvé dans votre espace ROMEO. Ce n’est pas votre identifiant SSH. Aucun compte de calcul n’est fourni par défaut.

doctor vérifie la configuration et la présence de la documentation sans connexion au cluster. Il doit indiquer account_configured: true et des pages documentaires disponibles. Il ne valide pas vos droits Slurm ni votre clé SSH.

Pour vérifier aussi les accès réels, utilisez le Python du même venv :

python -m romeo_mcp doctor --live

Ce diagnostic lit la connexion SSH, l’association au projet Slurm, les partitions et les quotas utilisateur et projet. Chaque contrôle fournit son état et une explication en cas d’échec. Il ne soumet aucun job. Options et résultats du diagnostic.

La commande configure conserve votre choix hors du dépôt :

Système

Configuration personnelle

Windows

%LOCALAPPDATA%\romeo-mcp\config.json

Linux / macOS

$XDG_CONFIG_HOME/romeo-mcp/config.json, sinon ~/.config/romeo-mcp/config.json

Un autre alias SSH ou une autre QOS peuvent être indiqués avec --host et --qos. Les variables d’environnement priment sur le fichier ; voir le guide de configuration.

3. Ajouter le MCP à votre assistant

Depuis la racine du dépôt, choisissez le client voulu :

# Windows : inspecter, puis enregistrer dans Codex
.\.venv\Scripts\python.exe tools/install_mcp.py --targets codex --dry-run
.\.venv\Scripts\python.exe tools/install_mcp.py --targets codex
# Linux / macOS : inspecter, puis enregistrer dans Codex
./.venv/bin/python tools/install_mcp.py --targets codex --dry-run
./.venv/bin/python tools/install_mcp.py --targets codex

Remplacez codex par claude-code, claude-desktop, ou une liste séparée par des virgules. --list affiche les emplacements détectés. L’installateur sauvegarde les fichiers modifiés et vérifie le démarrage du serveur. Relancer l’installation met à jour la même entrée.

Relancez ensuite le client concerné lorsque vos opérations en cours sont terminées. Le serveur apparaît sous le nom romeo. Pour un autre client stdio ou une configuration manuelle : exemples de configuration.

4. Faire un premier essai

Commencez par demander à l’assistant :

Cherche dans la documentation ROMEO comment lancer un calcul CPU. Puis consulte l’état du cluster et mes quotas, sans soumettre de job.

Puis préparez un petit calcul :

Prépare en simulation un job hello-romeo, sur un seul nœud x64cpu, avec un cœur, 1 Go de RAM et une minute. La commande est hostname. Montre-moi le script et les avertissements avant toute soumission.

L’appel correspondant à job_prepare est :

{
  "name": "hello-romeo",
  "command": "hostname",
  "arch": "x64cpu",
  "time_limit": "1m",
  "nodes": 1,
  "cpus_per_task": 1,
  "mem_gb": 1
}

La préparation retourne le script et un plan_id, sans soumettre le calcul. Elle peut consulter les chemins distants par SSH et conserve le plan localement pendant 24 heures. Après lecture, soumettez exactement ce plan avec job_submit({"plan_id": "IDENTIFIANT_RECU", "confirm": true}). Suivez ensuite le job reçu avec job_status, job_log_tail et job_efficiency.

Les tableaux et pipelines suivent le même parcours avec job_array_prepare / job_array_submit et job_pipeline_prepare / job_pipeline_submit. Les préparations ne prennent pas de paramètre confirm. Un aperçu hors ligne aux chemins illustratifs ne peut pas être soumis. D’autres actions, comme les transferts, l’annulation ou cluster_gpu_health_run, agissent directement.

Des demandes utiles

Situation

Exemple de demande

TP de calcul scientifique

« Prépare un tableau Slurm pour ces paramètres et explique les ressources choisies. »

Code MPI

« Trouve le logiciel via Spack, puis prépare un lancement MPI sur deux nœuds. »

Calcul GPU

« Vérifie la compatibilité ARM de mes dépendances avant de préparer ce job GPU. »

Job en échec

« Analyse le job indiqué et lis seulement les extraits de logs utiles au diagnostic. »

Optimisation

« Compare le temps et la mémoire réellement utilisés aux ressources réservées. »

Reproductibilité

« Exporte la fiche de ce job avec les empreintes de ces fichiers d’entrée. »

La page Tools propose un catalogue cliquable : chaque outil possède une fiche avec son rôle, ses paramètres et un exemple. Elle détaille aussi MPI, PyTorch, Apptainer, les transferts, le profilage et les limites de chaque mesure.

Un profil essentiel pour commencer

Le profil essential présente 22 outils : documentation, état du cluster, quotas, logiciels, soumission simple, suivi et diagnostic des jobs, transferts et export de fiches. Il réduit le catalogue envoyé au modèle.

Pour accéder aux tableaux de paramètres, aux pipelines, aux services interactifs ou au profilage, demandez à l’assistant :

Passe le profil d’outils ROMEO à full avec tool_profile_set.

Le client reçoit une notification de changement du catalogue. Le choix vaut pour le processus MCP actuel. Pour le conserver au prochain lancement :

python -m romeo_mcp configure --profile essential

Cette commande conserve votre projet et votre alias SSH. Sans choix explicite, le profil reste full, pour les opérations métier. Le profil expert ajoute les exécuteurs de commandes arbitraires. Les profils modifient la découverte des outils ; les autorisations restent celles du client et de ROMEO. Liste et configuration des profils.

Une fiche de reproductibilité par job

job_report_export crée un dossier privé avec report.json, report.md et script.sbatch.txt. Les nouveaux jobs conservent les ressources demandées et tentent de relever le commit Git, l’environnement chargé et les empreintes des fichiers choisis avant le calcul. job_report_collect(job_id) enregistre un relevé daté et rend report_id. job_report_export(report_id) exporte exactement ce relevé, sans SSH.

Pour choisir les entrées à relever au démarrage, ajoutez data_files à job_prepare. Les fichiers choisis seulement à la collecte sont datés comme observations après coup. Les captures sont limitées à 20 fichiers et 64 Mio par relevé.

Les exports restent hors des dépôts Git. Les secrets reconnaissables sont masqués ; le contenu des données, les variables d’environnement complètes et les adresses des dépôts Git ne sont pas exportés. Les informations absentes sont signalées. Exemples, protection des données et limites.

Documentation locale et contexte de l’IA

Le corpus ROMEO accompagne le dépôt et les paquets Python : 42 pages officielles, 21 images, un sommaire et un manifeste de provenance.

  • search_docs classe les sections par pertinence lexicale (BM25), avec prise en compte des titres et des accents.

  • Les extraits comportent les sources, les lignes, l’empreinte du document et les arguments de lecture.

  • read_doc permet de lire une section complète. Si elle dépasse le budget, next_call poursuit la lecture sans supprimer du texte du corpus.

  • Aucun service d’embeddings ni accès réseau n’est requis pour cette recherche.

Un extrait seul peut manquer de prérequis : l’assistant doit poursuivre la lecture de la section ou des sections parentes. Le corpus est daté ; romeo_status, romeo_quota et romeo_selfcheck renseignent l’état actuel du cluster.

Après déplacement du dossier, recréez le venv et relancez l’installateur pour mettre à jour les chemins du client. La documentation reste dans le projet. Fonctionnement et renouvellement du corpus.

Quotas : quel chiffre regarder ?

Limite

À quoi elle sert

Stockage utilisateur

Espace attribué à votre compte sur un système de fichiers

Stockage projet

Espace partagé, consommé collectivement par les membres du projet

Quota souple / strict

Seuil pouvant ouvrir une période de grâce / plafond bloquant les nouvelles écritures

Nombre de fichiers

Limite d’inodes ; beaucoup de petits fichiers peuvent l’atteindre avant le volume en Go

CPU / GPU / jobs Slurm

Ressources et nombre de jobs autorisés par les associations et QOS

Fairshare

Priorité influencée par l’usage passé du groupe ; ce n’est pas du stockage disponible

romeo_quota lit les quotas effectifs ; df indique la capacité du système de fichiers entier. Les plafonds Slurm dépendent du projet. Aucun quota personnel n’est présumé par défaut dans cette version.

En cas de problème

Symptôme

Vérification

« Projet ROMEO absent »

Exécuter configure --account …, puis relancer le processus MCP

SSH refusé ou bloqué

Tester ssh romeo1 dans le même environnement utilisateur ; vérifier clé et agent SSH

Le MCP n’apparaît pas

Relancer le client et vérifier le chemin du Python avec install_mcp.py --list

No module named romeo_mcp

Réinstaller avec le Python du venv utilisé par le client

Documentation introuvable

Exécuter doctor et retirer un ancien ROMEO_DOCS_DIR s’il n’est plus valable

Un job attend longtemps

Lire le motif Slurm ; vérifier disponibilité, compte, QOS, dépendances et limites

Binaire incompatible sur GPU

Recompiler pour aarch64 sur un nœud adapté avec compute_command_prepare

Contribuer et vérifier

python tests/run_all.py
python tools/check_privacy.py --history

Utilisez le Python du venv. Les suites par défaut sont hors ligne. Les tests --live demandent un accès ROMEO et peuvent soumettre de vrais jobs ; ils sont exclus de la CI.

La CI vérifie les suites hors ligne, la documentation, les métadonnées de commit et les fichiers publiés, puis construit le paquet. Les procédures de contribution et de signalement se trouvent dans CONTRIBUTING.md et SECURITY.md.

Code distribué sous licence MIT. La documentation et les visuels officiels conservent leurs droits et leur attribution : contenus tiers.


↑ Haut de page · Documentation · Catalogue Tools

Available Tools

47 tools
allocate_debug_nodeB

Reserve un noeud de calcul pour de la mise au point interactive (compilation, profilage, essais). Rend la commande a executer pour obtenir un shell dessus. Utile pour tester sur aarch64 sans passer par un aller-retour de job batch.

ParametersJSON Schema
NameRequiredDescriptionDefault
archNoarmgpu
confirmNo
minutesNo
time_limitNo
cpus_per_taskNo
gpus_per_nodeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations declaring readOnlyHint=false and destructiveHint=false, the safety profile is already covered. The description adds that it reserves a compute node and returns a command to obtain a shell, which is useful behavioral context, but it omits operational details such as whether confirmation is required (confirm parameter), duration limits (minutes, time_limit), or what happens if no node is available.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core purpose, followed by the return behavior and usage hint. Every sentence contributes meaning and there is no wasted wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Though annotations cover the safety profile and an output schema presumably documents the returned command, the description is inadequate for a six-parameter tool with 0% schema coverage. It does not explain the confirm flag, time limits, or resource sizing, which are essential for correct invocation of a reservation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description does not mention any of the six parameters (arch, confirm, minutes, time_limit, cpus_per_task, gpus_per_node). It only vaguely references 'aarch64', which does not clarify the arch parameter. The description fails to compensate for the complete lack of parameter documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ('Reserve') and resource ('un noeud de calcul') and scopes the purpose to interactive debugging (compilation, profiling, tests). It implies an alternative to batch jobs but does not name a specific sibling tool like submit_job or launch_interactive_service, so it falls just short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly states when the tool is useful: interactive debugging and testing on aarch64 without a batch-job round trip. It does not explicitly say when not to use it or name alternatives by tool name, but the context is clear enough to guide selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

audit_orphan_filesA
Read-only

Repere les fichiers volumineux abandonnes sous le scratch : anciens, sans job actif associe, ou typiques de fichiers temporaires de calcul (.rwf, .scr, .tmp). Ne supprime rien. Utile quand romeo_quota signale un depassement.

ParametersJSON Schema
NameRequiredDescriptionDefault
topNo
daysNo
pathNo
min_size_mbNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With readOnlyHint=true already in annotations, the description adds real value by reinforcing 'Ne supprime rien' and disclosing the heuristics used (age, job association, temp-file extensions). It omits any note on result ordering, limits, or cost, but the output schema covers the return shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the identification criteria and the safety assurance, with the usage trigger last. No padding, though the criteria list is slightly dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only audit tool with an output schema and read-only annotation, the description supplies the heuristics and the trigger. The remaining gap is parameter meaning, which is unaddressed in both schema and description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for 4 parameters. The description loosely implies them ('volumineux' → min_size_mb, 'anciens' → days, 'sous le scratch' → path) but never explains top, defaults, or units, so a caller still has to guess from the schema titles.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

It states a specific verb (Repere/identifies) and resource (large abandoned files under scratch) and spells out the detection criteria: old, no active associated job, or typical temp compute files (.rwf, .scr, .tmp). It also plainly distinguishes itself from the delete-action siblings by declaring 'Ne supprime rien'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit trigger condition: 'Utile quand romeo_quota signale un depassement', naming the sibling that signals the need. It does not say when NOT to use it or mention storage_cleanup_helper as the alternative for the actual cleanup, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

build_on_nodeA

Compile ou installe sur un noeud de calcul de l'architecture voulue, via srun. INDISPENSABLE pour cibler les noeuds GPU : ils sont en aarch64 alors que le noeud de login est en x86_64, donc tout artefact produit sur le login y est inutilisable. Synchrone, plafonne a 15 min.

ParametersJSON Schema
NameRequiredDescriptionDefault
archNoarmgpu
cpusNo
minutesNo
modulesNo
workdirNo
commandsYes
with_gpuNo
spack_packagesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and destructiveHint=false. The description adds valuable non-obvious behavior: execution is synchronous and capped at 15 minutes, and it runs remotely via srun. It stops short of stating what gets written where or failure semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, front-loaded with the action, then the critical arch rationale in caps, then the execution constraint. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists, so return values need no explanation, and the description covers purpose and key constraints well. However, for an 8-parameter tool with zero schema documentation, the description leaves the bulk of invocation semantics unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 8 parameters, so the schema explains nothing. The description only loosely gestures at the arch parameter ('architecture voulue') and the minutes cap; cpus, modules, workdir, commands, with_gpu and spack_packages get no meaning at all.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb pair (compile/install) and resource (compute node of the desired architecture) plus the mechanism (srun). It clearly separates itself from login-node siblings like run_login_command and build_wheel by naming the target environment.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a strong when-to-use rule: INDISPENSABLE for GPU nodes because they are aarch64 while the login node is x86_64, so login-produced artifacts are unusable. It implicitly contrasts with login-node work but never names an alternative tool explicitly to route the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

build_wheelA

Compile un paquet Python en roue binaire sur un noeud de la bonne architecture et la depose dans un depot local. Evite de recompiler les extensions C++/CUDA (deepspeed, flash-attn, bitsandbytes) a chaque nouvel environnement. Les roues ainsi produites sont ensuite trouvees automatiquement par romeo_pip_install.

ParametersJSON Schema
NameRequiredDescriptionDefault
archNoarmgpu
confirmNo
git_urlNo
minutesNo
packageYes
with_gpuNo
build_flagsNo
cpus_per_taskNo
spack_packagesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=false, covering the safety profile. The description adds substantive behavioral context (architecture-matched node, deposit into a local repo, downstream discovery by romeo_pip_install), but omits the presence of a confirmation gate (the 'confirm' param) and the fact that this is a long, resource-heavy compile job.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core action, then the rationale, then the integration point. No wasted text, though it could be slightly denser given the number of undocumented parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described, and annotations cover safety. However, for a 9-parameter mutating build tool with 0% schema coverage, the description leaves key operational parameters (confirm, resource limits) unexplained, so it is adequate but incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 9 parameters, so the description carries the full burden. It adds meaning for only 'package' and 'arch' (bonne architecture); the build-resource and control params (confirm, minutes, cpus_per_task, git_url, with_gpu, build_flags, spack_packages) are left entirely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+destination: 'Compile un paquet Python en roue binaire' on an architecture-matched node, deposited into a local repository. It also situates itself relative to the sibling romeo_pip_install, which consumes the produced wheels, so an agent can tell the two apart immediately.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear motivating context ('Evite de recompiler les extensions C++/CUDA... a chaque nouvel environnement'), implying when the tool is worth using. It names the downstream tool romeo_pip_install but does not state an explicit alternative to use instead or a when-not-to-use condition, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cancel_jobB

Annule un job en file ou en cours.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and destructiveHint=false, so the mutation/safety profile is partly covered. The description adds the useful constraint that only queued or running jobs can be cancelled, but says nothing about whether the cancellation is reversible, what happens to partial output, or that job_id must reference an active job.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short sentence, fully front-loaded, with no filler. It is efficient, though at the cost of omitting detail rather than saying too much.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and the tool has one required parameter. Still, for a state-changing action the description omits what the cancellation does to the job and its artifacts, leaving the agent slightly under-informed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the sole parameter job_id has no description in either the schema or the tool description. The description does not compensate for this gap; however, job_id is largely self-evident from context, so it is not a total void.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Annule un job') and adds scope ('en file ou en cours'), which is more precise than the name alone. It does not name or distinguish itself from siblings such as job_status or diagnose_job, but the action is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the scope clause — this is for jobs that are queued or running, not finished ones. There is no explicit when-to-use guidance, no mention of alternatives (e.g., checking job_status first), and no prerequisites stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

diagnose_jobA
Read-only

Autopsie d'un job en echec, en un seul appel : etat SLURM, fin des journaux, causes reconnues et remedes. Reconnait notamment le binaire x86 execute sur un noeud aarch64, la saturation de VRAM, le depassement de memoire, l'erreur de bus, les echecs NCCL et les quotas. En cas de depassement de temps, cherche aussi les points de reprise disponibles. A appeler des qu'un job echoue, avant toute autre investigation.

ParametersJSON Schema
NameRequiredDescriptionDefault
linesNo
job_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true, so the description carries most of the burden and does so well: it discloses that it aggregates SLURM state and log tails, performs pattern-based cause recognition, and additionally searches for resume points on timeout. It omits cost/latency and how deep it scans, but the behavioral picture is substantially richer than the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four compact sentences, front-loaded with the single-call promise and the list of recognized causes. Every clause adds signal; only the enumeration is slightly long, which is justified by the router value it provides.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return structure need not be explained, and the failure-mode coverage is thorough for a read-only diagnostic. The only material hole is parameter guidance, which the 0%-coverage schema pushes back onto the description and which is not addressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for two parameters, so the description must compensate — and it does not. The 'lines' parameter (log tail length, default 80) and the format/expectation of 'job_id' are never mentioned, leaving the agent to infer their meaning entirely from property names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (diagnose/autopsy a failed job) and enumerates the exact aggregates it produces: SLURM state, log tails, recognized causes and remedies. The named failure modes (x86 binary on aarch64 node, VRAM saturation, OOM, bus error, NCCL failures, quotas) clearly separate it from siblings like job_output or job_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit trigger and priority: 'call as soon as a job fails, before any other investigation,' plus a conditional extension for time-limit cases. It does not, however, name the alternative tools (job_output, job_status) an agent might otherwise reach for, so the routing is implied rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

download_from_romeoC

Rapatrie un fichier ou un repertoire depuis ROMEO vers la machine locale.

ParametersJSON Schema
NameRequiredDescriptionDefault
verifyNo
recursiveNo
local_pathYes
remote_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and destructiveHint=false, so the safety profile is covered. The description adds nothing beyond the basic operation: it does not say whether existing local files are overwritten, how the 'verify' integrity check behaves, or what happens on failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero waste. It is efficient, though its brevity contributes to the coverage gaps noted elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

While an output schema exists, the description is inadequate for a 4-parameter transfer tool: no usage guidance, no parameter semantics, and no behavioral context on overwriting or verification. The agent lacks what it needs to call this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for 4 parameters, so the description must compensate and does not. 'fichier ou repertoire' loosely hints at the recursive flag, but remote_path, local_path, and verify are left entirely undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource with direction: 'Rapatrie un fichier ou un repertoire depuis ROMEO vers la machine locale.' An agent can tell it moves data from remote to local, distinguishing it from upload_to_romeo by direction, though no sibling is named.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use context, no prerequisites, and does not mention alternatives such as read_remote_file or upload_to_romeo. The directionality is inferable from 'depuis ROMEO vers la machine locale', but nothing guides selection explicitly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

export_job_reportA

Exporte hors de Git une fiche JSON/Markdown d'un job soumis via ce MCP : script Slurm filtre, provenance, ressources mesurees et SHA-256 de fichiers distants explicitement choisis. Lecture seule sur ROMEO ; cree des fichiers locaux prives. live=false autorise l'export hors ligne.

ParametersJSON Schema
NameRequiredDescriptionDefault
liveNo
job_idYes
code_dirNo
data_filesNo
output_dirNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds real context beyond the annotations: the operation is read-only on ROMEO but creates private local files, and live=false enables offline export. The annotations (readOnlyHint=false, destructiveHint=false) alone would leave the agent unsure whether the cluster is mutated, so this clarification is genuinely valuable. It stops short of describing permissions or where files land.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence, front-loaded with the export purpose before the side-effects clause. No filler, though the run-on structure packs multiple ideas together.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values needn't be explained. But for a 5-param tool with 0% schema coverage, the description omits semantics for code_dir and output_dir and gives no hint of file layout or naming, leaving gaps an agent would need.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 5 params. The description explains the live flag's offline behavior and implies data_files are 'explicitly chosen' remote files, but code_dir, output_dir, and job_id receive no added meaning. It partially compensates but leaves several params unexplained in both schema and prose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (exporter) and resource (fiche JSON/Markdown d'un job soumis via ce MCP) and enumerates the report contents (script Slurm filtré, provenance, ressources mesurées, SHA-256). An agent can tell it apart from job_output or job_efficiency. It does not name a sibling alternative, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clarifies the offline case (live=false autorise l'export hors ligne) and the read-only-on-ROMEO constraint, which implies when the tool is applicable. However, it never states when to choose this over job_output or profile_report, nor any prerequisites, so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inject_io_stagingA
Read-only

Insere la mise en cache d'un jeu de donnees en memoire vive dans un script sbatch existant, que ce serveur n'a pas genere. Rend le script modifie sans rien ecrire : a toi de le relire puis de le deposer. Pour un job cree ici, prefere submit_job(stage_archive=...).

ParametersJSON Schema
NameRequiredDescriptionDefault
scriptYes
variableNoDATASET_DIR
dataset_archiveYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true, and the description reinforces that with 'Rend le script modifie sans rien ecrire' plus the follow-up step (you must re-read and submit it yourself). This is useful non-obvious behavioral context; it stops short of covering failure modes or what 'mise en cache' injects exactly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with the action first, the no-write caveat second, and the alternative routing last. Every clause earns its place; nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and the description covers scope, side-effect behavior, and alternatives. The only shortfall is the undocumented 'variable' parameter, which is minor against the rest of the coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for all three parameters, so the description carries the full burden. It obliquely covers 'script' (existing sbatch script) and 'dataset_archive' (the dataset), but the 'variable' parameter and its DATASET_DIR default are never mentioned, leaving a real gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: injecting in-memory dataset caching into an existing sbatch script that this server did not generate. It explicitly distinguishes itself from the sibling submit_job(stage_archive=...), so an agent can route correctly without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives both the when ('un script sbatch existant, que ce serveur n'a pas genere') and the when-not with a named alternative ('Pour un job cree ici, prefere submit_job(stage_archive=...)'). It also states the post-call obligation (re-read then deposit the script), leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

job_efficiencyA
Read-only

Efficacite reelle d'un job termine : CPU, memoire, GPU alloues, plus des recommandations de redimensionnement. Remplace seff, absent de ROMEO. A lire systematiquement apres un job pour calibrer le suivant.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation already establishes that this is a safe, non-mutating read, so the bar is lower. The description adds the meaningful precondition that the job must be finished (it will not work on running jobs) plus the fact that it returns sizing recommendations, which is genuine context beyond the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core purpose, then the equivalence to seff, then the usage instruction. Every sentence earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description needn't explain return values, and it correctly avoids doing so. For a single-parameter read-only tool it is nearly complete; the only gap is any hint about where job_id comes from or its expected form.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description says nothing about the single job_id parameter, so it does not compensate for the schema gap. The parameter name is largely self-explanatory, but no format or sourcing guidance is given, so it adds no meaning beyond the field title.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource: it reports the real efficiency (CPU, memory, GPU) of a finished job, and explicitly scopes it to completed jobs ('un job termine'). That scope qualifier alone distinguishes it from live-metric siblings such as job_live_metrics and job_status without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear usage condition ('A lire systematiquement apres un job pour calibrer le suivant'), telling the agent exactly when to call it. It does not, however, name an alternative tool or state when not to use it, which is what keeps this from a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

job_energy_footprintA
Read-only

Estime l'energie consommee et l'empreinte carbone d'un job. ATTENTION : ROMEO n'active aucun greffon de comptabilite energetique SLURM, le resultat est donc un MODELE et non une mesure. Sur un job en cours, la puissance GPU reelle est relevee, ce qui reduit fortement l'incertitude.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes
gpu_load_factorNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, it discloses a critical behavioral caveat: the output is a simulation, not a measurement, because ROMEO runs no SLURM energy-accounting plugin. It also explains exactly when uncertainty drops (running job → real GPU power sampled), which is precisely the kind of epistemic framing an agent needs before trusting the numbers.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core purpose and then the accuracy caveat. The capitalized ATTENTION warning is deliberate and earns its prominence because it prevents misuse of the result.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is unnecessary, and the description covers the accuracy model and the running-job nuance well. The one real gap is the unexplained gpu_load_factor parameter, which an agent must guess at.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description never mentions either parameter. In particular, gpu_load_factor (default 0.6) is a significant modeling knob whose meaning and effect on the estimate are left entirely undocumented, so the description fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource in its first clause ('estimates the energy consumed and carbon footprint of a job'), which is immediately distinguishable from adjacent tools like job_efficiency, profile_job or job_live_metrics. An agent can identify the tool's domain without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear usage context: results are a model when SLURM energy accounting is unavailable, and accuracy improves substantially for in-progress jobs where GPU power is actually sampled. It does not, however, name an alternative tool or state exclusions (e.g. when to prefer job_efficiency instead).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

job_live_metricsA
Read-only

Telemetrie instantanee d'un job EN COURS, sans lire de journal ni attendre la fin : occupation et memoire des GPU, temperature, puissance, et processus les plus actifs. Detecte le cas ou le calcul dort sur des entrees-sorties pendant que les GPU sont reserves, et la montee vers la saturation de VRAM avant qu'elle ne provoque un echec.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only assert readOnlyHint=true; the description adds real behavioral context beyond that — it is instantaneous/non-blocking, requires no log reading, and does not need the job to finish. It does not cover edge behavior such as what is returned when the job has already ended, so it is strong but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence captures the core purpose and scope, followed by a second sentence covering diagnostic value. Dense but every clause earns its place; the only mild cost is the length of the enumeration.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, and the description adequately conveys what is measured and why. Missing only operational caveats such as prerequisites (job must be running) or behavior for a completed job.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description never mentions the job_id parameter or its format. The single required parameter is largely self-explanatory from its name, so this is acceptable rather than damaging, but the description contributes no meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource and mode ('Telemetrie instantanee d'un job EN COURS') and enumerates exactly what it reports (GPU occupancy/memory, temperature, power, top processes). It also carves itself out from siblings by specifying 'sans lire de journal ni attendre la fin', which separates it from log-reading and wait-for-completion tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'sans lire de journal ni attendre la fin' implicitly routes the agent here instead of job_output or wait_for_job, and it names two concrete diagnostic situations (I/O-sleep while GPUs reserved, VRAM-saturation climb). It stops short of explicitly naming the alternative tools or stating prerequisites such as the job needing to be active.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

job_outputA
Read-only

Sortie d'un job, tronquee par defaut. N'affiche JAMAIS un log entier : utilise lines pour la fin du fichier et grep pour cibler. Le mode stream='auto' montre la sortie d'erreur si elle n'est pas vide.

ParametersJSON Schema
NameRequiredDescriptionDefault
grepNo
linesNo
job_idYes
streamNoauto
max_charsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true; the description adds the critical disclosure that output is truncated by default and that a full log is never returned, plus the stderr behavior of stream='auto'. That is genuine behavioral context beyond the annotation, though truncation limits themselves are not quantified.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the truncation default so the agent reads the most important constraint first. No filler; only `max_chars` is left unexplained.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return structure needn't be described, and the key behavioral caveats (truncation, stderr in auto mode) are covered. The unaddressed `max_chars` parameter is the one remaining gap for a 5-param tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 5 parameters, so the description must carry the load. It meaningfully explains `grep`, `lines`, and `stream`, but says nothing about `max_chars` (which directly controls the truncation it warns about) and only implicitly covers `job_id`.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Sortie d'un job') and immediately qualifies the default behavior (truncated). It is distinguishable from job_status/list_jobs, but it never names or contrasts those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete conditional guidance: use `lines` to see the tail, use `grep` to target, and stream='auto' surfaces stderr when non-empty. It lacks any when-not/alternative-tool routing (e.g. vs read_remote_file), but the operational context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

job_stack_traceA
Read-only

Capture la pile d'appels des processus d'un job EN COURS, pour diagnostiquer un blocage : interblocage MPI, noyau CUDA fige, attente sur verrou. Utilise pstack, avec repli sur gdb puis eu-stack. Ne s'attache pas de maniere interactive : il preleve une trace et rend la main, ce qui n'interrompt pas le calcul.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes
process_nameNo
max_processesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description discloses the fallback chain (pstack -> gdb -> eu-stack), that attachment is non-interactive, and crucially that the sampling does not interrupt the running computation. That is exactly the behavioral context an agent needs before touching a live job.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose, then the mechanism, then the non-intrusive guarantee in a few dense sentences with no filler. Slightly information-heavy for its length but every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and readOnlyHint covering the safety profile, the description is largely complete on purpose and behavior. The one real gap is parameter semantics, which nothing else fills since schema coverage is 0%.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for all three parameters (job_id, process_name, max_processes), so the description must carry the burden and does not. It never mentions that a running job_id is required, what process_name filters, or what max_processes limits (default 3).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Capture la pile d'appels des processus d'un job EN COURS') and scopes it to in-progress jobs, immediately separating it from the many job-status siblings. The stated diagnostic intent (interblocage MPI, noyau CUDA fige, attente sur verrou) makes the tool's niche unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear triggering context: use it when a job is blocked/hung, with three concrete hang categories enumerated. It does not name an alternative sibling (e.g., diagnose_job) or state when NOT to reach for it, so it stops short of explicit routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

job_statusB
Read-only

Etat d'un job : file d'attente puis historique si le job est termine. Inclut le demarrage estime pour un job en attente.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true already establishes the safe-read profile, so the description's added value is the state-machine detail: queued vs finished, and the estimated start for pending jobs. It says nothing about polling cadence, staleness of the estimate, or retention of finished-job history.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the state progression front-loaded and no filler; every clause carries information about what the caller receives.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be detailed, and the description still conveys the lifecycle semantics (queue then history) that matter for interpreting a status call. Only the absence of any usage routing among many job-related siblings keeps it from being fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single required job_id is undocumented, but it is a self-explanatory identifier with no ambiguity in type or role. Minimum-viable baseline is appropriate since the parameter needs little elaboration.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (job state) and enumerates what it returns: queue position, history once finished, and estimated start for pending jobs. That distinguishes it from job_output and list_jobs by implication, but it never explicitly names a sibling the way the top-tier example does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance and no mention of alternatives such as job_live_metrics, job_system_health, or wait_for_job, despite a crowded set of job-inspection siblings. Usage is only implied by the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

job_system_healthA
Read-only

Sante systeme d'un job EN COURS, GPU ou non : charge processeur face aux coeurs reserves, attente d'entrees-sorties, memoire utilisee face a la reservation. Detecte les trois gaspillages classiques : un code cense etre parallele qui tourne sur un seul coeur, un calcul qui passe son temps a attendre le systeme de fichiers, et une reservation memoire massivement surdimensionnee.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds substantive behavioral context beyond that: it discloses the exact comparison metrics and the three waste patterns it detects, which tells the agent what kind of analysis to expect. It does not discuss permissions, cost, or interpretation, but with annotations present the extra specificity is valuable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose and then efficiently lists the diagnostic signals and waste patterns in one dense sentence. There is no filler, though the colon-and-list structure is slightly long for the amount of routing information provided.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the description need not explain return values. It covers the tool's purpose, scope, and the specific diagnostics it performs, which is sufficient for a one-parameter read-only health check. The only missing piece is any parameter-level detail, but that is a small gap given the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there is one required parameter (job_id). The description never mentions the parameter, its expected format, or its role. While the name 'job_id' is largely self-explanatory, the description does not compensate for the missing schema documentation as required when coverage is low.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource ('santé système d'un job EN COURS') and enumerates the exact signals it inspects (CPU load vs reserved cores, I/O wait, memory used vs reserved) plus the three classic wastes it detects. It clearly distinguishes a running-job health check from a general status or efficiency report, though it does not explicitly contrast itself with sibling tools like job_efficiency or job_live_metrics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly scopes usage to a running job ('EN COURS'), which gives a clear context for when to call it. It adds that the job may be GPU or non-GPU. However, it offers no explicit when-not guidance or alternative routing against siblings such as job_efficiency or diagnose_job.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

launch_interactive_serviceB

Lance un service web sur un noeud de calcul (JupyterLab, TensorBoard, vLLM, MLflow) et rend la commande de pont SSH exacte a executer en local pour y acceder. Attend l'affectation du noeud pour pouvoir composer cette commande.

ParametersJSON Schema
NameRequiredDescriptionDefault
archNo
portNo
modelNo
logdirNo
confirmNo
minutesNo
serviceYes
workdirNo
local_portNo
time_limitNo
cpus_per_taskNo
gpus_per_nodeNo
spack_packagesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and destructiveHint=false, so the mutation/safety profile is partly covered. The description adds genuinely useful context beyond that: it waits for node allocation (blocking behavior) and produces a specific output artifact (the exact local SSH bridge command). It still omits resource consumption, time-limit implications, and whether the launch is cancellable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences that front-load the purpose and then the blocking/output behavior, with no filler. It is appropriately sized for the task, though it could have spent those same words clarifying parameters given the low schema coverage.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The presence of an output schema removes the need to explain return values, but with 13 parameters and 0% description coverage the definition is materially incomplete. An agent cannot know what minutes, time_limit, local_port, or confirm do, which is a serious gap for a tool that provisions compute resources.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 13 parameters, so the schema carries almost no semantic load. The description only hints at the "service" field via its examples and says nothing about port, minutes, time_limit, cpus_per_task, gpus_per_node, spack_packages, local_port, or the other numeric/resource parameters. This leaves most parameters undocumented in both structured data and prose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ("Lance") and resource ("service web sur un noeud de calcul") and lists concrete service examples (JupyterLab, TensorBoard, vLLM, MLflow). It also explains the side effect of returning an SSH bridge command, which clarifies the operation. It does not, however, explicitly differentiate itself from the closest sibling, spawn_remote_workspace.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage through the service examples and the note that the node must be allocated, giving an agent situational context. But it never says when to pick this tool over spawn_remote_workspace or other workspace tools, and offers no exclusions or prerequisites beyond the implicit allocation wait.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_dirC
Read-only

Contenu d'un repertoire distant (taille et date incluses).

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNo.
limitNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true already tells the agent this is a safe read. The description adds that size and date are returned per entry, but with an output schema present that information is already structurally available, so the added value is modest. It says nothing about recursion, error behavior, or remote path conventions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence that front-loads the resource. No padding, though it is arguably too terse given the undocumented parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema covers return values, but with 0% parameter description coverage the definition leaves the agent guessing about path resolution and limit semantics on a remote filesystem — core to calling this correctly. The French-only description amid an otherwise English toolset also adds a small comprehension cost.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and neither parameter is explained in the description. The two parameters (path with default '.', limit with default 100) are unintuitive in a remote-cluster context — relative to what root? does limit truncate or paginate? — and the description compensates for none of this.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource and operation: listing a remote directory's contents. The parenthetical clarifies the listing includes size and date, which distinguishes it from read_remote_file without needing the schema. It does not, however, name a sibling or state how it differs from other remote-file tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus read_remote_file, download_from_romeo, or audit_orphan_files. No prerequisites, no path format hints, no note about what happens with a missing directory.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_jobsA
Read-only

Tes jobs : ceux en file cote SLURM, et ceux soumis via ce serveur (avec leur repertoire de travail et le chemin de leurs logs, meme apres une perte de contexte).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true already establishes this as a safe read, but the description adds genuine behavioral context beyond the annotation: it discloses that submitted jobs are tracked with their working directory and log path, and that this record survives a context loss. That persistence guarantee is non-obvious and valuable, though pagination/limit behavior is not covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence that front-loads the two job categories and then appends the most useful distinguishing detail (working directory, log path, context-loss persistence). Nothing is wasted, though the parenthetical makes it slightly heavy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description adequately covers what the listing contains. The one gap is the undocumented limit parameter, which matters for an agent that needs to know whether results are truncated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single 'limit' parameter has a default of 15 with no explanation anywhere. The description never mentions the limit, pagination, or any bound on results, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Tes jobs') and enumerates the two sources it covers: SLURM-queued jobs and jobs submitted via this server. This clearly separates it from singular siblings like job_status or job_output, though those siblings are not named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the enumeration of what gets listed, giving an agent a reasonable sense of when to reach for it. However, there is no explicit when/when-not guidance, no mention of how it relates to job_status or job_output, and no prerequisites are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

profile_jobC

Profile un calcul GPU avec NVIDIA Nsight Systems et produit un rapport exploitable, au lieu de laisser deviner pourquoi un code est lent. La capture est fenetree pour ne pas produire une trace enorme. Lis ensuite le resume avec profile_report.

ParametersJSON Schema
NameRequiredDescriptionDefault
archNoarmgpu
nameNomcp-profile
commandYes
confirmNo
workdirNo
time_limitNo30m
warmup_stepsNo
cpus_per_taskNo
delay_secondsNo
gpus_per_nodeNo
profile_stepsNo
spack_packagesNo
duration_secondsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=false and destructiveHint=false, so the safety profile is largely covered. The description adds real context: it uses NVIDIA Nsight Systems and windows the capture to avoid an enormous trace. It still omits job-submission behavior, resource consumption, and what the confirm parameter changes, so it is only moderately transparent for a mutating tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core action and followed by the reason and the handoff to profile_report. The phrase 'au lieu de laisser deviner pourquoi un code est lent' is persuasive padding, but overall the description is efficient and easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 13-parameter mutating tool with an output schema, the description is too thin. The output schema means return values need not be explained, but the description should still cover the profiled command, why windowing exists, and what the agent must provide or confirm. It leaves the parameter surface and operational requirements almost entirely unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are 13 parameters with 0% schema description coverage, and the description explains none of them. The passing phrase 'La capture est fenetree' vaguely hints at the profiling-window parameters (profile_steps, warmup_steps, duration_seconds, delay_seconds), but it does not map to any argument, define units, defaults, or meaning. With this many undocumented parameters, the description fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb and resource: 'Profile un calcul GPU avec NVIDIA Nsight Systems et produit un rapport exploitable.' That is specific enough for an agent to know it launches a GPU profiling run. It does not differentiate itself from the sibling tool_profile, which appears to be a related profiling entry point, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a motivating context ('au lieu de laisser deviner pourquoi un code est lent') and a follow-up instruction ('Lis ensuite le resume avec profile_report'). However, it never says when to use this instead of alternatives such as tool_profile, nor does it state prerequisites or exclusions. Usage is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

profile_reportA
Read-only

Resume le rapport de profilage d'un job traite par profile_job : repartition entre calcul et transferts memoire, et noyaux les plus couteux. Condense une sortie nsys de plusieurs centaines de lignes en quelques constats.

ParametersJSON Schema
NameRequiredDescriptionDefault
topNo
job_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true already tells the agent this is a safe read. The description adds that the operation is a condensation of a several-hundred-line nsys output into a few findings, which hints at lossy aggregation, but it does not describe failure behavior when the job was never profiled or any rate/permission constraints. Adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, zero filler, with the core purpose front-loaded before the condensation benefit. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the description need not detail return values, and it correctly communicates the report's general content and the profile_job prerequisite. The one gap is the undocumented 'top' parameter, which leaves invocation details to inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry parameter meaning. It names no parameters: job_id is only implied, and the 'top' parameter (default 5) is never mentioned, even though 'noyaux les plus coûteux' gestures at a top-N concept without explaining that the count is configurable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Résume) and resource (le rapport de profilage d'un job traité par profile_job), and enumerates the actual content: répartition calcul/transferts mémoire and noyaux les plus coûteux. It clearly distinguishes itself from sibling profile_job, which produces the raw profiling rather than the report summary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'd'un job traité par profile_job' establishes a clear prerequisite: the job must first have been profiled. It gives strong usage context but does not state when NOT to use it or name explicit alternatives among the many sibling job tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_docA
Read-only

Lit le texte original d'une page locale ROMEO. Utiliser les read_args de search_docs ou choisir start_line/end_line (inclusives, base 1). max_chars borne la reponse. Si truncated=true, rappeler avec next_call pour recuperer la suite exacte ; offset est relatif a la plage choisie. expected_sha256 detecte un changement de page entre deux lectures.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageYes
offsetNo
end_lineNo
max_charsNo
start_lineNo
expected_sha256No

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation already establishes the safety profile, but the description adds substantial behavioral detail: max_chars bounds the response, truncated=true requires a follow-up with next_call, offset is relative to the selected range, and expected_sha256 detects page changes between reads. This is rich, useful behavioral context beyond the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with purpose, then usage, then pagination behavior. It is dense but well-structured, and every sentence carries actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists, the description need not explain return values, and it still covers the important pagination, continuation, and integrity-check behaviors. Together with the readOnlyHint annotation, an agent has enough to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema description coverage at 0%, the description must compensate, and it does for most parameters: start_line/end_line inclusivity and 1-based indexing, max_chars response limiting, offset relativity, and expected_sha256 change detection. The required page parameter's exact identifier semantics are left implicit rather than fully explained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: reads the original text of a local ROMEO page. It is clear what the tool does and references search_docs for range arguments, but it does not explicitly differentiate read_doc from other read-oriented siblings such as read_remote_file.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear operational guidance: use read_args from search_docs, or choose start_line/end_line. It also explains how to continue a truncated read. However, it does not explicitly state when not to use this tool or compare it against alternatives beyond the search_docs integration hint.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_remote_fileB
Read-only

Lit une tranche d'un fichier distant. Toujours borne : precise offset et limit plutot que de rapatrier un fichier entier.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
limitNo
offsetNo
max_charsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true already tells the agent this is a safe, non-mutating read. The description adds the genuinely useful constraint that reads are bounded by range, but says nothing about error behavior for out-of-range offsets, file-size limits, or remote access/permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and followed by the key constraint. Nothing is padded or redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value documentation is correctly omitted, and the slicing intent is communicated. However, with 0% schema coverage and four parameters, the silent max_chars default and the path format are gaps an agent would likely need when calling the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry parameter meaning, and it only covers offset and limit. The path parameter and especially max_chars (a non-obvious truncation cap defaulting to 8000) are left completely unexplained in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: reading a slice of a remote file, with the scoping intent ('toujours borne') made explicit. It distinguishes itself from whole-file retrieval, though it never names the obvious sibling (download_from_romeo) or other readers like read_doc.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives operational guidance ('precise offset et limit plutot que de rapatrier un fichier entier'), which implies the preferred usage pattern, but offers no when-to-use/when-not versus alternatives such as download_from_romeo, list_dir, or read_doc.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

romeo_fairshare_forecastB
Read-only

Estime l'effet d'une charge envisagee sur la part d'usage du compte, donc sur la priorite des jobs suivants de l'equipe. Lit l'usage courant via sshare et le compare a la consommation projetee.

ParametersJSON Schema
NameRequiredDescriptionDefault
duration_hoursNo
simulated_cpusNo
simulated_gpusNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation already establishes this is a safe read operation. The description adds useful mechanism context by stating it reads current usage via sshare and compares it with projected consumption, but it does not describe return values, limitations, or how the comparison is performed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences: the first states the purpose and downstream impact, the second states the mechanism. It is front-loaded and contains no redundant or filler language.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. The purpose and mechanism are covered, but with three completely undocumented parameters and no explicit routing among many sibling tools, the definition is only minimally complete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for all three parameters. The description mentions a 'charge envisagée' and projected consumption conceptually, but it never maps those concepts to duration_hours, simulated_cpus, or simulated_gpus, leaving parameter meanings to inference from names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: estimating the effect of a planned load on the account's usage share and therefore on team job priority. It clearly distinguishes the tool from generic job submission or monitoring siblings, though it does not explicitly name a sibling alternative such as romeo_quota.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The implied use case is clear: forecast priority impact before submitting a planned load. However, there is no explicit when-to-use or when-not-to-use guidance, and no alternative tool is named for comparison.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

romeo_modulesA
Read-only

Liste les modules Environment Modules, herites de l'ancien calculateur. Sur ROMEO 2025 la voie officielle est Spack : utilise romeo_software en premier lieu, et ne recours aux modules que si un logiciel n'existe pas dans le catalogue Spack.

ParametersJSON Schema
NameRequiredDescriptionDefault
searchNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so safety is covered. The description adds meaningful behavioral context beyond annotations: the modules are legacy/inherited, and the tool is a fallback rather than the official path. It does not describe output or search behavior, but the output schema covers returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences: the first front-loads what the tool lists, the second immediately routes the agent to the preferred alternative. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only listing tool with an output schema and readOnlyHint, the description is complete on purpose and usage. The only notable gap is the undocumented `search` parameter, which keeps it from a perfect score.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the single `search` parameter is never mentioned in the description. The parameter name is intuitive, but the description does not compensate for the missing schema documentation by explaining what search filters or how it works.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: it lists Environment Modules inherited from an older machine. It explicitly distinguishes itself from the sibling romeo_software by naming it as the official path.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when/when-not guidance: use romeo_software first, and only fall back to modules if a software package is absent from the Spack catalog. This is exactly the routing information an agent needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

romeo_pip_installA

Installe des paquets Python dans un environnement virtuel, sur un noeud de la bonne architecture, en privilegiant les roues deja compilees par build_wheel. Evite de recompiler les extensions natives et n'utilise jamais le noeud de login, dont les roues seraient en x86_64.

ParametersJSON Schema
NameRequiredDescriptionDefault
archNoarmgpu
confirmNo
minutesNo
env_pathYes
packagesYes
extra_flagsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this is a write operation (readOnlyHint=false) that is not destructive. The description usefully adds that execution is routed to a compute node of the matching architecture and never the login node, which is real operational context. However it says nothing about what happens on failure, whether the environment is mutated in place, or how confirmation/timeouts behave.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence that front-loads the action and the architecture constraint, with the login-node caveat at the end. Efficient and free of filler, though the multiple subordinate clauses make it slightly harder to scan than a two-sentence split would be.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. Still, for a mutating installer with 6 parameters at zero schema coverage, the description leaves the agent without semantics for confirm, minutes or extra_flags, which matters for calling it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 6 parameters, so the description must carry the burden and does not. It only indirectly conveys the packages and arch concepts ('noeud de la bonne architecture'); env_path, confirm, minutes and extra_flags are entirely undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (installs Python packages), the target scope (a virtual environment, on a node of the correct architecture), and names the sibling tool build_wheel as the source of preferred pre-built wheels. An agent can distinguish this from build_wheel, build_on_node and run_login_command without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear operational guidance: prefer wheels already compiled by build_wheel, avoid recompiling native extensions, and never use the login node because its wheels are x86_64. It names an alternative (build_wheel) but does not spell out the inverse condition (when to skip pip install and call build_wheel directly), so it stops short of explicit when/when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

romeo_quotaA
Read-only

Quotas de stockage reels, par espace (home, scratch, projet), lus avec mmlsquota. Signale les depassements de quota souple et l'expiration du delai de grace, qui bloque l'ecriture. A consulter avant de produire des sorties volumineuses.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so safety is covered; the description adds real behavioral context beyond that: it reports soft-quota exceedances and grace-period expiry, and notes that exceeding the grace period blocks writes. That write-blocking consequence is genuinely useful operational knowledge.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences with no filler, and the most decision-relevant content (what it reads and the write-blocking risk) is front-loaded. The usage cue is properly placed last as the call-to-action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-format explanation is not required, and the read-only nature is carried by annotations. The description covers what an agent needs to decide to call it; only the parameter semantics remain thin.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single `project` parameter is undocumented in the schema, so the description must compensate. It partially does by naming the space categories (home, scratch, projet), hinting at what `project` selects, but gives no format, optionality, or default semantics for the parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (real storage quotas per space: home, scratch, project) and a concrete mechanism (`mmlsquota`), which distinguishes it from generic cluster-status siblings like romeo_status. It does not, however, explicitly name a sibling it differs from or scope what it does NOT cover.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to consult it: "A consulter avant de produire des sorties volumineuses" (check before producing large outputs), which is a clear usage trigger. No explicit exclusions or alternative tools are named, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

romeo_selfcheckA
Read-only

Confronte le modele de cluster encode dans ce serveur a la realite de SLURM : partitions et limites de temps, comptes de noeuds par architecture, capacites d'un noeud, plafonds du compte, outils declares absents. Rend les ecarts, sans rien corriger. A lancer quand un refus de dimensionnement parait injustifie, ou apres une maintenance du cluster.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true; the description adds real behavioral context beyond that — it enumerates the categories of discrepancy it reports (partitions/limits, node counts per architecture, node capacities, account ceilings, absent tools) and explicitly states 'sans rien corriger' (it corrects nothing), confirming a pure diagnostic. No rate limits or auth notes, so not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded: the comparison action comes first, the enumeration of what is checked next, and the usage trigger last. Three dense sentences with no filler. Slightly list-heavy but every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained; the description instead conveys scope, non-mutating behavior, and invocation triggers. For a zero-parameter diagnostic this is nearly complete, only missing explicit sibling disambiguation against run_cluster_sanity_check.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters, so there is no parameter semantics to explain; the baseline of 4 applies. The description correctly implies the tool takes no inputs by describing a self-contained comparison.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource: it confronts the server's encoded cluster model against SLURM reality (partitions, time limits, node counts by architecture, account caps, absent tools). Clearly a diagnostic/audit operation rather than a query. However it does not distinguish itself from the sibling 'run_cluster_sanity_check', which sounds functionally adjacent, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit trigger conditions are given: run it when a sizing refusal looks unjustified, or after cluster maintenance. That is concrete when-to-use guidance. It names no alternative (e.g. romeo_status, run_cluster_sanity_check) and states no when-not, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

romeo_softwareA
Read-only

Cherche un logiciel dans le catalogue Spack de ROMEO, plusieurs centaines de paquets qui constituent la voie officielle de chargement des logiciels. ATTENTION : le catalogue differe selon l'architecture. Des outils absents du PATH du noeud de login (conda via anaconda3, apptainer, plusieurs versions de cuda) s'y trouvent. Les paquets trouves se passent ensuite a submit_job ou build_on_node via spack_packages.

ParametersJSON Schema
NameRequiredDescriptionDefault
archNoarmgpu
limitNo
searchNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true already marks this as a safe read, and the description adds real context beyond that: the catalogue varies by architecture, and tools missing from the login-node PATH (conda, apptainer, multiple cuda versions) are only reachable here. It could say more about result shape/pagination, but the architecture caveat is valuable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose, then the architecture warning, then the downstream handoff. Three focused sentences with little waste, though the parenthetical tool list is slightly dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description covers purpose, the key architecture caveat, and the follow-on tools. Only the parameter meanings remain thin for a 3-param tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for all three params (arch, limit, search). The description partially compensates by stressing that the catalogue depends on architecture, hinting at the `arch` parameter's importance, but it adds nothing about `limit` or `search`.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (recherche/cherche) and resource (catalogue Spack de ROMEO), scoping it as the official software-loading path. An agent can distinguish it from romeo_modules or build_on_node without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states the downstream workflow: found packages are passed to submit_job or build_on_node via `spack_packages`. It gives a clear context for use but does not explicitly contrast with the sibling romeo_modules tool, leaving some routing ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

romeo_statusA
Read-only

Etat de ROMEO en un appel : partitions, noeuds libres par architecture, tes jobs en cours et ta part d'ordonnancement. A appeler avant de dimensionner un job.

ParametersJSON Schema
NameRequiredDescriptionDefault
include_queueNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation already establishes that this is a safe read, so the description's job is lighter. It usefully frames the tool as a single aggregated snapshot rather than many calls, but says nothing about permissions, cost, or how include_queue alters behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, with the returned content front-loaded and the usage note last. No filler; the agent gets contents and timing immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no prose, and read-only safety is annotation-covered. The remaining gap is the undocumented include_queue parameter, which is the one thing an agent must decide before calling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the sole parameter include_queue (default true) is never mentioned in the description. The description lists result categories but leaves the one input entirely unexplained, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Etat de ROMEO en un appel') and enumerates the exact payload: partitions, free nodes per architecture, running jobs, and scheduling share. This makes the tool's output scope unambiguous, though it does not explicitly contrast itself with close siblings like romeo_quota or romeo_fairshare_forecast.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The final sentence ('A appeler avant de dimensionner un job') gives a concrete calling context, telling the agent when this tool is appropriate. It stops short of naming an alternative for other situations, so no explicit routing to siblings is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_cluster_sanity_checkA
Read-only

Verifie la sante des GPU sur un ou plusieurs noeuds : bridage thermique ou de puissance, erreurs memoire non corrigees, frequence anormalement basse. Un noeud degrade ne plante pas, il ralentit tout un job reparti sans erreur visible. Rend une clause --exclude prete a l'emploi pour les noeuds suspects.

ParametersJSON Schema
NameRequiredDescriptionDefault
nodesNo
minutesNo
max_nodesNo
check_typeNogpu

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true already establishes the safe-read profile, so the description is not obligated to restate safety. It adds genuine behavioral context: the failure mode is silent slowdown rather than a crash, and the output is a ready-to-use --exclude clause for suspect nodes. It does not mention runtime cost or how long a sanity sweep takes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with no filler: what is checked, why it matters, and what comes back. The important scoping information is front-loaded before the rationale.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value detail is not required, though the description helpfully previews the --exclude output. The gap is on the input side: with zero schema description coverage and undocumented parameters like check_type and minutes, an agent cannot confidently set non-default values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across four parameters, so the description must compensate and largely does not. 'Un ou plusieurs noeuds' loosely maps to the nodes parameter, but minutes, max_nodes, and check_type (whose enumeration of allowed values is undisclosed) are never explained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (verifie/verify) and a precisely scoped resource: GPU health on one or more nodes, enumerating the exact conditions checked (thermal/power throttling, uncorrected memory errors, abnormal frequency). This distinguishes it from general health siblings like job_system_health or romeo_selfcheck, and signals its Slurm-oriented output (--exclude clause).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the motivating scenario clearly: a degraded node does not crash but silently slows a distributed job, which tells the agent to run this before submitting spread work. No explicit 'use X instead of Y' routing versus siblings such as diagnose_job or job_system_health, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_login_commandA

Execute une commande COURTE et LEGERE sur le noeud de login (inspection, git, ls, grep). Les compilations, installations et calculs sont refuses par defaut : le noeud de login est partage par tout le laboratoire. allow_heavy=true leve ce refus pour les cas que la documentation ROMEO autorise explicitement, comme un pip install en environnement virtuel a destination du x86_64. Delai maximal 20 s.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNo
commandYes
allow_heavyNo
timeout_secondsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations specify readOnlyHint=false and destructiveHint=false. The description adds valuable operational context: shared login node constraints, the 20-second timeout, the allow_heavy flag mechanism, and the policy-based refusal behavior. However, it does not mention possible return values, error modes, or how the refusal manifests, which could be useful given the output schema exists but its contents are not described here.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core purpose and constraint, followed by the exception condition and the timeout. Every sentence adds essential information without waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (command execution on a shared node with policy restrictions) and the presence of an output schema, the description covers the key behavioral constraints, the allow_heavy mechanism, and the timeout. It could mention the default timeout (15s) to complement the stated maximum (20s), and briefly note output/error handling, but overall it is nearly complete for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description compensates well for the two non-obvious parameters: allow_heavy is explained in detail (lifts heavy-command refusal, only for documented cases), and timeout_seconds is implied by 'Delai maximal 20 s' although this actually overrides the default 15s, which could cause confusion. The command and cwd parameters are self-explanatory by name, but cwd semantics (default behavior) are not described. Overall solid compensation but with a minor gap on the timeout default.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (run) and resource (login node command) and names the command categories allowed (inspection, git, ls, grep). It clearly distinguishes from siblings like submit_job or run_cluster_sanity_check by emphasizing short/light commands on the shared login node, not compute jobs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says heavy operations (compilations, installations, calculations) are refused by default, and that allow_heavy=true lifts the refusal only for cases explicitly authorized by ROMEO documentation, giving a concrete example (pip install within a venv targeting x86_64). This tells the agent when to use this tool versus when to use submit_job or build_on_node.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sbatch_lintA
Read-only

Verifie un script de soumission AVANT de l'envoyer : fins de ligne Windows qui corrompent l'interpreteur, directives #SBATCH placees trop tard, --mem manquant que ROMEO exige, chemins inexistants, variables non definies, secrets ecrits en clair. Accepte le texte du script ou le chemin d'un script deja depose sur le cluster.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNo
scriptNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation already signals a safe, non-mutating operation. Beyond that, the description discloses exactly what the lint checks for (Windows line endings, late #SBATCH directives, missing --mem, nonexistent paths, undefined variables, plaintext secrets) and that it accepts either script text or a cluster path, adding concrete behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the purpose and timing constraint, followed by the specific validations and input modes. Every clause earns its place by listing concrete checks or clarifying input handling; no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only lint tool with an output schema, the description supplies the purpose, checks performed, and input modes, so return values need not be explained. It could mention error behavior or required cluster access, but otherwise it is complete enough for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description must carry parameter meaning, and it does: it maps 'script' to script text and 'path' to a script already deposited on the cluster, and signals that either can be supplied. It does not clarify precedence when both are provided or whether they are mutually exclusive, but the core semantics are covered.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Verifie) and resource (un script de soumission) and distinguishes itself from submission siblings via 'AVANT de l'envoyer'. The list of checks (line endings, #SBATCH placement, missing --mem, etc.) further sharpens the purpose beyond the name alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'AVANT de l'envoyer' clearly establishes the usage context: run this before submitting a job. It implies the alternative (submit_job) without naming it, and gives no explicit exclusions, but the timing condition is clear enough for an agent to select it correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_docsA
Read-only

Recherche locale classee par pertinence dans la documentation ROMEO livree avec le MCP. Accepte mots-cles ou questions, sans distinction d'accents/casse. mode='phrase' cherche un motif exact. page_prefix restreint une branche (ex. ressources/romeo_2025/). max_chars borne le total des extraits. Chaque resultat donne source, titres, lignes et read_args pour lire la section complete. Suivre next_call pour parcourir les autres resultats sans perte par troncature.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoterms
queryYes
offsetNo
max_charsNo
max_resultsNo
page_prefixNo
context_linesNo
expected_revisionNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so safety is covered. The description adds genuinely useful behavior beyond that: accent/case-insensitive matching, exact-pattern phrase mode, that max_chars bounds total extract size, and that truncation is avoidable by following next_call. Minor gap: no direct statement about scope of the doc set or result ordering beyond 'pertinence'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A compact, front-loaded paragraph where each sentence adds a distinct fact (scope, matching rules, phrase mode, branch restriction, size cap, result contents, pagination). No filler, though the density borders on terse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-shape explanation is not required, but the description still leaves four of eight parameters (offset, max_results, context_lines, expected_revision) unexplained for a non-trivial search tool. Adequate but with clear gaps for an 8-parameter definition.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 8 parameters, so the description must compensate. It meaningfully explains mode, page_prefix (with an example path), and max_chars, and implies query semantics, but leaves offset, max_results, context_lines, and expected_revision entirely undocumented in both schema and description. Partial compensation only.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: a local, relevance-ranked search over the ROMEO documentation bundled with the MCP. It distinguishes itself from read_doc implicitly by noting results carry read_args to read the full section. Clear enough to select without opening the schema, though read_doc is never named as the contrasting alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete context (accepts keywords or free-form questions, accents/case insensitive, mode='phrase' for exact patterns) and tells the agent to follow next_call to page through results. However it never states when to use this versus read_doc or other siblings, leaving the routing decision implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

secret_env_setupA
Read-only

Explique et prepare le stockage des variables sensibles (jetons, mots de passe) pour les jobs. Les valeurs ne transitent JAMAIS par ce serveur : tu les ecris toi-meme dans un fichier a droits restreints, que submit_job source au demarrage via secret_env_file. Elles n'apparaissent ainsi ni dans le script sbatch, ni dans le registre local, ni dans cette conversation.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNosecrets.env

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true is already given, but the description adds substantial context beyond it: secrets never traverse the server, the agent must write the restricted-permission file itself, and the values stay out of the sbatch script, local registry, and conversation. This is genuinely useful security-model disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with the core constraint (values never transit the server) front-loaded, and every sentence carries information. No filler or restatement of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the security concept is well covered. However, for a tool whose only parameter is undocumented in both schema and description, the definition leaves a real gap in how the agent should actually call it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has one parameter (`name`, default 'secrets.env') with 0% description coverage, and the description never mentions it at all. It references `secret_env_file`, but that is a submit_job parameter, not this tool's, so the sole argument is left entirely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a recognizable verb+resource combination ('explique et prepare le stockage des variables sensibles pour les jobs') that clearly demarcates it from siblings like write_remote_file or submit_job. The word 'prepare' is slightly vague about what the tool actually produces, but the security-centered intent is unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Tells the agent when this applies (handling tokens/passwords for jobs) and ties the workflow to submit_job's `secret_env_file` sourcing at startup. It does not name an explicit alternative or state when NOT to use it, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

spawn_remote_workspaceA

Ouvre un espace de travail interactif sur un noeud GPU et rend tout ce qu'il faut pour s'y connecter : commande de pont SSH a coller, et URL locale avec son jeton d'authentification. Contrairement a launch_interactive_service, le jeton est genere ici, donc l'URL est utilisable immediatement sans lire la sortie du job.

ParametersJSON Schema
NameRequiredDescriptionDefault
archNoarmgpu
gpusNo
portNo
confirmNo
workdirNo
local_portNo
time_limitNo2h
cpus_per_taskNo
spack_packagesNo
workspace_typeNojupyter

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=false, so the mutation/safety profile is covered. The description adds that the returned token makes the URL immediately usable, which is genuinely useful behavioral context, but it omits any mention of resource consumption (GPU/time_limit), confirmation semantics, or cleanup expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences: purpose and output first, then the differentiator against the sibling. Every clause earns its place with no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values needn't be restated, but for a 10-parameter tool with zero schema coverage, a mutating spawn that consumes GPU resources, the description leaves too much undocumented — notably the confirm flag, port/local_port, and resource sizing parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Ten parameters with 0% schema description coverage and the description adds essentially no parameter meaning: arch, gpus, port, confirm, workdir, local_port, time_limit, cpus_per_task, spack_packages and workspace_type are all unaddressed. Only a vague signal ('GPU node', 'interactive workspace') hints at the resource parameters, leaving confirm and local_port entirely opaque.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Ouvre un espace de travail interactif sur un noeud GPU') and enumerates what it produces (SSH bridge command, local URL with token). It explicitly contrasts itself with the sibling launch_interactive_service, so an agent can distinguish the two without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the alternative (launch_interactive_service) and the condition that favors this tool: the token is generated here so the URL works immediately without reading job output. That is clear routing guidance, though it doesn't state explicit prerequisites or when this tool should be avoided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stage_datasetB

Telecharge un jeu de donnees depuis un noeud de calcul plutot que depuis le noeud de login, dont la bande passante est partagee. Accepte une URL directe, un depot Hugging Face ou un depot git.

ParametersJSON Schema
NameRequiredDescriptionDefault
archNox64cpu
kindNoauto
sourceYes
confirmNo
minutesNo
time_limitNo
destinationYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and destructiveHint=false, covering the mutation/safety profile. The description adds useful context: the download occurs on a compute node rather than the login node, and supports URL/Hugging Face/git sources. However it says nothing about the confirm flag, timeout/minutes behavior, or prerequisites, which are exactly the behavioral traits an agent would need beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the key rationale (compute node vs shared-bandwidth login node) front-loaded and the accepted source types following. No filler, though the second sentence could be folded more economically.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter tool with 0% schema coverage, an output schema, and no parameter documentation, the description is not complete enough: the role of confirm, minutes, time_limit, and arch is unexplained, and there is no indication of what staging returns or how long it takes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 7 parameters, so the description must compensate, and it only partially does: it explains the source kinds (direct URL, Hugging Face repo, git repo), which maps to source/kind. It says nothing about destination, arch, confirm, minutes, or time_limit, leaving most parameters undocumented in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action (download/stage a dataset) and a specific resource, plus a distinctive rationale (run on a compute node instead of the bandwidth-shared login node) and the accepted source types. It is clear what the tool does, though it does not explicitly name the sibling it replaces (e.g. download_from_romeo/upload_to_romeo) to remove ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies the when-to-use case (you want staging to happen on a compute node because the login node's bandwidth is shared), but there is no explicit when-not-to-use guidance and no alternatives such as download_from_romeo or run_login_command are named. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

storage_cleanup_helperB
Read-only

Repere ce qui occupe le stockage : plus gros repertoires, journaux de jobs anciens, points de reprise volumineux, environnements virtuels dupliques. Ne supprime rien ; propose les commandes de menage. A utiliser quand romeo_quota signale un depassement.

ParametersJSON Schema
NameRequiredDescriptionDefault
topNo
pathNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true already tells the agent this is a safe, non-writing operation. The description usefully reinforces and extends this by stating it deletes nothing and instead proposes cleanup commands, which clarifies the output is advisory. It adds real value but no depth on scope, cost, or discovery behavior beyond the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences, front-loaded with what it detects before stating the non-destructive behavior and the trigger condition. No filler; only the missing parameter detail keeps it from being fully economical.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-format explanation is not required, and the readOnly safety profile is covered by annotations plus the 'supprime rien' sentence. However, with 0% schema description coverage and no scope guidance for path/top, the definition leaves an agent unable to invoke it precisely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for both parameters, so the description carries the full burden, yet it never mentions 'top' (how many results) or 'path' (what gets scanned). An agent cannot tell from the text whether path scopes the scan or what the default empty path means. This leaves both parameters effectively undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (spot/repere) and resource (storage) and enumerates concrete targets: biggest directories, old job logs, large checkpoints, duplicated virtual envs. This is far more informative than the generic name. It does not explicitly differentiate itself from close siblings such as audit_orphan_files, which keeps it from a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger: 'A utiliser quand romeo_quota signale un depassement', naming the sibling tool whose output routes the agent here. No when-not condition or alternative (e.g. audit_orphan_files) is offered, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

submit_array_jobA

Soumet un balayage parametrique en tableau SLURM : une tache par jeu de parametres, avec un plafond de taches simultanees. Chaque tache recoit sa ligne de parametres dans la variable $PARAMS, que ta commande peut interpoler. Ideal pour une recherche d'hyperparametres ou une evaluation sur plusieurs jeux de donnees. Simulation par defaut, comme submit_job.

ParametersJSON Schema
NameRequiredDescriptionDefault
archNo
nameYes
mem_gbNo
commandYes
confirmNo
modulesNo
workdirNo
parametersYes
time_limitNo1h
cpus_per_taskNo
gpus_per_nodeNo
max_concurrentNo
spack_packagesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare non-read-only and non-destructive behavior, and the description adds valuable detail: one task per parameter set, $PARAMS interpolation, concurrency cap, and simulation by default. It does not cover confirmation requirements or submission side effects, but adds substantial context beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, front-loaded with purpose, and each sentence adds distinct value (mechanics, variable interpolation, use case, simulation default). No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 13-parameter submission tool, the description covers purpose and some behavior but leaves most parameter semantics unexplained. Output schema exists to cover return values, but the parameter documentation gap makes it only adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description only explains the parameters array concept and the max_concurrent cap, leaving 11 other parameters (arch, mem_gb, modules, workdir, time_limit, cpus_per_task, gpus_per_node, spack_packages, confirm, etc.) completely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (soumet) and resource (balayage parametrique en tableau SLURM), with mechanics (one task per parameter set, cap on concurrency). It implicitly distinguishes from submit_job by being array-based and references submit_job for simulation default, making its niche clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context: ideal for hyperparameter search or evaluation on multiple datasets. However, it does not explicitly state when not to use it or name alternatives besides the simulation-default reference to submit_job.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

submit_jobA

Prepare et soumet un job SLURM. distributed='mpi' couvre le calcul parallele courant (prefixe srun, avec ou sans GPU) ; les familles ddp, accelerate, deepspeed et srun sont propres a PyTorch et ajoutent son point de rendez-vous. Options : container pour une image Apptainer, redirect_caches pour detourner les caches Python hors du home, stage_archive pour mettre un jeu de donnees en memoire vive. data_files choisit les entrees a empreinter au demarrage pour export_job_report (20 fichiers, 64 Mio). En simulation par defaut : rend le script sbatch genere, la partition et l'architecture deduites, et les avertissements de dimensionnement, SANS rien soumettre. Relance avec confirm=true pour soumettre reellement. La partition est deduite du temps demande et l'architecture du besoin en GPU : ne les force que si tu as une raison precise.

ParametersJSON Schema
NameRequiredDescriptionDefault
archNo
nameYes
arrayNo
nodesNo
mem_gbNo
commandYes
confirmNo
modulesNo
workdirNo
cpu_bindNo
containerNo
partitionNo
data_filesNo
job_tmpdirNo
nccl_debugNo
time_limitNo1h
distributedNo
cpus_per_taskNo
gpus_per_nodeNo
keep_patternsNo
stage_archiveNo
spack_packagesNo
ntasks_per_nodeNo
redirect_cachesNo
secret_env_fileNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and destructiveHint=false; the description adds substantial context beyond that – the safe dry-run-by-default flow, the confirm gate to trigger real submission, and the automatic partition/architecture deduction with a warning against overriding. This is exactly the kind of mutation-behavior disclosure annotations cannot carry. No contradiction with the write annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense paragraph that front-loads the core action and dry-run behavior. Every sentence carries information, though the parameter enumeration in the middle is slightly packed and the value would be higher if the most-used params led.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 25-parameter, write-capable tool the description covers the distinctive knobs and the simulation/confirm lifecycle well, and an output schema exists so return values need not be explained. But the majority of parameters are still undocumented given 0% schema coverage, leaving a meaningful gap for an agent configuring a job.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% across 25 parameters, so the description must compensate, and it partially does by explaining distributed, container, redirect_caches, stage_archive, data_files (with the 20-file/64 MiB cap), confirm, partition, and arch. Roughly two-thirds of the parameters (nodes, mem_gb, modules, time_limit, gpus_per_node, array, secret_env_file, etc.) remain undocumented in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Starts with a specific verb+resource ('Prepare et soumet un job SLURM') and characterizes the distributed families it orchestrates, so the agent knows exactly what it generates. It does not, however, draw the boundary against close siblings like submit_array_job, submit_resilient_job, or submit_pipeline, so differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear usage context: runs in simulation by default (returns the sbatch script, deduced partition/arch, sizing warnings) and requires confirm=true to actually submit. It also says not to force partition/arch unless there is a specific reason. It stops short of naming when to prefer an alternative sibling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

submit_pipelineA

Soumet un enchainement de jobs relies par des dependances SLURM : preparer, calculer, rassembler. Chaque etape decrit une intention (commande, temps, ressources) et les etapes dont elle depend ; le serveur ordonne, valide chacune comme submit_job le ferait, et pose les --dependency. L'architecture declaree pour l'enchainement est heritee par toutes les etapes, ce qui evite qu'une etape sans GPU parte sur x86_64 alors que les autres tournent en aarch64. SIMULATION PAR DEFAUT : rappelle avec confirm=true pour soumettre.

ParametersJSON Schema
NameRequiredDescriptionDefault
archNo
nameYes
stagesYes
confirmNo
modulesNo
spack_packagesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only say readOnlyHint=false and destructiveHint=false. The description adds substantial behavior: server-side ordering and per-stage validation, automatic placement of --dependency, arch inheritance across stages (with the x86_64/aarch64 failure case it prevents), and the critical dry-run-by-default gate. This is exactly the context annotations cannot supply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded, then mechanics, with the simulation warning placed last where it is most likely to be read before calling. Dense but each sentence adds a distinct fact; slightly longer than strictly necessary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. For a mutation tool, the simulate/confirm contract and the dependency/arch behaviors are covered well. The main gap is the three undocumented parameters (name, modules, spack_packages) that an agent may need to populate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden. It meaningfully documents stages (intention: command, time, resources, dependencies), arch (inherited across the chain), and confirm (the submit gate), covering half the parameters with real semantics. It leaves name, modules, and spack_packages undocumented, but the coverage it does provide is substantive rather than restating field names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource: submitting a chain of jobs linked by SLURM dependencies, with a clear lifecycle (prepare, compute, gather). It even references submit_job as the validation reference. It does not, however, disambiguate against submit_array_job or submit_resilient_job, which are the closest siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly explains the operational flow: simulation is the default, and the caller must re-invoke with confirm=true to actually submit. That is strong actionable guidance. It stops short of saying when to prefer this over submit_job or submit_array_job for dependency-free work.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

submit_resilient_jobA

Soumet un calcul long sous forme de chaine de segments reprenables, pour depasser la limite de temps d'une partition rapide. Chaque segment recoit SIGUSR1 avant son expiration pour sauvegarder, et le suivant demarre apres lui via une dependance, en reprenant du dernier point de sauvegarde. Un marqueur de fin fait sauter les segments restants si le calcul se termine avant terme. Ton code doit savoir reprendre depuis checkpoint_dir et, idealement, traiter SIGUSR1.

ParametersJSON Schema
NameRequiredDescriptionDefault
archNo
nameYes
mem_gbNo
commandYes
confirmNo
workdirNo
segment_timeNo1h
cpus_per_taskNo
gpus_per_nodeNo
signal_beforeNo
stage_archiveNo
checkpoint_dirNo
max_total_timeNo6h
spack_packagesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover non-readOnly and non-destructive status, while the description adds rich runtime behavior: segment chain execution, SIGUSR1 before expiration, dependency-based continuation, resume from checkpoint_dir, end-marker skipping, and the requirement that user code handle resumes and ideally SIGUSR1. This is exactly the operational detail annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, front-loaded with purpose followed by the execution mechanism, with no obvious filler. It is appropriately sized for the complex execution model, though it could better organize the parameter implications.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 14-parameter submission tool with no schema descriptions, the description omits most input parameters and any alternative-tool guidance. Output schema exists so return values need not be explained, but invocation-critical parameter context is largely absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 14 parameters, so the description must carry parameter meaning. It only mentions checkpoint_dir and links SIGUSR1 to the signal mechanism, while omitting name, command, segment_time, max_total_time, cpus_per_task, gpus_per_node, confirm, workdir, arch, mem_gb, stage_archive, and spack_packages.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: submits a long computation as a chain of resumable segments to exceed a fast partition's time limit. It does not name or contrast with sibling tools like submit_job or submit_array_job, so sibling differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for use: long computations that need to exceed a fast partition's time limit via checkpointed segments. No explicit when-not condition or alternative tool is named, but the intended scenario is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

suggest_submission_slotC
Read-only

Recommande ou soumettre en confrontant la file d'attente a l'etat du parc : quelle partition et quelle architecture demarreraient le plus vite pour la taille de job envisagee.

ParametersJSON Schema
NameRequiredDescriptionDefault
hoursNo
nodesNo
gpus_per_nodeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, yet the description says the tool may 'soumettre' (submit), implying a write/side-effecting action. This directly conflicts with the read-only annotation. No other behavioral context (cost of query, freshness of queue data) is disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence with the recommendation goal front-loaded and the comparison basis (queue vs. cluster state) appended. No filler, though the 'ou soumettre' clause adds ambiguity rather than value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, but with zero parameter documentation, no usage guidance, an annotation-contradicting verb, and a language mismatch (French description vs. English tool/params), the definition is not complete enough for reliable invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there are 3 undocumented parameters (hours, nodes, gpus_per_node). The phrase 'taille de job envisagee' loosely gestures at job size, but the description never maps to the individual parameters or explains defaults/units, so it does little to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description conveys a specific function: comparing the queue to cluster state to recommend which partition/architecture would start fastest for a given job size. However, 'Recommande ou soumettre' ('recommend or submit') blurs whether this is an advisory tool or an action tool, and it never names the sibling it complements (e.g. submit_job, romeo_status). The core idea is graspable but not crisply bounded.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is an implied context (use before submitting a job to pick the best slot), but no explicit when-to-use, when-not-to-use, or alternatives such as romeo_status or romeo_quota. An agent is left to infer the workflow position from the description alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tool_profileA

Consulte ou change le profil d'outils pour cette connexion : essential ou full. full reaffiche immediatement les outils avances. Aucun job ni fichier distant n'est modifie.

ParametersJSON Schema
NameRequiredDescriptionDefault
profileNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=false; the description adds real context by clarifying that no job or remote file is modified and that 'full' takes effect immediately. This usefully reassures the agent about side-effect scope, though it does not detail what switching profiles costs or affects beyond remote jobs/files.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero waste, with the action and the parameter domain front-loaded ahead of the behavioral reassurance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-optional-parameter tool with an output schema present, the description is nearly sufficient: it names the enum values, behavior of 'full', and the non-destructive blast radius. The only minor gap is not clarifying the omitted-parameter (consult) case explicitly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the load — and it does, defining the single optional parameter's domain ('essential' or 'full') and the effect of 'full'. It does not spell out what 'essential' hides or what a null/omitted value does, but the 'consult or change' framing implies the latter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: consult or change the tool profile ('profil d'outils') for this connection, and names the two possible profiles. This is distinguishable from sibling administrivia like profile_report or profile_job, though it never explicitly contrasts with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the 'consult or change' framing and the enumerated values, and it notes that 'full' immediately redisplays advanced tools. But it gives no explicit when-to-use guidance, no prerequisites, and no when-not-to-use or alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upload_to_romeoC

Envoie un fichier ou un repertoire local vers ROMEO.

ParametersJSON Schema
NameRequiredDescriptionDefault
verifyNo
local_pathYes
remote_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and destructiveHint=false; the description adds nothing about overwriting existing remote files, whether the transfer is resumable, size limits, or authentication requirements. For a mutation/transfer tool it does not enrich the behavioral picture beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence that is front-loaded with the core action and destination, with no filler. It is appropriately sized, though its brevity partly reflects the missing detail noted elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, but for a transfer tool with three undocumented parameters and no annotation coverage of overwrite/auth behavior, the description is too thin to call correctly without external assumptions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there are three parameters, yet the description mentions none of them. It does not clarify the meaning of local_path, remote_path, or the verify flag (default true), leaving the agent to guess parameter semantics entirely.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (envoie/upload) and resource (fichier ou repertoire local) with a clear destination (ROMEO). It implicitly distinguishes itself from download_from_romeo by naming direction, though it doesn't call out the sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as write_remote_file or download_from_romeo, nor any prerequisites (credentials, quota, target filesystem). The agent must infer usage purely from the verb.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wait_for_jobA
Read-only

Attend qu'un job se termine, avec un plafond strict (600 s). Utilise-le seulement pour un job court : sinon reviens interroger job_status.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes
poll_secondsNo
timeout_secondsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true (safe read), so the description adds real behavioral context by disclosing a hard 600 s cap and the intent to wait rather than poll repeatedly. It does not describe return shape, but an output schema exists. Minor tension: the cap (600 s) sits alongside an undocumented timeout_seconds default of 120, leaving their interaction unexplained.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the core action and the hard constraint front-loaded, and the fallback alternative appended. No filler, though it could have used one clause to name the key parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be described, and usage + the time cap are covered. However, for a 3-parameter tool with 0% schema coverage, the total silence on parameters (especially timeout vs. the stated 600 s cap) leaves a gap an agent must guess at.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description is the only source of parameter meaning, yet it mentions none of job_id, poll_seconds, or timeout_seconds and does not explain polling cadence or how timeout interacts with the cap. It fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Attend qu'un job se termine' / wait for a job to finish) and immediately bounds the behavior with a strict ceiling. It is clearly distinct from a plain status check, though the differentiation from siblings is carried mostly by the usage sentence rather than the purpose statement itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit when-to-use ('seulement pour un job court' / only for a short job) and an explicit alternative plus condition ('sinon reviens interroger job_status' / otherwise poll job_status). The routing decision is unambiguous and names the sibling directly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

write_remote_fileB

Ecrit ou remplace un fichier texte distant (script, configuration).

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
contentYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=false, so the mutation/safety profile is partly covered. The description usefully adds that existing files are replaced and that content is limited to text, but it omits whether parent directories must exist, permission requirements, or encoding behavior. Note: 'remplace' (replaces) has mild tension with destructiveHint=false, but replacing content is not clearly irreversible destruction, so this is not flagged as a contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence, front-loaded with the verb, with no filler. It is appropriately sized and wastes nothing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. However, for a mutating tool with 0% parameter description coverage and no usage guidance, the description should at least address path semantics and overwrite implications; it does not.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across both required parameters (path, content), so the description carries the full burden. It implies text content via 'fichier texte' but says nothing about path format (absolute/relative, remote convention) or size/encoding limits, leaving a real gap for a write tool with two required params.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb+resource ('écrit ou remplace un fichier texte distant') and clarifies scope with examples ('script, configuration') and a text-only constraint. It does not differentiate itself from close siblings like upload_to_romeo or read_remote_file, so it lands at 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus upload_to_romeo, download_from_romeo, or build_on_node, all of which touch remote files. The parenthetical examples describe file kinds, not selection criteria, so the agent is left to infer usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 47 tool updatesv1.4.0
    • First observedallocate_debug_node
    • First observedaudit_orphan_files
    • First observedbuild_on_node
    • First observedbuild_wheel
    • First observedcancel_job
    • First observeddiagnose_job
    • First observeddownload_from_romeo
    • First observedexport_job_report
    • First observedinject_io_staging
    • First observedjob_efficiency
    • First observedjob_energy_footprint
    • First observedjob_live_metrics
    • First observedjob_output
    • First observedjob_stack_trace
    • First observedjob_status
    • First observedjob_system_health
    • First observedlaunch_interactive_service
    • First observedlist_dir
    • First observedlist_jobs
    • First observedprofile_job
    • First observedprofile_report
    • First observedread_doc
    • First observedread_remote_file
    • First observedromeo_fairshare_forecast
    • First observedromeo_modules
    • First observedromeo_pip_install
    • First observedromeo_quota
    • First observedromeo_selfcheck
    • First observedromeo_software
    • First observedromeo_status
    • First observedrun_cluster_sanity_check
    • First observedrun_login_command
    • First observedsbatch_lint
    • First observedsearch_docs
    • First observedsecret_env_setup
    • First observedspawn_remote_workspace
    • First observedstage_dataset
    • First observedstorage_cleanup_helper
    • First observedsubmit_array_job
    • First observedsubmit_job
    • First observedsubmit_pipeline
    • First observedsubmit_resilient_job
    • First observedsuggest_submission_slot
    • First observedtool_profile
    • First observedupload_to_romeo
    • First observedwait_for_job
    • First observedwrite_remote_file

TDQS

B3.1/5.0

Scored across 47 tools

Disambiguation3/5

Descriptions are detailed and most tools target distinct actions, but several clusters overlap: launch_interactive_service vs spawn_remote_workspace, storage_cleanup_helper vs audit_orphan_files, and the multiple job-diagnostic/monitoring tools require careful reading to choose correctly. Boundaries are usually explained, yet the sheer number of near-neighbours invites misselection.

Naming Consistency3/5

snake_case is used throughout and the verb_noun pattern is common, but conventions are mixed: some tools carry a romeo_ prefix (romeo_status, romeo_quota) while analogous ones do not (search_docs, list_dir), and many tools are noun-first (job_output, job_efficiency, tool_profile). Readable but not fully predictable.

Tool Count2/5

At 47 tools this is far beyond the well-scoped 15-tool range; even for a complex HPC workflow it is heavy and necessitates profile gating via tool_profile (essential/full). Many tools are individually justified, but the volume itself is a significant usability cost.

Completeness4/5

The surface covers the full job lifecycle: submission variants (batch, array, resilient, pipeline), monitoring, diagnostics, profiling, file transfer, storage management, interactive services, builds, docs and quotas. Missing pieces such as remote file deletion/move or directory creation are minor and can be worked around via run_login_command.

Maintenance

ActivityMaintained
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables Claude Code to interact with a TACC or SLURM HPC cluster for bioinformatics pipelines, allowing job management, log reading, file browsing, remote script execution, and job submission through natural language.
    -
  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables AI assistants to manage SLURM HPC clusters via SSH. Supports job submission, resource monitoring, queue management, and file operations.
    8 npm
    4
    -
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables managing OpenI platform resources (login, query nodes, submit jobs, view logs) via natural language in Claude or Codex, following a kubectl/docker-style CLI.
    133
    MIT