Skip to main content
Glama

@imenam/mcp-feature-analyzer

Une table de relecture pour le code écrit par un agent IA, avec une interface graphique où un humain relit, trie et tranche.

L'agent qui vient de développer une feature ouvre une analyse : le serveur fige le diff git, l'agent y ajoute un résumé, des constats classés par gravité et des schémas. L'utilisateur relit le tout dans la GUI, fichier par fichier, ignore ce qui ne compte pas, annote, puis enregistre une décision. L'agent relit ensuite exactement ce qui a été retenu, sous la forme d'un prompt de correction prêt à suivre.

Agent ──create_analysis / add_findings / set_diagram──▶  Analyse  ◀──relecture, remarques──  Utilisateur (GUI)
  ▲                                                                                              │
  └──────────── get_review_feedback ──── décision, points retenus, prompt ◀── « Enregistrer le retour »

Le serveur n'appelle aucun LLM et n'écrit jamais dans le dépôt : il lit git, conserve l'analyse, et affiche.


Installation

npm install -g @imenam/mcp-feature-analyzer

Ou directement via npx, sans installation. Le serveur communique en JSON-RPC sur stdio : il est lancé par le client MCP (Claude Code, par exemple), pas à la main. Par défaut, il tourne sur la machine qui héberge le dépôt et y exécute git lui-même ; derrière un gateway, sur une autre machine, voir Derrière mcp-http-gateway.

Déclaration dans un projet

Depuis la racine du projet :

npx -y @imenam/mcp-feature-analyzer --claude-setup-mcp

La commande crée (ou complète) le .mcp.json du répertoire courant avec une entrée mcp-feature-analyzer dont les variables sont pré-remplies. Ajoutez --force pour remplacer une entrée existante ; les autres serveurs sont préservés. Un .mcp.json illisible fait échouer la commande, sans rien écrire.

Deux flags choisissent directement comment la GUI sera exposée :

# GUI derrière le proxy central
npx -y @imenam/mcp-feature-analyzer --claude-setup-mcp --proxy_url=http://localhost:3000

# GUI en accès direct, sans proxy : http://localhost:4600
npx -y @imenam/mcp-feature-analyzer --claude-setup-mcp --app_port=4600

Les deux flags sont exclusifs : la commande échoue si les deux sont passés. Sans l'un ni l'autre, PROXY_URL reçoit le modèle http://localhost: dont il reste à compléter le port.

Exemple d'entrée générée dans C:/dev/mon-projet avec --proxy_url=http://localhost:3000 :

{
  "mcpServers": {
    "mcp-feature-analyzer": {
      "command": "npx",
      "args": ["-y", "@imenam/mcp-feature-analyzer"],
      "env": {
        "PROXY_URL": "http://localhost:3000",
        "APP_PORT": "",
        "APP_GROUP": "Mon-projet",
        "APP_NAME": "Feature Analyzer Mon Projet",
        "APP_PATH": "/feature-analyzer-mon-projet",
        "MCP_FEATURE_ANALYZER_DATA_DIR": "C:/dev/mon-projet/.feature-analyzer-data",
        "MCP_LOG_DIR": ""
      }
    }
  }
}

Une variable laissée vide est inactive.

Variables d'environnement

Elles se renseignent dans le champ env de .mcp.json, ou dans un fichier .env à la racine du package, qui prime et est relu à chaque relance de la GUI.

Variable

Rôle

PROXY_URL

URL du proxy central : le proxy attribue le port et monte la GUI sous APP_PATH. Prime sur APP_PORT.

APP_PORT

Port local utilisé quand PROXY_URL est absent : la GUI est servie directement sur http://localhost:<port>, à la racine.

APP_PATH

Chemin de montage de la GUI derrière le proxy (défaut /feature-analyzer). Ignoré en mode APP_PORT.

APP_PATH_PREFIX

Préfixe de namespace appliqué au chemin par le client proxy partagé ; le <base href> de la GUI suit le chemin réellement enregistré.

APP_NAME

Nom affiché dans le dashboard du proxy (défaut Feature Analyzer).

APP_GROUP

Section repliable du dashboard du proxy (facultatif).

MCP_FEATURE_ANALYZER_DATA_DIR

Répertoire de stockage des analyses. À défaut : MCP_DATA_DIR, puis <package>/.feature-analyzer-data.

MCP_LOG_DIR

Répertoire des logs. Défaut : <données>/logs. Lu au démarrage, depuis la déclaration du serveur.

MCP_FEATURE_ANALYZER_ROLE

standalone (défaut) : git s'exécute sur la machine du serveur. remote : le serveur est derrière mcp-http-gateway et fait exécuter git sur la machine de l'agent. Une autre valeur fait refuser le démarrage. Lu au démarrage.

PROXY_URL et APP_PORT s'excluent : PROXY_URL prime si les deux sont renseignés. Si aucune des deux n'est définie, la GUI est désactivée ; les outils MCP continuent de fonctionner. Un APP_PORT qui n'est pas un entier entre 1 et 65535 désactive aussi la GUI : le worker s'arrête en écrivant la raison sur sa sortie d'erreur.

Derrière mcp-http-gateway

Placé dans un gateway (mcp-http-gateway, éventuellement piloté par le gateway manager), le serveur tourne sur la machine du gateway et n'y trouve pas les dépôts de l'agent. En rôle remote, create_analysis et refresh_analysis font exécuter git et la lecture des fichiers sur la machine de l'agent, par son relais mcp-http-gateway --mcp (convention _gateway_exec). Le snapshot figé est ensuite conservé côté gateway, et les autres outils comme la GUI n'ont plus besoin du dépôt.

Côté gateway, dans le bloc env du serveur :

{
  "MCP_SERVER_ROUTE": "/feature-analyzer",
  "MCP_FEATURE_ANALYZER_ROLE": "remote"
}

Côté agent, la route doit figurer dans la liste blanche d'exécution du relais :

{
  "mcpServers": {
    "gateway": {
      "command": "npx",
      "args": ["-y", "@imenam/mcp-http-gateway", "--mcp"],
      "env": {
        "GATEWAY_URL": "https://gateway.example.com",
        "GATEWAY_TOKEN": "…",
        "GATEWAY_EXEC_ROUTES": "/feature-analyzer"
      }
    }
  }
}
  • repo_path est le chemin absolu de la racine du dépôt sur la machine de l'agent. git et node doivent y être dans le PATH : node lit les fichiers de la copie de travail.

  • Sans GATEWAY_EXEC_ROUTES, create_analysis et refresh_analysis échouent en indiquant la route à ajouter ; aucun git n'est lancé côté gateway.

  • Le relais limite la sortie de chaque commande à 1 Mio. Le diff est donc demandé fichier par fichier :

    • le diff d'un seul fichier qui dépasse 1 Mio fait échouer l'analyse, avec le nom du fichier ;

    • un fichier modifié de plus de 1 Mio est conservé sans contenu, comme en local ;

    • un fichier non suivi de plus de 1 Mio fait échouer l'analyse en mode working_tree.

  • Les commandes sont envoyées au relais par paquets de 25. Le relais accepte 50 allers-retours par appel d'outil, soit environ 550 fichiers modifiés par analyse.

  • L'endpoint natif /mcp du gateway n'a pas de relais agent : create_analysis et refresh_analysis n'y fonctionnent pas en rôle remote.


Related MCP server: CodePeel MCP Server

Les outils

Quatorze outils. Les noms, descriptions et textes de résultat sont en anglais ; chaque résultat de modification se termine par le lien vers l'analyse dans la GUI (Open the review: <url>#/<id>) quand la GUI est active.

Analyses

Outil

Effet

list_analyses

Point d'entrée : les analyses regroupées par projet, chacune avec ses refs, ses constats ouverts par gravité, ses fichiers revus, son nombre de schémas et l'état de sa revue ; project limite la liste à un projet (un nom inconnu liste les projets existants). Invite l'agent à reprendre l'analyse existante d'une feature plutôt qu'à en créer une autre. Signale aussi les fichiers d'analyse illisibles.

create_analysis

Calcule et fige le diff : repo_path (racine absolue du dépôt), project (projet ou feature, dossier de rangement, requis), title, mode (branch ou working_tree), base (requis en branch), head (en branch, défaut HEAD), request_text et request_source (la demande initiale : ticket, prompt). Renvoie l'identifiant et la liste des fichiers.

get_analysis

Vue complète : projet, demande, vue d'ensemble, résumé, fichiers (revus ou non), constats avec statut et emplacement, remarques du relecteur, schémas (signalés quand leur tracé a des croisements), état de la revue.

update_analysis

Change title, project, la vue d'ensemble (objective, approach, attention_points), summary (les changements fonctionnels de la feature, une puce par changement, remplace toute la liste) ou la demande initiale (request_text, request_source ; null efface). Les champs absents sont conservés.

refresh_analysis

Recalcule le diff avec le même dépôt, le même mode et les mêmes refs, après correction. Voir Les deux modes de diff.

delete_analysis

Supprime l'analyse et son diff figé. Le dépôt n'est pas touché. Supprime aussi une analyse listée comme illisible.

La vue d'ensemble est l'analyse globale de l'agent : l'objectif fonctionnel de la feature, l'approche retenue (architecture, flux principal) et les points d'attention à vérifier en priorité. Sans vue d'ensemble, objective et approach sont requis ensemble pour la créer (attention_points vaut alors une liste vide s'il est omis) ; ensuite, chaque champ donné remplace le précédent. objective: null retire toute la vue d'ensemble. L'objectif figure aussi en tête du prompt de retour.

L'état de revue est dérivé des fichiers revus et de la soumission (shared/review-state.ts) : not_started (aucun fichier revu), in_progress (au moins un), files_reviewed (tous, revue non soumise), submitted (avec sa décision).

Diff

Outil

Effet

get_diff

Diff figé d'un fichier (path) ou de tous, avec pour chaque ligne son numéro ancien et son numéro nouveau. Les constats, remarques et nœuds s'ancrent toujours sur les numéros du côté nouveau (deuxième colonne).

Constats

Outil

Effet

add_findings

Ajoute un lot de constats. Chacun a une gravité (critical, major, minor, trivial), une nature (issue par défaut, ou requirement_gap pour un comportement demandé mais absent), un titre, un corps, et selon le cas path + start_line (+ end_line), une suggestion (texte de remplacement des lignes visées, conservé tel quel) et un prompt propre. Les lignes visées sont copiées du diff figé. Un seul constat invalide et rien n'est ajouté ; toutes les erreurs sont listées.

update_finding

Modifie les champs d'un constat. Un nouvel emplacement est vérifié contre le diff et rouvre un constat obsolète ; path: null retire l'emplacement (requirement_gap seulement). Le statut « ignoré » appartient au relecteur et ne se change pas ici.

delete_findings

Supprime des constats par identifiant et les retire de la sélection de la revue. Un identifiant inconnu fait échouer l'appel sans rien supprimer.

Un issue a toujours un emplacement ; un requirement_gap peut ne pas en avoir. Une suggestion n'est jamais appliquée : la GUI l'affiche (lignes actuelles barrées, lignes proposées) et le prompt la transmet.

Explications

Une explication décrit ce que fait un bloc de code ajouté ou supprimé, et comment il le fait. Elle aide le relecteur à comprendre le code avant de le juger : ce n'est pas un constat, elle ne demande rien et n'entre pas dans le retour de revue. Les instructions du serveur demandent à l'agent d'en écrire une pour chaque bloc long (une trentaine de lignes ou plus) ou complexe (algorithme, machine à états, concurrence, expression régulière délicate…), et jamais pour du code trivial.

Outil

Effet

add_explanations

Ajoute un lot d'explications : titre, corps, path + start_line (+ end_line) et side. side: "new" (par défaut) vise du code ajouté ou conservé, en numéros du côté nouveau ; side: "old" vise du code supprimé, en numéros du côté ancien (première colonne de get_diff), lignes toutes présentes dans le diff et dont au moins une est supprimée. Les lignes décrites sont copiées du diff figé. Une seule explication invalide et rien n'est ajouté.

update_explanation

Modifie le titre, le corps ou l'emplacement d'une explication. Un nouvel emplacement est vérifié contre le diff et rend son statut « à jour » à une explication obsolète.

delete_explanations

Supprime des explications par identifiant. Un identifiant inconnu fait échouer l'appel sans rien supprimer.

refresh_analysis traite les explications comme les constats : une explication dont les lignes ne correspondent plus au texte copié passe à « obsolète », et redevient à jour si elles correspondent de nouveau.

Schémas

Outil

Effet

set_diagram

Crée un schéma, ou remplace entièrement celui de diagram_id. L'agent ne donne que des nœuds et des liens, sans coordonnées ni couleurs : la GUI place et colore. Le résultat contrôle le tracé : No crossing., ou le nombre de croisements avec chaque paire de liens et chaque lien qui traverse un nœud, suivis de conseils (réordonner les nœuds, découper le schéma, changer de sorte). Le schéma est enregistré dans les deux cas.

delete_diagram

Supprime un schéma.

Trois sortes de schéma (kind) :

  • flow — un process, placé de gauche à droite le long des liens. Une branche secondaire sans jonction et le nœud terminal d'une chaîne descendent sous le nœud qui les précède, ce qui garde le flux compact ; les liens qui remontent le flux passent sous le schéma.

  • layers — une colonne par entrée de layers (GUI, API, cœur, stockage…), chaque nœud dans sa couche : le périmètre d'impact.

  • mindmap — un arbre radial autour de sa racine : la carte des concepts. Le schéma doit être un arbre (une seule racine, un seul parent pour chaque autre nœud, aucun cycle) ; sinon set_diagram le refuse en listant les problèmes.

Règles de lisibilité : un sujet par schéma, 5 à 12 nœuds, nœuds déclarés dans l'ordre de lecture (l'ordre de déclaration est l'ordre de départ de chaque rang ou colonne), aucun croisement attendu. Pour flow et layers, le placement part de l'ordre de déclaration et des ordres obtenus par balayages barycentriques, garde celui dont le tracé final a le moins de croisements et de traversées de nœuds, puis échange deux voisins d'un même rang tant que cela en retire. shared/diagram-quality.ts mesure ces croisements sur le tracé de layoutDiagram.

Chaque nœud a un id, un label, une forme shape (box par défaut, pill, decision en losange) et un statut status qui fixe sa couleur : new, modified, impacted, existing (défaut), finding, missing (dessiné en contour). detail s'affiche au survol ; path (fichier modifié de l'analyse) et line rendent le nœud cliquable vers la revue. Un lien a from, to et un label facultatif.

Retour de revue

Outil

Effet

get_review_feedback

Une fois la revue soumise : décision (approve, request_changes, reject), constats et remarques retenus avec leur détail, et le prompt généré, recopié tel quel. Avant la soumission : l'avancement (fichiers revus, constats ignorés, remarques). Lecture seule.

Le retour se lit avant refresh_analysis, qui le remplace par une revue vierge.

GUI

Outil

Effet

reconnect_gui

Relance le worker GUI et le réenregistre auprès du proxy, par exemple quand le proxy a démarré après le serveur. force: true redémarre même un worker qui semble vivant. En mode APP_PORT, il indique l'état de la GUI et la marche à suivre si elle ne tourne pas.


Boucle de travail type

list_analyses   { project: "mon-projet" }           ← reprendre une analyse existante de la feature ?
create_analysis { repo_path: "C:/dev/mon-projet", project: "mon-projet", title: "Rafraîchissement OAuth",
                  mode: "branch", base: "master", request_text: "…", request_source: "TM-142" }
get_diff        { analysis_id }                     ← numéros de ligne du côté nouveau
update_analysis { analysis_id, objective: "…", approach: "…", attention_points: ["…"],
                  summary: ["…", "…"] }             ← vue d'ensemble et changements fonctionnels
add_explanations { analysis_id, explanations: [...] } ← blocs longs ou complexes, ajoutés ou supprimés
add_findings    { analysis_id, findings: [...] }    ← y compris les requirement_gap
set_diagram     { analysis_id, kind: "flow", nodes: [...], links: [...] }

  … l'agent donne le lien à l'utilisateur, qui relit dans la GUI et enregistre son retour …

get_review_feedback { analysis_id }                 ← décision, points retenus, prompt
  … l'agent corrige le code …
refresh_analysis    { analysis_id }                 ← nouveau diff, nouveau tour de revue

L'interface

Thème sombre. La barre supérieure porte la marque (un clic ramène à l'accueil), le sélecteur d'analyse groupé par projet avec l'état de revue de chaque analyse (sa valeur est la ref de tête, ou « Copie de travail »), l'état de la connexion et, pendant la revue, le bouton « Terminer la revue ».

Accueil (#/)

Les analyses sont rangées par projet, dans des sections repliables. L'en-tête d'un projet compte ses analyses par état de revue. Chaque analyse affiche son titre (coupé par des points de suspension quand la place manque ; une bulle l'affiche en entier au survol), ses refs, la date de sa dernière modification, la progression des fichiers revus, ses constats ouverts par gravité (exigences manquantes comprises, ce que dit l'infobulle) et son état de revue : « Non commencée », « En cours », « Fichiers revus » ou « Soumise », suivi de la décision. Le filtre « À relire » (par défaut) masque les revues soumises, « Toutes » les montre. L'accueil n'ouvre aucune analyse d'office ; sans analyse, il explique que l'agent les crée avec create_analysis.

Écran Feature (#/<id>)

Titre, mode et refs (un sha complet est abrégé à 7 caractères), nombre de fichiers, lignes ajoutées et supprimées, bouton « Commencer la revue » qui ouvre le premier fichier non revu. Quatre onglets, avec leur compteur : fichiers modifiés, constats ouverts (exigences manquantes comprises ; l'infobulle donne aussi le total), schémas :

  • Résumé — la carte « Vue d'ensemble » (objectif, approche, points d'attention) quand l'agent l'a rédigée, « Ce que l'agent a fait », la demande initiale et sa source, la section « Constats ouverts » (compteurs par gravité, constats ouverts du plus grave au moins grave, lien « Voir les N constats, tous statuts confondus ») et « Copier le prompt pour l'agent » (tous les constats ouverts).

  • Fichiers — chaque fichier avec son statut, ses +/−, ses constats ouverts par gravité et son état de revue.

  • Constats — la liste complète, tous statuts, filtrable par gravité et par statut (ouvert, ignoré, obsolète) ; les puces comptent tous les statuts ; un clic ouvre le constat dans la revue.

  • Schémas — sélecteur de schéma, qui passe à la ligne quand les schémas sont nombreux (un schéma dont le tracé présente des croisements porte l'étiquette « croisements »), rendu SVG, légende des statuts présents et des gravités des points de constat dessinés (« Constat ouvert (couleur de la gravité la plus haute) »), contrôles « − », « + », « Ajuster » (recentre et adapte le zoom), « Plein écran », et zoom à la molette. Un clic sur un nœud rattaché à un fichier ouvre la revue à ce fichier et à cette ligne. Un schéma qui ne peut pas être placé (carte mentale qui n'est pas un arbre, par exemple) affiche l'erreur à la place du dessin.

Le plein écran passe par l'API Fullscreen du navigateur : le schéma occupe l'écran, légende et contrôles restent visibles, le zoom est réajusté à l'entrée et à la sortie, Échap ou « Quitter le plein écran » en sortent. Si le navigateur refuse, un message l'indique sous le schéma.

L'ajustement à la fenêtre ne descend jamais sous l'échelle 0,85 : en dessous, les libellés deviennent illisibles. Un schéma plus grand s'ouvre sur son début, à cette échelle, et se parcourt en le faisant glisser (curseur main, indication « Glisser pour parcourir le schéma ») ; la vue reste bornée au schéma.

Écran Revue (#/<id>/review/<chemin>)

  • Panneau des fichiers — progression, fichiers groupés par dossier, état revu, pastilles des constats ouverts par gravité. Un nom trop long est tronqué au milieu (reminder-sch…uler.ts), l'extension reste visible et le chemin complet est en infobulle. Chaque pastille est un lien : elle ouvre le fichier sur son premier constat ouvert de cette gravité (?line=N) ; son infobulle le dit (« 3 constats mineurs ouverts — aller au premier »). Le bouton en tête de l'en-tête du fichier masque ou réaffiche le panneau, pour toute la session : le diff s'élargit vers la gauche, la colonne de droite garde sa largeur. Il reste affiché quand le chemin demandé ne fait pas partie de l'analyse.

  • En-tête du fichier — bascule du panneau des fichiers, chemin, +/−, statut, fichier précédent et suivant, « Remarque sur le fichier », « Marquer comme revu ».

  • Diff — code coloré selon le langage du fichier, déduit de son extension ou de son nom (grammaires de Shiki, chargées à la demande) : le côté nouveau à partir du contenu complet du fichier quand le snapshot le conserve, sinon bloc par bloc ; les lignes supprimées bloc par bloc ; chaque correctif proposé séparément ; les lignes barrées par un correctif restent sans couleur. Un fichier dont le langage n'est pas reconnu s'affiche sans coloration ; une grammaire qui ne se charge pas est signalée dans la barre du diff. Numéros du côté nouveau et repère vertical à la couleur de la gravité le long des lignes visées par un constat. La barre du diff porte trois réglages, communs à tous les fichiers : chacun est un bouton compact qui affiche sa valeur courante, et un clic ouvre la liste de ses choix, chacun avec son icône et une description d'une ligne (un clic dehors, Échap ou un défilement la referme). « Afficher » choisit entre « Modifications » (les blocs du diff git, entourés de leurs lignes de contexte) et « Fichier entier » (tout le fichier en un seul bloc : les lignes inchangées, tirées du contenu conservé, s'intercalent entre les modifications, ce qui donne à lire en entier une fonction modifiée en son milieu ; les deux autres sélecteurs s'y appliquent de la même façon, un correctif dont les lignes chevauchaient deux blocs s'y affiche, et « Copier une demande d'explication » porte alors sur tout le fichier). En « Fichier entier », une mini-carte de 22 px longe le bord droit du diff et reste à l'écran pendant le défilement : toute sa hauteur représente tout le fichier affiché. Sa piste de gauche place les lignes ajoutées (vert), supprimées (rouge), supprimées et ajoutées en vis-à-vis en vue « Côte à côte » (moitié rouge, moitié verte) ou seulement réindentées (violet) ; sa piste de droite porte un point par constat, à la couleur de sa gravité (atténué s'il n'est plus ouvert), et un point par ligne portant une remarque. Un cadre bleuté suit la portion visible. Le survol d'un repère en donne la description (« 12 lignes ajoutées · lignes 120 à 131 », « Constat 3 · Majeur » et son titre) ; un clic ou un glissé sur la carte fait défiler le diff jusqu'à l'endroit visé. « Fichier entier » est inactif pour un fichier dont le contenu n'est pas conservé ; la liste en donne la raison à la place de sa description. « Code » choisit entre « Unifié » (une seule colonne : lignes supprimées sur fond rouge au-dessus des lignes ajoutées sur fond vert), « Côte à côte » (ancien code à gauche avec ses numéros du côté ancien, nouveau code à droite : lignes de contexte des deux côtés, lignes supprimées et ajoutées appariées ligne à ligne, case hachurée en face d'une ligne sans vis-à-vis ; remarques, repères et lignes proposées par un correctif sont du côté droit, le bandeau d'un correctif sur toute la largeur) et « Nouveau code » (le code du côté nouveau seul, dans les mêmes blocs : les lignes supprimées ne sont pas affichées et les lignes ajoutées portent un fin filet vert à la place de leur fond et de leur marque ; une explication de code supprimé s'aligne alors sur la ligne voisine du bloc). « Correctifs » choisit entre « Sans correctif », « Avant / après » et « Code corrigé ». Le correctif proposé s'affiche à l'endroit où il s'applique : bandeau « Correctif proposé par l'agent — remplace les lignes X à Y » avec le numéro du constat et « Masquer le correctif », lignes visées barrées sur fond ambre, lignes proposées juste en dessous sur fond violet (marque « › », sans numéro). Le correctif est affiché d'office pour un constat ouvert ; masqué, ou pour un constat sans correctif, une pastille numérotée marque la première ligne visée. Un filet sarcelle longe les lignes décrites par une explication. En « Modifications », quand le contenu du fichier est conservé, l'en-tête de chaque bloc porte un réglage « −5 · N lignes · +5 » : « +5 » ajoute cinq lignes de contexte au-dessus et cinq en dessous du bloc, « −5 » les retire, et le compteur donne le nombre de lignes de contexte au-dessus de la première modification. Les lignes ajoutées s'insèrent sous l'en-tête et après le bloc, si bien que le réglage ne bouge pas à l'écran ; elles se distinguent par un fond plus sombre et un filet gris (« Contexte ajouté » dans la légende). Les lignes entre deux blocs vont d'abord au bloc du dessus ; quand plus aucune ligne masquée ne les sépare, les deux blocs n'en forment plus qu'un, sous un seul en-tête dont « +5 » et « −5 » s'appliquent à chacun des blocs réunis. « +5 » s'éteint quand le bloc atteint les bords du fichier. Le contexte choisi est propre à chaque fichier et dure le temps de la session. L'en-tête de chaque bloc porte aussi « Copier une demande d'explication » (icône du robot) : le bouton copie un prompt à coller dans la conversation de l'agent, qui situe le bloc (analyse, dépôt et refs, fichier, lignes des deux côtés, code tel que figé) et lui demande de l'expliquer en enregistrant ses explications avec add_explanations, ou update_explanation si une explication couvre déjà ces lignes. Deux correctifs dont les lignes se recouvrent ne s'affichent pas ensemble. Le fond ambre des lignes visées se distingue du rouge des lignes − supprimées par la feature. Une légende suit le diff, limitée aux fonds que la vue et le mode affichent : « Ajouté par la feature » (fond ou filet selon la vue), « Supprimé par la feature » (Unifié et Côte à côte), « Lignes remplacées par le correctif » (Avant / après), « Correctif proposé » ou « Correctif appliqué », « Code décrit par une explication ». Un clic sur un numéro de ligne ouvre la saisie d'une remarque sur cette ligne. Les exigences manquantes sans emplacement s'affichent en bandeau au-dessus du diff de chaque fichier ; le bandeau « N exigences manquantes » se replie, et reste replié pour la session.

  • Affichage des correctifs — le sélecteur « Correctifs », au-dessus du diff, propose trois modes, conservés d'un fichier à l'autre pendant la session :

    • « Sans correctif » : le diff seul, sans aucun correctif ; les interrupteurs des cartes sont inactifs.

    • « Avant / après » (par défaut) : le rendu décrit ci-dessus, lignes visées barrées puis lignes proposées.

    • « Code corrigé » : le code tel qu'il se lirait avec les correctifs appliqués. Les lignes visées disparaissent ; les lignes proposées prennent leur place, sur fond violet avec la marque « › » et sans numéro, sous un bandeau « Correctif N appliqué — remplace les lignes X à Y ». Un lien ?line=N vers une ligne retirée met ce bandeau en évidence ; les remarques ne se posent que sur les lignes du diff, pas sur les lignes proposées.

    Dans les deux derniers modes, l'interrupteur de chaque carte (« Afficher le correctif dans le code » ou « Appliquer le correctif dans le code ») choisit les correctifs dessinés.

  • Colonne « Constats et explications » — l'en-tête compte les constats ouverts, et le total quand il diffère (« 4 constats ouverts · 5 au total »). Deux boutons, « Constats » et « Explications », affichent ou masquent chacun leur type, indépendamment l'un de l'autre et pour tous les fichiers de la session ; chacun indique le nombre d'éléments du fichier. Masquer les constats retire aussi leurs repères, pastilles et correctifs du diff et le bandeau des exigences manquantes : on parcourt alors le code avec ses seules explications, avant de rétablir les constats. Masquer les explications retire leurs cartes et leurs filets. Dessous, un index compact des constats (numéro à la couleur de la gravité et première ligne, ignorés et obsolètes atténués) : un clic fait défiler jusqu'aux lignes et à la carte, mises en évidence comme avec ?line. Tant que des cartes de constat sont sous la partie visible, une indication collée en bas de la colonne les compte (« 3 constats plus bas ↓ ») ; un clic amène la suivante. Les constats, numérotés dans l'ordre de leur première ligne, sont des cartes alignées sur leur ligne et reliées au diff ; des cartes voisines s'empilent sans se chevaucher et défilent avec le diff. Chaque carte montre numéro, gravité, lignes, titre, explication, l'interrupteur « Afficher le correctif dans le code » (constat avec correctif) et les boutons « Copier le prompt », « Ignorer » / « Rouvrir », « Remarque ». Les constats ignorés ou obsolètes restent visibles, atténués. Les explications sont des cartes sarcelle alignées de la même façon sur la première ligne qu'elles décrivent, une ligne supprimée pour du code supprimé ; à hauteur égale, l'explication précède les constats. Chaque carte montre les lignes (côté ancien et mention « Code supprimé » pour du code supprimé), le titre et le texte, replié au-delà de huit lignes avec « Lire la suite ». Une explication obsolète est atténuée et signale que le code décrit a changé. Les remarques de ligne sont des cartes alignées de la même façon, datées au format court (date complète en infobulle), avec « Modifier », qui rouvre le texte sur place dans la même saisie que la création, et « Supprimer » ; modifier une remarque retenue dans une revue déjà soumise recalcule le prompt enregistré ; les remarques sur le fichier entier et ce qui vise des lignes hors du diff affiché sont en tête de colonne. Sous 1200 px de large, la colonne passe sous le diff et ses cartes s'empilent.

Un fichier binaire, supprimé ou de plus de 1 Mo n'a pas de contenu conservé : on ne peut pas y ancrer de constat ni de remarque de ligne, seulement une remarque sur le fichier entier.

Écran Fin de revue (#/<id>/finish)

Décision (Approuver, Demander des corrections, Rejeter ; « Demander des corrections » est proposée d'emblée dès qu'il y a des points), points à transmettre (constats ouverts et remarques, tous cochés pour une revue en attente, la sélection enregistrée pour une revue déjà soumise), aperçu du prompt qui se met à jour, « Copier le prompt », « Retour à la revue » et « Enregistrer le retour », qui soumet la revue lue ensuite par get_review_feedback.

Toute modification faite par l'agent pendant que la GUI est ouverte y apparaît en direct, via SSE. Chaque action de la GUI porte la révision affichée : si l'analyse a changé entre-temps, l'action est refusée, l'analyse rechargée, et un message invite à refaire l'action.


Architecture

Deux processus, conformément au standard de l'écosystème (@imenam/mcp-gui-interface) :

  • Maître MCP (src/index.ts) — JSON-RPC sur stdio, outils (src/tools/, une fonction pure par outil), logique métier (src/core/), stockage, cycle de vie du worker via GuiLauncher.

  • Worker GUI (src/gui-worker.ts) — serveur Hono, enregistrement auprès du proxy via ProxyClient ou écoute sur APP_PORT, API REST et SSE (/api/events), fichiers de la SPA avec la balise <base> du chemin de montage.

Le worker ne lit ni n'écrit aucune donnée : chaque requête REST est transmise au maître par IPC (messages {type, correlationId, data, error, timestamp} validés par zod) et le maître pousse les changements, que le worker relaie en SSE. Le worker s'arrête quand le maître disparaît (canal IPC fermé, PID parent absent) ou quand une autre instance est déjà enregistrée, ce qui évite les GUI orphelines.

shared/ est importé à la fois par le serveur et par la GUI : schémas zod (shared/schemas/), prompt de retour (prompt.ts, utilisé par l'aperçu de la GUI et par get_review_feedback), placement des schémas (diagram-layout.ts, fonctions pures testées sous Node) et contrôle de leurs croisements (diagram-quality.ts), état de revue dérivé (review-state.ts), libellés et couleurs des gravités et statuts (labels.ts, severity.ts). Ce que l'utilisateur voit en aperçu est donc exactement ce que l'agent reçoit.

Persistance

Un fichier JSON par analyse dans <données>/analyses/<id>.json, et son diff figé dans <données>/diffs/<id>.json. Les écritures sont atomiques (fichier temporaire puis rename), sérialisées entre processus par un verrou .lock, et le cache mémoire est invalidé dès que la date ou la taille du fichier change. Tout document est validé par zod à la lecture et à l'écriture ; src/core/migrate.ts met à niveau les anciens formats à la lecture : une analyse sans projet reçoit le nom du dossier du dépôt (basename(repoPath)), une analyse sans vue d'ensemble reçoit overview: null, une analyse sans explications reçoit une liste vide.

Invariants

Toute modification, qu'elle vienne d'un outil ou de la GUI, passe par AnalysisStore.mutate(id, editor, fn, { baseRevision }), qui vérifie le schéma et les invariants avant d'écrire et incrémente la révision :

  • refs cohérentes avec le mode (branch : ref et commit de tête ; working_tree : base HEAD, sans tête) ;

  • identifiants uniques (fichiers, constats, explications, remarques, schémas, nœuds d'un schéma) ;

  • un constat ouvert vise un fichier de l'analyse dont le contenu est conservé, avec 1 ≤ startLine ≤ endLine ≤ nombre de lignes ; un issue a un emplacement ;

  • une explication à jour vise un fichier de l'analyse ; côté nouveau, son contenu est conservé et ses lignes existent ;

  • les liens d'un schéma relient des nœuds existants ; en layers, chaque nœud est dans une couche déclarée ;

  • la revue ne sélectionne que des constats et des remarques existants ;

  • les champs issus du diff (dépôt, refs, commits, liste des fichiers hors « revu ») ne changent que par refresh_analysis.

Les deux modes de diff

  • branch — git diff --find-renames <base>...<head> : ce que la branche apporte depuis son point de départ. head vaut HEAD par défaut. Le contenu des fichiers est lu dans le commit de tête : en local, en deux commandes git quel que soit le nombre de fichiers (ls-tree, puis cat-file --batch) ; en rôle remote, par un cat-file blob <commit>:<chemin> par fichier.

  • working_tree — git diff --find-renames HEAD : les modifications non commitées, index compris, plus les fichiers non suivis (hors .gitignore), présentés comme ajoutés. Le contenu est lu sur disque.

Le diff est figé à la création. refresh_analysis le recalcule : un fichier revu dont le diff n'a pas changé reste revu, les autres repassent à revoir ; un constat ouvert dont les lignes ne correspondent plus au texte copié passe à « obsolète », un constat obsolète dont les lignes correspondent de nouveau redevient ouvert ; la revue repart vierge.

Rien n'est écrit dans le dépôt

Le serveur lance git sans shell, uniquement en lecture (rev-parse, diff, ls-tree, cat-file, ls-files), et lit les fichiers de la copie de travail. En rôle remote, ce sont les mêmes commandes git, plus un node qui écrit un fichier sur sa sortie standard, exécutées par le relais agent dans repo_path, avec GIT_LITERAL_PATHSPECS=1 pour qu'un nom de fichier ne soit jamais lu comme un motif. Les correctifs proposés ne sont qu'affichés et transmis dans le prompt. repo_path doit être la racine absolue d'un dépôt git ; tout autre chemin est refusé avec un message explicite.

Exposition réseau de la GUI

La GUI montre le code des dépôts analysés et ses retours sont lus par l'agent : elle n'est joignable que depuis la machine elle-même, avec la garde de requêtes de @imenam/mcp-gui-interface (honoGuiGuard).

  • Le worker écoute sur 127.0.0.1 uniquement (GUI_LISTEN_HOST). L'accès depuis un autre poste passe par le proxy ou le gateway manager et leur authentification ; tous deux relaient vers localhost:<port>.

  • L'en-tête Host doit désigner la boucle locale : une page tierce qui fait pointer son propre domaine vers 127.0.0.1 est refusée (403).

  • Une requête de modification venant d'un autre site est refusée (403) : Sec-Fetch-Site doit valoir same-origin ou none et, en son absence, l'hôte de l'Origin doit être la machine ou le proxy. Une page web ouverte dans le navigateur ne peut pas soumettre de revue à la place du relecteur.


Développement

npm install
npm run build   # version, backend (tsc -> dist/), puis GUI (gui/dist)
npm test        # node --test sur dist/test
cd gui && npm run lint   # oxlint

Dossier

Contenu

src/index.ts

Maître MCP : déclaration des outils, instructions, IPC, lancement de la GUI.

src/gui-worker.ts

Worker GUI : Hono, REST, SSE, fichiers statiques, proxy.

src/tools/

Un fichier par outil, (store, args) => résultat.

src/core/

Stockage, git (local : git.ts ; machine de l'agent : agent-repo.ts, gateway-exec.ts), rôle, parseur de diff, emplacements, invariants, recalcul, actions du relecteur.

src/cli/setup-mcp.ts

Commande --claude-setup-mcp.

shared/

Code commun au serveur et à la GUI.

gui/

Application Vite + React 19 + zustand, CSS écrit à la main.

test/

Tests node --test.

La suite comprend des tests unitaires (parseur de diff, git sur de vrais dépôts temporaires, en local et à travers un relais agent simulé, store, invariants, emplacements, recalcul, prompt, placement des schémas), des tests de chaque outil, et deux suites qui lancent un vrai serveur sur stdio :

  • test/gui-api.test.ts — la GUI en mode standalone sur un port libre : API REST et boucle agent → GUI → agent. Elle ne dépend de rien d'extérieur.

  • test/e2e.test.ts — l'enregistrement auprès du proxy (E2E_PROXY_URL, défaut http://localhost:3000), la balise <base>, le SSE et reconnect_gui. Sans proxy joignable, elle est ignorée plutôt qu'en échec.

Pour inspecter le worker sans traverser le proxy, ajoutez ?__direct=1 à l'URL de son port local : les assets ne sont alors pas préfixés par le chemin de montage.

Available Tools

17 tools
add_explanationsA

Explain blocks of added or removed code to the reviewer, as a batch: what each block does and how it works. Write one for every long (about 30 lines or more) or complex block; skip trivial code. An explanation judges nothing: problems go in findings. Each location is checked against the frozen diff and its lines are copied; if any explanation is invalid, nothing is added and every problem is reported.

ParametersJSON Schema
NameRequiredDescriptionDefault
analysis_idYesId of the analysis, as returned by create_analysis or list_analyses.
explanationsYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses batch/atomic semantics ('if any explanation is invalid, nothing is added and every problem is reported'), validation against a frozen diff, and line-copying behavior. It is missing only surrounding operational context such as permission or auth requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four compact sentences, front-loaded with the core action and scope, then guidance, then failure behavior. It is dense but every sentence carries usable information; the single long paragraph is only mildly heavy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description covers purpose, inclusion criteria, output effect (lines copied onto the diff), and atomic failure behavior, which is enough for correct invocation. Return-value details are unnecessary given no output schema, leaving only minor gaps like auth context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Top-level coverage is 50%, but the nested item schema documents body, path, side, title, and line fields in detail. The description still adds semantics the schema does not spell out — that locations are checked against the frozen diff and their lines copied — which explains why exact path/line/side values matter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource ('Explain blocks of added or removed code to the reviewer, as a batch') and immediately clarifies scope by contrasting with the sibling concept of findings ('problems go in findings'). An agent can tell this apart from add_findings without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states when to create one ('every long (about 30 lines or more) or complex block') and when not to ('skip trivial code'), plus routing away from findings. It stops short of naming the update/delete siblings (update_explanation, delete_explanations) as alternatives for existing explanations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_findingsA

Record findings for the reviewer, as a batch. Each location is checked against the frozen diff and its lines are copied; if any finding is invalid, nothing is added and every problem is reported. Include requirement_gap findings for requested behaviour that is missing.

ParametersJSON Schema
NameRequiredDescriptionDefault
findingsYes
analysis_idYesId of the analysis, as returned by create_analysis or list_analyses.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses the atomic transaction semantics ("if any finding is invalid, nothing is added and every problem is reported") and the validation against the frozen diff with line copying. It omits auth/permission needs and return shape, but the failure behavior is unusually well spelled out.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with the core action and its batching nature front-loaded, followed by the failure contract and one kind-specific rule. Dense but every sentence earns its place; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a batch-mutation tool with no annotations and no output schema, the description covers the essential behaviors an agent needs: batching, atomicity on invalid input, diff validation, and the requirement_gap case. Only the success return and any permission requirements are left unstated, which is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Top-level schema coverage is 50% (analysis_id documented, the findings array is not), but the nested item schema is richly documented for severity, kind, location, and suggestion. The description's only parameter-level addition is the requirement_gap hint, which largely repeats the kind enum's own description, so it adds little beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ("Record findings") and scope ("as a batch"), which cleanly separates it from single-item siblings like update_finding and from add_explanations. It stops short of naming which sibling to use instead, but the purpose is unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives one concrete usage directive ("Include requirement_gap findings for requested behaviour that is missing"), which implies the when for that kind. However, there is no guidance on when to batch-add here versus updating an existing finding (update_finding) or when this is preferred over other entry points.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_analysisA

Create an analysis once a feature is developed: computes and freezes the git diff of the repository for human review. In branch mode the diff is base...head (head defaults to HEAD); in working_tree mode it is the uncommitted changes, index included, against HEAD. Returns the id and the changed files; then read them with get_diff.

ParametersJSON Schema
NameRequiredDescriptionDefault
baseNoBranch mode only, required: the ref the feature branched from, e.g. "master". Not accepted in working_tree mode.
headNoBranch mode only: the ref holding the feature (default HEAD). Not accepted in working_tree mode.
modeYesbranch: diff between two refs; working_tree: uncommitted changes against HEAD.
titleYesShort title of the feature, e.g. "OAuth token refresh".
projectYesProject the analysis belongs to: the product or feature name, used as its folder in the GUI. Reuse the exact name of an existing project (see list_analyses) for another analysis of it.
repo_pathYesAbsolute path of the root of the git repository.
request_textNoInitial request the feature answers (ticket, user prompt), shown to the reviewer.
request_sourceNoWhere the initial request comes from, e.g. a ticket id or URL. Needs request_text.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses that the diff is 'frozen' (a persisted snapshot, not a live read), that working_tree includes the index, and what the return contains. It omits any permission/auth requirements or idempotency notes, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the purpose, then the mode semantics, then the return/follow-up. No filler; each clause conveys a distinct fact.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers mode behavior and return values even though no output schema exists, and notes the get_diff follow-up. Only minor gaps remain around permissions and long-term lifecycle, which siblings like get_analysis/delete_analysis imply.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds value by explaining the parameter interactions: base...head semantics with head defaulting to HEAD, and mode-dependent acceptance. It reinforces the schema rather than merely repeating it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create an analysis') plus the concrete effect of the operation ('computes and freezes the git diff of the repository for human review'). It also distinguishes itself from the sibling get_diff by routing the agent there for reading the results.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear trigger ('once a feature is developed') and points to the follow-up tool ('then read them with get_diff'). It stops short of stating when NOT to use it or naming alternatives beyond get_diff.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_analysisA

Delete an analysis and its frozen diff. The repository is not touched. Also removes an analysis listed as unreadable.

ParametersJSON Schema
NameRequiredDescriptionDefault
analysis_idYesId of the analysis, as returned by create_analysis or list_analyses.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and does disclose what is destroyed (the analysis and its frozen diff) and what is deliberately spared (the repository). It covers the edge case of removing an analysis flagged as unreadable. It stops short of stating irreversibility, confirmation requirements, or any permission/auth constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with zero filler, and the core action plus its blast radius are front-loaded before the edge-case sentence. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter delete tool with no output schema and no annotations, the description supplies the key information: destructive scope and the untouched boundary. It lacks only irreversibility/permission context, which is a minor gap given the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single analysis_id parameter is already documented in the schema as coming from create_analysis or list_analyses. The description adds nothing about the parameter, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('delete') and resource ('analysis'), plus the coupled artifact ('its frozen diff'), which cleanly separates it from siblings like delete_findings or delete_explanations. The scope note that the repository is not touched further pins down what is and isn't affected.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the destructive verb rather than stated: there is no explicit 'use this when you want to permanently remove an analysis' or routing to alternatives such as update_analysis or refresh_analysis. The repository-untouched note gives partial scope guidance but is not a when-to-use rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_diagramC

Delete a diagram from an analysis.

ParametersJSON Schema
NameRequiredDescriptionDefault
diagram_idYesId of the diagram, as shown by get_analysis.
analysis_idYesId of the analysis, as returned by create_analysis or list_analyses.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden for a destructive operation. It does not say whether deletion is permanent, whether it cascades to findings or explanations attached to the diagram, or what permissions are required. Only the bare mutation intent is conveyed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler or redundancy. It is efficient, though it is arguably too terse for a destructive tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a delete operation with no annotations and no output schema, the description should at least address irreversibility or side effects. It leaves an agent unable to judge the consequences of invoking it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already explains that diagram_id comes from get_analysis and analysis_id from create_analysis/list_analyses. The description adds no parameter meaning beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (delete a diagram) and scopes it to an analysis, so the action is unambiguous. It does not explicitly distinguish itself from similarly named siblings like delete_analysis or delete_findings, but the resource noun carries the differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool, no prerequisites, and no mention of alternatives such as delete_analysis or set_diagram. The agent must infer usage entirely from the name and schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_explanationsA

Delete explanations by id, e.g. when the code they describe is gone. If any id is unknown, nothing is deleted.

ParametersJSON Schema
NameRequiredDescriptionDefault
analysis_idYesId of the analysis, as returned by create_analysis or list_analyses.
explanation_idsYesIds of the explanations to delete.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it usefully discloses all-or-nothing semantics ('If any id is unknown, nothing is deleted') — non-obvious transactional behavior. However, it says nothing about irreversibility, permission requirements, or whether deletions cascade to related records, which matters for a destructive operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero filler, with the operation front-loaded and the edge-case behavior immediately after. Nothing could be cut without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter destructive tool with no output schema, the description is close to sufficient: it covers scope and failure semantics. It falls short only on irreversibility and permission/confirmation expectations, which arguably an agent should still know before invoking a delete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters are documented in the schema, so the baseline is 3. The description adds no syntax, format, or batching nuance beyond 'by id' for either analysis_id or explanation_ids.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Delete explanations by id') that distinguishes it from the sibling deletes (delete_findings, delete_analysis, delete_diagram). The agent can identify exactly what gets removed without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete triggering context ('when the code they describe is gone'), which is more than most delete tools offer. It does not name an alternative tool or state when-not to use it, but the resource-scoped siblings make the choice fairly unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_findingsA

Delete findings by id, typically once refresh_analysis shows them solved. If any id is unknown, nothing is deleted.

ParametersJSON Schema
NameRequiredDescriptionDefault
analysis_idYesId of the analysis, as returned by create_analysis or list_analyses.
finding_idsYesIds of the findings to delete.

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose a genuinely non-obvious transactional trait: 'If any id is unknown, nothing is deleted' (all-or-nothing atomicity). However, it never states that deletion is permanent/irreversible or what permissions are required, so the destructive profile is only partially covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero waste, with the primary action front-loaded and the atomicity caveat as a tight trailing clause. Nothing redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity, no-output-schema two-param tool this is close to sufficient, and the atomicity note is valuable. But with no annotations the description should also say the deletion is irreversible and whether it is a hard or soft delete — an agent cannot infer that from the structured fields.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both required params (analysis_id, finding_ids) are documented in the schema, so baseline is 3. The description adds no syntax, ID-format, or batch-limit detail beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Delete findings by id') and the sibling-set context (findings vs. explanations/analysis) makes the target clear. It is distinguishable from delete_analysis and delete_explanations without opening a schema, though it does not explicitly name the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides real workflow context: 'typically once refresh_analysis shows them solved,' which names a sibling tool and the condition that motivates the call. No explicit when-not guidance or alternatives are given, so it stops short of 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_analysisA

Read a whole analysis: project, request, overview, summary, files (reviewed or not), findings with their status and location, reviewer notes, diagrams (flagged when their layout has crossings) and review state. Use it to check what is already recorded before adding more.

ParametersJSON Schema
NameRequiredDescriptionDefault
analysis_idYesId of the analysis, as returned by create_analysis or list_analyses.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does much of it: it discloses the rich return shape (files reviewed or not, findings with status/location, diagrams flagged on layout crossings) and the read-only nature implied by 'Read'. It stops short of explicitly stating there are no side effects or any permission requirements, so it is strong rather than complete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with what the tool reads, and no filler. The long enumeration is dense but each item is load-bearing given there is no output schema to describe the return value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description must define the return contract itself, and it enumerates the payload comprehensively (project through review state, including the diagram-crossing flag). Nothing an agent needs in order to call and interpret this tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is exactly one parameter at 100% schema coverage, and the schema already documents it ('as returned by create_analysis or list_analyses'). The description adds no format, constraint, or edge-case detail beyond the schema, which is the expected baseline when the schema does the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb and resource with explicit scope; clearly separable from its read-only siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use it to check what is already recorded before adding more' gives a clear read-before-write context, which is genuinely actionable. It stops short of naming alternatives or exclusions, but the whole-analysis scope implicitly separates it from the narrower reads in the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_diffA

Read the frozen diff of one file, or of every file when path is omitted, with the old and the new line number of each line. Use the NEW-side numbers (second column) for every finding location and diagram node line.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoChanged file to show (default: all files).
analysis_idYesId of the analysis, as returned by create_analysis or list_analyses.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose meaningful traits the schema cannot: the diff is 'frozen' (an immutable snapshot rather than live state) and it carries old and new line numbers per line. It omits error behavior for an invalid analysis_id and any size/truncation limits, which keeps it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences, front-loaded with the resource and scope, with zero filler. Every clause adds information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema and no annotations, so the description must describe the return shape, which it does (old and new line numbers per line). It leaves gaps around failure modes and behavior on very large diffs, but is otherwise complete for a read-only inspection tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline would be 3; the description still earns extra credit by explaining the two-column line-number output and which column to use for downstream locations. The path default it mentions is largely redundant with the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Read') and resource ('the frozen diff'), plus the exact scope: one file, or all files when path is omitted. An agent immediately knows what this returns and how scope is controlled, without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear conditional guidance ('or of every file when path is omitted') and, unusually, cross-tool guidance that the NEW-side column must be used for finding locations and diagram node lines. It does not name alternatives or exclusions (e.g. when to prefer get_analysis), so it stops short of 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_review_feedbackA

Read the reviewer's feedback once the review is submitted: decision, selected findings and notes, and the prompt to act on. Before submission it reports the review progress. Read-only. Always read the feedback before calling refresh_analysis, which discards it.

ParametersJSON Schema
NameRequiredDescriptionDefault
analysis_idYesId of the analysis, as returned by create_analysis or list_analyses.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and mostly delivers: it declares 'Read-only' (substituting for readOnlyHint), describes two behavioral modes (pre- vs post-submission), and flags that refresh_analysis discards the data. It stops short of covering error cases, permissions, or response shape when feedback does not yet exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the primary purpose, then the pre-submission variant, then the ordering hazard. Every sentence carries distinct information with no repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema and no annotations, so the description is the only source of return-value and safety information — and it covers both the post-submission payload and the pre-submission progress state adequately. A brief note on what happens with an invalid or unknown analysis_id would make it fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the single analysis_id parameter is documented in the schema as coming from create_analysis or list_analyses. The description adds no extra meaning about the identifier, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb (read) and resource (the reviewer's feedback) and enumerates the payload: decision, selected findings and notes, and the prompt to act on. It also distinguishes itself from get_analysis by describing state-dependent content, and explicitly names refresh_analysis as the contrasting sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when: after submission you get the final feedback, before submission you get progress. It also gives a hard ordering rule — 'Always read the feedback before calling refresh_analysis, which discards it' — naming the alternative and the consequence of sequencing incorrectly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_analysesA

List the analyses grouped by project, with their refs, open findings by severity, reviewed files, diagram count and review state (not started, in progress, all files reviewed, submitted). Start here: when the project already has an analysis of the feature, continue it (update_analysis to rewrite the overview and summary, refresh_analysis after changing the code) instead of creating a duplicate; create a new one with create_analysis for another feature or another diff.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectNoOnly list the analyses of this project (exact name). An unknown name lists the existing projects.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden, and it does disclose the full return shape including the enumeration of review states, which is genuine behavioral context. It stops short of stating read-only safety explicitly or any pagination/permission behavior, but 'List' strongly implies a read.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, well front-loaded with the return content before the routing advice. The second sentence is dense but every clause earns its place by naming a sibling and its condition; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by enumerating the returned fields, and with no annotations it supplies routing context for a read-only list tool. It is essentially complete for invocation, lacking only minor operational details like pagination or error behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already explains the project filter, including the useful 'unknown name lists existing projects' fallback. The description adds no parameter meaning beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (analyses), then enumerates exactly what the listing contains: refs, open findings by severity, reviewed files, diagram count, and review state with its four values. This is clearly separable from get_analysis and the other sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly frames itself as the entry point ('Start here') and names the alternatives with their triggering conditions: update_analysis to rewrite the overview, refresh_analysis after code changes, create_analysis for a different feature or diff. Explicit when-to-use and when-not-to-duplicate guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reconnect_guiA

Relaunch the GUI worker and re-register it with the proxy server. Use this when the proxy server was not running when the MCP server started (so the GUI registration failed), to redeploy the GUI without restarting the MCP server. Does nothing if the GUI is already registered, unless force=true.

ParametersJSON Schema
NameRequiredDescriptionDefault
forceNoRestart the GUI worker even if it is already running and registered (default: false)

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the no-op behavior when already registered, the force override, and that it avoids restarting the MCP server. It does not mention permissions or side effects on an existing GUI worker, which keeps it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the action, then the trigger condition, then the edge case. No filler and every sentence carries distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-optional-parameter operation with no output schema and no annotations, the description covers purpose, trigger, and the already-registered edge case. Nothing required to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the force parameter is already documented. The description still adds practical meaning by framing force as the override for the otherwise-no-op path ('Does nothing... unless force=true'), complementing the schema's phrasing rather than merely repeating it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb pair (relaunch, re-register) with clear resources (GUI worker, proxy server). There are no functionally similar siblings in the list, and the tool's unique identity as the GUI/proxy reconnection mechanism is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the triggering condition: 'use this when the proxy server was not running when the MCP server started (so the GUI registration failed)'. It also states the negative case ('does nothing if the GUI is already registered, unless force=true'), which is exactly the when/when-not guidance an agent needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

refresh_analysisA

Recompute the frozen diff with the same repository, mode and refs, after fixing the code. Files whose diff changed become unreviewed, findings whose lines changed become outdated, and the review restarts as pending: the previous feedback is discarded, so read it with get_review_feedback first.

ParametersJSON Schema
NameRequiredDescriptionDefault
analysis_idYesId of the analysis, as returned by create_analysis or list_analyses.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses that changed files become unreviewed, changed findings become outdated, the review reverts to pending, and prior feedback is discarded. It omits idempotency, permission requirements, and cost, which keeps it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the action, then the state consequences, then the prerequisite tool. Every clause earns its place, though the semicolon-chained structure is dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations and no output schema, the description covers the critical side effects an agent must know before calling. It does not describe what the call returns, but that gap is minor given the destructive-behavior disclosure is thorough.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter and schema coverage is 100%, so the schema already documents analysis_id fully (including its origins in create_analysis/list_analyses). The description adds no syntax or format meaning beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Recompute the frozen diff') and immediately pins the scope ('same repository, mode and refs'), which distinguishes it from create_analysis and update_analysis. It also names the sibling get_review_feedback, so an agent can route without inspecting schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear triggering condition ('after fixing the code') and a prerequisite action ('read it with get_review_feedback first'), which is strong context. It stops short of stating when-not to use it or naming a full alternative, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_diagramA

Create a diagram, or fully replace the one with diagram_id, when a picture helps the reviewer (impact perimeter, process flow, concept map). Give nodes and links only, with no coordinates or colours: the GUI lays the diagram out and colours nodes by status. flow = a process with steps and decisions, laid out left to right along the links, where a side branch without a join and the last node of a chain hang below the node before them; layers = the impact perimeter, one column per entry of layers (e.g. GUI, API, core, storage), each node in its layer; mindmap = a tree of concepts drawn radially around its root: exactly one root, one incoming link for every other node, no cycle (refused otherwise). Rules: one subject per diagram, 5 to 12 nodes; declare nodes in reading order, since the declaration order is the initial order of every rank or column; zero crossings expected. The result reports every link crossing and every link through a node, with advice: reorder the nodes, split the diagram or change its kind, then send it again.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindYesflow = a process with steps and decisions, laid out left to right along the links, where a side branch without a join and the last node of a chain hang below the node before them; layers = the impact perimeter, one column per entry of `layers` (e.g. GUI, API, core, storage), each node in its layer; mindmap = a tree of concepts drawn radially around its root: exactly one root, one incoming link for every other node, no cycle (refused otherwise).
linksNoDirected links between nodes (default: none). In a mindmap, from parent to child.
nodesYesNodes in reading order (5 to 12): the declaration order is the initial order of every rank or column.
titleYesTitle of the diagram.
layersNoRequired when kind is layers, forbidden otherwise: column names from left to right.
diagram_idNoId of the diagram to replace, or of the new one; generated when omitted.
analysis_idYesId of the analysis, as returned by create_analysis or list_analyses.

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so: it discloses the destructive 'fully replace' semantics of passing diagram_id, the division of labour (agent supplies nodes/links only, GUI handles layout and status colours), composition rules (one subject, 5-12 nodes, zero crossings), refusal conditions (mindmap with a cycle or multiple roots), and what the response reports back with remediation advice.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is front-loaded with purpose and the three kind behaviors, which is good, but it is very dense and a large share of it duplicates the `kind` and `nodes` schema descriptions word-for-word. Those sentences do not earn their place given they are already structured-field content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 7-parameter tool with no annotations and no output schema, the description covers kinds, composition constraints, replacement semantics, and the shape of the returned crossing/advice feedback, which compensates for the missing output schema. Minor gaps remain around permissions and the mismatch between the schema's minItems and the stated 5-12 node rule.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description largely restates schema content -- the `kind` enum text and the nodes reading-order/5-to-12 wording appear verbatim in the schema -- and adds only marginal extras such as 'no coordinates or colours' and 'declare nodes in reading order'. It does not compensate for the schema's misleading minItems of 1 versus the stated 5-12 range.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first clause names a specific verb and resource ('Create a diagram, or fully replace the one with diagram_id') and adds the reviewer-facing purpose. An agent can immediately tell this is the diagram-creation/replacement tool, distinct from delete_diagram in the sibling list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear when-to-use trigger ('when a picture helps the reviewer') with concrete examples (impact perimeter, process flow, concept map) and explains which `kind` fits which situation. It stops short of explicit when-not guidance or naming alternatives such as emitting findings or explanations instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_analysisA

Change the title, the project, the overview, the summary or the initial request of an analysis. Right after reading the diff, write the overview (objective, approach, attention_points: your global analysis of the feature) and the summary (the functional changes of the feature, one bullet per change, not one bullet per file). Give at least one field; fields left out are kept.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleNoNew title.
projectNoNew project. Project the analysis belongs to: the product or feature name, used as its folder in the GUI. Reuse the exact name of an existing project (see list_analyses) for another analysis of it.
summaryNoFunctional changes of the feature, one bullet per change (e.g. "Expired tokens are refreshed once before the request fails"), not one bullet per file. Replaces the whole list; [] clears it.
approachNoOverview: how the feature achieves it (architecture choices, main flow). Replaces the current approach.
objectiveNoOverview: what the feature lets users do, from a functional point of view. When the analysis has no overview yet, give objective and approach together to create it. null removes the whole overview (approach and attention points included).
analysis_idYesId of the analysis, as returned by create_analysis or list_analyses.
request_textNoInitial request the feature answers (ticket, user prompt), shown to the reviewer. null removes the request.
request_sourceNoWhere the initial request comes from, e.g. a ticket id or URL. null clears it.
attention_pointsNoOverview: risks and points the reviewer should check first, one per string. Replaces the whole list; [] clears it. Defaults to an empty list when the overview is created without it.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses partial-update semantics ('fields left out are kept') and content expectations, but says nothing about permissions, whether the analysis must exist, error behavior, or the response. Adequate but incomplete for a 9-parameter mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the field list and then content guidance. The middle sentence is dense and slightly overloaded with parentheticals, but every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

All 9 parameters are fully documented in the schema and there is no output schema to explain, so the description only needs the workflow and patch semantics it already provides. Missing only permission/error context, which is minor here.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds conceptual meaning the schema lacks: that 'overview' is composed of objective+approach+attention_points, and that summary is one bullet per change, not per file. This mapping value exceeds the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (change/update) plus resource (analysis) and enumerates the updatable fields (title, project, overview, summary, initial request). Purpose is unmistakable, though it does not explicitly differentiate itself from siblings like update_finding or update_explanation beyond the resource noun.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives real workflow context ('Right after reading the diff, write the overview... and the summary') plus the patch constraint ('Give at least one field; fields left out are kept'). It does not name alternatives, but the timing guidance is concrete and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_explanationA

Change one explanation: title, body or location. A new location (path and start_line, with side for removed code) is checked against the diff and makes an outdated explanation current again; rewrite the body too when the code changed.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyNoWhat the code does (inputs, outputs, effects), then how it works, step by step in the order of the code. Describe, never judge. Plain text; line breaks are kept.
pathNoPath of a changed file of the analysis, relative to the repository root, as listed by get_diff.
sideNoSide of the diff the lines belong to (default "new"): "new" for added or kept code, with NEW-side numbers (second column of get_diff); "old" for removed code, with OLD-side numbers (first column of get_diff), all shown in the diff and at least one removed.
titleNoNew title.
end_lineNoLast described line, inclusive (default: start_line).
start_lineNoFirst described line, on the side given by side.
analysis_idYesId of the analysis, as returned by create_analysis or list_analyses.
explanation_idYesId of the explanation, as shown by get_analysis.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose a genuinely non-obvious behavior: location updates are validated against the diff and can make an outdated explanation current again. However it says nothing about partial-update semantics (whether omitted fields are left unchanged), failure modes for an invalid location, or required permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly packed sentences, front-loaded with the core action and then the diff-validation consequence. Dense but every clause carries information; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter mutation tool with no annotations and no output schema, the description covers the key behavioral twist (diff-checked location) but leaves partial-update semantics, error behavior, and return expectations unstated. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all eight parameters thoroughly. The description adds the useful framing that path/start_line/side form a single 'location' concept tied to the diff check, but does not extend the per-parameter meaning beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Change one explanation') and enumerates the mutable facets (title, body, location), which cleanly separates it from the sibling update_analysis and update_finding tools. It does not explicitly name those siblings, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives conditional usage guidance: a new location is diff-checked and can revive an outdated explanation, and the body should be rewritten when the code changed. This is actionable context, though it never states when-not to use the tool or names an alternative path.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_findingA

Change the fields of one finding: severity, kind, title, body, location, suggestion or prompt. A new location is checked against the diff and reopens an outdated finding; path null removes the location. Whether a finding is ignored is the reviewer's decision and cannot be changed here.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyNoNew explanation.
kindNoissue or requirement_gap.
pathNoPath of a changed file of the analysis, relative to the repository root, as listed by get_diff. Give it with start_line to move the finding; null removes the location (requirement_gap only).
titleNoNew title.
promptNoPrompt to hand back to the agent for this finding, replacing the generated one (body, current lines, replacement) in the review feedback. null clears it.
end_lineNoLast targeted line, inclusive (default: start_line).
severityNoSeverity: critical, major, minor, trivial (from most to least serious).
finding_idYesId of the finding, as shown by get_analysis.
start_lineNoFirst targeted line. Line numbers are NEW-side numbers, as shown in the second column of get_diff.
suggestionNoReplacement text proposed for the targeted lines, kept verbatim (indentation included); an empty string proposes deleting them. Displayed to the reviewer only: nothing is written to the repository. Needs a location. null clears it.
analysis_idYesId of the analysis, as returned by create_analysis or list_analyses.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the side effect that a new location reopens an outdated finding, the null-clears semantics, and that the ignored flag is immutable here. It does not state permission/auth requirements or failure behavior for invalid locations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the action and the field list, with side-effect caveats packed into the second sentence. Dense but no wasted filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 11-parameter mutation tool with no annotations and no output schema, the description covers the important side effects (reopening, null removal, ignored immutability) that an agent cannot infer from the schema. Remaining gaps are minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents every parameter including the null/removal semantics for path and prompt. The description's field enumeration largely restates what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Change the fields of one finding') plus the exact field set (severity, kind, title, body, location, suggestion, prompt), so an agent can distinguish it from update_explanation or add_findings. It does not explicitly name or contrast a sibling tool, which keeps it short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete usage conditions: a new location is validated against the diff, path null removes the location (requirement_gap only), and ignored status cannot be changed here — a clear exclusion that routes the agent elsewhere. It stops short of naming which alternative tool handles the excluded case.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 17 tool updatesv1.0.13
    • First observedadd_explanations
    • First observedadd_findings
    • First observedcreate_analysis
    • First observeddelete_analysis
    • First observeddelete_diagram
    • First observeddelete_explanations
    • First observeddelete_findings
    • First observedget_analysis
    • First observedget_diff
    • First observedget_review_feedback
    • First observedlist_analyses
    • First observedreconnect_gui
    • First observedrefresh_analysis
    • First observedset_diagram
    • First observedupdate_analysis
    • First observedupdate_explanation
    • First observedupdate_finding

TDQS

A3.8/5.0

Scored across 17 tools

Disambiguation5/5

Each tool targets a distinct resource and action (analysis lifecycle, findings, explanations, diagrams, diff, review feedback), with detailed descriptions clarifying boundaries. Potential confusion between add/update/delete variants is mitigated by singular/batch distinctions.

Naming Consistency4/5

All tools use snake_case with a verb_noun pattern (e.g., create_analysis, get_diff). Minor inconsistency: plural in add_findings/add_explanations/delete_* but singular in update_finding/update_explanation.

Tool Count4/5

17 tools cover a complex workflow with CRUD for analyses, findings, explanations, and diagrams. Slightly above the ideal 15 but each tool appears necessary and well-scoped.

Completeness4/5

Covers full analysis lifecycle, diff retrieval, findings/explanations/diagrams CRUD, review feedback, and GUI reconnect. Minor gap: no explicit read tool for explanations (though get_analysis may include them) and no standalone list for findings/explanations.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    A collaborative code and markdown review tool that bridges human reviewers and AI agents, enabling both to browse files, inspect git diffs, leave structured comments, and save a final review report from the same UI in real time.
    41 PyPI
    2
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables coding agents and human reviewers to share a single live view of code, letting the agent open and annotate specific lines while the human marks ranges to ask about, all without any write access to the underlying files.
    1
    MIT