Skip to main content
Glama
didou92i

lmstudio-local

by didou92i

LM Studio Local MCP — modèles, connecteurs et RAG, bannière vert et crème

LM Studio ↔ Codex — MCP 2.1

Serveur MCP indépendant pour administrer LM Studio et utiliser des modèles locaux depuis un client MCP. Il utilise l’API native /api/v1, l’API compatible OpenAI, la commande officielle lms et le SDK Python officiel. Il expose 28 outils : administration, inférence, réglages, documentation, connecteurs et RAG. Les versions Python sont verrouillées dans uv.lock.

Installation

Prérequis : Python 3.11 ou supérieur, uv, Git et LM Studio avec son serveur local activé. Plateforme LM Studio validée : macOS. Windows n’est pas pris en charge nativement ; le connecteur utilise des verrous POSIX et des chemins macOS par défaut.

Cloner ce dépôt, ouvrir un terminal dans son dossier, puis :

uv sync --frozen
cp .env.example .env
./run.sh --diagnose

Adapter .env uniquement si le port, l’authentification ou les chemins diffèrent. Le serveur LM Studio écoute par défaut sur http://127.0.0.1:1234.

Pour Codex CLI, depuis le dossier cloné :

codex mcp add lmstudio-local -- "$PWD/.venv/bin/python" -m lmstudio_mcp.server

Dans les paramètres MCP du client, prévoir 45 secondes pour le démarrage et jusqu’à 1 000 secondes par appel long. Pour les autres clients, adapter le chemin absolu de examples/mcp-client.json. L’installation crée un environnement local ; aucun modèle ni moteur n’est téléchargé automatiquement.

Related MCP server: LM Studio MCP Server

Utilisation dans Codex

Après installation, enregistrer le serveur sous lmstudio-local dans le client MCP. Le programme reste dans le dossier cloné ; après un déplacement, adapter le chemin enregistré. Recharger le client si les outils ne sont pas encore visibles.

Exemples de demandes :

  • « Vérifie LM Studio et liste mes modèles. »

  • « Estime la mémoire nécessaire pour Gemma avant de le charger. »

  • « Interroge Gemma avec cette consigne, puis libère son instance. »

  • « Indexe ces documents et réponds avec les citations sources. »

  • « Diagnostique mon connecteur MCP et vérifie son programme de démarrage. »

  • « Enregistre un profil de chargement et compare les réglages effectifs. »

  • « Vérifie les mises à jour des moteurs, sans les installer. »

Par défaut, l’inférence s’exécute sur la machine locale. Les résultats demandés par Codex sont ensuite transmis à Codex. Les autres logiciels nécessitent leurs propres connexions autorisées.

Vérification à chaque utilisation

À l’ouverture d’une connexion MCP et avant chaque outil :

  1. Lecture de la version de l’application installée et des moteurs présents.

  2. Vérification réelle de /api/v1/models et de sa structure.

  3. Synchronisation automatique du dépôt documentaire complet si son cache a expiré (une heure par défaut). Les sources hors ligne sont signalées comme périmées.

  4. Contrôle de la fraîcheur des informations officielles : changelog et pages API, chargement, conversation. Cache d’une heure ; lm_status(refresh_updates=true) force la lecture distante. LM_MCP_UPDATE_TTL=0 vérifie en ligne à chaque appel.

  5. Comparaison de l’empreinte application/moteurs et du code/dépendances du connecteur avec le dernier test réel. Un changement invalide la validation et demande un nouveau test.

  6. Contrôle des capacités utiles avant l’action, puis vérification de l’instance après chargement/déchargement. Les paramètres de chargement non appliqués produisent une erreur MCP avec l’état réellement obtenu.

Chaque réponse comprend ok, data ou error, et diagnostics. Une erreur réseau ne devient jamais une déclaration « à jour ». Une mutation en timeout a un résultat inconnu : consulter l’état avant de la répéter. Les POST ne sont pas réessayés automatiquement.

Une détection de version ne prouve pas que toutes ses nouveautés fonctionnent. Les nouveautés inconnues nécessitent une adaptation du code et les tests correspondants. Les changements de documentation restent signalés dans le cache jusqu’à revue. Aucun téléchargement de mise à jour logicielle n’est lancé à la connexion.

Outils

Outil

Usage

lm_status

Version, API, moteurs, dernières versions et validation

lm_models

Modèles, variantes, instances, capacités et configurations

lm_load

Chargement natif : contexte et options GGUF documentées

lm_load_advanced

GPU, parallélisme, TTL, décodage spéculatif via lms ; identifier requis

lm_estimate

Estimation mémoire sans chargement

lm_unload

Déchargement d’une instance précise avec vérification

lm_download

Téléchargement depuis le catalogue/Hugging Face

lm_download_status

Suivi du job de téléchargement

lm_chat

Texte, images, paramètres, raisonnement, conversation persistante et statistiques

lm_openai_chat

Historique explicite, sortie structurée JSON, définitions de fonctions

lm_embeddings

Vecteurs, en lot ou à l’unité

lm_server_control

État, démarrage sur loopback, arrêt et redémarrage du serveur

lm_runtime

Moteurs, matériel, disponibilité, simulation ou installation explicite d’une mise à jour stable

lm_integrations

MCP configurés, programme disponible, permissions

lm_diagnose

API, matériel, disque, instances, erreurs et signaux des journaux

lm_connections

Enregistrer, tester et sélectionner des serveurs LM Studio

lm_docs

Documentation complète, recherche, procédures, couverture, changements et sources

lm_model_config

Schéma SDK installé, inspection et chargement avancé

lm_profiles

Enregistrer et réutiliser des configurations de modèles

lm_mcp_config

Prévisualiser/appliquer une configuration avec sauvegarde et contrôle de conflit

lm_mcp_probe

Connexion réelle et découverte des outils d’un MCP

lm_mcp_call

Appel d’un outil MCP autorisé avec validation de ses arguments

lm_rag_index

Index local de TXT, MD, PDF texte et DOCX sélectionnés

lm_rag_search

Recherche avec extraits, chemins, pages et empreintes

lm_rag_ask

Réponse locale avec contrôle des identifiants et citations exactes

lm_rag_manage

Liste des collections ou suppression d’un index

lm_link

État et administration LM Link exposés par le CLI installé

lm_import_model

Simulation/import par copie d’un GGUF local

lm_chat prend un texte ou une liste {"type":"text","content":"..."} / {"type":"image","data_url":"data:image/png;base64,..."}. Avec options.store=true, transmettre le response_id reçu dans options.previous_response_id pour continuer. Répéter le system_prompt souhaité. Les valeurs de raisonnement sont contrôlées contre celles du modèle.

lm_openai_chat retourne les appels de fonctions ; il ne les exécute pas. lm_mcp_call appelle directement un outil autorisé ; lm_chat.integrations permet sa délégation à un modèle local, selon les permissions de LM Studio.

Réglages et administration

Consulter lm_model_config(action="schema") avant de configurer le SDK : contexte, GPU, quantification KV, mmap, RoPE, seed et autres options réellement exposées par la version installée. Les options supplémentaires GGUF sont refusées sur MLX. Utiliser lm_load_advanced pour les flags CLI disponibles : parallélisme, TTL, offload GPU, assistant de décodage.

lm_profiles stocke les paramètres dans ce dossier. Un chargement SDK exige un identifiant d’instance neuf ; l’inspection ne charge jamais implicitement un modèle. Les paramètres ignorés produisent une erreur avec l’instance réellement créée, permettant de la décharger.

lm_diagnose fournit des pistes de réparation fondées sur les erreurs, sans renvoyer les lignes brutes des journaux. Ces pistes ne sont pas une preuve de cause. Il vérifie aussi l’existence du programme des connecteurs configurés. lm_server_control pilote le serveur local ; lm_runtime gère les moteurs. L’installation de l’application et la connexion au compte LM Link restent dans les interfaces officielles.

Connecteurs MCP

  1. Lister les entrées via lm_mcp_config(action="list").

  2. Prévisualiser une modification (upsert, remove, authorize) avec apply=false.

  3. Appliquer en reprenant expected_digest de cette prévisualisation : le MCP conserve les autres entrées et crée une sauvegarde.

  4. Utiliser lm_mcp_probe pour établir la connexion et découvrir les outils réels.

  5. Autoriser leurs noms exacts avec allowed_tools, puis appeler lm_mcp_call.

La prévisualisation et l’application peuvent être enchaînées par Codex dans la demande autorisée. Le contrôle de conflit évite d’écraser une modification concurrente. Changer une configuration invalide ses anciennes permissions. Aucune commande shell générique n’est proposée ; un connecteur stdio exécute néanmoins le programme configuré lors de sa connexion. Une configuration réussie dans mcp.json ne prouve pas son activation dans l’interface LM Studio, qui peut nécessiter un rechargement.

Les transports stdio, HTTP Streamable et SSE sont pris en charge. Les commandes, arguments et secrets complets ne sont pas retournés par l’inventaire. Les sauvegardes conservent la configuration originale et doivent rester privées.

RAG documentaire

Déposer les documents dans documents/, puis utiliser lm_rag_index avec une liste explicite de fichiers et le modèle d’embeddings. Une collection conserve textes, vecteurs, fichiers, pages PDF et empreintes SHA-256 dans .state/rag.sqlite3.

lm_rag_search retourne les extraits ; lm_rag_ask interroge le modèle et vérifie les identifiants de sources et citations exactes. Une réponse sans citations vérifiables est remplacée par answer=null. Le format JSON est imposé ; les repères de sources manquants sont ajoutés après validation des citations structurées. Ces contrôles ne prouvent pas que toutes les conclusions du modèle sont exactes. Les documents sont des données, jamais des instructions exécutables.

Les fichiers modifiés, supprimés ou sortis des répertoires autorisés sont écartés. Changer de modèle d’embeddings, de moteur ou de serveur impose une réindexation complète avec replace=true. Les formats acceptés sont TXT/MD UTF-8, PDF texte et DOCX avec tableaux ; les PDF scannés nécessitent un OCR préalable. Limites : 25 Mo/fichier, 50 fichiers/appel, 500 pages/PDF, 2 500 fragments/appel, 10 000/collection.

Cet index appartient au connecteur ; il ne pilote pas les pièces jointes ni les index internes des conversations de l’application LM Studio. Les embeddings et questions utilisent le serveur sélectionné : choisir un profil distant lui transmet les extraits nécessaires. lm_rag_manage(action="delete") supprime l’index, jamais les documents sources.

Serveurs et documentation officielle

lm_connections conserve des profils local, réseau privé ou HTTPS, teste /api/v1/models, puis sélectionne la connexion. Les jetons restent dans des variables LM_REMOTE_* ou LM_STUDIO_* référencées par le profil. Un profil distant désactive les commandes CLI locales ; il n’administre pas le système d’exploitation distant. Les opérations de retour au profil local restent accessibles même si la connexion sélectionnée est défaillante.

La documentation complète du dépôt officiel est automatiquement récupérée et indexée : 284 fichiers au snapshot validé, avec provenance et empreintes. lm_docs fournit la recherche français/anglais, la lecture intégrale paginée, douze guides opérationnels, un dossier de références par demande (brief), la couverture outils/tests/limites et les changements amont. scope="all" permet aussi de consulter les brouillons et fichiers de support, clairement identifiés.

Codex reçoit les instructions de consultation à la connexion. Une ressource d'orientation, des ressources de pages et le prompt lmstudio_workflow complètent les outils. La synchronisation vérifie la fraîcheur à chaque connexion/utilisation et utilise un cache d'une heure ; sync force le contrôle. En cas de réseau indisponible, la copie précédente reste accessible avec un avertissement. Les médias externes sont référencés par leurs liens.

La disponibilité d'une documentation ne prouve pas qu'une fonction est pilotable. Chaque page dispose d'une correspondance explicite, invalidée si son contenu change. Voir la documentation opérationnelle et ses limites.

Transport et n8n

Codex utilise stdio. ./run.sh --http fournit également http://127.0.0.1:8765/mcp pour les clients locaux, avec les protections Host/Origin du SDK. Il ne publie pas le service sur Internet.

L’exemple examples/n8n-mcp-workflow.json est fourni à titre indicatif, désactivé par défaut et sans validation dans une instance n8n.

Configuration

Copier .env.example vers .env si nécessaire. Les variables d’environnement du processus ont priorité. Ne pas placer de vrai jeton dans Git.

  • LM_STUDIO_BASE_URL : origine HTTP locale, par défaut http://127.0.0.1:1234.

  • LM_STUDIO_API_TOKEN : jeton LM Studio si son authentification est activée.

  • LMS_PATH : chemin du binaire lms, par défaut ~/.lmstudio/bin/lms.

  • LM_MCP_TIMEOUT : délai API, 300 secondes par défaut.

  • LM_MCP_UPDATE_TTL : durée du cache changelog/pages API en secondes, 3600 par défaut.

  • LM_MCP_DOCS_AUTO_SYNC : synchronisation documentaire automatique, true par défaut.

  • LM_MCP_DOCS_TTL : fraîcheur du dépôt documentaire en secondes, 3600 par défaut ; 0 force les contrôles distants.

  • LM_MCP_STATE_DIR : cache et preuves, .state/ dans ce dossier par défaut.

  • LM_MCP_ALLOWED_INTEGRATIONS : ancienne autorisation globale par ID de plugin ; préférer les permissions exactes via lm_mcp_config.

  • LM_MCP_CONFIG_PATH : configuration MCP de LM Studio, par défaut ~/.lmstudio/mcp.json.

  • LM_MCP_RAG_ROOTS : répertoires autorisés séparés par : sur macOS, par défaut documents/.

Les jetons ne sont pas envoyés aux sites de vérification des versions. Le dossier .state contient notamment les textes/vecteurs du RAG, les sauvegardes de configuration, les profils et les preuves de test. Il reste privé et hors Git. store=true permet cependant à LM Studio de conserver les échanges côté serveur.

Validation et maintenance

./run.sh --diagnose
.venv/bin/pytest -q
.venv/bin/ruff check src tests scripts
.venv/bin/python scripts/docs_smoke.py
.venv/bin/python scripts/smoke_test.py
.venv/bin/python scripts/operations_smoke.py
.venv/bin/python scripts/server_smoke.py
# Ajout facultatif du test de redémarrage LM Studio :
.venv/bin/python scripts/server_smoke.py --restart
# Serveur HTTP lancé dans un autre terminal :
.venv/bin/python scripts/smoke_test.py --http http://127.0.0.1:8765/mcp

Le test réel utilise les deux modèles déjà téléchargés, effectue des inférences synthétiques puis libère uniquement les instances qu’il a créées. Il peut consommer plusieurs minutes et environ 7 Go pour Gemma. Il faut le relancer après changement d’application ou de moteurs ; ses limites détectées restent dans le rapport .state/smoke-stdio.json.

Pour une nouvelle installation : uv sync --frozen. Pour mettre à jour les dépendances du connecteur : uv lock --upgrade, uv sync --frozen, puis tous les tests. La mise à jour de l’application LM Studio se fait via son interface officielle ; lm_runtime(action="update") concerne seulement les moteurs stables et nécessite un nouveau test d’inférence.

Voir les résultats et limites pour les résultats et limites constatés. operations_smoke.py utilise une configuration MCP isolée et un document fictif ; il supprime les instances, profils et index créés pour le test.

Sources

API et aide CLI locales inspectées le 13 septembre 2026. Les API stables exposées ici ne donnent pas accès à tous les réglages internes de l’interface, ni à l’ensemble des fonctionnalités de Bionic. Les possibilités futures sont détectées, mais leur implémentation reste un travail explicite.

Available Tools

28 tools
lm_chatC
Destructive

Native chat/vision/stats. store=true returns response_id for continuation. Optional explicitly enabled LM Studio MCP plugins.

ParametersJSON Schema
NameRequiredDescriptionDefault
inputYes
modelYes
optionsNo
integrationsNo

TDQS

C2.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this as destructive, open-world, and not read-only. The description adds useful context about the store=true continuation behavior and optional MCP plugins, but it does not explain what gets destroyed, how plugins are authorized, or what side effects to expect from a destructive chat tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loaded, which is good, but it is fragmented and under-specified. The first sentence is vague, and the remaining two sentences cover only narrow aspects of the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with nested chat options, required model and input parameters, optional integrations, and no output schema, the description is substantially incomplete. It does not explain input structure, available options, expected return values, or integration requirements.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain the four parameters. It only touches on store=true and optional MCP plugins while saying nothing about model, input format (string vs array), or the many options fields like temperature, context_length, and previous_response_id.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase 'Native chat/vision/stats' gives a general idea that this is a local model chat tool with vision and stats support, but it lacks a clear verb and does not distinguish itself from siblings like lm_openai_chat or lm_rag_ask. An agent must guess what 'native' means and whether this is the right chat tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use lm_chat versus lm_openai_chat, lm_rag_ask, or other siblings. It mentions store=true and MCP plugins, but these are behavioral notes, not usage conditions or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lm_connectionsB

Manage/test/select LM Studio connection profiles (local, private LAN, HTTPS). Secrets via LM_REMOTE_* env only. Remote profile disables local CLI.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNo
nameNolocal
actionYes
token_envNo

TDQS

B3.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile, and the description adds genuinely useful context beyond them: secrets are supplied only via LM_REMOTE_* env vars (including what token_env likely refers to) and selecting a remote profile disables the local CLI. It omits what 'delete' or 'save' actually do to existing configuration.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short clauses, front-loaded with the core verb and resource, no filler. It is arguably over-compressed given the gaps in parameter and action semantics, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A multi-action mutation tool with 0% parameter coverage and no output schema should explain what each action does and which parameters each action requires. The description leaves the agent guessing whether url is needed for test/save or how select differs from save.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for 4 parameters. The description only indirectly hints at token_env via the LM_REMOTE_* mention; url, name, and especially the action enum values (list/save/test/select/delete) get no explanation of their effects or required combinations.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb set (manage/test/select) and resource (LM Studio connection profiles) with a clarifying taxonomy of local, private LAN, and HTTPS targets. It does not differentiate itself from the sibling lm_profiles, which an agent could plausibly confuse with connection-profile management.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance or routing against alternatives like lm_profiles, lm_status, or lm_server_control. The only context given is a side-effect note ('Remote profile disables local CLI'), which is behavior, not usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lm_diagnoseB
Read-only

Diagnose API, server, disk, loaded models, error history and redacted log signals. Returns evidence and repair steps.

ParametersJSON Schema
NameRequiredDescriptionDefault
include_hardwareNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered by structured data. The description does add value by saying it gathers error history and redacted log signals and returns 'evidence and repair steps', which tells the agent what to expect back; however it says nothing about cost, latency, or whether log collection touches external systems.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences with zero waste. The scope enumeration comes first and the return-value note second, so the most decision-relevant content is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only diagnostic with no output schema, the description adequately states the subject areas and the shape of the result ('evidence and repair steps'). It nonetheless leaves the include_hardware flag and the depth/expense of the diagnosis unexplained, which is a noticeable gap for a tool that probes many subsystems.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter include_hardware has 0% schema description coverage and is not mentioned anywhere in the description. With no zero-param baseline to fall back on, the one undocumented flag leaves the agent unable to tell what enabling it changes about the diagnostic scope.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Diagnose') and enumerates the resources examined: API, server, disk, loaded models, error history and redacted log signals. That is far better than a tautology, but it never distinguishes itself from the many adjacent diagnostic/status siblings such as lm_status, lm_runtime and lm_connections, so an agent cannot confidently pick it over them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to reach for this tool versus lm_status or lm_runtime, nor any stated prerequisites or exclusions. The diagnostic intent is only implied by the verb 'Diagnose', leaving the agent to guess the trigger condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lm_docsC

Complete official repo knowledge. brief: French/English task references; guide: procedures; coverage: tools/tests/gaps; catalog/search/read: paginated sources; changes: latest delta; sync: force refresh. scope=all includes labeled drafts, configs and source scripts (never executed). Offset paginates results or characters within a read line range. Follow returned cursors. Media links are references, not downloaded/interpreted.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNo
limitNo
queryNo
scopeNopublished
actionYes
offsetNo
sectionNo
start_lineNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, which is a significant gap: the description mentions 'sync: force refresh' implying state mutation, but nonspecific. The description does not explain that sync force-refreshes and therefore may change data, nor does it clarify whether any action is destructive or requires permissions. With readOnlyHint=false, the description carries a higher burden and only lightly touches on sync's side effect.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a dense, semicolon-separated list that is front-loaded with a purpose statement but becomes a run-on inventory of actions. It is efficient in word count but sacrifices readability; an agent must parse multiple concepts in one sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with nine actions, eight parameters, and no output schema, the description is somewhat complete regarding scope, pagination, and media handling, but it omits key behavioral details like return formats, the effect of sync on system state, and which parameters pair with which actions. It is a minimum-viable description for a complex tool and would benefit from structured action-to-parameter mapping.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does explain that offset paginates results or characters, that scope=all includes drafts/configs, that limit controls pagination implicitly, and that path/section are used by read/catalog. However, it does not map parameters to specific actions, leaving ambiguity about which of the eight parameters apply to which of the nine actions. The description adds partial meaning but not enough to fully document eight parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description packs nine action verbs into a list of what the tool can do, making the tool's overall purpose (documentation/knowledge access) reasonably clear but each action's specific purpose is buried in dense latch-on phrases. It is distinguishable from siblings like lm_rag_search or lm_diagnose by its documentation focus, though it doesn't explicitly say 'this is the documentation tool versus X'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use lm_docs versus alternatives like lm_rag_ask or lm_status. The brief parentheticals hint at what each action does but do not explain when an agent should choose one action over another, or when to use this tool instead of a sibling RAG tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lm_downloadB
Destructive

Download a requested catalog model or Hugging Face URL. Returns job ID/status; can consume substantial disk.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelYes
quantizationNo

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (destructiveHint=true, openWorldHint=true), and the description adds real value beyond them: it discloses the disk-space cost and that it returns a job ID/status rather than the artifact itself. That async, resource-consuming nature is important operational context not present in structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the action front-loaded and the side effect (disk consumption) attached immediately. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully notes the job ID/status return, and it flags the resource cost. The only gap is the quantization parameter's meaning and whether polling via lm_download_status is expected, but the core operational picture is complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It roughly conveys what 'model' accepts (catalog name or HF URL) but says nothing about 'quantization', leaving one of two parameters undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Download) and resource scope (catalog model or Hugging Face URL), which distinguishes it reasonably well from lm_import_model and lm_load. However it does not explicitly contrast with the closest sibling, lm_import_model, leaving some inference required.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus lm_import_model, lm_load, or lm_download_status. The description never states prerequisites or the alternative path for loading a model already present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lm_download_statusB
Read-only

Check actual download progress; a submitted job is not a completed download.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds a genuinely useful behavioral caveat about job submission vs. completion, but says nothing about polling cadence, terminal states, or what happens for an invalid job_id.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with the key concept front-loaded ('check actual download progress'). No padding, though the semicolon clause is a caveat rather than an actionable instruction.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description would need to convey what a status response looks like (in-progress vs. complete vs. failed) — it only gestures at completion. For a simple single-parameter poll tool with annotations covering safety, this is minimally adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is exactly one parameter (job_id) with 0% schema description coverage, so the description carries the full burden of explaining it — and it says nothing about where the job_id comes from or its format. Since the description ignores the sole parameter entirely, it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (check download progress) and implicitly distinguishes from the lm_download sibling by warning that a submitted job is not a completed download. It never names the sibling tool explicitly, so sibling differentiation is inferred rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The clause 'a submitted job is not a completed download' implies this tool is used after lm_download to poll for completion, but it never says when to call it, how often, or which alternative to use. Usage is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lm_embeddingsC

Compute vectors with an embedding model for RAG/search. May auto-load the model.

ParametersJSON Schema
NameRequiredDescriptionDefault
inputYes
modelYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, destructiveHint=false, openWorldHint=false, covering the safety profile. The description adds one genuinely useful trait not in the annotations — 'May auto-load the model' — which explains the non-read-only hint and warns of a side effect. It still omits return format, batch/size limits, and any cost or latency notes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core purpose and no wasted words. Efficient, though the brevity is partly the source of the missing detail rather than pure tightness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Two required parameters with zero documentation, no output schema, and a thin two-sentence description leave the agent under-informed about how to call the tool or what it returns. The 'May auto-load the model' note helps but does not close the gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and there is no output schema, so the description carries the full burden — yet it says nothing about the 'model' or 'input' parameters. It does not clarify that model takes an identifier or that input accepts a single string or an array of strings for batch embedding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (compute) and resource (vectors) plus a purpose (RAG/search), which is enough for an agent to grasp the tool. It does not explicitly distinguish itself from related siblings like lm_rag_index or lm_rag_search, but the embedding-generating role is clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for RAG/search' implies a use context but gives no explicit when-to-use, when-not-to-use, or alternative routing. An agent cannot tell from this description whether to call lm_embeddings directly versus letting lm_rag_index or lm_rag_search produce embeddings internally.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lm_estimateB
Read-only

Estimate model memory needs without loading it, using installed CLI estimate-only support.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelYes
optionsNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, covering the safety profile. The description adds genuinely useful behavioral context by confirming the model is not actually loaded, i.e. a side-effect-free dry run. It says nothing about estimate accuracy or how the many options affect the result.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the operation front-loaded and zero filler. It is efficient, though arguably too terse for a tool exposing a large nested options object.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 0% schema description coverage on a tool whose 'options' object contains a dozen undocumented tuning fields, the description is not complete enough for an agent to invoke it correctly beyond passing a model name. No output schema exists, so return values need not be explained, but the parameter guidance gap is significant.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%: neither 'model' nor the deeply nested 'options' object (gpu, ttl, context_length, speculative_draft_*) is documented anywhere. The description mentions no parameter at all and does not compensate for the coverage gap, leaving the substantial options surface opaque.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Estimate model memory needs') and adds the key qualifier 'without loading it,' which cleanly distinguishes it from write-oriented siblings like lm_load and lm_load_advanced. It stops short of naming an alternative explicitly, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: you gather that this is a pre-load check, and 'using installed CLI estimate-only support' hints at a prerequisite. There is no explicit statement of when to prefer this over lm_load/lm_load_advanced or what to do if the estimate-only support is missing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lm_import_modelA

Preview/import a local GGUF model by copying it into LM Studio. Never moves the original file.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
dry_runNo
user_repoNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, destructiveHint=false and openWorldHint=false, so the safety envelope is already given. The description adds genuinely new behavioral context by clarifying the operation copies rather than moves, meaning the source file survives — something the annotations do not express. It still omits overwrite/collision and destination behavior, so not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and resource, and the non-destructiveness guarantee immediately follows. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating tool with no output schema and zero schema description coverage, the description should explain dry_run vs commit behavior and the user_repo/namespace parameter. It covers the core operation but leaves key parameter semantics and any return/error behavior undocumented.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 3 parameters, so the description carries the full documentation burden. It only loosely implies dry_run semantics via 'Preview' and says nothing about where the model lands, and user_repo is entirely unexplained in both description and schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (preview/import), the resource (local GGUF model), and the mechanism (copying it into LM Studio). This is readily distinguished from lm_download (remote fetch) and lm_load (loading into memory), so an agent can pick it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Preview/import' combined with the dry_run default of true implies the preview-then-commit workflow, but the description never states when to use this over lm_load or lm_download, nor any prerequisites (e.g., the file must already exist locally). Usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lm_integrationsA
Read-only

List LM Studio MCP plugin names and connector allowlist without exposing commands or secrets.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds genuine behavioral context beyond that: it discloses what is NOT returned (commands and secrets), which tells the agent this call is safe to surface and will not leak sensitive material.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the verb and resource, with the privacy guarantee appended. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-param, read-only listing tool with no output schema, the description conveys both what is returned and what is withheld. Only the lack of routing guidance relative to sibling tools keeps it from being fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. There is nothing further the description could add about parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and a concrete resource (LM Studio MCP plugin names and connector allowlist). This is distinguishable from LM siblings such as lm_mcp_config or lm_connections, though the description never explicitly names or contrasts with any sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus the many adjacent LM tools (lm_mcp_config, lm_mcp_probe, lm_connections). Usage is only inferable from the resource named in the sentence.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lm_loadC

Native load: context, flash attention, batch, experts, GPU KV cache. Verifies state and config.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelYes
optionsNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations mark this as non-read-only but non-destructive, framing it as a mutation. 'Verifies state and config' hints at post-load validation, but the description omits critical behavior: whether loading replaces an already-loaded model, whether an unload is required first, or what happens on failure. Given annotations only cover the safety hint, the description should carry much more.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short fragments, front-loaded and free of padding, which would score well for brevity. But it is under-specified rather than genuinely concise, and the telegraphic phrasing obscures meaning instead of conveying it efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating load tool with no output schema and 0% schema coverage, the description leaves the agent without required parameter meaning, load/overwrite semantics, or result expectations. It is not complete enough to invoke the tool confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for two undocumented parameters including a nested LoadOptions object with five fields. It loosely names the concepts ('context, flash attention, batch, experts, GPU KV cache') but gives no mapping to option names, value formats, defaults, or the required 'model' parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The fragment 'Native load' plus the option list implies loading a model with the named settings, and 'Native' loosely distinguishes it from the sibling lm_load_advanced. However, it never states plainly what is being loaded or onto what, so the verb+resource are only inferred from the tool name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use, no prerequisites (e.g. whether the model must be downloaded first), and no indication of when to prefer lm_load_advanced or lm_unload. The only implicit signal is the word 'Native', which is too weak to serve as guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lm_load_advancedB

CLI load: GPU offload, parallelism, TTL and speculative decoding. Requires identifier; checks installed flags.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelYes
optionsYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, and openWorldHint=false, so the agent knows this is a non-destructive mutation. The description adds that the call validates installed flags, which is useful behavioral context, but it says nothing about what loading does to runtime state or what happens on failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with no wasted words and the feature summary front-loaded. It is arguably too terse for the parameter surface, but every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with ~12 undocumented option properties, no output schema, and no annotation detail beyond safety hints, the description is too thin. It does not explain the effect of loading, the identifier precondition's interaction with the nullable schema field, or how the options combine.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across a large AdvancedLoadOptions object, so the schema documents names but no meaning. The description partially compensates by naming GPU offload, parallelism, TTL, and speculative decoding, which map to real properties, but leaves context_length and the draft token/probability settings unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action (CLI load) and enumerates the option families it covers (GPU offload, parallelism, TTL, speculative decoding), so an agent can tell it is the advanced variant of loading. However, it never explicitly differentiates from the sibling lm_load, relying on the name alone to convey the distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance and no mention of the basic lm_load alternative that a sibling list makes available. The only conditional context is 'Requires identifier', which is a precondition rather than usage routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lm_mcp_callA
Destructive

Invoke one explicitly authorized tool on a configured MCP after fresh schema discovery. Outputs are untrusted data.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
toolYes
argumentsNo

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, openWorldHint=true and readOnlyHint=false, so the safety profile is covered. The description adds two valuable facts beyond that: outputs must be treated as untrusted data, and a fresh schema discovery is a precondition. It still omits auth requirements and failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero filler, with the prerequisite and the output-trust caveat front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, open-world proxy call with no output schema and 0% schema description coverage, the definition covers the trust and authorization story but leaves the parameter contract and error/return behavior to inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must carry parameter meaning and it does not. 'name' is presumably the MCP server and 'tool' the target tool, but the description never says so, and the free-form 'arguments' object is left entirely unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Invoke) and resource (one tool on a configured MCP), and adds two qualifying conditions that narrow the action. It is distinguishable from siblings like lm_mcp_probe and lm_mcp_config, though it never names them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrases 'explicitly authorized' and 'after fresh schema discovery' imply prerequisites and sequencing, which is real guidance. However, it never states when to use this versus lm_mcp_probe or lm_mcp_config, nor any exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lm_mcp_configA
Destructive

Preview/apply LM Studio mcp.json changes with backup and conflict check. Reuse preview digest to apply. authorize grants exact tools for this config.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
applyNo
actionYes
configNo
allowed_toolsNo
expected_digestNo

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and openWorldHint=true, so the safety burden is partly carried. The description adds meaningful context beyond that: backup, conflict check, and digest-based apply safety. It still does not say exactly what gets overwritten or what permissions are required.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, zero waste, and the core purpose is front-loaded ahead of the workflow and authorize details. Every sentence carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Annotations cover the destructive/open-world profile and there is no output schema, so the description needn't explain returns. But for a complex 6-parameter, 0%-coverage mutation tool, it leaves several parameters and the list/remove actions opaque, making it only adequately complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 6 parameters, so the description must compensate but only partially does. It implies apply (preview/apply), expected_digest ('reuse preview digest'), and allowed_tools ('authorize grants exact tools'), but leaves the action enum values, config, and name unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb pair (preview/apply) and resource (LM Studio mcp.json changes), plus a safety qualifier (backup and conflict check). It is clearly distinct from siblings like lm_mcp_probe and lm_mcp_call, though it does not explicitly name alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an operational sequence ('Reuse preview digest to apply') and a scoping rule for authorize, which is useful implied guidance. However, it never states when to choose this tool over lm_mcp_probe/lm_mcp_call or when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lm_mcp_probeA
Destructive

Start/connect a configured MCP and list its real tools. Stdio executes its configured program; no tool invocation.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true and openWorldHint=true, which is a serious safety burden. The description is the only place clarifying that Stdio 'executes its configured program' as a side effect, while reassuring that no tool is actually invoked. That contradiction-resolution detail (executes a program but doesn't invoke a tool) is exactly the kind of behavioral context annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the primary action and followed by the critical constraint. No filler; every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a destructive, open-world operation with no output schema and an undocumented parameter, the description should at minimum clarify what 'name' refers to and what the probe returns. It covers the execution side-effect well but leaves the input and return shape unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single required 'name' parameter. The description mentions only 'configured MCP' obliquely and never explains what 'name' should contain (a config profile name, a connection identifier, etc.). With low coverage the description is required to compensate and it does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: connect a configured MCP and list its real tools. The phrase 'real tools' meaningfully distinguishes it from lm_mcp_config (which likely manages configuration) and lm_mcp_call (which invokes a tool). It does not explicitly name those siblings, keeping it just short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage comes before lm_mcp_call via the clarifying 'no tool invocation', which tells the agent this is for discovery rather than execution. However, it never explicitly states when to prefer this over lm_mcp_config or lm_diagnose, so the guidance is inferred rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lm_model_configA

Official SDK advanced load config (GPU, KV quantization, mmap, RoPE, seed). schema first; inspect compares SDK and REST; load requires new instance ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNo
actionYes
configNo
instance_idNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, destructiveHint=false, and openWorldHint=false, so the agent already knows this mutates but is not destructive. The description adds a useful behavioral detail — that 'load requires new instance ID' — implying a side effect of creating a new instance. However, it says nothing about permissions, reversibility, or what happens to running instances, so it adds only partial context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense, semicolon-separated clauses with no filler; the config scope is front-loaded ahead of the action routing. It is appropriately sized for a compact multi-action tool, though the terse phrasing borders on cryptic.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and a 0%-coverage schema containing a free-form config object, the description carries substantial burden. It covers action semantics and the instance_id requirement but leaves the config object's shape and the model parameter unaddressed, so an agent still has to inspect the schema to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does explain the three enum values of the required 'action' parameter and mentions that 'load' needs an instance_id, and it names config areas (GPU, KV quantization, mmap, RoPE, seed) that map to the opaque config object. But 'model' and the internal structure of the free-form config object remain undocumented, leaving meaningful gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the resource (official SDK advanced load config) and enumerates the config dimensions it governs (GPU, KV quantization, mmap, RoPE, seed), and it names the three actions (schema, inspect, load). It is clear enough to distinguish from the sibling lm_load_advanced by emphasizing the SDK config surface, though the boundary between the two is not fully drawn.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The second sentence routes between the tool's own actions: 'schema first; inspect compares SDK and REST; load requires new instance ID.' That is explicit when-to-use guidance at the action level. It does not, however, compare against sibling tools like lm_load or lm_load_advanced, so it falls short of full alternatives coverage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lm_modelsB
Read-only

List downloaded models, loaded instance IDs, effective configs and supported capabilities.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description usefully adds what content the response contains (models, instance IDs, effective configs, capabilities), which goes beyond the annotations, but it says nothing about pagination, filtering behavior, or cost of a large model inventory.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, listing the returned artifacts efficiently. Nothing needs trimming and nothing is buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description's enumeration of returned data is helpful, and for a read-only listing tool the missing safety details are not critical. However, the undocumented 'model' parameter and the lack of differentiation from similar sibling tools leave an agent without enough to decide when and how to call it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single 'model' parameter, so the description carries full responsibility — but it never explains what 'model' does. It is unclear whether it filters the listing to one model, selects a default, or is ignored; the description merely says 'models' generically.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a clear verb ('List') and enumerates the specific resources returned: downloaded models, loaded instance IDs, configs, and capabilities. It gives a concrete picture of the tool's output, though it never distinguishes itself from overlapping siblings like lm_status, lm_runtime, or lm_model_config.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus alternatives, no prerequisites, and no exclusions. With siblings such as lm_status, lm_runtime, lm_model_config, and lm_profiles that plausibly overlap in scope, the absence of routing guidance is a real gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lm_openai_chatA

OpenAI-compatible chat, explicit message history, JSON schema and function definitions. Returns tool calls; never executes them.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelYes
toolsNo
messagesYes
max_tokensNo
temperatureNo
tool_choiceNo
response_formatNo

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The final clause 'Returns tool calls; never executes them' is a critical behavioral disclosure that the annotations do not provide. It tells the agent that tool calls are returned for external execution, preventing a dangerous assumption that tools are executed automatically. Annotations declare non-readOnly and non-destructive, which aligns with a call that returns data without side effects, but the description adds unique operational context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence that is front-loaded with the core purpose and ends with the most important behavioral caveat. Every clause earns its place, with no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 7 parameters, 0% schema description coverage, no output schema, and no annotations indicating safety, the description is thin. It covers the key behavior (tool calls are returned, not executed) but doesn't explain the parameters or the expected response format beyond that. It is minimally adequate but has clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 7 parameters, so the schema itself provides no semantic help. The description mentions message history, JSON schema, and function definitions, which loosely maps to 'messages' and 'tools'/'response_format', but it doesn't explain parameters like model, max_tokens, temperature, or tool_choice. It partially compensates but leaves many parameters undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific operation: an OpenAI-compatible chat endpoint accepting explicit message history, JSON schema, and function definitions. This clearly separates it from siblings like lm_chat and lm_embeddings. However, it doesn't explicitly state what distinguishes it from lm_chat, which is a likely alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no when-to-use or when-not-to-use guidance. There is no mention of when to prefer this over lm_chat or other chat-like siblings. It provides implied usage through the feature list, but no explicit routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lm_profilesC

Save/reuse named model configurations in this project. load uses the SDK with postcondition verification.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
modelNo
actionYes
configNo
instance_idNo

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, destructiveHint=false, openWorldHint=false, so the agent knows this mutates local state. The description adds nothing about what save/delete actually do, whether delete is reversible, or what persists. With 4 actions including a destructive-sounding 'delete' (yet destructiveHint=false), the description should clarify but doesn't.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Very short (two sentences), so no bloat, but it is under-specified rather than concise. The second sentence is oddly scoped to one action out of four rather than front-loading the overall behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A 5-parameter, multi-action mutation tool with 0% schema coverage, no output schema, and only two thin sentences of description. Essential context about each action, parameter roles, and persistence/verification behavior is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description explains none of the 5 parameters. The only enum (action) has four values whose meaning (list/save/load/delete) is left entirely to inference, and name/model/config/instance_id receive no explanation in either schema or description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States the general purpose (save/reuse named model configurations) with a specific verb+resource, but the second sentence about 'load' is confusing and the description never explains the multi-action 'action' enum (list/save/load/delete). It doesn't clearly distinguish this from siblings like lm_model_config or lm_load.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use or when-not-to guidance, no mention of alternatives. The fragment 'load uses the SDK with postcondition verification' hints at a load path but doesn't tell the agent when to choose this tool over lm_load or lm_load_advanced.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lm_rag_askC

Retrieve then answer locally. Returns no answer without valid source IDs and exact source quotes; semantic correctness still needs review.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelYes
queryYes
top_kNo
min_scoreNo
collectionYes
embedding_modelYes

TDQS

C2.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations present (readOnlyHint=false, openWorldHint=false, destructiveHint=false), the description adds genuinely non-obvious behavior: it will return NO answer unless valid source IDs and exact source quotes exist, and semantic correctness still requires human review. That prevents an agent from treating an empty result as a failure and flags the trust boundary. It stops short of explaining why readOnlyHint is false for what sounds like a read operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the core behavior and the failure condition. No filler, though the extreme compression trades away detail an agent would need.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter, 4-required tool with no output schema and no annotation coverage of the parameters, the description is far too thin: it omits parameter meanings, required inputs, and return shape. The behavioral caveats are useful but insufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Six parameters with 0% schema description coverage, so the description carries the full burden — yet it explains none of them (collection, query, model, embedding_model, top_k, min_score). An agent cannot know what a 'collection' or 'min_score' means here from either structured fields or the text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States the core activity (retrieve then answer locally) which implies RAG question answering, and is distinguishable from lm_rag_search/lm_rag_index by the 'answer' verb. However it never names what is being queried (a collection) or explicitly contrasts with lm_rag_search, leaving the purpose only partially pinned down.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus lm_rag_search, lm_rag_index, or lm_chat, nor any prerequisites. The reader must infer usage from the name and the fragmentary description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lm_rag_indexA

Index selected TXT/MD/PDF/DOCX files with local embeddings and source/page/hash metadata. Roots set by LM_MCP_RAG_ROOTS. replace rebuilds collection.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathsYes
replaceNo
collectionYes
embedding_modelYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare write/non-destructive profile. The description adds real operational context beyond that: roots limited by LM_MCP_RAG_ROOTS env var, metadata captured (source/page/hash), and that 'replace rebuilds collection'. That's meaningful behavioral disclosure for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with what gets indexed and how, followed by the environment constraint and the flag behavior. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-param indexing tool with no output schema and zero schema coverage, the description should explain what 'collection' and 'embedding_model' mean and what happens on success/failure. It covers replace, roots, and metadata but leaves the two most consequential required params opaque.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden. It explains 'replace' semantics and partially clarifies 'paths' by listing file types, but 'collection' and 'embedding_model' — both required — get no explanation. Partial compensation only.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Index') and resource (TXT/MD/PDF/DOCX files) with the storage mechanism (local embeddings) and metadata captured. Clear enough to distinguish from lm_rag_ask/lm_rag_search, though it doesn't explicitly name those siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The mention of 'replace rebuilds collection' gives condition-level guidance for that flag, and the roots constraint implies a precondition. But there's no explicit when-to-use versus lm_rag_search/lm_rag_ask, nor guidance on incrementally adding vs rebuilding.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lm_rag_manageB

List RAG collections/document roots, or delete only an index. Original documents are retained.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYes
collectionNo

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the safety profile (readOnlyHint=false, destructiveHint=false), and the description adds genuinely useful scope information: only the index is removed and the original documents are retained, which explains why the delete is not flagged destructive. It does not cover permissions or whether the index can be rebuilt afterward.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence that front-loads the two operations and ends with the key retention guarantee. Nothing is padded, though the retention clause could be tightened.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and 0% parameter coverage, the description should explain what 'list' returns and how 'collection' scopes each action. Those gaps leave the delete path ambiguous, which matters for a mutation-capable tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry both parameters. It mentions 'collections/document roots' generically but never explains the 'collection' parameter — whether it is required for delete, optional for list, or what identifiers are accepted.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource for each mode: listing RAG collections/document roots and deleting an index. It clarifies scope ('only an index') but never names or contrasts with the obvious siblings lm_rag_index and lm_rag_search, so an agent must infer the boundary itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied through the two verbs, which mirror the action enum (list/delete). There is no explicit when-to-use statement, no condition for choosing delete over re-indexing, and no mention of alternatives such as lm_rag_index or lm_rag_search.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lm_runtimeC
Destructive

Inspect engines/hardware or check stable updates (dry run). update explicitly installs stable engines; retest inference afterwards.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the mutation risk is broadly known. The description usefully differentiates actions by noting that check_updates is a dry run and that 'update' installs stable engines, and warns that inference must be retested afterwards — real added context. It still omits permission requirements, disruption scope, and download/size implications for a destructive open-world operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly packed sentences with the read/inspect path front-loaded before the mutating 'update' caveat. Minimal waste, though the phrasing is compressed to the point of slight ambiguity about which action maps to which verb.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter, multi-action tool with no output schema, the description covers the destructive path and its follow-up but does not explain what 'list', 'hardware', or 'available' return, nor the format of results. Adequate but with visible gaps given zero schema coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the single 'action' enum is undocumented, and the description must compensate. It maps part of the enum ('inspect engines/hardware', 'check_updates', 'update') but leaves 'list' and 'available' undefined, so the enum is only partially disambiguated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names verbs and resources for most actions ('Inspect engines/hardware', 'check stable updates', 'installs stable engines'), which is more than a tautology. However, two of the five enum values ('list', 'available') are never explained, and the description does not distinguish this tool from similarly-named siblings such as lm_status, lm_diagnose, or lm_models that plausibly also report engine/hardware state.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'check stable updates (dry run)' implies that one action is safe to probe, and 'retest inference afterwards' gives an internal follow-up step. But there is no guidance on when to pick lm_runtime over lm_diagnose/lm_status, no prerequisites, and no when-not-to-use statement, leaving tool selection to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lm_server_controlA
Destructive

Check/start/stop LM Studio HTTP server. Starts bound to 127.0.0.1, never to the LAN.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and openWorldHint=false, so the mutating and local-only nature is partly covered. The description adds a genuinely useful detail not in the annotations: the server binds to 127.0.0.1 and never the LAN. It does not say what stopping the server does to in-flight inference or loaded models, which is the most consequential missing behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero filler, with the primary capability stated first and the network-binding constraint second. Nothing could be removed without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter control tool with no output schema, the description is nearly sufficient: it covers the operations and the security-relevant binding behavior. It is only slightly thin on what 'status' returns and on the restart action.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and there is a single enum parameter whose values are largely self-describing. The description maps three of the four values (check=status, start, stop) but leaves 'restart' unexplained, so it adds only partial meaning over the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource (the LM Studio HTTP server) and the operations performed on it (check/start/stop), which is more than a restatement of the name. It stops short of distinguishing itself from adjacent siblings such as lm_status or lm_runtime, so an agent must still infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the enumerated operations rather than stated: there is no explicit guidance on when to control the server versus inspecting it via lm_status, and no prerequisites or exclusions. The description also omits 'restart', one of the four enum values, so the usage picture it paints is incomplete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lm_statusB
Read-only

Live API/app/runtime status, compatibility diagnostics and official update check (cached 1h).

ParametersJSON Schema
NameRequiredDescriptionDefault
refresh_updatesNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, non-destructive, closed-world), so the bar is lower. The description adds one genuinely useful behavioral fact - the 1h cache - but says nothing about what happens on a cache miss, whether refresh_updates bypasses the cache, or what the diagnostic output contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence with no filler; the cache caveat is placed at the end where it is least disruptive. Every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and one undocumented parameter, the description should carry more burden than it does. It tells the agent what domains are covered but not what is actually returned or how this differs from lm_diagnose/lm_runtime, which matters in a 28-tool namespace.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single refresh_updates parameter. The description's mention of an 'official update check (cached 1h)' implicitly gestures at the refresh behavior, which is marginal added meaning, but it never states that refresh_updates forces a fresh check rather than using the cache.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names three concrete content areas (live status, compatibility diagnostics, official update check), which is more informative than a tautology. But it is a noun-phrase list with no verb and gives no signal to distinguish it from close siblings like lm_diagnose, lm_runtime, or lm_connections.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use or when-not-to-use guidance. With 28 sibling tools including the obviously overlapping lm_diagnose and lm_runtime, the absence of routing guidance leaves the agent to guess which status/diagnostic tool to call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lm_unloadA
Destructive

Unload precisely one instance and verify absence. May interrupt its active generations.

ParametersJSON Schema
NameRequiredDescriptionDefault
instance_idYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true, and the description adds crucial extra context: it interrupts active generations and verifies absence afterward. This goes beyond the annotation by naming the side effect and the confirmation behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, front-loaded sentences with zero filler. Every clause adds new information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter destructive tool with no output schema, the description covers the key risks (interrupting generations) and the verification step. Missing only explicit guidance on when to use it versus alternatives.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one required string parameter (instance_id) with 0% schema description coverage. The description's 'precisely one instance' emphasizes that a single instance_id is expected and that multi-unload is not supported, adding useful constraint.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Unload) and resource (one instance), and specifies cardinality ('precisely one instance'). Does not differentiate from sibling lm_load, though they are natural opposites.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus lm_load or lm_load_advanced, and no preconditions stated. Usage is only weakly implied by the verb.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 28 tool updatesv2.1.0
    • First observedlm_chat
    • First observedlm_connections
    • First observedlm_diagnose
    • First observedlm_docs
    • First observedlm_download
    • First observedlm_download_status
    • First observedlm_embeddings
    • First observedlm_estimate
    • First observedlm_import_model
    • First observedlm_integrations
    • First observedlm_link
    • First observedlm_load
    • First observedlm_load_advanced
    • First observedlm_mcp_call
    • First observedlm_mcp_config
    • First observedlm_mcp_probe
    • First observedlm_model_config
    • First observedlm_models
    • First observedlm_openai_chat
    • First observedlm_profiles
    • First observedlm_rag_ask
    • First observedlm_rag_index
    • First observedlm_rag_manage
    • First observedlm_rag_search
    • First observedlm_runtime
    • First observedlm_server_control
    • First observedlm_status
    • First observedlm_unload

TDQS

B3/5.0

Scored across 28 tools

Disambiguation3/5

Many tools are distinct, but several clusters overlap: lm_load/lm_load_advanced/lm_model_config/lm_profiles all configure loading; lm_status/lm_diagnose/lm_runtime overlap in inspection/update; lm_chat/lm_openai_chat and lm_rag_ask/lm_rag_search differ mainly by interface. Descriptions clarify intent but misselection risk remains.

Naming Consistency5/5

All 28 tools use the lm_ prefix and snake_case, with mostly predictable noun or verb_phrase patterns. lm_load_advanced is the only compound variation but remains readable and consistent with the set.

Tool Count2/5

28 tools is heavy for a single MCP server and exceeds the 16-25 borderline range. While the broad LM Studio domain justifies many operations, the set could be consolidated, especially around load/config/profile tools.

Completeness4/5

The set covers runtime, models, loading, downloads, chats, RAG, MCP, docs, diagnostics, and server control, which is broad for the domain. Minor gaps like deleting models or more direct MCP server lifecycle management exist, but core local lifecycle coverage is mostly complete.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    Unified MCP server for managing local model runtimes (Ollama, LM Studio, etc.), enabling provider-agnostic discovery, lifecycle management, hardware-fit checks, and delegated inference.
    16
    18 npm
    Creative Commons Attribution Non Commercial No Derivatives 4.0 International
  • F
    license
    A
    quality
    C
    maintenance
    MCP server that connects LLM agents to a local LM Studio instance, enabling model management, OpenAI-compatible chat completions, text completions, and embeddings through a set of tools.
    9
    1
    -