Skip to main content
Glama
RiasJ1Dar

glm-orchestrator

by RiasJ1Dar

glm-orchestrator

Відкритий MCP-оркестратор для GLM та інших моделей із OpenAI-compatible API. Він підтримує потокові відповіді, фонові запуски, скасування, локальний журнал, web UI та опціональний міст до інших MCP-серверів.

Цей публічний репозиторій має окрему чисту історію і не містить робочих промптів, журналів, локальних конфігів або даних доступу.

Інструкції для AI coding agents: AGENTS.md. Інтерактивна архітектурна схема: architecture.html.

Вимоги

  • Node.js 20.3 або новіший;

  • OpenAI-compatible endpoint із /v1/models, /v1/chat/completions або /v1/responses;

  • API key, якщо його вимагає провайдер.

Related MCP server: Agentforce MCP Integration Server

Встановлення

git clone https://github.com/RiasJ1Dar/glm-orchestrator-public.git
cd glm-orchestrator-public
npm install

Задайте конфігурацію через змінні процесу:

GLM_BASE_URL=https://your-openai-compatible-endpoint.example
GLM_API_KEY=your-api-key
GLM_MODEL=glm-5.2

GLM_API_KEY можна не задавати для локального endpoint без авторизації. Не записуйте ключ у git, MCP-конфіг або командний рядок.

Запуск

npm start

MCP-сервер працює через stdio. Web UI запускається автоматично на http://127.0.0.1:9751/. Автозапуск можна вимкнути через GLM_UI=0, а порт змінити через UI_PORT.

Приклад реєстрації в MCP-клієнті:

{
  "mcpServers": {
    "glm": {
      "command": "node",
      "args": ["/absolute/path/to/glm-orchestrator-public/src/server.mjs"],
      "env": {
        "GLM_BASE_URL": "https://your-endpoint.example",
        "GLM_MODEL": "glm-5.2"
      }
    }
  }
}

Передавайте GLM_API_KEY безпечним механізмом вашого середовища, а не у файлі, який комітиться.

Основні інструменти

  • glm_models — моделі, які повертає endpoint;

  • glm_build_prompt — підготовка структурованого prompt;

  • glm_run — синхронний або фоновий запуск;

  • glm_status, glm_result, glm_wait, glm_cancel — керування запуском;

  • glm_list_runs — локальна історія;

  • mcp_status — стан опціонального MCP-моста.

Додаткові MCP-сервери налаштовуються локальним mcp.json. За основу можна взяти mcp.example.json; справжній mcp.json ігнорується.

Безпека

  • ключ береться лише з GLM_API_KEY і не записується в журнал;

  • UI та proxy слухають тільки 127.0.0.1;

  • runtime-дані в state/ і logs/ не відстежуються Git;

  • запис файлів обмежений коренем репозиторію та GLM_WRITE_ROOTS;

  • читання й запис .env, ключів, SSH/AWS-конфігів та .git блокуються;

  • перед публікацією запускайте secret scanner для всього нового history.

Перевірка

npm test

Стандартний набір тестів не потребує мережі або справжнього API key.

Ліцензія

ISC — див. LICENSE.

Available Tools

20 tools
debug_tailНалагоджувальний журналA

Сирий обмін зі шлюзом: тіла запитів, відповіді, виклики інструментів. Пишеться лише коли сервер запущено з GLM_DEBUG=1. Тут видно те, чого немає у звичайному журналі: биті arguments, роздуті промпти, справжній текст відповіді шлюзу.

ParametersJSON Schema
NameRequiredDescriptionDefault
grepNoЛишити лише записи, що містять цей текст
blocksNoСкільки останніх записів (типово 3)

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full burden of disclosing behavior. It reveals that the tool reads raw gateway traffic, that its availability depends on a server-side debug flag, and that it exposes content not present in normal logs. It stops short of explicit verb/read-only semantics, but the context strongly implies a non-destructive tail operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences, each carrying distinct information: what the log contains, when it is written, and what unique value it offers. No redundancy or filler; the most essential fact is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter, no-output-schema tool, the description provides essential operational context: the raw content type, the conditional availability, and the extra diagnostic value. The only gap is that it does not explicitly describe the return format or what happens when debug mode is off, but the schema and the condition imply enough for a competent agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents both optional parameters (grep and blocks) with full coverage. The description does not add any parameter-level detail, so the baseline 3 for high schema coverage is appropriate; no extra meaning is necessary or missing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool exposes a raw gateway exchange log: request bodies, responses, and tool calls. It clearly distinguishes what this log contains (broken arguments, bloated prompts, actual gateway response text) and when it is available (GLM_DEBUG=1). This is far more informative than the title alone and differentiates it from the sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear condition for use: the log is written only when the server is started with GLM_DEBUG=1, so the agent knows to rely on this tool only in debug sessions. It also positions it as a complement to the 'usual log' by highlighting what is uniquely visible here. However, it stops short of explicitly naming alternative tools or stating when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

git_commitЗакомітитиA

Комітить зміни в D:/GLM-5.2. Файли, схожі на креденшели, відхиляє. Push не робить — це окреме рішення людини.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathsNoОбмежити коміт цими шляхами
messageYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals two non-obvious behaviors: automatic rejection of credential-like files and absence of push. This goes well beyond a generic 'commit' description, though it omits details about staging behavior and error handling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two terse sentences with no filler. The core action is front-loaded, and both behavioral caveats (credential filtering, no push) earn their place without repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter commit tool, the description adequately covers purpose, repository scope, and key behavioral caveats. The main gap is the undocumented required message parameter and the absence of guidance on pre-commit checks such as git_status.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: only paths has a description ('Обмежити коміт цими шляхами'), while the required message parameter is undocumented. The tool description adds no meaning for either parameter, so an agent gets no help understanding required message content or default paths behavior beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Комітить') and resource ('зміни в D:/GLM-5.2'), and adds two differentiating constraints: credential-like files are rejected and push is not performed. This clearly distinguishes the tool from read-only siblings like git_status, even though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies use for creating a local commit in the specified repository and explicitly warns that push is a separate human decision ('Push не робить — це окреме рішення людини'). However, it does not state when to prefer this tool over alternatives or suggest preconditions like checking git_status first.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

git_statusСтан репозиторіюA

Гілка, останній коміт, незакомічені зміни.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the transparency burden. The fragment implies a read-only inspection by enumerating status fields, but it does not explicitly state that the tool makes no changes, nor does it mention error behavior or whether a valid git repository is required.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single compact phrase containing only the three relevant data categories: branch, last commit, and uncommitted changes. There is no filler, repetition, or unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple parameterless read-only status tool, the description covers the main output content without requiring an output schema. It omits usage guidance and return formatting, but the low complexity and obvious semantics make this an acceptable gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and the input schema fully reflects that. The description correctly adds no unnecessary parameter details, and the baseline of 4 applies for a parameterless tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description lists concrete elements of what the tool exposes (branch, last commit, uncommitted changes), and the title clarifies that this is repository status. However, it lacks an explicit verb such as 'show' or 'report', and it does not explicitly differentiate from status-like siblings such as glm_status or mcp_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool instead of alternatives like git_commit, glm_status, or mcp_status. No context is given such as checking state before committing, and no exclusions or prerequisites are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glm_applyЗаписати FILE-блоки з відповіді GLMA

Бере результат запуску (типово останній done) і записує блоки FILE: шлях. Секрети й шляхи поза GLM_WRITE_ROOTS відхиляє.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idNo

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and discloses useful behavior: it consumes a run result, writes FILE: blocks, and rejects secrets or paths outside GLM_WRITE_ROOTS. It does not mention overwrite behavior or failure semantics, but the stated security constraints go beyond a generic 'writes files' statement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences contain the core action, default source, and key constraints with no redundant filler. The most important information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter tool, the description covers the main action, default selection, and safety boundaries. However, there is no output schema and the description does not indicate what the tool returns or how it reports rejected blocks, which leaves some ambiguity for an agent needing confirmation of success.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema only lists run_id as an optional string with no description. The description adds meaning by explaining that the run result is used and that the default is the most recent done run, effectively documenting run_id as optional and clarifying its default behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the verb (writes/applies) and resource (FILE: blocks from a GLM run result), and specifies the default source as the last done run. It is distinguishable from siblings like glm_result because it performs file writes rather than just returning the result.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is used after a run to apply its FILE outputs, but it does not explicitly state when to use it instead of glm_result, glm_run, or plan_append, nor does it mention conditions like 'use only after a completed run'. The only guidance is the default selection of the last done run.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glm_build_promptСкласти промптA

Складає промпт під конкретний інструмент за правилами скіла prompt-master. Складає glm-5.2. Для випадків, коли оркестратор не має власного скіла. Інші моделі шлюзу платні — не підставляти без потреби.

ParametersJSON Schema
NameRequiredDescriptionDefault
ideaYesСира ідея або завдання своїми словами
modelNoХто складає промпт (типово glm-5.2)
constraintsNoДодаткові межі
target_toolYesДля кого промпт: Claude Code, Codex CLI, GLM, Midjourney…

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden. It discloses that it uses glm-5.2 and that other models are paid, which is helpful cost behavior. It doesn't mention side effects or output format, but as a prompt builder it's likely safe. Still, it could disclose that the result is a prompt text or that it doesn't execute anything.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (three sentences) and front-loaded with the primary function. It includes usage condition and cost warning, all relevant. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and no annotations, the description should explain what the tool returns and any side effects. It doesn't mention that the result is a prompt string, nor how to handle errors. It's adequate for a simple tool but lacks detail about the return value.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all parameters have descriptions in the schema. The tool description adds no additional meaning to the parameters; it only repeats the general purpose. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action: composing a prompt for a specific tool using the prompt-master skill, and specifies the model (glm-5.2). It clearly distinguishes this from running a prompt (sibling tools like glm_run). However, it doesn't explicitly name an alternative tool, so it's clear but not maximally differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides a usage condition: 'for cases when the orchestrator does not have its own skill', and adds a warning about not substituting paid models unnecessarily. This gives context for when to use it, though it doesn't name an explicit alternative tool to use instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glm_cancelСкасувати запускB

Обриває запуск, що виконується.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of disclosing side effects. It only says the run is aborted; it does not reveal whether cancellation is irreversible, what happens to results already produced, or whether it errors on non-running runs. This is a meaningful gap for a destructive-ish operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short active sentence with no filler. The core verb and target are front-loaded, making it easy to parse quickly. All the conciseness concerns are separate from the lack of detail, which is captured in other dimensions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter cancel tool, the basic call shape is inferable and the 'running' qualifier is a useful precondition. However, with no annotations and no output schema, the description should ideally mention what happens after cancellation, whether it is idempotent, and whether it only applies to running runs. The gap is moderate but not severe.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema only provides a required string run_id, and schema description coverage is 0%. The tool description does not explain what run_id refers to, how to obtain it, or any format expectations. The parameter name is self-descriptive, but the description itself adds no parameter-level meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('обриває' / aborts) applied to a specific resource ('запуск' / run), and qualifies it as a running run. This clearly identifies the tool as the cancel counterpart to glm_run or glm_wait, though it does not explicitly name those siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The qualifier 'що виконується' implies the tool should be used on in-progress runs, which is a useful usage clue. However, there is no explicit guidance about when not to use it, what alternatives exist (e.g., glm_wait or glm_status), or what preconditions apply.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glm_list_runsОстанні запускиC

Перелік запусків, найновіші згори.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It only states the action 'list' which implies a read-only operation, but it does not mention pagination, limit semantics, response format, or any side effects. The description is too sparse to convey behavior beyond the obvious.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise in length but under-specified. It consists of a single short sentence that provides minimal information. There is no front-loading of key details, and it lacks any structure to aid comprehension. The brevity is not a strength here because it omits essential context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has no output schema, no annotations, and an undocumented parameter, the description is grossly incomplete. An agent cannot determine what the response looks like, how 'limit' behaves, or any error conditions. The description fails to provide the necessary context for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has one parameter, 'limit', with no description in the schema (0% coverage). The tool description does not mention this parameter at all. An agent cannot infer the meaning or usage of 'limit' from the description, making the parameter semantics completely unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'Перелік запусків' means 'List of runs', and it specifies ordering with 'найновіші згори' (newest at top). This is a specific verb-resource combination that distinguishes it from siblings like glm_run (which executes a run) and glm_status (which shows status). The purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention any context, prerequisites, or exclusion criteria. The only implicit cue is that it lists runs, but there is no explicit direction or comparison to sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glm_modelsСписок моделейB

Моделі, які повертає налаштований OpenAI-compatible endpoint.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It only states that models come from the configured endpoint; it does not describe whether the call is live, cached, ordered, or what the return shape looks like. This is minimal behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence with no filler. It front-loads the core subject and adds the useful source detail about the OpenAI-compatible endpoint. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is low in complexity and requires no parameters, but there is no output schema and no description of the return format. An agent trying to pick a model for glm_run may need to know whether the result is a list of IDs, names, or objects; that is currently unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and the input schema is an empty object, so schema coverage is effectively 100%. The description adds no parameter details, but none are needed; with no parameters, the baseline of 4 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the tool as returning the models provided by the configured OpenAI-compatible endpoint, and the title 'Список моделей' reinforces it as a listing operation. It is clearly distinct from siblings like glm_run or glm_apply. It lacks an explicit imperative verb, but 'returns' plus the title makes the purpose clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives. It does not mention related tools, preconditions, or any scenario where this list should be consulted, so an agent receives no help in selecting it over other GLM tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glm_resultРезультат запускуC

Повний результат разом із промптом, яким його отримано.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYes
with_promptNoДодати текст промпту

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, but it only mentions the returned contents and not read-only behavior, prerequisites, output format, or error conditions. It also implies the prompt is always included ('разом із промптом'), while the schema indicates the prompt is only added when with_prompt is true.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler, and the key idea is front-loaded. It loses a point because the noun-phrase style and ambiguity about the prompt weaken its structural clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema and no annotations, the description is too thin to support correct invocation. It omits how run_id is supplied, whether the response is raw output or a structured object, and whether the prompt inclusion is conditional, leaving an agent to guess about the with_prompt behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%, so the description must compensate for the undocumented run_id, but it does not explain how to obtain or format run_id. The with_prompt parameter is adequately described in the schema, and the description only loosely echoes that concept without clarifying the optionality.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states that the tool exposes the full run result and the prompt that produced it, so the resource and payload are clear. It lacks an explicit action verb like 'retrieve' and does not directly differentiate from sibling tools such as glm_status or glm_list_runs, though 'full result' hints at that distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use glm_result versus glm_status, glm_wait, or glm_list_runs. The description only defines what the result is, not the context in which an agent should call it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glm_runПоставити задачу GLMB

Виконує готовий промпт. Промпт MUST бути складений викликаючою стороною за скілом prompt-master — сира постановка дає помітно гірший результат. За замовчуванням чекає результат; з background: true повертає run_id одразу.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskYesГотовий промпт, складений за prompt-master
filesNoАбсолютні шляхи файлів, які додати у промпт
modelNoМодель (типово glm-5.2)
skillNoСкіл, правила якого вкласти у промпт — свій або Grok-івський (див. skills_list)
thinkNoВнутрішнє міркування моделі. false пришвидшує втричі і звільняє бюджет на саму відповідь — для генерації документів ставити false
refineNoСкласти промпт іншою моделлю — запасний шлях для клієнта без скіла (типово вимкнено)
use_mcpNoДати моделі інструменти MCP-серверів: граф коду, канал до інших агентів
done_whenNoКритерій готовності
backgroundNoНе чекати результат, віддати run_id
max_tokensNoСтеля вихідних токенів (типово 32000)
constraintsNoЧого робити не можна
mcp_serversNoОбмежити перелік серверів, напр. ["codebase-memory"]
refine_modelNoХто складає промпт при refine (типово glm-5.2)
output_formatNoЯкої форми чекаємо результат

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that it waits by default and returns run_id immediately with background:true, and notes the prompt must be prepared with prompt-master. However, it omits other behaviors like token consumption, error handling, or whether the operation is destructive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loaded, stating the core purpose and the key constraint first. Every sentence adds relevant information, and it avoids redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 14 parameters and no annotations or output schema, the description is insufficient. It doesn't describe the return format (other than run_id for background), error handling, or how to retrieve results later. The tool is complex, and the description leaves many operational gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all parameters. The description adds some context (e.g., the prompt requirement, background behavior) but doesn't significantly go beyond the schema's own descriptions. It meets the baseline but adds little extra value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool executes a ready prompt, using a specific verb and resource. It differentiates from siblings like glm_build_prompt by emphasizing the prompt must already be composed, but it doesn't explicitly name alternative tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides context on when to use it (with a ready prompt composed via prompt-master skill) and mentions a fallback via refine, but it doesn't explicitly list alternative tools or state when not to use it. The guidance is implicit rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glm_statusСтан запускуB

Статус запуску за run_id, без повного тексту результату.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It only mentions that the full result text is not included, which is a useful behavioral trait, but it omits other relevant aspects such as read-only nature, whether the tool blocks or returns immediately, or any authentication requirements. This is minimal coverage for a status-checking tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, direct sentence with no filler or redundant phrasing. It front-loads the core purpose and the key distinction (excluding full result text), making it efficient for an agent to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one string parameter, no output schema), the description is minimally adequate but not complete. It lacks information about the response format, possible status values, or how the status should be interpreted. It also does not clarify the relationship to sibling tools like glm_wait or glm_cancel, leaving some ambiguity for the agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It mentions run_id but only by name, without explaining its format, origin, or how it relates to the run lifecycle. The schema itself only defines it as a string, so the description adds no meaningful semantic value beyond restating the parameter name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool retrieves the status of a run by run_id and explicitly excludes the full result text, which clearly distinguishes it from a result-fetching tool like glm_result. It is specific about the resource (run) and identifier, but does not name the alternative sibling tool directly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this tool is for checking status without pulling the full result, which gives some context. However, it does not explicitly state when to choose this over alternatives like glm_wait or glm_result, nor does it provide any exclusion criteria or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glm_waitДочекатись запускуA

Чекає завершення фонового запуску і віддає результат. Поки чекає — шле клієнту progress-нотифікації, щоб оркестратор бачив хід, а не смикав glm_status наосліп. На таймауті запуск НЕ обривається: повертає стан, а робота йде далі.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYes
timeout_secNoСкільки чекати (типово 600)

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it discloses key behavioral traits: it blocks until completion, emits progress notifications while waiting, and does not abort the run on timeout. It does not go into return-shape details, but the behavior most relevant to an agent is clearly stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, all load-bearing: what it does, what progress signals it sends, and what happens on timeout. The most important fact is front-loaded and nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple blocking-wait tool, the description covers the workflow, timeout semantics, and relation to glm_status. Without an output schema or annotations, a bit more detail about the returned result/state shape would make it fully complete, but not much is missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: timeout_sec already has a description with a default, and run_id is self-explanatory from the run-management sibling set, but the description itself adds no parameter-level detail. This is acceptable but not additive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('Чекає завершення фонового запуску') and states the output ('віддає результат'). It differentiates the tool from glm_status by explicitly rejecting blind polling, so an agent can select it correctly among the run-management siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It names the alternative glm_status and tells the orchestrator to use this tool to get progress rather than poll status blindly. It also sets expectations for timeout behavior ('на таймауті запуск НЕ обривається'), which is important usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mcp_statusСтан мосту до MCP-серверівA

Які сервери з mcp.json піднялись, які їхні інструменти доступні моделі і що не піднялось — з причиною.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden of behavioral disclosure. It usefully states the output content categories: available servers, accessible tools, and failures with reasons. However, it does not explicitly confirm the operation is read-only, nor describe output format, timeout, or network-check behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that immediately conveys purpose and key output categories. Every part earns its place, and there is no redundant wording or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-parameter status tool, the description covers the essential return-value categories: successful servers, available tools, and failures with reasons. It does not describe exact output formatting, but given the low complexity and absence of an output schema, this is a minor gap rather than a critical omission.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, so the baseline for this dimension is 4. The description correctly focuses on the tool's behavior rather than parameters, and no parameter documentation is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb-resource pairing: it reports which MCP servers from mcp.json are up, which of their tools are available, and which failed with reasons. This clearly distinguishes it from sibling status tools like git_status and glm_status by naming the MCP server context explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is used to check MCP server and tool availability, but it does not explicitly say when to prefer this over alternatives or mention any exclusions. It gives clear context but leaves the when-to-use decision to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_addЗапам'ятати фактA

Дописує факт у пам'ять проєкту. Годиться для того, що має пережити сесію: рішення, обмеження середовища, домовленості. Не для того, що видно з коду.

ParametersJSON Schema
NameRequiredDescriptionDefault
factYesОдин факт, одним реченням
tagsNoМітки, напр. ["середовище"]

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It does disclose that this is an append operation and that stored facts persist beyond the session. However, it does not mention return behavior, errors, permissions, or how facts relate to memory_read, leaving some behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two compact sentences: the first states the operation, and the second gives use cases and an exclusion. Every sentence earns its place, with no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter append tool, the description and schema together cover everything needed to call it correctly: what it does, when to use it, when not to use it, and how to format the fact and tags. No output schema is needed for this kind of operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters: fact is one sentence and tags are labels like ["середовище"]. The description adds no parameter-specific meaning beyond what the schema provides, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Дописує факт у пам'ять проєкту' (appends a fact to project memory). It also differentiates the tool from code-related siblings by explicitly saying it is not for what is visible from the code, making its scope unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear usage context: use it for decisions, environment constraints, and agreements that must survive the session, and do not use it for code-visible facts. It provides both positive and negative guidance, though it does not explicitly name alternative tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_forgetЗабути фактA

Прибирає з пам'яті рядки, що містять указаний текст. Факт, що став неправдою, шкідливіший за його відсутність.

ParametersJSON Schema
NameRequiredDescriptionDefault
containsYesФрагмент рядка, який прибрати

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must fully disclose behavioral traits. It states that rows containing the text are removed, but does not mention that this is permanent/irreversible, that all matching rows are affected, or any side effects. The philosophical note adds no operational transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, with the operational purpose front-loaded. The second sentence is a philosophical justification that is not needed for function but is short. Overall concise and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter destructive tool without annotations, the description is adequate but lacks explicit statements about permanence, effect on all matches, or return value. It implies destructiveness but does not state it clearly, which is a gap given the absence of annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description for 'contains' is minimal ('fragment of row to remove'), but the tool description clarifies that matching is substring-based ('rows that contain the specified text'), which adds semantic meaning beyond the schema. This is useful for the agent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action: 'removes from memory rows that contain the specified text'. The verb and resource are specific, and it is distinct from siblings like memory_add and memory_read. No ambiguity about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies a use case ('a fact that has become false is more harmful than its absence') but does not explicitly state when to use this tool versus alternatives like memory_add or memory_read. There is no mention of conditions where this tool should not be used.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_readПрочитати пам'ятьA

Факти проєкту, що пережили сесію. Ті самі факти автоматично йдуть у system кожного запуску, тож окремо вкладати їх у промпт не треба.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden. It does reveal that the facts persist across sessions and are automatically included in the system prompt each run万里. It does not, however, explicitly state that the tool returns these stored facts or that it has no side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single compact sentence with no filler. Each clause contributes useful information: what the facts are, that they persist, and that they are already available in the system prompt.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only tool with no output schema, the description provides sufficient context: the nature of the memory, its persistence, and its automatic availability in the system prompt. A minor gap is that it does not explicitly describe the return value as the set of stored facts.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and an empty input schema, so there is nothing for the description to explain about parameters. The baseline of 4 applies because no parameter documentation is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The tool name and title provide the verb ('read') and resource ('memory'), and the description defines the content as 'project facts that survived the session'. It is clear enough, though the description itself does not explicitly state the retrieval action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives useful operational context: the facts are automatically injected into the system prompt on every run, so the agent should not redundantly add them to prompts. However, it does not specify when to call memory_read versus memory_add, memory_forget, or other siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

plan_appendДописати до плануA

Додає розділ у кінець PLAN.md. Наявний текст не переписує.

ParametersJSON Schema
NameRequiredDescriptionDefault
markdownYesГотовий фрагмент Markdown

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It does disclose that the operation appends and does not overwrite, which is a useful trait. However, it does not mention whether PLAN.md must already exist, whether the tool creates it if missing, or any error conditions, leaving the agent with an incomplete picture for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with zero waste, stating the primary action first and the key behavioral guarantee second. It is well-structured and easy to parse, meeting the standard for conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple append tool with one parameter and no output schema, it covers the essential operation, but it lacks context such as whether the plan file must exist prior to calling, what happens on failure, or the effect of repeated appends. Given that annotations are absent, the description could be more complete about the tool's side effects and expectations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already fully describes the parameter ('Готовий фрагмент Markdown' – ready-made Markdown fragment) with 100% coverage. The description adds no extra meaning about the parameter beyond what the schema provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Adds a section to the end of PLAN.md') and the target resource, and explicitly notes it does not rewrite existing text, which distinguishes it from any overwrite operation. This is specific and unambiguous, and the sibling plan_read clearly contrasts in function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives. The description only states what it does, without mentioning context such as 'use plan_read to view the plan before appending' or any prerequisite conditions. The agent must infer usage from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

plan_readПрочитати планB

PLAN.md — стан проєкту і наступні кроки.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full responsibility for behavioral disclosure. It does not explicitly state that the tool is read-only, has no side effects, or what it returns. The read behavior is only implied by the tool name, not the description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise phrase that front-loads the key resource (PLAN.md) and its purpose. It is appropriately sized, though it could be slightly more explicit by including the verb 'read' without adding much length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with no output schema and no annotations, the description should explicitly state that the tool reads and returns the contents of PLAN.md. It instead only describes the file's content, leaving the agent to infer the tool's behavior from its name and title. This is a minor but real gap in completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline is 4. The description adds no parameter semantics because there are none to document, but it provides context about what the tool operates on (PLAN.md).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the resource (PLAN.md) and what it contains (project state and next steps), which strongly implies the read action. It distinguishes from the sibling plan_append by contrasting with the plan file itself, though the verb 'read' is not explicitly stated in the description and relies on the name/title.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives like plan_append or other project tools. The description simply describes the file without stating when it should be invoked or under what conditions it is preferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

skill_readПрочитати скілB

Повний текст скіла за назвою.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It only states that the full text is returned and does not disclose error behavior, output format, or whether the operation is read-only. Some value exists in specifying the return content, but it is minimal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence with no filler. It is front-loaded with the key fact and every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read operation this is nearly sufficient, but the lack of annotations and output schema means the agent gets no information about return format or failure modes. It is simple enough to call correctly, but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has one string parameter 'name' with 0% description coverage. The description adds that the name identifies the skill, but it does not specify exact naming format or how to obtain valid names. For a single simple parameter this is adequate but not rich.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool returns the full text of a skill by name. It identifies the resource (skill) and the action (read/get full text), but it does not explicitly differentiate it from sibling tools like skills_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance or alternatives are provided. The description implies that a skill name is needed, but it does not mention skills_list for discovering skill names or clarify when to prefer this tool over similar read tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

skills_listСписок скілівA

Скіли, доступні на цій машині: свої (Claude) і Grok-івські. Назву звідси можна передати у glm_run параметром skill.

ParametersJSON Schema
NameRequiredDescriptionDefault
ownerNoФільтр: claude, grok, extra

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses that the tool returns both Claude and Grok skills and implies the output contains names usable in glm_run, which is useful behavioral context. However, it does not explicitly state that the operation is read-only, what the exact return structure is, or whether any side effects occur, leaving some gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with no fluff. The first sentence states the core purpose, and the second adds a useful operational hint about using output names in glm_run. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with one optional parameter and no output schema, the description gives enough context: it states what is listed and how the result can be used. However, it omits explicit mention of the 'extra' filter option and does not clearly describe the returned format beyond implying names, leaving slight ambiguity for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents the 'owner' parameter. The description adds semantic context by mentioning Claude and Grok categories, which align with filter values, but it does not explain the 'extra' value or the filtering mechanism beyond what the schema states. It adds marginal meaning, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the resource (available skills on this machine) and distinguishes between Claude's and Grok's skills. It also hints at the tool's output being usable in glm_run, which separates it from execution or reading tools. However, it lacks an explicit action verb like 'list' or 'returns', though the title 'Список скілів' compensates.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a downstream usage hint: names from here can be passed to glm_run as the skill parameter. This implies the tool is for discovering skill names before running a skill, but it does not explicitly state when to prefer this over siblings like skill_read or how the owner filter relates to alternatives. No exclusions or contrasting with other tools are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 20 tool updatesv0.1.0
    • First observeddebug_tail
    • First observedgit_commit
    • First observedgit_status
    • First observedglm_apply
    • First observedglm_build_prompt
    • First observedglm_cancel
    • First observedglm_list_runs
    • First observedglm_models
    • First observedglm_result
    • First observedglm_run
    • First observedglm_status
    • First observedglm_wait
    • First observedmcp_status
    • First observedmemory_add
    • First observedmemory_forget
    • First observedmemory_read
    • First observedplan_append
    • First observedplan_read
    • First observedskill_read
    • First observedskills_list

TDQS

B3.3/5.0

Scored across 20 tools

Disambiguation4/5

Most tools are clearly separated by resource and action (plan, memory, git, skills, GLM runs). The only mild ambiguity is among glm_status, glm_result, and glm_wait, but their descriptions make the distinction reasonably explicit.

Naming Consistency4/5

The snake_case verb_noun pattern is dominant and recognizable across all tool groups. Minor deviations like glm_models (noun-only) and the skills_list vs skill_read singular/plural mismatch prevent a perfect score.

Tool Count3/5

Twenty tools is on the heavy side for a single server, even though the grouping into GLM lifecycle, memory, plan, git, and skills makes the set navigable. It falls into the 16–25 range where the tool count starts to feel like a burden rather than a curated surface.

Completeness4/5

Core workflows are well covered: GLM runs have create/status/result/cancel/wait/list/apply, memory has full lifecycle, skills are readable, and plan/git basics exist. Minor gaps include no plan editing/removal and no git diff/log, but these are workable or intentionally left to the user.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    A Model Context Protocol server that bridges MCP clients with local LLM services, enabling seamless integration with MCP-compatible applications through standard tools like chat completion, model listing, and health checks.
    -
  • A
    license
    A
    quality
    A
    maintenance
    Unified MCP server for managing local model runtimes (Ollama, LM Studio, etc.), enabling provider-agnostic discovery, lifecycle management, hardware-fit checks, and delegated inference.
    16
    46 npm
    Creative Commons Attribution Non Commercial No Derivatives 4.0 International