Agile Today
Server Details
Project management methodology graph: route planning, stage-gate decisions and plan checks.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 8 tools
Most tools have clearly distinct roles: retrieval (list_context, search_graph, explain_node, when_to_use), generation (plan_project), and evaluation (check_gate, validate_plan, check_mix). However check_mix and validate_plan overlap—both detect antipatterns and redundant alternatives on a set of ids—which could cause misselection between them.
Almost all tools follow a consistent snake_case verb_noun-ish pattern (check_gate, list_context, plan_project, search_graph, validate_plan, explain_node, check_mix). The lone deviation is when_to_use, which is not verb_noun and breaks the otherwise predictable pattern.
Eight tools is well-scoped for a methodology/knowledge-graph advisory server. Each tool earns its place: retrieval, search, node explanation, fit assessment, plan generation, and multiple validation checks without redundancy. No filler or missing obvious helpers.
The surface covers onboarding (list_context), discovery (search_graph, explain_node, when_to_use), planning (plan_project), and validation (validate_plan, check_gate, check_mix), forming a coherent lifecycle. Minor gap: no tool to persist, compare, or retrieve prior project/plan state, so the workflow is read/evaluate-only.
Available Tools
8 toolscheck_gateCheck a gateARead-onlyIdempotentInspect
Evaluates gate g.g0 … g.g10 against your evidence. Outcome precedence: Stop > Recycle > Hold > Conditional go > Go, with the rule that fired (outcome_reason). Critical evidence missing or stale → Hold; required non-critical missing, or a waiver without a reason or an approver who decides at this gate → Conditional go; a waiver with both → Go with the Waiver modifier (under regulation a waived critical artifact stays Conditional go until the deviation is recorded with the quality approval); optional gaps are notes. Modifiers: Work started without the gate, Overdue, Stale, Waiver. gate.deciders decide, gate.advisors take part. A disproved key product hypothesis is a Recycle, not a Stop: pass earlier_stage_invalid:true with invalid_stage:'st.discovery'. At G8 (go / no-go), work_started_beyond:true means production was changed before the decision: Hold (R3), not a Go with a modifier. Call without evidence to list what the gate needs. Returns a receipt hash of the inputs.
| Name | Required | Description | Default |
|---|---|---|---|
| gate | Yes | Gate id, e.g. g.g5 | |
| lang | No | Answer language, default en | |
| context | No | Project context. Unknown keys and values are rejected in the answer envelope with did_you_mean. Defaults: software, hybrid, complicated, tm, single, medium size, ui, data. | |
| evidence | No | Map artifact id → state or {state, reason, ref, by, date}. A waiver is {state:'waived', reason, by}: by is the role id of a decider of this gate (see gate.deciders, e.g. r.client_owner) or 'name, role'. Unknown ids are reported as warnings with a suggestion. | |
| grace_days | No | Overrides both limits. Default: 80% of planned_stage_days, else by approach (predictive/hybrid 10, agile/flow 5); required evidence then waits twice as long | |
| days_in_stage | No | ||
| invalid_stage | No | With earlier_stage_invalid: the stage to return to, if known | |
| invalid_artifact | No | With earlier_stage_invalid: the artifact that is no longer true, to find where to return | |
| planned_stage_days | No | Planned length of the stage in days. With it, missing critical evidence is overdue after 80% of it and missing required evidence after all of it | |
| business_case_valid | No | false → Stop | |
| work_started_beyond | No | True if work of later stages already started (adds the 'work started without the gate' modifier; at G8 it is a Hold) | |
| earlier_stage_invalid | No | true → Recycle; the answer names the stage to return to |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare the safety profile (readOnly, idempotent, non-destructive); the description adds substantial behavioral context the annotations cannot: outcome precedence (Stop > Recycle > Hold > Conditional go > Go), the rule that fired, waiver handling, Hold-vs-Conditional-go triggers, modifier semantics, and the receipt-hash return. This is exactly the beyond-annotations disclosure the dimension rewards.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose followed by the outcome precedence, then edge cases. It is dense and long but nearly every sentence carries decision-relevant information (precedence, waiver rules, G8 exception). Slightly prose-heavy for a single paragraph but no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter tool with nested context, no output schema, and high branching logic, the description covers the decision model, modifiers, and even mentions the return value ('receipt hash of the inputs'). Complete enough to call correctly; only the exact receipt contents and warning envelope details are left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 92% schema coverage the baseline is 3, but the description adds real meaning for several parameters: business_case_valid=false → Stop, earlier_stage_invalid=true → Recycle with a named return stage, and work_started_beyond at G8 meaning Hold rather than a modifier. These are semantic rules not captured by the schema, lifting it above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource+scope: 'Evaluates gate g.g0 … g.g10 against your evidence.' An agent knows exactly what the tool consumes and produces. It does not, however, explicitly differentiate itself from siblings like validate_plan or check_mix, so it lands at 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides concrete usage guidance: 'Call without evidence to list what the gate needs,' and tells when to pass earlier_stage_invalid, work_started_beyond, and invalid_stage. It never states when to prefer a sibling tool (check_mix, validate_plan) over this one, so it lacks explicit alternatives but the in-tool context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_mixCheck a mix of practicesCRead-onlyIdempotentInspect
For a set of ids: antipatterns in the set and detected by rules, conflicts, redundant alternatives (e.g. seven estimation methods), items to avoid or not applicable here, risks, missing inputs and suggested additions. ok = no antipatterns, conflicts, avoid or not-applicable items; a detected pattern of severity warning (e.g. output over outcome on a client contract) does not change ok.
| Name | Required | Description | Default |
|---|---|---|---|
| ids | Yes | ||
| lang | No | Answer language, default en | |
| context | No | Project context. Unknown keys and values are rejected in the answer envelope with did_you_mean. Defaults: software, hybrid, complicated, tm, single, medium size, ui, data. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so the safety profile is covered. The description adds genuine behavioral context by defining the 'ok' semantics and clarifying that a severity=warning pattern (e.g. output over outcome) does not flip ok. It omits anything about response shape beyond that flag or about how context alters detection.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
All content is in one dense run-on sentence that front-loads the output categories before revealing the purpose, mixing result fields with the ok-flag rule. It is not padded, but the structure makes it hard to scan and buries the tool's actual function behind an output enumeration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description reasonably enumerates the return categories and explains the ok flag, which partially fills the gap. However, it fails to explain how the context object affects results or what form the findings take, leaving the response semantics only half-covered for a fairly complex nested-input tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only the required 'ids' parameter is even implied ('For a set of ids'), and nothing is said about the rich 'context' object or the 'lang' enum. At 67% schema description coverage the description needs to carry some weight for the nested context fields, yet it contributes none — 'not applicable here' gestures at context without explaining it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete operation on a specific resource: given a set of practice ids, it reports antipatterns, conflicts, redundant alternatives, risks and missing inputs. The detection categories make the tool's output intent legible. It does not, however, distinguish itself from siblings like check_gate or validate_plan, which an agent must infer from name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance, no prerequisites, and no mention of the alternative tools (check_gate, validate_plan, when_to_use) that could be selected instead. The agent is left to infer that this is the multi-practice consistency check versus a single-gate check, which is a plausible but unstated distinction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
explain_nodeExplain a nodeBRead-onlyIdempotentInspect
One node with its description, conditions, source, level and its links grouped by type, each with a link id (Lxxxx) to cite. Filter with link_types, direction and limit.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| lang | No | Answer language, default en | |
| limit | No | ||
| direction | No | ||
| link_types | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds genuine value by disclosing the return shape (node fields plus links grouped by type with citable ids), but says nothing about truncation, the limit default, or behavior when the node is missing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, output shape front-loaded before the filtering sentence. Dense but no filler; the only weakness is that the second sentence reads as a parameter list rather than guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of explaining return values and does describe node fields and link grouping adequately. Still, for a 5-parameter tool at 20% schema coverage, gaps remain around id semantics, limit defaults, and missing-node behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 20% (just lang), so the description must compensate and partially does: it names link_types, direction and limit and their filtering role. However, the id parameter is unexplained and no value semantics (e.g., default limit, valid link type strings) are added beyond the enums already in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the specific resource (one node) and enumerates the returned fields — description, conditions, source, level — plus grouped links with citable Lxxxx ids. This makes it clearly distinguishable from search_graph or list_context, though it never names a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only usage guidance is 'Filter with link_types, direction and limit,' which is parameter mechanics rather than when to use this tool. There is no statement of prerequisites (e.g., needing a node id from search_graph) or when a different sibling should be chosen instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_contextList context optionsARead-onlyIdempotentInspect
Call first. Returns the context dimensions and flags (with defaults and how many rules each flag drives), stages, gates, evidence states and levels, the gate outcome precedence and the suggested workflow.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Answer language, default en |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish the safety profile (readOnly, idempotent, non-destructive, closed-world), so the description's value-add is the disclosure that this is the mandatory first call and an enumeration of what the response contains. With no output schema, spelling out the returned vocabulary is meaningful behavioral context. It omits any note on size, caching, or whether the returned flags are defaults vs. configured values.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, with the imperative 'Call first' front-loaded before the payload inventory. The dense enumeration is efficient rather than padded, and nothing could be cut without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-required-param discovery tool with no output schema, the description adequately covers purpose, ordering, and return contents while annotations cover safety. Minor gap: it never states that the result is static reference data the agent can reuse rather than re-fetch, which would help an agent avoid redundant calls.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'lang' enum parameter is fully documented in the schema (including its 'default en' note), so the schema already carries the burden. The description says nothing about the lang parameter or its effect on the returned content, leaving it at the baseline for full schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the exact resource set returned: context dimensions, flags (with defaults and rule counts), stages, gates, evidence states and levels, gate outcome precedence, and the suggested workflow. That is far more specific than a generic 'get context' and lets an agent predict the payload. It does not, however, contrast itself against any sibling like when_to_use or explain_node.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Call first' gives an explicit ordering instruction that positions this as the bootstrap/discovery step before other tools, which is genuine when-to-use guidance. It stops short of naming alternatives or stating when this call is unnecessary, so it is clear context without exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
plan_projectPlan a project routeARead-onlyIdempotentInspect
Route from presale to closure for a project context. Default is compact (ids by strength + a names dictionary, core items and gate evidence for all stages, about 6–10k characters). Ask for one stage with stages:['st.design'] and format:'full' for reasons, avoid texts and gate details. 'core' = do it here (context rule, flag or gate-critical artifact), at most 8 work items per stage (producers of the gate's critical artifacts first, then context rules, then flags).
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Answer language, default en | |
| format | No | Default compact | |
| stages | No | Stage ids or codes, e.g. ['st.analysis','design']. Default: all | |
| context | No | Project context. Unknown keys and values are rejected in the answer envelope with did_you_mean. Defaults: software, hybrid, complicated, tm, single, medium size, ui, data. | |
| include | No | Default all four | |
| min_strength | No | Default: core for all stages, all when stages are given |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the bar is lower. The description adds real behavioral context beyond them: expected response size (~6–10k characters), what compact returns versus full, and the composition rule for 'core'. It omits any mention of cost, latency, or failure modes, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded and only three sentences, but the final sentence is a dense run-on packed with nested parentheticals and internal shorthand ('core' = ...), which is hard to parse on first read. It is information-dense rather than wasteful, but structure suffers.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and a large nested context object, the description usefully compensates by describing the default response shape and size, and the schema itself documents all context keys at 100% coverage. It stops short of describing the conceptual shape of the returned route (stages/gates/work items) or validation failure behavior, which the schema only partially implies.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description genuinely extends it by explaining the 'core' strength tier (context rules, flags, gate-critical artifacts; max 8 items per stage; producer-first ordering) and the stage+format interaction — semantics not present in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence establishes scope ('route from presale to closure for a project context') and the rest details what the route contains (ids by strength, names dictionary, core items, gate evidence per stage), so an agent can tell it produces a routed project plan. It never differentiates itself from siblings like validate_plan or when_to_use, and 'route' as a verb is jargon-heavy, keeping it below 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives invocation guidance for modes (default compact vs format:'full' with a single stage for reasons/avoid texts, and what 'core' means), which implies how to call it. However there is no when-to-use vs alternatives guidance — nothing tells the agent why to call plan_project rather than validate_plan, check_gate or when_to_use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_graphSearch the graphBRead-onlyIdempotentInspect
Search names, ids, descriptions and aliases in English and Russian (word stems, small typos tolerated). Returns total and the best matches.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Answer language, default en | |
| type | No | ||
| limit | No | ||
| query | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, not openWorld, non-destructive). The description adds a useful trait beyond that – fuzzy matching (word stems, typo tolerance) and multilingual coverage. It doesn't describe result ranking beyond 'best matches' or pagination/limit behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with the search subject and capabilities, then the return shape. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers what is searched and the matching behavior, but with no output schema and only 25% param coverage, it leaves the return structure ('total and the best matches') vague and doesn't explain how type, limit, or lang refine the search.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 25% (only lang is described). The description says queries may be in English or Russian, which hints at the lang parameter but doesn't explain its default or semantics. The type and limit parameters receive no description coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search) and resource (the graph), and names the searchable fields (names, ids, descriptions, aliases). The sibling names (check_gate, explain_node, etc.) suggest a graph-oriented toolset, but the description doesn't explicitly differentiate this lookup tool from them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use or when-not-to-use guidance. It doesn't say whether search_graph is the right tool to resolve an entity before calling explain_node, or how it relates to list_context. No alternatives named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_planValidate a planARead-onlyIdempotentInspect
Checks an ordered list of steps (stages, gates, events, practices, techniques, notations, artifacts). Errors: empty plan, unknown ids, non-step types, stage order (analysis and design work planned early is only a warning 'early'; building before G3–G5 is an order error), contract after delivery gates, a gate after the step it authorises (G8 after the cutover), antipatterns, avoid or not-applicable items. Warnings: duplicates, redundant alternatives, early work, a cutover with no G8 or go / no-go meeting before it, antipatterns of severity warning. 'Shadow start' looks at the order: G4 or G5 must come before the first build step. 'expected' lists gate evidence not yet produced by earlier steps.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Answer language, default en | |
| steps | Yes | ||
| context | No | Project context. Unknown keys and values are rejected in the answer envelope with did_you_mean. Defaults: software, hybrid, complicated, tm, single, medium size, ui, data. | |
| evidence | No | Artifacts that already exist |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive behavior, so the safety profile is covered. The description goes well beyond that by disclosing the exact error/warning taxonomy, the severity distinction (analysis-before-design is a warning 'early' while building before G3-G5 is an order error), and the semantics of 'shadow start' and 'expected' — genuinely useful operational context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded and every clause carries information, but the body is a dense run-on enumerating errors and warnings with no separation or formatting, which hurts scannability. It is information-rich rather than padded, but structure is weak.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must describe the result — and it does, effectively documenting the error and warning categories the validator returns. Given the complex nested context object, it is largely complete, though it says nothing about the context/evidence inputs that shape validation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75% and the nested context object documents most fields, so the schema does the heavy lifting. The description never mentions the context, evidence, or lang parameters and adds no syntax or format detail for steps beyond naming the element types, so it does not compensate for the remaining gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (checks/validates) against a specific resource (an ordered list of steps), and enumerates the element types it understands (stages, gates, events, practices, etc.). It is clear what the tool does, though it never explicitly contrasts itself with the similarly named siblings check_gate, check_mix, or plan_project.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the enumeration of error and warning categories, which tells the agent the tool is for pre-flight validation of a planned sequence. However there is no explicit 'use this when / not when' statement and no routing guidance relative to check_gate or check_mix, so the agent must infer the boundary.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
when_to_useWhen to use an itemARead-onlyIdempotentInspect
For one event, practice, technique, notation, tool, artifact or antipattern: does it fit this context (fits / avoid / not_applicable / antipattern), its strength and why, alternatives and their group advice, risks, and which gates require it.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | e.g. t.planning_poker, p.stage_gating, x.velocity_kpi | |
| lang | No | Answer language, default en | |
| context | No | Project context. Unknown keys and values are rejected in the answer envelope with did_you_mean. Defaults: software, hybrid, complicated, tm, single, medium size, ui, data. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds useful behavioral return context: the verdict categories, rationale, alternatives, risks, and gate requirements. It does not add operational details like rate limits, but the annotations carry that burden.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single dense sentence that front-loads the input scope ('For one ... item') and then enumerates the returned information. There is no filler or redundancy, though the long comma-separated clause list makes it slightly less scannable than ideal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema, the description usefully outlines the response content: fit verdict, strength and why, alternatives and their advice, risks, and required gates. The nested context schema is richly documented elsewhere. The main missing piece is sibling routing guidance, which keeps it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with id, lang, and context parameters fully documented, including enum values and nested context field descriptions. The description adds only the conceptual idea of evaluating an item against 'this context,' not any parameter syntax or meaning beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific evaluation task: for one named item (event, practice, technique, etc.), decide whether it fits the given context and return fit, rationale, alternatives, risks, and required gates. This is clear enough to distinguish from siblings like search_graph or validate_plan, though it does not explicitly name or compare against any sibling tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It identifies the input scope ('For one ... item') but gives no explicit when-to-use, when-not-to-use, or alternative-tool guidance. With siblings such as check_gate, check_mix, and explain_node, an agent must infer when this tool is preferable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
- First observed
check_gate - First observed
check_mix - First observed
explain_node - First observed
list_context - First observed
plan_project - First observed
search_graph - First observed
validate_plan - First observed
when_to_use
Related MCP Connectors
Draft-safe project planning: tasks, dependencies, resources, schedules and cost metrics.
Project controls for construction schedules: approved changes, look-aheads, briefs, signed records.
Agent Pay and planning graph for stablecoin commerce.
Transform project ideas into paint-by-numbers development plans with phases, tasks, and subtasks.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables AI orchestrators to manage hierarchical implementation roadmaps with phases, tasks, and plan-change tracking.Apache 2.0
- FlicenseAqualityNot gradedmaintenanceProvides structured project management through phase-based workflows, enabling planning, execution tracking, compliance checking, and automated documentation for software development projects.12-
- FlicenseNot gradedqualityCmaintenanceEnforces work-readiness gates by tracking project plots, plans, stakeholder views, slots, stops, and journal entries via REST and MCP, with LLM-based plan agreement and result verification.-
- AlicenseNot gradedqualityBmaintenanceHelps AI agents decompose goals into costed execution plans, checkpoint progress with deviation tracking, and learn from past project data.31 npmMIT
Glama MCP Gateway
Add one secure layer between your agents and this server.