io.github.koten-ai/zeus-dev-helper
OfficialThis stdio MCP server coaches a coding agent and human through Zeus first-app onboarding and smoke tests, without running data-plane Explore/Verify verbs or Hub admin mutations.
Run health and environment checks with
doctor(health,env,compat,cache,all)Start or reset a first-app checklist and get the current blocker/next recommended tool (
start_project,next_step)Store non-secret prereqs like Zeus URL, bucket, scope, and auth presence flags (
set_prereq)Check live Zeus readiness: healthz/readyz/version, auth, bootstrap, chat_request (
readiness_check)Scaffold a Python ZeusRuntime app: CLI or FastAPI
POST /turn(scaffold_app)Use UI samples: travel clone, beer catalog UI, or yelp demo clone (
use_sample)Copy a stamped
contract.hashonly; refuse empty or locally computed hashes (bind_contract)Recommend Client surface/verb and a do-not list from intent (
recommend_surface)Run a no-LLM smoke test: readiness plus read-only describe (
smoke_test_zeus)Run one Client
run_turnagent smoke test with an LLM key (smoke_test_agent)Diagnose HTTP/body/error codes with failure class and Detective URL templates (
diagnose_error)Read
zeus-helper://resources for checklist, glossary, verbs, policies, and catalog modesUse prompts
first_green,smoke_question, andsupport_packEnable opt-in toolsets:
catalog,lint,travel,support,handoff, orallIt enforces boundaries: public Zeus API on port 8080 only, no invented contract hashes, no secrets in results, semantic cache off
Clones public sample and template repositories (demo_travel_sample, demo_yelp, and the zeus_chat_request catalog templates) from GitHub when they are not already present locally, letting the coach path pull ready-made UI samples and catalog templates for the first-app walkthrough.
Provides a Yelp demo UI path: start_project(sample=demo_yelp) plus use_sample(sample=demo_yelp) clone the demo_yelp app when missing, point at the yelp-demo bucket / _default scope, and set DEMO_YELP_SAMPLE_DIR so a Yelp-style demo UI can be stood up and smoke-tested. Bare sample=yelp is reserved as the multi-agent handoff.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@io.github.koten-ai/zeus-dev-helperWalk me through a first Zeus Client app turn, from onboarding to smoke tests."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Zeus Dev Helper MCP
Stdio MCP server that coaches a coding agent and a human to a first successful Zeus Client app turn.
This is not a data-plane MCP. It does not run Explore/Verify verbs on your behalf, invent contract hashes, or perform Hub admin mutations. After the first-green smokes pass, data-plane and multi-agent work are handoffs only.
Package |
|
Registry name |
|
Transport | stdio |
Python | 3.11+ |
MCP SDK |
|
What it does
The server walks a first-app checklist: prereqs, live readiness on the public Zeus API, catalog templates, contract bind (copy a stamped hash only), surface/verb coaching, config lint, and smoke tests. Prefer a live Zeus stamp for catalogs. fetch_chat_request is always template only.
Hard constraints the tools enforce:
Public Zeus API on port 8080 only (never Hub 9091 from the app path)
Never invent
contract_hashNo secrets in tool results, checklist evidence, or support packs
Semantic cache stays off
Related MCP server: Planwright
Install
pip install zeus-dev-helper-mcp
# or
uvx zeus-dev-helper-mcpOptional extra for smoke_test_agent (pulls the Zeus Client package):
pip install "zeus-dev-helper-mcp[agent]"Run
zeus-dev-helper-mcp
# or
python -m zeus_dev_helper_mcpPrefer the published console script (uvx / pip install) so hosts do not need a source checkout.
Host install
Set ZEUS_URL to your Zeus public API on port 8080 (never Hub :9091). Prefer the remote host you actually use. Use http://localhost:8080 only when Zeus runs on the same machine as the MCP host.
# Remote Zeus (typical lab / shared engine) — put your host here
export ZEUS_URL=http://192.168.0.219:8080
# or: http://<zeus-host>:8080
#
# Same-machine Zeus only:
# export ZEUS_URL=http://localhost:8080If the user names a URL or sample in chat, call set_prereq with that zeus_url / bucket / scope (do not keep a stale localhost default). doctor reports stored vs effective URL routing (see ZDM-3).
Grok Build
grok mcp add treats flags like -m as its own unless they come after --. The uvx argument is the PyPI package zeus-dev-helper-mcp, not the MCP server id zeus-dev-helper.
grok mcp add zeus-dev-helper \
-e ZEUS_URL=http://192.168.0.219:8080 \
-- uvx zeus-dev-helper-mcpFrom a local checkout after pip install -e ".[dev]", point command at this tree’s venv so the host can start the server even when it was launched without the venv activated:
grok mcp add zeus-dev-helper \
-e ZEUS_URL=http://192.168.0.219:8080 \
-- "$(pwd)/.venv/bin/python" -m zeus_dev_helper_mcpEquivalent config:
[mcp_servers.zeus-dev-helper]
command = "uvx"
args = ["zeus-dev-helper-mcp"]
env = { ZEUS_URL = "http://192.168.0.219:8080" }
enabled = trueThen refresh MCP servers, or grok mcp doctor zeus-dev-helper.
Common failures:
unexpected argument '-m'— missing--before the python commanduvx zeus-dev-helper/No solution found— wrong package name; usezeus-dev-helper-mcpNo module named 'zeus_dev_helper_mcp'/python: No such file or directory— the host did not inherit the venv; use the.venv/bin/pythonpath aboveNo module named 'mcp.server.fastmcp'— mcp 2.x renamed FastMCP; use Helper 0.6.0+ (mcp>=1.8.0,<3)
Claude Code / Claude Desktop
{
"mcpServers": {
"zeus-dev-helper": {
"command": "uvx",
"args": ["zeus-dev-helper-mcp"],
"env": {
"ZEUS_URL": "http://192.168.0.219:8080"
}
}
}
}Cursor
Add (or merge) .cursor/mcp.json in the project (or use Cursor’s global MCP settings):
{
"mcpServers": {
"zeus-dev-helper": {
"command": "uvx",
"args": ["zeus-dev-helper-mcp"],
"env": {
"ZEUS_URL": "http://192.168.0.219:8080"
}
}
}
}Other stdio hosts
Point the host’s MCP stdio entry at uvx zeus-dev-helper-mcp (or python -m zeus_dev_helper_mcp from a venv) with the same env vars. Replace the sample IP with your Zeus host.
Day-one coach path
doctor → (if user named URL/sample) set_prereq → start_project → next_step
→ readiness_check
→ use_sample | scaffold_app → bind_contract → recommend_surface
→ smoke_test_zeus → smoke_test_agent → diagnose_errorPrefer next_step over dumping the full checklist. Two first-green paths (TravelPlan is not the only path):
Travel + LLM (UI default):
start_project(sample=travel)→use_sample, which clones publicdemo_travel_samplewhen missing (optionalproject_namefor the directory) and setsDEMO_TRAVEL_SAMPLE_DIR. Checklist 1.2 stays open until this process hasLLM_API_KEY,XAI_API_KEY, orOPENAI_API_KEY.smoke_test_agentreads that variable. Standalone clones pinkotenai-zeus-client>=2.4.0,<2.5sorun_turncan recover a fenced pipeline inside the SDK.Beer catalog UI:
start_project(sample=beer)→use_sample(sample=beer)writesdemo_beer_sample(FastAPI BFF + static page). Search matches travel:rt.agent.run_turnwithchat_requestomitted socatalog.load_for_turnmerges the live SCOPE BRIEF and MINI-SCHEMA. The key is required even whenhas_llm_key=false. Put it in the app.env. Leaveconfig.jsonllm.api_key_envas that variable name.verify_local_setupkeeps checklist 3.2 open until that file check passes. The BFF does not build a pipeline body. Do not clonedemo_travel_sample.Yelp demo UI:
start_project(sample=demo_yelp)→use_sample(sample=demo_yelp)clonesdemo_yelpwhen missing (utterancesyelp-demo/demo_yelp) and setsDEMO_YELP_SAMPLE_DIR.set_prereqbucket isyelp-demo, scope_default. Do not clonedemo_travel_sample. Baresample=yelpstays the multi-agent handoff.API-only: user asks for an API/REST app →
start_project(sample=api)→scaffold_app(app_kind=api, coding_language=python)(FastAPIPOST /turnonkotenai-zeus-client). Other languages not scaffolded yet.Credentials from chat → process env / gitignored
.env;set_prereqpresence flags only.has_llm_key=truedoes not count as the key. Pass the user’s Zeus URL intoset_prereq(zeus_url=…). Do not paste the secret intollm.api_key_env.Integrating into an arbitrary existing repo is out of scope.
Read zeus-helper:// resources for glossary, verbs, policies, and catalog modes. Hosts can pick prompts first_green, smoke_question, and support_pack.
Default tools (core)
Live tools/list is the call contract. Default surface is 12 tools (ZEUS_DEV_HELPER_TOOLSETS=core).
Tool | Job |
| Health. |
| Init checklist; |
| Current item plus recommended tools and resource links |
| Store non-secret prereqs (presence flags only; |
| Live gates: healthz / readyz / version, auth, bootstrap. Marks 1.2 done only when the URL is set and, on a |
| CLI or FastAPI ( |
| Travel UI clone, beer catalog UI ( |
| Copy a stamped |
| Intent → Client surface + do-not list |
| No LLM: readiness plus a read-only describe |
| One Client |
| Map HTTP / body / error codes to a failure class |
Opt-in toolsets (static, comma-separated): catalog, lint, travel, support, handoff. all enables every set. Full when/args/side-effects map: docs/TOOLS.md.
Resources (always on): zeus-helper://checklist, zeus-helper://glossary/{topic}, zeus-helper://verbs/{name}, zeus-helper://policy/hash-boundary, zeus-helper://policy/req-id, zeus-helper://catalog/modes.
Environment
Secrets stay in the process environment. set_prereq stores presence flags only. Tool results redact secret values.
Variable | Purpose |
| Public Zeus API base URL (port 8080). Remote host first; |
| Scope for bootstrap and auth probes |
| Default catalog mode ( |
| Auth mode ( |
| Basic auth (never logged) |
| Bearer auth (never logged) |
| Process key for checklist 1.2, |
| Local directory of min catalog templates; auto-set when |
| Local sample directory for |
| Local |
| Checklist, prereqs, and local metrics (default |
| Static toolsets: |
Boundaries
This MCP | Not this MCP |
Onboarding coach to first green | Data-plane Explore/Verify tools |
Catalog templates plus readiness and smoke | Inventing or locally computing |
Verb explain / lint / draft ( | POSTing |
Detective URL templates | Hub scrape or Hub admin mutations |
Multi / data-plane handoffs | Multi-agent job runtime |
Local checklist and metrics | Shipping secrets in evidence or support packs |
Dev install
From a local checkout:
python3 -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
# optional agent smoke:
pip install -e ".[agent]"
export ZEUS_URL=http://localhost:8080pytest -qMCP Registry
Official registry name: io.github.koten-ai/zeus-dev-helper. The registry hosts metadata only; the install artifact is the PyPI package zeus-dev-helper-mcp.
License
BSD-3-Clause — see LICENSE.
Available Tools
12 toolsbind_contractARead-onlyIdempotent
Extract stamped contract.hash only. Refuse placeholders / compute_local (ZDH-22).
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | ||
| path | No | ||
| scope | No | ||
| bucket | No | ||
| json_text | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds a meaningful behavioral trait beyond annotations: it refuses certain inputs (placeholders/compute_local) and only handles stamped contracts. This is useful context that annotations do not provide. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with zero wasted words. The primary action is front-loaded, and the refusal condition is stated in a compact second sentence. It is exemplary in brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
While an output schema exists (which may cover return values), the description fails to explain the purpose of any parameter, the meaning of 'stamped contract', or the expected input format. For a tool with five optional parameters and no schema descriptions, this leaves agents under-equipped. The refusal condition is useful, but the absence of parameter semantics makes the tool contextually incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage for its five parameters (mode, path, scope, bucket, json_text), so the description must compensate, but it does not mention any parameter or its meaning. Agents receive no guidance on what values to provide, making correct invocation nearly impossible without external knowledge.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Extract') and resource ('stamped contract.hash'), and specifies a constraint that it only operates on stamped contracts and refuses placeholders/compute_local. This clearly distinguishes the tool's purpose and is unlikely to be confused with any sibling, which are unrelated (e.g., smoke_test_agent, diagnose_error).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit context on when to use the tool (to extract stamped contract.hash) and explicitly lists refusal conditions (placeholders/compute_local), which serves as a negative usage guideline. It does not name an alternative tool for those cases, but the guidance is clear enough for an agent to decide against using this tool for disallowed inputs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
diagnose_errorARead-onlyIdempotent
Map error signals to failure_class + errors.md anchor (ZDH-7 / ZDH-19).
Includes Detective URL templates when req_id/chat_id are present (folded detective_links). Does not scrape Hub.
| Name | Required | Description | Default |
|---|---|---|---|
| body | No | ||
| req_id | No | ||
| status | No | ||
| chat_id | No | ||
| message | No | ||
| turn_id | No | ||
| zeus_url | No | ||
| error_code | No | ||
| session_id | No | ||
| error_class | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the read-only, idempotent annotations, the description adds real behavioral context: it conditionally folds Detective URL templates when req_id/chat_id are present and explicitly states that Hub scraping is not performed. This is meaningful and does not contradict the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three terse, front-loaded sentences cover the core mapping behavior, the conditional link behavior, and a key exclusion. There is no filler, repetition, or unnecessary detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 10 optional parameters, the description explains the central behavior and one important conditional, and an output schema exists to document return values. Still, it does not clarify which parameters should be populated or how the mapping consumes them, so an agent must infer a significant part of invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only names and empty-string defaults, so the description's clarification that req_id/chat_id trigger detective_links is a genuine addition. But the remaining eight parameters are only loosely implied to be 'error signals,' with no specifics about how they are used, leaving a notable coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete action and object ('Map error signals to failure_class + errors.md anchor'), so an agent can tell what the tool accomplishes. The notes about Detective URL templates and not scraping Hub further sharpen scope, but it does not explicitly distinguish itself from sibling diagnostic tools like doctor.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The purpose sentence implies this should be used when error signals need mapping to a failure class and errors.md anchor, and 'Does not scrape Hub' is a useful exclusion. However, there is no explicit 'use when...' or 'instead of...' guidance, and no alternative sibling is named, so routing is largely left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
doctorBRead-onlyIdempotent
Health / doctor. detail: health | env | compat | cache | all.
env/compat/cache fold validate_env, compat_check, and semantic_cache_status (those names stay on the lint toolset).
If the user already gave a Zeus URL or sample name, next call set_prereq with those values (not Helper localhost defaults), then start_project / next_step. Do not grep the Zeus engine tree for first green.
| Name | Required | Description | Default |
|---|---|---|---|
| detail | No | health |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds context that env/compat/cache fold related validation tools and includes an operational warning ('Do not grep the Zeus engine tree for first green'), which is useful, but it does not clarify what the output looks like or what 'first green' means.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loads the detail values early. It avoids fluff, and each section adds some information, though the phrasing 'fold' and the telegraphic 'Health / doctor.' make it less polished than ideal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple optional-parameter tool with read-only and idempotent annotations plus an output schema, the description is mostly sufficient. The main gaps are ambiguity about what 'fold' means and the unexplained reference to 'first green,' as well as no explicit differentiation from the diagnostic sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter detail has no schema description and no enum, but the description supplies the allowed values: health | env | compat | cache | all. It also maps env/compat/cache to validate_env, compat_check, and semantic_cache_status, giving the parameter more meaning than the schema alone. It does not spell out what each value returns, though an output schema exists.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a noun fragment, 'Health / doctor,' which mostly restates the tool name and the domain rather than stating a specific verb and resource. It does enumerate detail options (health | env | compat | cache | all), which give partial purpose, but it does not distinguish doctor from sibling tools like diagnose_error, readiness_check, or smoke_test_agent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a follow-up workflow condition ('If the user already gave a Zeus URL or sample name, next call set_prereq...'), but it does not say when to use doctor versus the diagnostic sibling tools. The line about env/compat/cache folding validate_env, compat_check, and semantic_cache_status hints at coverage but never explicitly states selection criteria or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
next_stepARead-onlyIdempotent
Return the single current checklist blocker + recommended Helper tools (coach).
Call only the recommended tool next. If the user named a URL/bucket and prereqs are unset, recommend set_prereq first. Prefer zeus-helper:// + readiness_check / smoke_test_zeus over hand-rolled curl.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds a behavioral contract beyond the annotations: the instruction to call only the recommended tool next and the conditional recommendation logic. It does not add side-effect details because none exist, but it does add operational behavior that is not in the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short paragraphs with no filler. The first sentence front-loads the purpose, and the second paragraph delivers precise usage rules. Every sentence earns its place, and the structure is easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has zero parameters, an output schema exists, and annotations cover safety, the description is complete. It states what the tool does, when to use it, how to sequence calls, and which sibling alternatives to prefer. There are no missing pieces that an agent would need to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
This tool has zero parameters and the input schema is empty, so there is nothing to explain. The baseline for 0 params is 4, and the description does not need to compensate for any schema gaps. It mentions external context (URL/bucket, prereqs) but that is not parameter input, so no further semantic detail is required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') and resource ('single current checklist blocker + recommended Helper tools'), and it distinguishes itself from siblings by framing itself as a coach that recommends the next step. It clearly tells the agent what the tool produces without needing to open the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: 'Call only the recommended tool next', a conditional rule for set_prereq when URL/bucket is named and prereqs are unset, and a preference for zeus-helper:// + readiness_check / smoke_test_zeus over hand-rolled curl. This is rich, actionable guidance that selects among siblings (set_prereq, readiness_check, smoke_test_zeus).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
readiness_checkBIdempotent
Live Zeus platform gates: healthz/readyz/version, auth, bootstrap, chat_request.
Never returns secret values. Emits failure_class + next_action on red paths.
| Name | Required | Description | Default |
|---|---|---|---|
| update_checklist | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide idempotentHint and destructiveHint, and the description adds useful behavioral context: 'Never returns secret values' and 'Emits failure_class + next_action on red paths.' This goes beyond the annotations. However, it does not disclose any side effects associated with the update_checklist parameter, despite readOnlyHint being false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences: the first names the core scope, and the second adds safety and failure-output behavior. There is no filler, and the main purpose is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema covers return structure, and annotations cover idempotency and destructiveness. Still, the definition omits usage guidance relative to several similar siblings and leaves the update_checklist parameter semantically ambiguous, so the overall context is not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention update_checklist at all. The name and default true imply some update behavior, but the meaning of 'checklist' and the effect of setting it to false remain unexplained. With low schema coverage, the description needed to compensate and did not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names specific resources and scopes: 'healthz/readyz/version, auth, bootstrap, chat_request.' This clearly identifies what the tool checks. However, it does not explicitly differentiate readiness_check from siblings like smoke_test_zeus or diagnose_error, so it stops short of full sibling distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use readiness_check versus the many sibling tools such as smoke_test_zeus, doctor, or diagnose_error. There are no conditions, exclusions, or alternative recommendations, leaving the agent to guess based on the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recommend_surfaceCRead-onlyIdempotent
Pick Direct vs agent surface + Trace-Class (ZDH-18). Does not call Zeus.
| Name | Required | Description | Default |
|---|---|---|---|
| qps | No | ||
| intent | Yes | ||
| needs_llm | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds a meaningful external behavior—'Does not call Zeus'—which is not in the annotations, but it omits other behavioral context like prerequisites or side effects. Given the annotation coverage, the added value is modest.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with no padding, and the core purpose is front-loaded. Each sentence earns its place, though the terseness borders on under-specification. Overall, it is well-structured and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With three parameters, no schema descriptions, and only this terse description, the tool is under-specified for an agent to call it correctly. The presence of an output schema helps with return values but does not compensate for missing parameter semantics. The description reads as an internal note rather than a complete tool contract.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the tool description provides no explanation of the three parameters (qps, intent, needs_llm). An agent cannot infer how to set 'intent' or when to supply 'qps' or 'needs_llm' from either source. This is a critical gap for a tool with a required parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Pick') and resource ('Direct vs agent surface'), and names a trace class, which gives a clear sense of the tool's decision-making role. It also explicitly notes 'Does not call Zeus,' differentiating it from Zeus-related siblings. However, 'Trace-Class (ZDH-18)' is unexplained and may confuse an agent unfamiliar with the domain.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus alternatives. The only hint is 'Does not call Zeus,' which implies it is not for workflows requiring Zeus, but it does not name sibling tools or provide selection criteria. This is insufficient for effective routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scaffold_appADestructive
Write a ZeusRuntime middle-man.
app_kind=cli (default) or api (FastAPI POST /turn). coding_language=python only today. UI demos use use_sample / demo_travel_sample, not this tool. When no Zeus URL is stored, the MCP call asks before writing files. Do not pass a password.
| Name | Required | Description | Default |
|---|---|---|---|
| force | No | ||
| app_kind | No | cli | |
| target_dir | Yes | ||
| project_name | No | zeus_first_app | |
| coding_language | No | python |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate destructiveHint=true and readOnlyHint=false, and the description adds useful behavior beyond that: writing files, asking before writing when no Zeus URL is stored, and a security warning not to pass a password. It does not explain force/overwrite behavior, but it provides meaningful context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short, scannable, and every line adds actionable information: purpose, variants, language restriction, UI-demo exclusion, confirmation behavior, and a security warning. There is no filler or repetition of schema defaults.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 5 parameters, a required target_dir, and no enum constraints, and the description covers the main variants and caveats. However, it leaves gaps: force behavior is not explained, target_dir expectations are not stated, and there is no mention of how the generated scaffolding interacts with the Zeus URL state beyond asking before writing. An output schema exists, so return values are less of a concern, but the missing parameter semantics prevent a higher score.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must add meaning for parameters. It does explain app_kind values and restricts coding_language, which helps. However, it does not describe the required target_dir parameter, project_name, or force. The names are somewhat self-explanatory, but force in particular lacks semantic guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Write a ZeusRuntime middle-man.' It then clarifies the supported variants (cli or api) and explicitly separates itself from UI demo tooling ('UI demos use use_sample / demo_travel_sample, not this tool'), which distinguishes it from a relevant sibling. This gives an agent a clear idea of the tool's job.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives direct usage constraints: app_kind can be cli or api, coding_language is python-only today, and UI demos should use other tools. It also flags that confirmation may be required when no Zeus URL is stored. It does not explicitly contrast against start_project or other siblings, but the guidance is clear enough for typical selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_prereqAIdempotent
Store non-secret prereqs for readiness (does not store password/token values).
Put real secrets in environment variables (ZEUS_PASSWORD, ZEUS_BEARER_TOKEN, LLM_API_KEY). When the user named a Zeus URL or sample in chat, pass that zeus_url / bucket / scope here (not Helper localhost defaults). If the user did not name a URL, leave zeus_url empty and wait for the form. Do not copy the host ZEUS_URL. Persisted values override MCP host ZEUS_* env defaults so readiness/doctor hit the user's cluster (ZDM-3). Presence flags only for credentials and LLM key.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | ||
| role | No | ||
| scope | No | ||
| bucket | No | ||
| zeus_url | No | ||
| auth_mode | No | ||
| collection | No | ||
| has_bearer | No | ||
| has_llm_key | No | ||
| has_password | No | ||
| has_username | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide idempotentHint=true and destructiveHint=false, so the safety profile is already declared. The description adds genuinely useful behavior beyond that: values are persisted ('Persisted values override MCP host ZEUS_* env defaults'), the tool explicitly excludes secrets, and presence flags only apply to credentials/LLM key. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, followed by dense, purposeful guidance on secrets, URL handling, and persistence. Every sentence earns its place and there is no fluff, though six sentences make it slightly long; it could be tightened without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter tool with zero schema description coverage, the description covers the core flow well: secrets handling, URL-passing rules, persistence override, and presence-flag semantics. Output schema exists so return values need no explanation. But mode, role, auth_mode, and collection remain undefined, leaving meaningful gaps for such a complex, poorly-schematized tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does for the most important parameters: zeus_url, bucket, scope (URL/sample passing rules) and the has_* presence flags ('Presence flags only for credentials and LLM key'). However, mode, role, auth_mode, and collection are never mentioned across 11 total parameters, leaving roughly a third of the surface unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Store non-secret prereqs for readiness' and immediately negates what it does not do ('does not store password/token values'). This distinguishes it from the secret-handling context and from siblings like readiness_check and doctor, giving an agent a precise mental model of the tool's scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides exceptionally explicit when/when-not conditions: pass zeus_url/bucket/scope when the user named a URL or sample in chat, leave zeus_url empty and wait for the form otherwise, and 'Do not copy the host ZEUS_URL.' The secret-vs-env-var routing ('Put real secrets in environment variables') tells the agent what must not go through this tool. No sibling is named, but the conditional guidance is detailed enough to drive correct invocation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
smoke_test_agentC
One ZeusRuntime run_turn (requires kotenai-zeus-client>=2.3.0 + LLM key).
If the client is missing and demo_travel_sample documents Docker install, returns guide-only docker compose next_action instead of only pip install.
| Name | Required | Description | Default |
|---|---|---|---|
| question | No | In one short sentence, what data is available in this scope? | |
| update_checklist | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, destructiveHint=false. The description adds some behavioral context: it requires a specific client version and LLM key, and it may return a guide-only docker compose next_action. However, it does not explain side effects, failure modes, or what 'run_turn' entails. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded with the core dependency, but the second sentence is dense and somewhat cryptic. It earns its place by adding conditional behavior, but the phrasing is awkward and could be clearer. Not overly verbose, but not well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is an output schema, the description needn't explain return values, but it still lacks essential context: what the tool actually does, when to use it, and what the parameters mean. The dependency and conditional behavior are useful but incomplete. An agent would struggle to invoke this correctly without more information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does not explain the 'question' parameter or 'update_checklist' parameter at all. The description's mention of 'run_turn' and 'next_action' hints at behavior but not parameter meaning. With 0% coverage and no param explanation, this is a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description says 'One ZeusRuntime run_turn' but does not clearly state what the tool does. It mentions a dependency and a conditional behavior about docker compose, but the core action is vague. It does not distinguish itself from siblings like smoke_test_zeus or diagnose_error.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a conditional behavior (if client missing, return guide-only docker compose next_action) but does not explicitly state when to use this tool versus alternatives. No clear context for when an agent should invoke smoke_test_agent over smoke_test_zeus or diagnose_error.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
smoke_test_zeusB
Smoke Zeus without LLM: readiness + POST /v2/{bucket}/{scope}/describe.
| Name | Required | Description | Default |
|---|---|---|---|
| update_checklist | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate readOnlyHint=false, idempotentHint=false, and destructiveHint=false, which is a mixed safety profile. The description adds 'without LLM' to clarify that no language model is invoked, but doesn't disclose what happens during the smoke test (e.g., whether it creates or modifies resources), or what the readiness check entails. There is no contradiction, but the description is thin on behavioral details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with essential information: it identifies the tool's purpose and the endpoint. It is front-loaded with the key action 'Smoke Zeus without LLM' and includes a specific HTTP endpoint, which is efficient. No fluff or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has a simple signature (1 optional param), annotations provide some safety context, and an output schema exists, so the description does not need to detail return values. However, given the existence of a sibling 'smoke_test_agent', the description should clarify the difference, and it doesn't explain what 'readiness' means in this context or what 'smoke' entails beyond an HTTP call. It is adequate but with gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has one optional parameter 'update_checklist' with 0% description coverage in the schema itself. The description does not explain it either, but with only one parameter and no required params, the burden is low. The parameter name is self-explanatory, and the description's focus on 'readiness' and 'describe' implies the checklist is updated, but explicit clarification would be better.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool performs a 'smoke test' on Zeus and mentions a specific endpoint (POST /v2/{bucket}/{scope}/describe), which is specific. However, it does not differentiate from sibling tools like 'smoke_test_agent' – the name 'smoke_test_zeus' is a tautology, and the description doesn't clarify what makes it distinct from the agent variant.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention when to choose smoke_test_zeus over smoke_test_agent or other diagnostic tools. The context signals show no explicit usage context, so an agent would have to infer from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_projectADestructive
Start or reset first-app coaching checklist (single-agent default).
sample:
travel (default) — UI path via demo_travel_sample / use_sample
beer — beer-sample catalog UI; search is run_turn + MINI-SCHEMA (use_sample sample=beer)
api — API-only FastAPI scaffold (scaffold_app app_kind=api)
yelp / multi — gated until single-agent smokes green unless force_multi
Multi-agent goals (goal=multi or sample=yelp) are gated until single-agent smokes are green, unless force_multi=true (ZDH-11).
When no Zeus URL is stored, the MCP call asks for the public :8080 URL before the checklist starts. Leave zeus_url empty when the user did not name one. Do not copy the host ZEUS_URL. Do not pass a password.
| Name | Required | Description | Default |
|---|---|---|---|
| goal | No | single-agent | |
| sample | No | travel | |
| force_multi | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the reset behavior aligns. The description adds valuable behavioral details beyond the annotations: the zeus_url prompt when not stored, instructions to leave it empty, not to copy the host ZEUS_URL, and not to pass a password. It also discloses the gating logic and the need for force_multi to override, which an agent would not infer from annotations alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is structured into a main purpose line followed by bullet-like sample explanations and gating rules. It is front-loaded with the core action and then provides necessary details. Each line serves a purpose, though the formatting is loose (plain text line breaks) and could be more compact or use clearer separators. Still, it is efficient and readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that an output schema exists, the description does not need to explain return values. It covers the main usage context, parameter nuances, gating, and credential handling. It could mention what happens on a reset (e.g., does it overwrite existing state) more explicitly, but the destructiveHint already flags that. Overall, it is complete enough for an agent to invoke this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden. It explains all three parameters: goal (default single-agent, multi-agent gated), sample (with mappings to travel, beer, api, yelp/multi and their underlying tools), and force_multi (overrides gating). It gives concrete value semantics and edge cases, fully compensating for the lack of schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Start or reset first-app coaching checklist'. It clearly differentiates from sibling tools by explicitly referencing them (use_sample, scaffold_app) and explaining how they fit into each sample path. An agent can immediately understand what this tool does and how it relates to siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool, including gating conditions (multi-agent goals restricted until single-agent smokes pass unless force_multi=true) and specific sample-derived behaviors (e.g., 'api' uses scaffold_app, 'beer' uses use_sample). It does not explicitly say 'use this instead of X' but implicitly routes to alternatives, which is adequate for these sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
use_sampleCIdempotent
UI sample: travel clone or beer Direct template.
sample=travel — locate/clone demo_travel_sample; set DEMO_TRAVEL_SAMPLE_DIR. sample=beer — write demo_beer_sample catalog UI. Search is rt.agent.run_turn with chat_request omitted (catalog.load_for_turn merges MINI-SCHEMA). LLM key required. No pipeline body. project_name = directory name (defaults: demo_travel_sample / demo_beer_sample). Extra travel-only phases stay on travel_golden_path (travel toolset). When no Zeus URL is stored, the MCP call asks before writing or cloning. Do not pass a password.
| Name | Required | Description | Default |
|---|---|---|---|
| sample | No | travel | |
| parent_dir | No | ||
| sample_dir | No | ||
| project_name | No | ||
| clone_if_missing | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, destructiveHint=false, openWorldHint=true, idempotentHint=true. The description adds that it may ask before writing or cloning, and that no password should be passed. This is useful, but the description is fragmented and doesn't fully explain side effects or requirements like the LLM key mentioned. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single block of dense, jargon-heavy text that lacks clear structure. It's not front-loaded with a clear purpose; instead, it lists technical details and internal identifiers. It could be more concise and organized with paragraphs or bullets.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 5 optional parameters, complex behavior (cloning vs writing, travel-specific phases), and an output schema. The description is incomplete: it doesn't cover all parameters, doesn't explain return values (though output schema exists, it doesn't), and leaves many ambiguities for an agent to resolve.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate by explaining parameters. It explains 'project_name = directory name' and 'sample=travel/beer', but it doesn't clarify 'parent_dir', 'sample_dir', or 'clone_if_missing' semantics. It omits details like when to set sample_dir vs parent_dir, and what clone_if_missing does beyond the name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description mentions two samples ('travel' and 'beer') and actions like 'locate/clone' and 'write catalog UI', but the phrasing is cryptic and uses jargon ('rt.agent.run_turn', 'catalog.load_for_turn merges MINI-SCHEMA') that obscures the tool's actual function. It is not a clear, plain-language statement of what the tool does, and it doesn't differentiate from siblings like 'scaffold_app' or 'start_project'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives some conditional hints ('travel-only phases stay on travel_golden_path', 'When no Zeus URL is stored, the MCP call asks before writing or cloning'), but it doesn't clearly state when to use this tool versus siblings. It lacks explicit when-to-use/when-not-to-use guidance and doesn't mention alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
12 tool updates
v0.7.3- First observed
bind_contract - First observed
diagnose_error - First observed
doctor - First observed
next_step - First observed
readiness_check - First observed
recommend_surface - First observed
scaffold_app - First observed
set_prereq - First observed
smoke_test_agent - First observed
smoke_test_zeus - First observed
start_project - First observed
use_sample
TDQS
Scored across 12 tools
Tools are mostly distinct: smoke_test_agent vs smoke_test_zeus differentiate by LLM usage, and doctor/readiness_check/diagnose_error have different scopes (health, live gates, error mapping). Minor overlap exists but descriptions clarify boundaries.
Most tools follow a verb_noun snake_case pattern (start_project, set_prereq, scaffold_app, use_sample, bind_contract, diagnose_error, recommend_surface). Two exceptions: 'doctor' is noun-like and 'next_step' is adjective_noun, which slightly breaks the pattern but overall is readable and predictable.
12 tools for a developer-helper/coaching server is well-scoped. Each tool has a clear purpose and covers the domain without excessive overlap or unnecessary bloat.
The surface covers smoke testing, scaffolding, samples, prereq setting, health checks, readiness, error diagnosis, and step-by-step coaching. Minor gaps like explicit update/delete operations are not needed for this domain, so coverage is adequate.
Maintenance
Related MCP Connectors
One message in, a full agentic application out: website and MCP app, live. Built from any AI client.
Free enterprise-grade due diligence for your app, run by your coding agent. Then the full SDLC.
1Outcome-as-a-Service commerce for AI agents: discover, hire, settle on proof. Live on devnet.
Deterministic AI code review, with an audit record. Governance inside the agent loop.
Related MCP Servers
- AlicenseNot gradedqualityFmaintenanceEnables AI agents to autonomously request services from other specialized agents and compensate them via x402 micropayments. Demonstrates a Machine-to-Machine economy using A2A protocol for agent communication, MCP for context management, and blockchain-based payments on Base network.30 npm2MIT

Planwrightofficial
FlicenseNot gradedqualityCmaintenanceEnables orchestration of autonomous coding agents (Claude Code, Cursor, etc.) through an objective-native planning board with hash-chained audit trail. Humans define outcomes, agents claim and execute tasks via MCP.7-- AlicenseBqualityAmaintenanceLocal-first Agent OS that wraps Claude Code, Codex CLI, and other coding agents in a replayable Seed → Ledger → Runtime contract, driven by an interview → seed → execute → evaluate → evolve workflow loop.3416,680 PyPI6,154MIT
- FlicenseAqualityCmaintenanceDemonstrates agent-native platform onboarding with human-in-the-loop API key provisioning, plus live usage and rate-limit queries via Anthropic Admin APIs.3-