aethis-mcp
OfficialClick on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@aethis-mcpcheck if a Vogon qualifies for spacecraft crew certification"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
aethis-mcp
MCP server for the Aethis decision engine. Compile legislation, policy, contracts, and regulation into deterministic logic — same input, same answer, every time, with a full audit trail.
Install · Skills · Quick start · Tools · Setup · Authoring · DSL · Troubleshooting
Install
Authoring is in private beta. Decision tools (
aethis_decide,aethis_schema,aethis_explain,aethis_next_question) are public — no key required. Authoring tools (rule generation, test refinement, publishing) require an invite. Request access at aethis.ai/developer-access.
Recommended — one command via aethis-cli:
uv tool install aethis-cli
aethis mcp install --target allWires the server into claude-code, cursor, claude-desktop, or windsurf. Idempotent. Restart your editor to pick up the change. Re-run after aethis account generate rotates a key. Full options: aethis mcp install --help.
Manual install:
claude mcp add aethis -- npx -y aethis-mcpFor Cursor / Claude Desktop / Windsurf manual config, see Setup.
Onboarding an AI coding agent end-to-end? See docs.aethis.ai/agents/onboarding — install + verify + auth + workflow patterns in one page.
Related MCP server: decide
Skills
After the MCP server is installed, add reusable agent workflows with aethis-skills:
npx skills add Aethis-ai/aethis-skillsThe skills package provides workflows for policy-to-ruleset authoring, test/refine/publish loops, decisions with trace, and regression comparison. It calls the MCP tools in this package; it does not replace the MCP server.
Quick start
aethis_decide({
ruleset_id: "aethis/spacecraft-crew-certification",
field_values: { "space.crew.species": "Vogon" },
include_trace: true
}){
"decision": "not_eligible",
"fields_provided": 1,
"fields_evaluated": 11,
"trace": {
"species_check": "FAIL — species is 'Vogon' (disqualifying, Section 3)"
}
}Public rulesets work without a key. Browse: aethis_discover_rulesets({}) or docs.aethis.ai.
Engine determinism + accuracy benchmarks: Aethis-ai/confidently-wrong-benchmark.
Tools
35 tools across six groups.
Group | Access | Tools |
Decision | public |
|
Discovery — public catalogue | public |
|
Discovery — your tenant | private beta |
|
Authoring — rulebooks | private beta |
|
Authoring — sections & fields | private beta |
|
Authoring — generation | private beta |
|
Management | private beta |
|
aethis_graph is public for a public showcase ruleset (ruleset_id) and tenant-scoped for a rulebook (rulebook_id) — it returns the ruleset-map graph ({nodes, edges, sections, stats}, each node's display.sentence/display.routes/display.expr) plus a ready-to-render mermaid diagram string. Pass include_graph_overlay: true to aethis_decide to get that same graph back with a specific decision's per-criterion status (satisfied/not_satisfied/pending) stamped onto it (graph_overlay in the response) — a "you are here" map for those inputs.
aethis_create_rulebook / aethis_update_rulebook manage a Rulebook's identity (name/domain/slug/description) and robot_hints — beat-keyed natural-language guidance for the conversational agent. Active beats: general_context, preamble, session_start, postamble, session_end, stuck. Reserved (accepted, not yet acted on): persona, conversational_style, section_transition. Composition (bridging rulesets via outcome_logic) is a separate, larger surface not covered by these two tools yet.
Workflows
Evaluate eligibility (2 calls):
aethis_schema(ruleset_id) → fields needed
aethis_decide(ruleset_id, fields) → eligible / not_eligible / undeterminedPass include_trace: true for the per-criterion evaluation trail. Pass include_explanation: true for human-readable rule descriptions.
aethis_decide accepts either ruleset_id (single ruleset, may be public) or rulebook_id (composed multi-ruleset rulebook) — the two are mutually exclusive. Rulebook decide always requires an API key (AETHIS_API_KEY); anonymous callers get HTTP 401. aethis_graph follows the same ruleset_id/rulebook_id split for the underlying map.
Conversational eligibility (next-question routing):
aethis_next_question(ruleset_id, field_values)Returns the most informative remaining question and the optimal_path of remaining questions. Call again after each answer; the engine recomputes from the updated state. Stops when a decision is reachable.
Authoring (private beta): see Authoring.
Prompts
Prompt | Description |
| Step-by-step TDD authoring workflow |
| Decision workflow guide; accepts optional |
Setup
Decision tools work with no key. For invited authoring access, run aethis login, then install with aethis mcp install --target <client>. The installer references a saved profile so the host configuration does not contain the API key.
Claude Code
# Decision tools only
claude mcp add aethis -- npx -y aethis-mcp
# With authoring access
claude mcp add aethis -e AETHIS_PROFILE=default -e XDG_CONFIG_HOME=/absolute/path/to/config -- npx -y aethis-mcpClaude Desktop
Add to ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows):
{
"mcpServers": {
"aethis": {
"command": "npx",
"args": ["-y", "aethis-mcp"]
}
}
}For authoring, add "env": { "AETHIS_PROFILE": "default", "XDG_CONFIG_HOME": "/absolute/path/to/config" }. Use your saved profile name and the absolute directory containing aethis/credentials (normally your home directory’s .config).
Cursor / Windsurf
Add to ~/.cursor/mcp.json or ~/.codeium/windsurf/mcp_config.json (same JSON shape).
Keys
AETHIS_PROFILE— non-secret saved profile name. It pins the account and endpoint used by this registration, even if the CLI’sactive_profilelater changes.XDG_CONFIG_HOME— absolute config directory containingaethis/credentials. Relative values and credentials symlinks escaping your home/config directory are refused; credential files must have no group/other permission bits (normally0600).AETHIS_API_KEY— optional deliberate process-environment override for the platform key. Prefer saved-profile references when installing; avoid putting raw keys in host config or command arguments. The host must securely supply the process environment; it may not inherit your shell environment.AETHIS_ANTHROPIC_KEY_ENV— the name of the env var holding the Anthropic key Aethis authoring tools may use (e.g.AETHIS_ANTHROPIC_KEY). The server never reads a provider key from the environment unless you set this. See Passing your Anthropic key safely.Rotate via
aethis account generate+aethis account revoke <key_id>. Mint one key per machine for surgical revocation.
Credential precedence and restart behavior
MCP parses the CLI credentials file as YAML. AETHIS_PROFILE selects a named profile; otherwise the file’s active_profile (or default) selects it. The profile supplies both its API key and base_url, with https://api.aethis.ai as the default endpoint. With an explicit AETHIS_PROFILE, AETHIS_API_KEY and AETHIS_BASE_URL deliberately override their respective values. Without an explicit profile selector, AETHIS_API_KEY uses AETHIS_BASE_URL or the default endpoint, ignoring the implicitly active profile (including implicit anonymous). Missing or malformed explicitly selected profiles fail visibly, including when environment overrides are present. Explicit AETHIS_PROFILE=anonymous always stays unsigned. After anonymous setup, run aethis login and install again to reference the saved authoring profile, then restart the host.
A saved profile outranks old macOS Keychain entries. If no profile is configured and no explicit name is selected, MCP can use the legacy default keychain entry, then the older flat credentials.yaml file. Flat api_key/base_url files at aethis/credentials remain supported. A configured profile awaiting login stays unsigned instead of borrowing another stored key.
The server keeps its authenticated startup key and endpoint paired until restart. If it started without a key, an authenticated tool can pick up a later login for the same endpoint. A changed endpoint causes a visible refusal: restart the MCP host to load the new pair. Startup stderr reports only the credential source, never key values.
Passing your Anthropic key safely
Authoring tools (aethis_generate_and_test, aethis_refine, aethis_discover_fields, aethis_refine_fields, aethis_discover_sections, aethis_refine_sections) need an Anthropic API key per call. Three accepted forms — listed in preferred order:
The server sends a provider key to Aethis only when you configured one for Aethis. It never picks up ANTHROPIC_API_KEY (or any other variable) from your environment on its own, and an env-var name supplied in a tool call by the host model is refused unless it is the one you configured. Only Anthropic keys (sk-ant-…) are accepted; anything else is refused locally and not sent.
AETHIS_ANTHROPIC_KEY_ENV(recommended). In the MCP server config, put the key in a dedicated env var and name that var inAETHIS_ANTHROPIC_KEY_ENV. Tools then use it automatically. The raw value never appears in the tool call payload, so it does not land in the MCP host's session transcript on disk.// claude_desktop_config.json { "mcpServers": { "aethis": { "command": "npx", "args": ["aethis-mcp"], "env": { "AETHIS_PROFILE": "default", "XDG_CONFIG_HOME": "/absolute/path/to/config", "AETHIS_ANTHROPIC_KEY_ENV": "AETHIS_ANTHROPIC_KEY", "AETHIS_ANTHROPIC_KEY": "sk-ant-..." // never echoed back to the LLM } } } }aethis_generate_and_test({ project_id }) // uses the configured keyanthropic_key_keychain(macOS). A keychain reference — either"account"(service defaults toaethis-anthropic-key) or"service:account". Store the key once withsecurity add-generic-password -U -s aethis-anthropic-key -a my-anthropic -w 'sk-ant-...', then call:aethis_generate_and_test({ project_id, anthropic_key_keychain: "my-anthropic" })anthropic_key(deprecated). Pass the raw key as a tool argument. Accepted for backwards compatibility, but the raw value is written verbatim to the host's session transcript JSONL on disk. If a key was ever passed this way, rotate it before relying on the safer forms.
Authoring (private beta)
Authoring requires an invite. Request access. Decision tools (above) are public.
Three-phase workflow. Phases 1–2 are for multi-section domains; skip them for single-section rules and go straight to Phase 3.
Phase 1 — Section discovery
aethis_discover_sections({ domain, sources: [{ name, content }, ...] })
aethis_validate_sections({ domain, expected_sections, discovered_sections })
aethis_refine_sections({ domain, feedback, sources })Phase 2 — Field vocabulary
aethis_set_field_spec({
project_id,
expected_fields: [{ key, sort, enum_values?, notes?: [{ note_text, source?, metadata? }] }, ...]
})
aethis_discover_fields({ project_id }) // auto-validates against the spec if set
aethis_refine_fields({ project_id, feedback })
aethis_validate_fields({ project_id, expected_fields })Phase 3 — Generate, test, publish
aethis_create_ruleset({
name, section_id, domain?, source_text,
test_cases: [{ name, field_values, expected_outcome, expectations? }, ...],
contract_version?: 1,
expected_review_bindings?: { field_id: { token: true | false | null } }
})
aethis_generate_and_test({ project_id })
aethis_refine({ project_id, feedback }) // iterate until tests pass
aethis_publish({ project_id }) // refuses if tests fail; returns ruleset_id on successWhen a test carries expectations, set contract_version: 1. The optional
binding catalogue is generic authoring metadata: omit it when no binding
assertion is needed, or pass {} to assert that no review bindings exist.
The server must confirm the complete stored contract before generation begins.
aethis_create_ruleset creates a project; it does not append or replace tests
on an existing project. Use aethis_set_tests with the complete version-1
contract for a later replacement, then generate or refine that same project.
Legacy updates cannot discard stored assertions. Object keys named __proto__
are rejected in field-value and binding maps rather than silently discarded.
If generation polling times out, call aethis_generation_status({ project_id })
before retrying: use its telemetry_availability, server-authoritative
worker_lifecycle, and retry_readiness, and retry only when readiness is
ready. An old heartbeat alone is not proof that the worker died. Call
aethis_cancel_generation({ project_id, job_id, confirm_job_id }) only after
showing the observed job_id and receiving explicit confirmation to abandon
that active run. It releases the project's job ownership, but worker shutdown
may be cooperative rather than immediate; inspect the returned detail. It is a
destructive, API-key-protected mutation. The response distinguishes a new
cancelled transition from the idempotent already_cancelled result.
Guidance
Targeted hints without regenerating, plus cross-section principles for a domain:
aethis_add_guidance({ project_id, guidance_text, process_type })
aethis_list_guidance({ project_id })
aethis_add_domain_guidance({ domain, guidance_text, process_type, notes? })
aethis_list_domain_guidance({ domain })process_type is rule_generation (default) or field_extraction.
Diagnose a failing test
aethis_explain_failure({
ruleset_id, field_values, expected_outcome, test_name
})
// Returns criterion statuses, the failing rule, and a targeted fix hint.Tests are the publish gate. aethis_publish refuses to publish a ruleset with a failing test. SMEs write the tests; the LLM generates the rules from source text + guidance; the platform refuses to ship rules that don't satisfy the tests. Better tests = faster convergence.
Anthropic key required for authoring. ConfigureAETHIS_ANTHROPIC_KEY_ENV or use anthropic_key_keychain (macOS keychain ref) rather than the raw anthropic_key argument — see Passing your Anthropic key safely. Used per-request, never stored server-side; the raw form, however, lands in the MCP host's session transcript on disk.
DATE fields use integer ordinals (date.toordinal()), not ISO strings. 2025-04-13 = 739354. Quick conversion: python3 -c "from datetime import date; print(date(2025,4,13).toordinal())".
Field types
Type | Description |
| True / false |
| Integer (counts, money as pence, percentages as integers) |
| Closed set of named values |
| Integer ordinal — |
| Integer days |
| Free text — prefer |
Operators
Category | Operators |
Logic |
|
Comparison |
|
Membership |
|
Arithmetic |
|
Aggregation |
|
Helpers
days_between(date_a, date_b)→Intyears_between(date_a, date_b)→Int— completed whole years between the two dates (leap-correct). Use this for age from a date-of-birth field; never derive age asdays_between(...) / 365.min(a, b, ...),max(a, b, ...)→IntConstant arithmetic folded at authoring time (
5 * 365→1825)
Not supported
Division between runtime field values
Weighted scoring or probabilistic outcomes
Lists as field values (use pre-aggregated
Int/Bool)More than 3 outcome tiers (
eligible/not_eligible/undetermined)
Troubleshooting
Error | Cause | Fix |
|
| Configure in MCP client settings, not shell profile |
| Missing Anthropic key | Set |
| Wrong ID or archived |
|
| Daily limit | Client retries automatically. eng@aethis.ai for higher tier |
| Tests don't pass |
|
Generation timeout (504) | Server still generating (5–15 min normal) | Wait, then |
| DATE field passed as ISO string | Use |
Related
aethis-cli — Python CLI; file-based authoring with YAML test cases
aethis-examples — runnable rulesets (spacecraft, construction-CAR, consumer credit) and benchmark scenarios
confidently-wrong-benchmark — paper, 225-scenario benchmark, LegalBench harness
Development
git clone https://github.com/Aethis-ai/aethis-mcp.git
cd aethis-mcp && npm install && npm test && npm run buildLicense
MIT
Available Tools
35 toolsaethis_add_domain_guidanceA
Add a guidance hint at domain level — applies to ALL projects in the domain, not just one project. Use for cross-section principles: solicitor navigation, discretion model, raw-facts principle. These hints are retrieved automatically during generation for any project in the domain. Use adherence='exact' with process_type='section_discovery' to specify exactly which sections the SME wants — the LLM will follow them precisely.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | SME commentary or legislation provenance. Never sent to LLM. | |
| domain | Yes | Domain identifier (e.g. 'uk_citizenship') | |
| adherence | No | How strictly the LLM must follow this hint. 'exact' = must follow, produce nothing beyond what is specified (use for SME-defined section lists); 'guided' = strong preference, may adapt if source text requires (default); 'loose' = soft suggestion. | guided |
| process_type | No | Which authoring phase this hint targets — rule_generation (default), field_extraction, or section_discovery | rule_generation |
| guidance_text | Yes | The guidance hint text |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare write (readOnlyHint=false), non-idempotent, open-world. The description adds real behavioral context beyond that: hints are auto-retrieved during generation for any project in the domain, and adherence='exact' forces the LLM to follow precisely. It does not address idempotency (what happens if the same hint is added twice), a notable gap for a non-idempotent write.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the crucial domain-vs-project distinction, and each sentence carries information. The final adherence/process_type sentence is useful but slightly ancillary to the core purpose, adding mild redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description covers purpose, scope, retrieval behavior, and key parameter usage, which is sufficient for a write tool. Remaining gaps (idempotency semantics, permission requirements) are minor but real.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description goes beyond the schema by prescribing a concrete parameter pairing (adherence='exact' with process_type='section_discovery') and its effect, which is a usage pattern the schema does not spell out.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (add a guidance hint at domain level) and immediately disambiguates scope: 'applies to ALL projects in the domain, not just one project.' This clearly separates it from the sibling aethis_add_guidance, which operates at project level.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete context for when to use it (cross-section principles like solicitor navigation, discretion model, raw-facts principle) and notes hints are retrieved automatically during generation. It implies the project-level alternative via 'not just one project' but never names aethis_add_guidance explicitly as the alternative to pick instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_add_guidanceA
Add a guidance hint to a project. Use for domain knowledge not in the source text. Then call aethis_generate_and_test to regenerate.
| Name | Required | Description | Default |
|---|---|---|---|
| adherence | No | How strictly the LLM must follow this hint. 'exact' = must follow precisely, produce nothing beyond what is specified; 'guided' = strong preference, may adapt if source text requires (default); 'loose' = soft suggestion. | guided |
| project_id | Yes | The project ID | |
| process_type | No | Which authoring phase this hint targets. Use 'field_extraction' for field design principles (e.g. raw-facts principle, solicitor navigation). Defaults to 'rule_generation'. | rule_generation |
| guidance_text | Yes | Domain knowledge or correction not present in the source text |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, destructiveHint=false. The description adds a behavioral note: the guidance applies to domain knowledge absent from source text, and implies adding it affects future generation steps. It does not cover authentication, side effects, or whether existing guidance is replaced. With annotations carrying the safety profile, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action, followed by usage scope and the next step. No waste; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema and annotations are moderate. The description gives a basic workflow but omits important details like whether the call is synchronous, whether it replaces existing guidance, and what happens on failure. For a mutation tool with openWorldHint, more context would be expected.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the description need not explain parameters. It hints at the purpose of guidance_text ('domain knowledge not in the source text') but adds no syntax, format, or default information beyond what the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Add a guidance hint to a project.' This is clear and actionable. It does not differentiate itself from the sibling aethis_add_domain_guidance, which sounds like it may overlap, but the core operation is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to use this for 'domain knowledge not in the source text' and directs the agent to call aethis_generate_and_test afterward. This provides clear context and a follow-up action, though it offers no exclusions or when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_archive_projectADestructiveIdempotent
Archive a project. Archived projects are preserved but excluded from listing. This is permanent.
| Name | Required | Description | Default |
|---|---|---|---|
| project_id | Yes | The project ID to archive |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, so the safety profile is covered. The description adds genuinely useful context beyond that: archived projects are 'preserved but excluded from listing,' which clarifies this is not a hard delete, and flags permanence. It stops short of saying whether the action can be undone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and followed by the two facts that matter (what happens to the object, and that it is irreversible). No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter mutation whose annotations already carry the destructive/idempotent profile, the description covers the essential outcome. It is slightly thin on whether archiving is reversible or requires specific authorization, but nothing critical for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter (project_id) and the schema already documents it at 100% coverage. The description adds no format, validation, or lookup guidance beyond what the schema states, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Archive) and resource (a project), and immediately clarifies the semantics of 'archive' as preserved-but-hidden. It does not explicitly differentiate itself from the sibling aethis_archive_ruleset, but the resource noun makes the target unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to archive versus using alternatives (e.g. aethis_list_projects, aethis_review_project) or any stated prerequisites such as required permissions or project state. Usage is only implied by the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_archive_rulesetADestructiveIdempotent
Archive a rule ruleset. Archived rulesets are preserved but excluded from /decide resolution. This is permanent.
| Name | Required | Description | Default |
|---|---|---|---|
| ruleset_id | Yes | The ruleset ID to archive |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, so the safety profile is covered. The description still adds real value beyond them by stating the post-condition (preserved but excluded from /decide) and the irreversibility ('This is permanent').
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and immediately followed by the effect and the permanence warning. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter mutation with no output schema and annotations that already carry the destructive/idempotent flags, the description covers action, consequence, and irreversibility. Only the lack of explicit usage routing keeps it short of full completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single ruleset_id parameter is fully documented in the schema, so the description adds no format or semantic detail beyond it. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('archive') and resource ('ruleset'), letting an agent distinguish it from siblings such as aethis_archive_project and aethis_list_rulesets without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when this is appropriate by explaining that archived rulesets are excluded from /decide resolution, but it never states explicit when-to-use/when-not conditions or names an alternative (e.g., deleting or updating instead).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_cancel_generationADestructiveIdempotent
Request cancellation of one observed generation job and release only its project ownership. First call aethis_generation_status and bind the exact job_id; confirm_job_id protects against accidental target mismatch but does not itself prove human approval. MCP hosts should require destructive-action approval, and agents must obtain a fresh explicit user reply before calling. The response outcome is cancelled or idempotent already_cancelled. Cancellation may be cooperative rather than immediate. It is a destructive mutation and requires an API key.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | The exact job ID returned by aethis_generation_status | |
| project_id | Yes | The project ID containing the observed generation job | |
| confirm_job_id | Yes | Repeat job_id to bind the cancellation target; host/user approval is a separate requirement |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true and openWorldHint=true, so the description goes further by disclosing that cancellation 'may be cooperative rather than immediate', that only project ownership is released, that outcomes are 'cancelled or already_cancelled', and that an API key is required. This is meaningful added context, though the destructive/idempotent framing partly restates the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and the required workflow order, and each sentence carries useful guidance. It runs somewhat long, and the closing 'It is a destructive mutation' sentence largely duplicates the destructiveHint annotation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation with no output schema, the description covers the workflow, approval requirement, possible outcomes and API-key need. It stops short of naming when an agent should NOT cancel or listing alternative tools, but the essential invocation context is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters are already documented. The description reinforces the binding intent ('bind the exact job_id', confirm_job_id 'protects against accidental target mismatch') but adds little syntax or format detail beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('cancellation of one observed generation job') and adds scope ('release only its project ownership'), which lets an agent distinguish this from sibling mutations like aethis_archive_project or aethis_publish.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the prerequisite and ordering ('First call aethis_generation_status and bind the exact job_id') and states the required human-approval condition ('agents must obtain a fresh explicit user reply before calling'), leaving little to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_create_rulebookA
Create a new Rulebook — the composed-whole execution unit that bridges multiple rulesets (the parts) via outcome_logic. Created empty: no rulesets, no field vocabulary, no tests, status='draft'. Populate afterwards with aethis_create_ruleset for each section, then wire up the field vocabulary and composition logic before publishing. Requires an API key.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Human-readable name for the rulebook (e.g. 'UK FSM') | |
| slug | No | Stable human-readable alias (e.g. 'aethis/uk-fsm'). Globally unique when set; recommended for any rulebook referenced from outside this session. | |
| domain | No | Domain hint, lower-snake (e.g. 'uk_fsm') | |
| description | No | Optional description | |
| robot_hints | No | Assistant guidance for the conversational agent, keyed by conversational beat. Natural language only — no rule syntax, no field keys. Active beats: general_context, preamble, session_start, postamble, session_end, stuck. Reserved (accepted, not yet acted on): persona, conversational_style, section_transition. An unknown beat key is rejected. Omit for no hints (create), or to leave existing hints unchanged (update). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover safety (readOnlyHint=false, destructiveHint=false, non-idempotent). The description adds real behavioral context beyond them: the object is created empty (no rulesets, no field vocabulary, no tests) with status='draft', and it requires an API key. It does not describe the returned object's shape, but with annotations carrying the safety profile this is a solid addition.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then creation-state semantics and the population workflow. Four compact sentences with little waste; the trailing auth note is useful but slightly bolted-on.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a creation tool with 100% schema coverage, no output schema, and annotations carrying the safety profile, the description covers the essential extras: post-creation state and the follow-up workflow. The only minor gap is not describing what the create call returns.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters including the nested robot_hints object are already documented in the schema. The description adds no parameter-level syntax or format detail beyond that, making the baseline 3 correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a new Rulebook') and defines it as the composed-whole execution unit that bridges rulesets via outcome_logic. This conceptual framing distinguishes it from aethis_create_ruleset (the parts) without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear next-step workflow ('Populate afterwards with aethis_create_ruleset for each section, then wire up the field vocabulary and composition logic before publishing') and names a sibling. It lacks an explicit when-not or alternative-selection condition, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_create_rulesetA
Create a new rule ruleset with source text and test cases (TDD). Test cases are required. After creation, call aethis_generate_and_test.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Human-readable name for the rule ruleset | |
| domain | No | Domain hint (e.g., 'uk_immigration') | |
| section_id | Yes | Unique section identifier (e.g., 'flight_readiness') | |
| test_cases | Yes | Test cases with optional strict acceptance expectations. | |
| source_text | Yes | The source legislation, policy, or specification text | |
| contract_version | No | Required when acceptance expectations or expected_review_bindings are supplied. | |
| expected_review_bindings | No | Optional review-binding catalogue; omit for no assertion or use {} to assert zero bindings. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already disclose the write profile (readOnlyHint=false, idempotentHint=false, openWorldHint=true, destructiveHint=false), so the safety burden is partly carried elsewhere. The description adds the required-test-cases constraint and the follow-up generation step, but says nothing about duplicate-creation risk from non-idempotency, permissions, or what the created artifact looks like.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and immediately followed by the precondition and the next tool. Every clause carries information and nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a fairly complex tool (7 params, nested test_cases, optional contract_version and expected_review_bindings) with no output schema, the description covers the creation-to-generation flow but omits how the response should be interpreted and when the optional expectation-related parameters matter. It is minimally adequate given the strong schema coverage rather than complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all seven parameters including the nested test_cases shape are already documented in the schema; baseline 3 applies. The description echoes that source_text and test_cases are inputs but adds no format, sizing, or conditional-requirement detail (e.g., contract_version being required when expectations are supplied).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a new rule ruleset with source text and test cases') and flags the TDD framing, so an agent knows this bootstraps a ruleset from source text plus tests. It does not, however, differentiate itself from the sibling aethis_create_rulebook, so the boundary between the two creation tools is left implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Adds real workflow guidance: 'Test cases are required' and 'After creation, call aethis_generate_and_test' tells the agent both a hard precondition and the mandatory next step. It stops short of naming when to prefer this over aethis_create_rulebook or aethis_set_tests, so no exclusions or alternatives are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_decideARead-only
Evaluate eligibility against either a single published ruleset (ruleset_id) or a composed rulebook (rulebook_id). Provide exactly one. A rulebook composes multiple rulesets via outcome_logic — use it for the whole-form decision (e.g. aethis/uk-fsm). A ruleset is one section in isolation (e.g. aethis/uk-fsm/child-eligibility). Returns eligible/not_eligible/undetermined with optional trace and explanation. When undetermined, includes next_question and optimal_path. Rulebook evaluation always requires an API key; ruleset evaluation can be anonymous against public rulesets.
| Name | Required | Description | Default |
|---|---|---|---|
| ruleset_id | No | The ID or slug of a single published ruleset. Mutually exclusive with rulebook_id. | |
| rulebook_id | No | The ID or slug of a composed rulebook (e.g. `aethis/uk-fsm`). Mutually exclusive with ruleset_id. Requires an API key — anonymous callers get HTTP 401. | |
| field_values | Yes | Input field values (see aethis_schema for required fields) | |
| include_trace | No | Include the full evaluation trace showing how each rule was evaluated | |
| include_explanation | No | Include human-readable rule explanations with source citations | |
| include_graph_overlay | No | Stamp this decision's per-criterion outcome (satisfied/not_satisfied/pending) onto the ruleset-map graph and return it as graph_overlay — the same {nodes, edges, sections, stats} shape as aethis_graph, letting a caller render a 'you are here' map for these specific inputs. Off by default; the response is byte-identical to a call without the flag. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare read-only/non-destructive/open-world; the description goes further by disclosing the three outcome states (eligible/not_eligible/undetermined), the undetermined payload (next_question, optimal_path), optional trace and explanation, and the API-key requirement per mode. With no output schema, this return-value disclosure is the primary behavioral value and it is covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and the either/or constraint, then builds outward to mode semantics, return states, and auth. Every sentence carries distinct information; nothing is repeated or padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a six-parameter, nested-object tool with no output schema, the description supplies the missing pieces: mutual exclusivity of the two ID modes, the auth difference between them, and the shape of the non-success outcome. Nothing an agent needs to call it correctly is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3; the description adds conceptual semantics beyond the schema's field-level text — that a rulebook composes rulesets via outcome_logic and that exactly one of the two IDs must be supplied. It does not add syntax or format detail for field_values beyond deferring to aethis_schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (evaluate eligibility) and immediately disambiguates the two modes of operation: single published ruleset vs composed rulebook. The examples (`aethis/uk-fsm` vs `aethis/uk-fsm/child-eligibility`) make the distinction concrete and let an agent separate this from sibling tools like aethis_schema or aethis_next_question.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing guidance: 'Provide exactly one,' plus the decision rule — use a rulebook for the whole-form decision, a ruleset for one section in isolation. It also states the auth precondition that selects between modes (rulebook always needs an API key; ruleset can be anonymous against public rulesets).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_discover_fieldsARead-only
Discover input fields from the project's source text. Returns field names, types, descriptions, and completeness assessment. Run this BEFORE writing test cases to ensure field names are consistent. Call repeatedly with aethis_refine_fields to improve completeness.
| Name | Required | Description | Default |
|---|---|---|---|
| openai_key | No | Retired and refused: Aethis LLM tools use Anthropic models only. | |
| project_id | Yes | The project ID | |
| anthropic_key | No | An Anthropic API key the user explicitly provided for this call. [sensitive — do not echo or log] Deprecated: the raw value is written verbatim to the host's session transcript. Never fill this from the environment. | |
| anthropic_key_env | No | Optional. Only honoured when it equals the env var the user configured via AETHIS_ANTHROPIC_KEY_ENV in this MCP server's config; that configured key is used automatically, so this can be omitted. Do not guess a variable name: the server refuses any name the user did not configure. | |
| anthropic_key_keychain | No | macOS keychain reference the user created for Aethis: either 'service:account' or just 'account' (service defaults to 'aethis-anthropic-key'). The server reads it via the `security` command at call time. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false and idempotentHint=false; the description usefully confirms the iteration model ('call repeatedly ... to improve completeness'), which matches the non-idempotent hint. However, it never discloses that the tool invokes an LLM and may require an Anthropic key or incur cost — that requirement is buried in the schema's key parameters rather than surfaced behaviorally.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three front-loaded sentences: purpose, return payload, then workflow sequencing. Nothing is wasted, though the return-value sentence could arguably be dropped given how terse the rest is.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by naming the returned artifacts (names, types, descriptions, completeness) plus the recommended call sequence. It is largely complete for an extraction tool, with the only real gap being the unstated LLM/key dependency.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters (including the retired openai_key and the three auth-key variants) are documented in structured data. The description only implies that project_id selects the source-text scope and adds no syntax or format detail beyond the schema — baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Discover input fields from the project's source text') and enumerates the outputs (field names, types, descriptions, completeness). It clearly differentiates itself from the sibling refine_fields, but stays silent on how it differs from validate_fields or set_field_spec, which also touch fields.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit sequencing rule ('Run this BEFORE writing test cases') and names the companion tool to pair with ('Call repeatedly with aethis_refine_fields'). It lacks a negative case (when not to use it vs validate_fields), so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_discover_rulesetsARead-only
List public showcase rulesets across all tenants. No authentication required. Use this for first-time discovery, demos, or whenever the user asks 'what rulesets are available?' without referencing a specific project. Returns slug, ruleset_id, name (the human-readable section title), description, field_count, rule_count for each — pass the slug or ruleset_id to aethis_decide / aethis_schema / aethis_explain to interact with one. Distinct from aethis_list_rulesets, which is tenant-scoped.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum rulesets to return (default 20, max 50). | |
| offset | No | Pagination offset (default 0). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds genuinely useful behavioral context beyond that: 'No authentication required' and the exact returned field list. It lacks any statement about pagination behavior or whether results are cached/stable, but it adds real value over the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and scope, then usage triggers, return shape, and routing. Every sentence carries information, though the middle clause enumerating return fields and downstream tools makes it dense; a tighter phrasing would read better without losing content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description compensates by naming the exact returned fields (slug, ruleset_id, name, description, field_count, rule_count) and the next tools to pass them to. For a two-parameter browse tool with full annotation coverage, nothing an agent needs is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both limit and offset fully documented in the schema itself. The description adds no parameter syntax or semantics beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List public showcase rulesets across all tenants') plus scope, and explicitly distinguishes itself from the tenant-scoped aethis_list_rulesets. An agent can select between the two without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit triggering conditions ('first-time discovery, demos, or whenever the user asks what rulesets are available without referencing a specific project') and names the sibling it is not (aethis_list_rulesets, tenant-scoped). The when-to-use and the alternative are both spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_discover_sectionsARead-only
Discover the logical sections of source legislation for a domain. Provide the raw text of your source documents (legislation, guidance notes, form instructions). The service analyses the content and identifies which sections should be authored as separate rule rulesets. Run BEFORE creating projects — you need to know the sections before you can create one. Call aethis_refine_sections if sections are missing or incorrectly split.
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | Domain identifier, e.g. 'uk_citizenship' | |
| sources | Yes | Source documents to analyse. Provide the actual text content. | |
| openai_key | No | Retired and refused: Aethis LLM tools use Anthropic models only. | |
| anthropic_key | No | An Anthropic API key the user explicitly provided for this call. [sensitive — do not echo or log] Deprecated: the raw value is written verbatim to the host's session transcript. Never fill this from the environment. | |
| anthropic_key_env | No | Optional. Only honoured when it equals the env var the user configured via AETHIS_ANTHROPIC_KEY_ENV in this MCP server's config; that configured key is used automatically, so this can be omitted. Do not guess a variable name: the server refuses any name the user did not configure. | |
| anthropic_key_keychain | No | macOS keychain reference the user created for Aethis: either 'service:account' or just 'account' (service defaults to 'aethis-anthropic-key'). The server reads it via the `security` command at call time. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=true and idempotentHint=false, so safety and mutability are covered. The description usefully discloses that the service analyses raw text (implying LLM-driven, non-deterministic output, consistent with idempotentHint=false), but it says nothing about cost, latency, or whether repeated calls vary. Adequate but not rich beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four front-loaded sentences that each carry signal: purpose, input, outcome, ordering, and fallback. Slightly marred by the 'rule rulesets' typo and the ordering/fallback points being split across two trailing sentences, but there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only analysis tool with no output schema, the description conveys the purpose, required input, sequencing, and the corrective sibling. It hints at the return ('which sections should be authored as separate rulesets') but never describes the shape of the returned sections, which is the one gap for a tool whose whole value is its output.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six parameters are documented in the schema itself. The description only reinforces that raw source text must be supplied (mapping to 'sources'), adding no format or constraint detail beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('discover the logical sections of source legislation') and goes further by explaining the outcome ('identifies which sections should be authored as separate rule rulesets'). This clearly separates it from siblings like aethis_refine_sections and aethis_validate_sections without opening their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit ordering constraint ('Run BEFORE creating projects') and names the alternative for the failure case ('Call aethis_refine_sections if sections are missing or incorrectly split'). Both the when-to-use and the redirect condition are stated, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_explainARead-only
Get human-readable descriptions of the rules in a ruleset, including criteria groups, requirements, and exception paths.
| Name | Required | Description | Default |
|---|---|---|---|
| ruleset_id | Yes | The ID of the published rule ruleset |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=true and idempotentHint=false, so the safety profile is covered. The description still adds real value by disclosing what the explanation contains (criteria groups, requirements, exception paths) — important because there is no output schema to describe the return shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the verb and resource, with the content enumeration appended rather than padded. No redundant restatement of the title or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, read-only explain tool this is close to sufficient: the agent knows the input and what the output narrative will cover. It does not note whether the output is scoped/truncated for large rulesets, which is the only notable gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter, ruleset_id, and schema description coverage is 100% ('The ID of the published rule ruleset'), so the schema fully documents it. The description adds no format, sourcing, or lookup guidance for the ID, which is the expected baseline-3 outcome when the schema does the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: it returns human-readable descriptions of the rules inside a ruleset, and it names the sub-components covered (criteria groups, requirements, exception paths). The 'human-readable' framing implicitly separates it from the schema-oriented siblings (aethis_schema, aethis_rulebook_schema), but no sibling is named explicitly, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: an agent can infer this is the tool to call when it needs to understand what a published ruleset actually enforces. There is no explicit when-to-use, when-not-to-use, or pointer to aethis_explain_failure / aethis_schema for related needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_explain_failureARead-only
Diagnose why a ruleset produced an unexpected outcome for specific test inputs. Use during rule authoring when a test fails — returns the diagnosis, criteria with DSL metadata (waivable, review_required), and a targeted hint for fixing the rule.
| Name | Required | Description | Default |
|---|---|---|---|
| test_name | No | Name of the failing test case (included in the diagnosis for context) | |
| ruleset_id | Yes | The ID of the rule ruleset to diagnose | |
| field_values | Yes | The test input values that produced the unexpected outcome | |
| expected_outcome | Yes | The outcome you expected from this input |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, destructiveHint=false and openWorldHint, so the safety profile is covered. The description adds real value beyond that by disclosing the return payload — diagnosis, criteria with DSL metadata (waivable, review_required), and a fix hint — which is important since no output schema exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the purpose before the usage and return details; every clause carries information. The return-value clause is slightly list-heavy but earns its place given there is no output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only diagnostic tool with full schema coverage, the description supplies the missing piece — what the diagnosis returns — which compensates for the absent output schema. Only the differentiation from aethis_explain is left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (ruleset_id, field_values, expected_outcome, test_name) are already documented in the schema. The description adds no format or syntax detail beyond it, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Diagnose why a ruleset produced an unexpected outcome') and scopes it to test inputs, which is clear. It does not explicitly distinguish itself from the sibling aethis_explain, so an agent must infer the boundary between 'explain' and 'explain_failure'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear triggering context — 'Use during rule authoring when a test fails' — which tells the agent when this tool is appropriate. It stops short of naming an alternative (e.g. aethis_explain) or stating exclusions, so it is context without routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_generate_and_testA
Generate rules from source text and run all test cases. Triggers generation, polls until complete, then runs tests. Returns pass/fail with regression detection. Usually takes 60-120 seconds; if polling times out, use aethis_generation_status before retrying, and aethis_cancel_generation only when the caller wants to stop the run.
| Name | Required | Description | Default |
|---|---|---|---|
| openai_key | No | Retired and refused: Aethis LLM tools use Anthropic models only. | |
| project_id | Yes | The project ID | |
| anthropic_key | No | An Anthropic API key the user explicitly provided for this call. [sensitive — do not echo or log] Deprecated: the raw value is written verbatim to the host's session transcript. Never fill this from the environment. | |
| anthropic_key_env | No | Optional. Only honoured when it equals the env var the user configured via AETHIS_ANTHROPIC_KEY_ENV in this MCP server's config; that configured key is used automatically, so this can be omitted. Do not guess a variable name: the server refuses any name the user did not configure. | |
| anthropic_key_keychain | No | macOS keychain reference the user created for Aethis: either 'service:account' or just 'account' (service defaults to 'aethis-anthropic-key'). The server reads it via the `security` command at call time. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false and openWorldHint=true, so the safety profile is covered. The description adds genuinely non-obvious behavior: a 60-120 second runtime, synchronous polling until completion, and the timeout recovery path. It stops short of disclosing cost implications of an LLM-backed run or whether a partial/timed-out generation leaves state behind.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with no filler, front-loading what the tool does and then the operational caveats. Every clause (duration, timeout escape hatch, cancel condition) carries information an agent needs at call time.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly takes on return semantics (pass/fail plus regression detection) and runtime/recovery behavior, which is the right coverage for a long-running compound operation. It could be more complete about prerequisites (a valid Anthropic key is mandatory per the schema) or what happens on a failed run.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the parameter docs are unusually rich (key-source precedence, deprecation and refusal notes). The description adds no parameter-level meaning of its own, so the baseline 3 for a fully documented schema is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence names a specific compound action (generate rules from source text + run all test cases), and the second clarifies the mechanism (trigger, poll, run). It reads distinctly from siblings like aethis_generation_status and aethis_cancel_generation rather than blurring into them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit routing conditions: on polling timeout use aethis_generation_status before retrying, and reserve aethis_cancel_generation for a caller who wants to stop the run. Both the alternative tool and the selecting condition are stated, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_generation_statusARead-only
Check the current generation job for a project without changing it. Returns generation_contract_version, telemetry_availability, server-authoritative worker_lifecycle, retry_readiness, and the active or most recent job's progress and safe failure diagnostics. Retry only when retry_readiness is ready; an old heartbeat alone does not prove worker death. Tenant-scoped — requires an API key.
| Name | Required | Description | Default |
|---|---|---|---|
| project_id | Yes | The project ID whose generation status to inspect |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, and the description adds genuinely new context: tenant scoping plus an API key requirement, the server-authoritative nature of worker_lifecycle, and the caveat that a stale heartbeat alone does not imply worker death. That is exactly the kind of operational nuance annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the read-only purpose, then the return contract, then the retry caveat and auth requirement. No sentence is redundant, and the field enumeration is justified because no output schema exists.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by naming the returned fields (generation_contract_version, telemetry_availability, worker_lifecycle, retry_readiness, progress, safe failure diagnostics), and it covers auth and tenant scoping. Nothing an agent needs to invoke and interpret this call is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With a single required parameter and 100% schema description coverage, the schema already documents project_id fully; the description adds only that the lookup is tenant-scoped, not any syntax or format detail. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Check the current generation job for a project') and explicitly marks it as non-mutating ('without changing it'), which cleanly separates it from siblings like aethis_cancel_generation, aethis_generate_and_test, and aethis_refine. The enumerated return fields further pin down what the agent gets back.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete conditional guidance: 'Retry only when retry_readiness is ready; an old heartbeat alone does not prove worker death,' which tells the agent when a retry is justified. It does not name the sibling retry tool or state when-not-to-call this status tool, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_graphARead-only
Get the ruleset-map graph for a single published ruleset (ruleset_id) or a composed rulebook (rulebook_id) — provide exactly one. Returns {ruleset_id|rulebook_id, slug, name, graph: {nodes, edges, sections, stats}, mermaid}: each node's display.sentence / display.routes / display.expr shows how that branch composes, and mermaid is a ready-to-render diagram string. Use this to visualise or explain a ruleset's/rulebook's structure before or instead of aethis_explain. Ruleset graphs may be public (no auth for public showcase rulesets); rulebook graphs always require an API key.
| Name | Required | Description | Default |
|---|---|---|---|
| ruleset_id | No | The ID or slug of a single published ruleset. Mutually exclusive with rulebook_id. | |
| rulebook_id | No | The slug (e.g. `aethis/uk-fsm`) or opaque id (`rb_*`) of a composed rulebook. Mutually exclusive with ruleset_id. Requires an API key — anonymous callers get HTTP 401. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=true, destructiveHint=false), and the description adds material context the annotations lack: public showcase rulesets need no auth while rulebook graphs always require an API key (401 for anonymous callers). It does not, however, explain why idempotentHint is false for a pure read, which is the one behavioral oddity left unaddressed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense paragraph with the core purpose and the either/or constraint front-loaded, followed by return shape and then auth caveats. It is efficient and well-ordered, though the return-shape sentence is long and could be trimmed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing the return payload, and it does so concretely ({ruleset_id|rulebook_id, slug, name, graph: {nodes, edges, sections, stats}, mermaid}) plus what display.sentence/routes/expr convey. Combined with usage and auth guidance, an agent has everything needed to call and interpret this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters are fully documented there, including mutual exclusivity and the rb_* / slug formats. The description's "provide exactly one" restates the schema rather than adding syntax or format detail beyond it, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (get) and resource (ruleset-map graph) and explicitly names both input modes: a published ruleset via ruleset_id or a composed rulebook via rulebook_id. It also names the sibling it competes with (aethis_explain), so an agent can distinguish it without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use this to visualise or explain a ruleset's/rulebook's structure before or instead of aethis_explain" gives an explicit when-to-use and names the alternative tool. It also states the exclusivity rule ("provide exactly one") and the auth conditions per mode, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_list_domain_guidanceBRead-only
List all active guidance hints for a domain. Returns cross-section hints that apply to all projects in the domain.
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | Domain identifier (e.g. 'uk_citizenship') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is covered. The description adds that only 'active' hints are returned and that they are cross-section, which is useful context, but it omits the notable idempotentHint=false implication and any return-format detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the action and scope. Every clause earns its place with no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does carry the burden of hinting at return content, and it does so briefly ('cross-section hints that apply to all projects'). For a single-parameter read-only list tool this is nearly sufficient, though the shape of an individual hint is left unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single 'domain' parameter, which already carries the example 'uk_citizenship'. The description adds no format or constraint detail beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('List all active guidance hints for a domain') and clarifies scope as 'cross-section hints that apply to all projects in the domain.' It implicitly separates itself from the sibling aethis_list_guidance by specifying domain-level scope, but never names that sibling or contrasts them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no alternatives named. The second sentence describes what is returned rather than when an agent should reach for this tool over aethis_list_guidance or aethis_add_domain_guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_list_guidanceARead-only
List all guidance hints accumulated for a project. Shows the text, source, and active status of each hint. Use before adding new guidance to avoid duplicates.
| Name | Required | Description | Default |
|---|---|---|---|
| project_id | Yes | The project ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=true and idempotentHint=false, covering the safety profile. The description adds the returned fields, which is useful, but says nothing about pagination, ordering, or result size. Adequate but not rich against the annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: what it does, what it returns, when to use it. Front-loaded with the core action and zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description compensates by naming the returned fields (text, source, active status). For a single-parameter read tool with full annotation coverage, this is close to complete, though ordering/scope of 'all' remains unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a single documented project_id parameter, so the schema does the heavy lifting. The description only reinforces 'for a project' without adding format, ID-source, or scoping detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List all guidance hints accumulated for a project') and adds what each item exposes (text, source, active status). It does not explicitly differentiate itself from the sibling aethis_list_domain_guidance, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use before adding new guidance to avoid duplicates' gives a clear, actionable condition for when to call it, tied to the sibling aethis_add_guidance. No explicit when-not or named alternative is given, keeping it below a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_list_projectsBRead-only
List all projects in the current tenant. Returns project IDs, names, domains, and latest ruleset information.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=true, so the safety profile is covered. The description adds useful scope context ('current tenant') and enumerates the returned fields, but says nothing about pagination, result limits, or the fact that idempotentHint is false. Adds some value beyond annotations without being rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, purpose first and return contents second, with no filler. Both sentences earn their place, though the second is essentially a return-value preview that could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully previews the return fields (project IDs, names, domains, latest ruleset info). For a zero-parameter, annotation-covered list tool this is close to complete; only pagination/scope caveats are absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is no parameter semantics for the description to carry; the baseline of 4 applies. Nothing in the description misrepresents the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource ('List all projects') scoped to 'the current tenant', which cleanly distinguishes it from siblings like aethis_list_rulesets and aethis_list_rulebooks. It does not explicitly name or contrast those siblings, but the resource is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisite (e.g. tenant selection), and no routing to alternatives such as aethis_review_project or the ruleset-listing tools. Usage is only implied by the verb 'List'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_list_rulebooksARead-only
List rulebooks (composed wholes that bridge multiple rulesets) in the current tenant. Returns rulebook_id, slug (e.g. aethis/uk-fsm), name, domain, status (draft/active/archived), version, outcome_logic (the composition Expr AST), ruleset_refs, and timestamps. Use this when the user asks 'what rulebooks exist?' or to disambiguate whether several <ns>/<x>/* rulesets are bridged into one parent rulebook. Tenant-scoped — requires an API key. Pass a returned rulebook_id or slug to aethis_decide (rulebook_id arg) or aethis_rulebook_schema.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly, destructive, idempotent, and openWorld. The description adds what annotations do not: tenant scoping and the API-key requirement. It does not discuss pagination or volume limits, but with annotations carrying the safety profile this is solidly above baseline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then returns, then usage, then auth and hand-off. The long field enumeration is justified because no output schema exists, though it is denser than strictly necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description supplies the return fields, the scope/auth constraint, the triggering intent, and the follow-up tools that consume rulebook_id/slug. An agent has everything needed to call it and use the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so there is nothing for the description to disambiguate; baseline 4 applies. The enumerated return fields add value but belong to output semantics, not parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('List rulebooks') and immediately defines the resource ('composed wholes that bridge multiple rulesets'), which distinguishes it from the sibling aethis_list_rulesets without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger ('when the user asks "what rulebooks exist?"') and a second disambiguation scenario (checking whether several `<ns>/<x>/*` rulesets are bridged into one parent). It also routes forward to aethis_decide and aethis_rulebook_schema.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_list_rulesetsARead-only
List all rule rulesets for a project, including version history. Shows ruleset ID, human-readable name (the section title the ruleset covers, e.g. 'Knowledge of language and life in the UK'), status (active/archived), version, field count, and rule count.
| Name | Required | Description | Default |
|---|---|---|---|
| project_id | Yes | The project ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered. Because there is no output schema, the description carries the return-value burden and does so concretely by naming every returned field, including the non-obvious 'version history' and the active/archived status values. It leaves the unusual idempotentHint=false unaddressed, and mentions no pagination behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence states the action and scope, followed by a compact enumeration of returned fields. The inline UK citizenship example is mildly verbose but genuinely clarifying for the 'name' field, so it largely earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does the necessary work of enumerating return fields, which makes it adequate for a one-parameter read tool. The gaps are the lack of differentiation from aethis_discover_rulesets and any statement about result size or pagination.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter with 100% schema description coverage, so the schema already documents project_id fully. The description adds only the loose phrase 'for a project' and contributes no format or sourcing detail beyond the schema. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (list rulesets) and enumerates exactly what each entry contains — ID, name, status, version, field count, rule count. It does not, however, distinguish itself from the sibling aethis_discover_rulesets, which an agent could easily confuse with this tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description says the tool operates 'for a project' but gives no when-to-use guidance, no prerequisites, and no routing away from alternatives such as aethis_discover_rulesets or aethis_list_rulebooks. The agent must infer the choice from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_next_questionARead-only
Get the optimal next question for a conversational eligibility check. Call with empty field_values for the first question, then add answers and call again until decision is reached. When the ruleset author attached notes to a question (e.g. why it is asked, or legal background), they are surfaced under a Notes block.
| Name | Required | Description | Default |
|---|---|---|---|
| ruleset_id | Yes | The ID of the published rule ruleset | |
| field_values | Yes | Answers collected so far (empty dict for first question) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=false; the description adds genuine behavioral context beyond them — a stateful iterative loop keyed on accumulating answers, and the surfacing of author notes under a Notes block. No contradiction with the non-idempotent hint (each call reflects new state). It stops short of describing auth, error behavior, or the response shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with what the tool returns and then the calling protocol; the Notes-block detail is placed last as secondary information. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries more burden but handles the key termination signal ('until decision is reached') and the Notes block. It still doesn't describe the response structure (question text, options, how a final decision is represented), which an agent driving the loop would benefit from knowing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description goes further by defining the semantics of field_values over time (empty dict for the first call, then collected answers) — information the schema's field description only partially conveys. ruleset_id is left to the schema, which is acceptable given full coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource ('Get the optimal next question') and frames it within a conversational eligibility check, which separates it from decision/explanation siblings. However, it never names an alternative tool or explicitly contrasts itself with siblings like aethis_decide or aethis_explain, so sibling differentiation remains implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly states the usage pattern: call with empty field_values for the first question, then accumulate answers and call again until a decision is reached. This is strong procedural context, but it offers no explicit when-not-to-use or named alternatives, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_publishA
Publish the latest rule ruleset. Runs tests first and refuses if they fail unless force=true. Auto-deprecates previous active ruleset.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Override the human-readable section name for this ruleset. When omitted, the ruleset keeps the name set at generation time (a titlecase of section_id, e.g. 'english_language' → 'English Language'). Section names are surfaced in rulebook responses so end users can see which sections compose a rulebook. | |
| force | No | Publish even if tests are not all passing | |
| label | No | Human-readable label for this ruleset version, e.g. 'v5 — raw facts, date arithmetic'. Stored on the ruleset and shown in aethis_list_rulesets. | |
| project_id | Yes | The project ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare non-read-only, non-idempotent, non-destructive, but the description adds materially more: tests run first, failure blocks publication, force bypasses, and the previous active ruleset is auto-deprecated. The auto-deprecation side effect is exactly the kind of context annotations do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, zero waste, with the core action and its gating condition front-loaded. Every sentence carries distinct information (action, test gate, side effect).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description covers the key risk behaviors an agent needs: test gating, force override, and deprecation of the prior ruleset. What publish returns and whether the deprecation is reversible is left unstated, a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all four parameters, including the force semantics the description echoes. The description reinforces force but adds nothing about name or label beyond what the schema states, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Publish) and resource (the latest ruleset) with clear scope, and separates itself from create_ruleset/aethis_update_rulebook by operating on the 'latest' existing ruleset. It stops short of naming a sibling, so sibling differentiation is only implied.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives one concrete conditional (tests must pass unless force=true), which implies when the tool will refuse, but it never says when an agent should choose publish over create_ruleset or the rulebook update tools. Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_refineA
Refine an existing published ruleset: add optional feedback, then make the MINIMAL edit to fix failing test cases while keeping passing tests green, and re-run the full suite (seed-from-existing incremental re-authoring). Use this to fix a specific finding without re-authoring the whole section; use aethis_generate_and_test for a from-scratch rebuild.
| Name | Required | Description | Default |
|---|---|---|---|
| feedback | No | Optional correction or domain knowledge to add before regenerating | |
| openai_key | No | Retired and refused: Aethis LLM tools use Anthropic models only. | |
| project_id | Yes | The project ID | |
| anthropic_key | No | An Anthropic API key the user explicitly provided for this call. [sensitive — do not echo or log] Deprecated: the raw value is written verbatim to the host's session transcript. Never fill this from the environment. | |
| anthropic_key_env | No | Optional. Only honoured when it equals the env var the user configured via AETHIS_ANTHROPIC_KEY_ENV in this MCP server's config; that configured key is used automatically, so this can be omitted. Do not guess a variable name: the server refuses any name the user did not configure. | |
| anthropic_key_keychain | No | macOS keychain reference the user created for Aethis: either 'service:account' or just 'account' (service defaults to 'aethis-anthropic-key'). The server reads it via the `security` command at call time. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, destructiveHint=false. The description adds real behavioral context beyond that: the edit is deliberately minimal, passing tests must stay green, the full suite is re-run, and feedback is applied before regeneration. It does not say whether the call blocks or returns a job handle, which matters given the generation-status/cancel siblings.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with verb+resource and the edit constraint, then the routing rule. No filler; every clause conveys an operational fact (minimality, test-preservation, suite re-run, alternative tool).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating, open-world tool with no output schema and no idempotency guarantee, the description covers what changes but not what comes back or whether the operation is synchronous. Given sibling tools like aethis_generation_status and aethis_cancel_generation, the absence of any note about job/async behavior or result reporting leaves a real gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and all six parameters (including the auth-key variants) are documented in the schema itself, so the baseline is 3. The description only adds that feedback is optional and is applied 'before regenerating'; it says nothing about the key-resolution parameters, which is acceptable since the schema handles them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Refine an existing published ruleset') plus the exact mechanism: add optional feedback, make the MINIMAL edit to fix failing tests, re-run the full suite. The parenthetical 'seed-from-existing incremental re-authoring' further pins the operation. It is clearly distinguishable from the from-scratch sibling it names.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing guidance: 'Use this to fix a specific finding without re-authoring the whole section; use aethis_generate_and_test for a from-scratch rebuild.' That names an alternative and the condition selecting it. It does not, however, disambiguate from the very close siblings aethis_refine_sections and aethis_refine_fields, which an agent must also choose between.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_refine_fieldsA
Add guidance to improve field discovery, then re-discover. Use when fields are missing, misnamed, or enum values are incomplete. Adds a field_extraction guidance hint and re-runs discovery.
| Name | Required | Description | Default |
|---|---|---|---|
| feedback | Yes | Guidance about missing or incorrect fields (e.g., 'Section 7 implies a criminal record check') | |
| openai_key | No | Retired and refused: Aethis LLM tools use Anthropic models only. | |
| project_id | Yes | The project ID | |
| anthropic_key | No | An Anthropic API key the user explicitly provided for this call. [sensitive — do not echo or log] Deprecated: the raw value is written verbatim to the host's session transcript. Never fill this from the environment. | |
| anthropic_key_env | No | Optional. Only honoured when it equals the env var the user configured via AETHIS_ANTHROPIC_KEY_ENV in this MCP server's config; that configured key is used automatically, so this can be omitted. Do not guess a variable name: the server refuses any name the user did not configure. | |
| anthropic_key_keychain | No | macOS keychain reference the user created for Aethis: either 'service:account' or just 'account' (service defaults to 'aethis-anthropic-key'). The server reads it via the `security` command at call time. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false and destructiveHint=false, covering the mutation/safety profile. The description adds that the operation appends a 'field_extraction' guidance hint and re-runs discovery, which is useful behavioral context, but it discloses no auth requirements or reversibility detail beyond that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with the trigger condition front-loaded. The first and third sentences overlap somewhat ('Add guidance to improve field discovery' vs 'Adds a field_extraction guidance hint and re-runs discovery'), a minor redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a non-read-only, non-idempotent mutation tool with no output schema, the description conveys purpose, trigger, and the resulting side effect adequately. Remaining gaps around auth handling are covered by the well-documented key parameters in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all six parameters including the feedback example. The description adds no syntax or format detail for parameters beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: adding field-discovery guidance and re-running discovery. It reads clearly and implies a distinct role from siblings like aethis_discover_fields or aethis_add_guidance, though it does not explicitly contrast itself with those tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a clear trigger condition: 'Use when fields are missing, misnamed, or enum values are incomplete.' However, it names no alternative tool or when-not condition, so routing against siblings like aethis_validate_fields is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_refine_sectionsA
Add guidance to improve section discovery, then re-discover sections. Use when sections are missing, incorrectly split, or named differently than expected. Saves the feedback as a domain-level guidance hint and immediately re-runs discovery so you can see the effect. Repeat until the section list matches your expectations.
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | Domain identifier, e.g. 'uk_citizenship' | |
| sources | Yes | The same source documents used in the initial aethis_discover_sections call | |
| feedback | Yes | What was wrong and how to fix it, e.g. 'The english language and life in the UK test should be separate sections' | |
| openai_key | No | Retired and refused: Aethis LLM tools use Anthropic models only. | |
| anthropic_key | No | An Anthropic API key the user explicitly provided for this call. [sensitive — do not echo or log] Deprecated: the raw value is written verbatim to the host's session transcript. Never fill this from the environment. | |
| anthropic_key_env | No | Optional. Only honoured when it equals the env var the user configured via AETHIS_ANTHROPIC_KEY_ENV in this MCP server's config; that configured key is used automatically, so this can be omitted. Do not guess a variable name: the server refuses any name the user did not configure. | |
| anthropic_key_keychain | No | macOS keychain reference the user created for Aethis: either 'service:account' or just 'account' (service defaults to 'aethis-anthropic-key'). The server reads it via the `security` command at call time. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare a non-read-only, non-idempotent, open-world operation, but the description adds the crucial detail that feedback is persisted as a domain-level guidance hint and that discovery re-runs immediately. This explains the non-idempotent persistence and the iteration model, which the annotations alone do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, front-loaded with the action and then the when-to-use condition. Slight redundancy between 'immediately re-runs discovery so you can see the effect' and 'Repeat until the section list matches your expectations,' but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter mutation tool with no output schema, the description explains the iteration loop and what the re-run produces. The sensitive API-key parameters are handled thoroughly in the schema itself, so the description need not repeat them, though it could note that key configuration is required to run.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents domain, feedback, sources, and the key parameters in detail. The description restates that feedback is a 'guidance hint' and that sources are the same documents, but adds no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb sequence (add guidance, then re-discover sections) on a specific resource (sections), and the 'refine' scope is distinguishable from the sibling aethis_discover_sections. An agent can tell it apart from a plain discovery call without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit trigger conditions: 'when sections are missing, incorrectly split, or named differently than expected,' plus a termination condition ('repeat until the section list matches your expectations'). It doesn't name aethis_discover_sections as the alternative explicitly, but the context is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_review_projectARead-only
Review an authoring project against the deterministic authoring-coach rubric and get skill-building feedback. Returns a score, per-check evidence across grounding / process / lifecycle, strengths, and the single highest-leverage next improvement. Advisory only — it never blocks publishing. The deterministic report needs no LLM key; set coach=true (with an Anthropic key) to add an LLM-synthesised coaching narrative on top of the computed checks.
| Name | Required | Description | Default |
|---|---|---|---|
| coach | No | Add an opt-in LLM-synthesised coaching narrative on top of the deterministic rubric. Requires an Anthropic key the user configured (AETHIS_ANTHROPIC_KEY_ENV, anthropic_key_keychain, or anthropic_key). Off by default — the deterministic report needs no key. | |
| openai_key | No | Retired and refused: Aethis LLM tools use Anthropic models only. | |
| project_id | Yes | The project ID to review | |
| anthropic_key | No | An Anthropic API key the user explicitly provided for this call. [sensitive — do not echo or log] Deprecated: the raw value is written verbatim to the host's session transcript. Never fill this from the environment. | |
| anthropic_key_env | No | Optional. Only honoured when it equals the env var the user configured via AETHIS_ANTHROPIC_KEY_ENV in this MCP server's config; that configured key is used automatically, so this can be omitted. Do not guess a variable name: the server refuses any name the user did not configure. | |
| anthropic_key_keychain | No | macOS keychain reference the user created for Aethis: either 'service:account' or just 'account' (service defaults to 'aethis-anthropic-key'). The server reads it via the `security` command at call time. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, non-destructive, open-world, and non-idempotent. The description adds real behavioral context beyond them: it is advisory-only, the deterministic path needs no LLM key, and coach=true requires a configured Anthropic key. It does not explain why the read is non-idempotent (the LLM narrative can vary), a minor omission.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with purpose and return shape before the advisory/key caveats; no filler. Slightly dense but every sentence carries information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so: it enumerates score, per-check evidence dimensions, strengths, and the next improvement. Combined with the advisory-only note and key requirements, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema itself documents every parameter in detail, including the key-resolution chain and the retired openai_key. The description only restates the coach/key relationship, adding framing but no semantics beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (review) and resource (authoring project) plus the rubric it applies, and describes the distinctive output (score, per-check evidence across grounding/process/lifecycle, strengths, next improvement). This separates it cleanly from the sibling validate/refine/explain tools, which do not produce skill-building feedback.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Tells the agent this is advisory and 'never blocks publishing', which contrasts usefully with the gating validate_* siblings, and explains the condition for opting into coach=true. It stops short of an explicit 'use this instead of X when Y' routing statement, so it is clear context rather than full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_rulebook_schemaARead-only
Get the composition + aggregated input fields for a rulebook. Returns the outcome_logic Expr AST (how the bridged rulesets compose, e.g. A AND (B OR C)), the list of bridged rulesets (ruleset_name, ruleset_id, slug, status), and the union of all required input fields. Use this BEFORE aethis_decide on a rulebook_id to know what field_values to supply, or to inspect how a rulebook is wired. Pass a rulebook slug (e.g. aethis/uk-fsm) or opaque id (rb_*).
| Name | Required | Description | Default |
|---|---|---|---|
| rulebook_id | Yes | The slug (e.g. `aethis/uk-fsm`) or opaque id (`rb_*`) of the rulebook |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered by structured data. The description adds return-structure detail but discloses nothing behavioral beyond the annotations (no auth requirements, rate limits, or error behavior), so a baseline 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three well-ordered sentences: purpose first, return contents second, usage and parameter format third. Appropriately sized and front-loaded, with only minor redundancy in the third sentence where the slug/id format restates the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the full burden of explaining return values and does so thoroughly (AST expression, ruleset fields, union of required inputs). Combined with complete parameter documentation, an agent has everything needed to call and interpret the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the single rulebook_id parameter is already fully documented with the same slug/opaque-id examples the description repeats. The description adds no format or semantics beyond what the schema provides, so the baseline 3 holds.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (get the composition + aggregated input fields for a rulebook) and enumerates exactly what is returned (outcome_logic AST, bridged rulesets list, union of required fields), which helps separate it from siblings like aethis_list_rulebooks or aethis_schema. Sibling differentiation is present but indirect rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear condition and names an alternative: 'Use this BEFORE aethis_decide on a rulebook_id to know what field_values to supply,' plus a secondary use 'to inspect how a rulebook is wired.' No explicit when-not guidance, but the intended call context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_schemaARead-only
Get the input fields required for an eligibility check. Returns field names, types, descriptions, and allowed values. Use this before calling aethis_decide.
| Name | Required | Description | Default |
|---|---|---|---|
| ruleset_id | Yes | The ID of the published rule ruleset |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered; the description adds useful context by enumerating the shape of the returned data (names, types, descriptions, allowed values). It stops short of describing error behavior or what happens for unpublished rulesets.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the purpose, then the return shape, then the workflow instruction. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly compensates by summarizing the returned fields, which is the key information an agent needs to chain into aethis_decide. Only minor omissions remain (e.g., behavior for missing or unpublished ruleset_id).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single ruleset_id parameter is documented in the schema itself ('The ID of the published rule ruleset'). The description adds no syntax, sourcing, or format guidance beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource: it fetches the input fields required for an eligibility check, and explicitly states what it returns (field names, types, descriptions, allowed values). It is clearly tied to aethis_decide, though it does not differentiate itself from the similarly named sibling aethis_rulebook_schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit sequencing guidance: 'Use this before calling aethis_decide,' which tells the agent exactly when in the workflow to invoke it. It does not, however, state any exclusions or conditions under which this tool should be skipped.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_set_field_specAIdempotent
Store the expected field specification for a project. Once set, every aethis_discover_fields call automatically validates discovered fields against this spec. Mismatches (missing fields, wrong types, wrong enum values) generate guidance hints automatically and appear in the validation_result block. Call this BEFORE running aethis_discover_fields when the SME has already defined the field vocabulary. The spec is persisted on the project and survives across sessions. Optionally provide ordered notes for a field. Omit notes to leave existing note guidance unchanged; pass an empty list to clear it.
| Name | Required | Description | Default |
|---|---|---|---|
| project_id | Yes | The project ID | |
| expected_fields | Yes | The fields the SME expects to be discovered for this project |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the annotations by disclosing automatic downstream validation ('every aethis_discover_fields call automatically validates discovered fields against this spec'), the resulting mismatch hints in the validation_result block, persistence across sessions, and precise note-handling behavior ('Omit notes to leave existing note guidance unchanged; pass an empty list to clear it'). These are concrete behavioral consequences not captured by readOnlyHint, idempotentHint, or destructiveHint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with purpose and then adds usage and behavioral detail in a logical order. It is somewhat verbose: the automatic validation and mismatch-hint sentences overlap, and the note-handling sentence duplicates schema text, but no sentence is entirely wasted for a setter with nested configuration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a two-parameter setter with nested expected_fields and no output schema, the description covers purpose, usage timing, downstream effects, persistence, and note semantics. One notable gap remains: it does not specify whether expected_fields replaces the entire existing spec or merges with it, leaving replacement semantics ambiguous.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters and their nested fields in detail. The description's note-handling sentence largely restates the schema's own notes description ('Omit to leave existing notes unchanged; pass [] to clear them') without adding new semantic detail, so it meets the baseline for high-coverage schemas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Store the expected field specification for a project.' It clearly identifies the action and object, but does not directly differentiate the tool from sibling setters like aethis_validate_fields or aethis_refine_fields within the purpose statement itself; that differentiation comes later as an ordering instruction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says when to call: 'Call this BEFORE running aethis_discover_fields when the SME has already defined the field vocabulary.' That gives a clear context for use, but it does not state when not to use it or name alternative tools for other scenarios, so it falls short of full when/when-not/alternatives guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_set_testsADestructive
Replace the complete reviewed test suite for an existing project after field discovery. Requires 1 to 100 legacy cases, or up to 500 with contract_version: 1, and replaces prior tests without creating a project or changing its sources, fields, or guidance. This is destructive. The target API must advertise replacement support before any write. If the response is interrupted, inspect the project before approving another replacement.
| Name | Required | Description | Default |
|---|---|---|---|
| project_id | Yes | Existing project ID whose complete test suite will be replaced | |
| test_cases | Yes | The complete authoritative reviewed suite (1-500 cases); this replaces existing tests | |
| contract_version | No | Required when acceptance expectations or expected_review_bindings are supplied. | |
| expected_review_bindings | No | Optional review-binding catalogue; omit for no assertion or use {} to assert zero bindings. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=false, but the description adds what that means operationally: prior tests are replaced wholesale, cardinality limits (1-100 legacy, up to 500 with contract_version: 1), a capability gate on the target API, and a recovery instruction after interruption. This is meaningful context beyond the structured hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and scope, followed by limits, the destructive warning, the capability gate, and the recovery note in descending order of importance. Dense but each sentence carries a distinct constraint; slight clause-stacking in the second sentence keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, open-world mutation with nested objects and no output schema, the description covers prerequisites, scope of destruction, limits, and post-interruption recovery. It does not describe permission or auth requirements beyond the capability advertisement, and gives no hint about the response shape, which is a modest remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds semantics the schema does not: the 1-100 vs up-to-500 split tied to contract_version. That nuance clarifies how the two caps interact, which is not derivable from maxItems: 500 alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (replace) and resource (the complete reviewed test suite) scoped to an existing project, and explicitly bounds what it does NOT do: 'without creating a project or changing its sources, fields, or guidance.' That negation cleanly separates it from siblings like aethis_refine_fields or aethis_generate_and_test.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives real preconditions: 'after field discovery' and 'The target API must advertise replacement support before any write.' It also handles the failure path ('If the response is interrupted, inspect the project before approving another replacement'). It stops short of naming an alternative tool when replacement support is absent, so no explicit routing for that case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_update_rulebookAIdempotent
Update a Rulebook's name, description, slug, or robot_hints (assistant guidance for the conversational agent). Provide at least one field to change; omitted fields are left as-is. Requires an API key.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | New human-readable name | |
| slug | No | New stable alias | |
| description | No | New description | |
| robot_hints | No | Assistant guidance for the conversational agent, keyed by conversational beat. Natural language only — no rule syntax, no field keys. Active beats: general_context, preamble, session_start, postamble, session_end, stuck. Reserved (accepted, not yet acted on): persona, conversational_style, section_transition. An unknown beat key is rejected. Omit for no hints (create), or to leave existing hints unchanged (update). | |
| rulebook_id | Yes | The slug (e.g. `aethis/uk-fsm`) or opaque id (`rb_*`) of the rulebook to update |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true and destructiveHint=false, so the safety profile is covered. The description adds two genuine facts beyond that: it requires an API key (auth requirement) and it behaves as a PATCH, leaving omitted fields untouched. It does not say what happens on a missing rulebook_id or whether the updated rulebook is returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the mutable fields and followed by the two operational constraints. No filler; each sentence carries information, though the structure is plain prose rather than optimized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-param partial-update tool with full schema coverage and a safety annotation set, the description covers what can change, that at least one field is required, that omissions are preserved, and the auth requirement. No output schema exists, so return values need not be described; only error behavior on an invalid rulebook_id is left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented in the schema, including the detailed robot_hints beat list. The description only restates the same field names, adding the parenthetical 'assistant guidance for the conversational agent' which the schema already states. Baseline 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (update) and resource (Rulebook) and enumerates the exact mutable fields: name, description, slug, robot_hints. An agent can tell this is the mutation counterpart to aethis_create_rulebook, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Provide at least one field to change; omitted fields are left as-is' communicates partial-update semantics, which is useful invocation guidance. However, it never says when to reach for this tool versus aethis_create_rulebook or the various refine/validate tools, and gives no exclusions or preconditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_usageARead-only
Show the caller's rate-limit budget per operation class over the rolling 24h window: for each of decide / generate / author / read / keys / admin, the used count, limit, remaining, and reset time. generate (LLM rule generation) is the scarce class; browsing and status polling (read) are effectively unlimited-but-metered. Check this before a large authoring run — and report remaining generate budget to the user — so a 429 is never the first signal. Tenant-scoped — requires an API key.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/openWorld/non-destructive, and the description adds meaningfully beyond them: tenant-scoped, requires an API key, the rolling 24h window, and the relative scarcity of `generate` vs metered-but-unlimited `read`. These are behavioral facts the agent could not infer from the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded: the purpose leads, followed by field detail, then the operational trigger and constraint. No sentence is filler, despite the length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description supplies the return shape (per-class used/limit/remaining/reset), the auth requirement, and the scoping. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so the baseline of 4 applies. The description instead documents the response fields (used, limit, remaining, reset time) which is useful but not a parameter concern.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (show the caller's rate-limit budget per operation class) and enumerates the exact operation classes and fields returned. It is unmistakably distinct from siblings like aethis_generation_status or aethis_list_projects.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly prescribes when to call it ('before a large authoring run') and the downstream behavior (report remaining `generate` budget to the user) so a 429 is never the first signal. That is a clear triggering condition plus rationale.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_validate_fieldsARead-only
Assert that the discovered fields match an expected field specification. Returns a structured diff: missing fields, type mismatches, enum value mismatches, and extra fields. all_match=true only when there are no missing fields and no type or enum mismatches. Extra discovered fields do not affect all_match. Run after aethis_discover_fields to verify field coverage before writing test cases. If all_match=false, call aethis_refine_fields with guidance about the missing or incorrect fields.
| Name | Required | Description | Default |
|---|---|---|---|
| project_id | Yes | The project ID | |
| expected_fields | Yes | The fields you expect to find in the discovered field set |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnly, non-destructive, openWorld), but with no output schema the description carries the return-value burden and does so thoroughly: it enumerates the diff categories (missing, type mismatch, enum mismatch, extra) and defines the all_match condition including the edge case that extra fields do not affect it. This is substantive behavior disclosure beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Six tightly packed sentences, front-loaded with what the tool asserts and the all_match contract, followed by sequencing guidance. Each sentence carries distinct information (return shape, success condition, edge case, precondition, failure path) with no repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read-only assertion tool with no output schema, the description supplies everything an agent needs: the precondition (after discovery), the success criterion, the meaning of extra fields, and the remediation call on failure. Nothing material is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so expected_fields, key, sort, and enum_values are all documented in the schema itself. The description references an 'expected field specification' but adds no syntax or value-format guidance beyond that, so the baseline 3 for fully-covered schemas applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Assert that the discovered fields match an expected field specification') and clearly positions the tool as a verification step distinct from discovery and refinement siblings. The phrase 'discovered fields' and the field-spec language differentiate it from aethis_validate_sections without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('Run after aethis_discover_fields to verify field coverage before writing test cases') and explicit next step when it fails ('call aethis_refine_fields with guidance about the missing or incorrect fields'). Both the trigger and the alternative are named, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aethis_validate_sectionsARead-only
Compare discovered sections against an expected specification. Returns missing sections (expected but not found) and extra sections (found but not expected). Call after aethis_discover_sections to check whether the LLM found all sections the SME expects. If sections are missing, call aethis_add_domain_guidance with adherence='exact' to enforce them.
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | Domain identifier, e.g. 'uk_citizenship' | |
| expected_sections | Yes | Section names/IDs the SME expects (snake_case, e.g. ['english_language', 'residence', 'good_character']) | |
| discovered_sections | Yes | Section names/IDs returned by aethis_discover_sections |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description goes beyond that by disclosing the shape of the result (missing vs extra) and the follow-up action, though it says nothing about the odd idempotentHint=false or any cost/size behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: purpose/return shape first, then when to call, then the conditional next step. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description takes on the burden of explaining the return value and does so (missing and extra sections). Combined with annotations covering safety and the workflow branch covering next steps, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents domain, expected_sections, and discovered_sections with formats and examples. The description only implies the expected-vs-discovered distinction and adds no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (compare/validate) and resource (discovered sections vs expected specification), and names exactly what the result contains (missing and extra sections). This cleanly distinguishes it from aethis_validate_fields and from aethis_discover_sections, which it explicitly references.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit sequencing ('Call after aethis_discover_sections') and an explicit remediation branch ('If sections are missing, call aethis_add_domain_guidance with adherence="exact"'), including the argument value to pass. No inference is required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
35 tool updates
v0.22.0- First observed
aethis_add_domain_guidance - First observed
aethis_add_guidance - First observed
aethis_archive_project - First observed
aethis_archive_ruleset - First observed
aethis_cancel_generation - First observed
aethis_create_rulebook - First observed
aethis_create_ruleset - First observed
aethis_decide - First observed
aethis_discover_fields - First observed
aethis_discover_rulesets - First observed
aethis_discover_sections - First observed
aethis_explain - First observed
aethis_explain_failure - First observed
aethis_generate_and_test - First observed
aethis_generation_status - First observed
aethis_graph - First observed
aethis_list_domain_guidance - First observed
aethis_list_guidance - First observed
aethis_list_projects - First observed
aethis_list_rulebooks - First observed
aethis_list_rulesets - First observed
aethis_next_question - First observed
aethis_publish - First observed
aethis_refine - First observed
aethis_refine_fields - First observed
aethis_refine_sections - First observed
aethis_review_project - First observed
aethis_rulebook_schema - First observed
aethis_schema - First observed
aethis_set_field_spec - First observed
aethis_set_tests - First observed
aethis_update_rulebook - First observed
aethis_usage - First observed
aethis_validate_fields - First observed
aethis_validate_sections
TDQS
Scored across 35 tools
Most tools target distinct actions within the authoring/decision lifecycle, and descriptions often explicitly disambiguate similar-sounding pairs (e.g. discover_rulesets vs list_rulesets, schema vs rulebook_schema). However, with 35 tools there are several easily confused clusters: discover/refine/validate for sections and fields, add_guidance vs add_domain_guidance, and explain vs graph vs explain_failure.
All tools use the same aethis_ snake_case prefix and are mostly verb_noun (e.g. create_ruleset, list_projects, archive_project). Minor deviations include noun-only or verb-only names like aethis_schema, aethis_graph, aethis_decide, and aethis_refine, but the overall convention is predictable.
35 tools is heavy for the apparent domain, well above the 3–15 sweet spot. Many operations are split into discover/refine/validate triples, suggesting consolidation could reduce surface area without losing capability.
The surface covers a full authoring lifecycle: discovery, field/section validation, generation, testing, refinement, publishing, archiving, guidance, usage, and review. Minor gaps remain, such as no rulebook archival/deletion, no direct ruleset update tool, and no get_project or list_tests operation.
Maintenance
Related MCP Connectors
- DatagoatOAuthio.datagoat
Governed decision engine: yes/no, score, choice and rank answers about cases, from past outcomes.
Governed data discovery, exact queries, decisions, simulations, and runtime utilities over MCP.
Author rules from policy docs, then decide: a Rete engine gives the verdict, an LLM explains why.
Append-only decisions with provenance, supersession, retrieval, and audited MCP actions.
Related MCP Servers
AlicenseNot gradedqualityAmaintenanceDeterministic decision engine with DAG-based receipts. Build entity graphs, query with MCP, get auditable proof.17Apache 2.0- AlicenseNot gradedqualityAmaintenanceDeterministic MCP notary suite for subscription policies: refund, cancellation, return, and trial terms. Stateless, read-only rules engine with auditable verdicts and hosted remote endpoints.14 npm1MIT
- AlicenseNot gradedqualityBmaintenanceEnables users to trace regulatory rule changes to affected parties and required actions through deterministic safety gates, returning dated action plans and hash-linked evidence records. Supports 12 MCP tools over stdio or streamable HTTP for source comparison, obligation decomposition, scope assessment, planning, and evidence validation across multiple domain packs.MIT
- AlicenseNot gradedqualityCmaintenanceEnables agents to evaluate single records or batches with dynamically authored typed questions, returning structured decisions and probabilities.MIT