Skip to main content
Glama
Aethis-ai

aethis-mcp

Official
by Aethis-ai

aethis_refine

Fix a specific failing test case in a published ruleset with minimal edits, keeping other tests green and re-running the suite—no full re-authoring needed.

Instructions

Refine an existing published ruleset: add optional feedback, then make the MINIMAL edit to fix failing test cases while keeping passing tests green, and re-run the full suite (seed-from-existing incremental re-authoring). Use this to fix a specific finding without re-authoring the whole section; use aethis_generate_and_test for a from-scratch rebuild.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
feedbackNoOptional correction or domain knowledge to add before regenerating
openai_keyNoRetired and refused: Aethis LLM tools use Anthropic models only.
project_idYesThe project ID
anthropic_keyNoAn Anthropic API key the user explicitly provided for this call. [sensitive — do not echo or log] Deprecated: the raw value is written verbatim to the host's session transcript. Never fill this from the environment.
anthropic_key_envNoOptional. Only honoured when it equals the env var the user configured via AETHIS_ANTHROPIC_KEY_ENV in this MCP server's config; that configured key is used automatically, so this can be omitted. Do not guess a variable name: the server refuses any name the user did not configure.
anthropic_key_keychainNomacOS keychain reference the user created for Aethis: either 'service:account' or just 'account' (service defaults to 'aethis-anthropic-key'). The server reads it via the `security` command at call time.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.22.0

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, destructiveHint=false. The description adds real behavioral context beyond that: the edit is deliberately minimal, passing tests must stay green, the full suite is re-run, and feedback is applied before regeneration. It does not say whether the call blocks or returns a job handle, which matters given the generation-status/cancel siblings.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with verb+resource and the edit constraint, then the routing rule. No filler; every clause conveys an operational fact (minimality, test-preservation, suite re-run, alternative tool).

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating, open-world tool with no output schema and no idempotency guarantee, the description covers what changes but not what comes back or whether the operation is synchronous. Given sibling tools like aethis_generation_status and aethis_cancel_generation, the absence of any note about job/async behavior or result reporting leaves a real gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and all six parameters (including the auth-key variants) are documented in the schema itself, so the baseline is 3. The description only adds that feedback is optional and is applied 'before regenerating'; it says nothing about the key-resolution parameters, which is acceptable since the schema handles them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Refine an existing published ruleset') plus the exact mechanism: add optional feedback, make the MINIMAL edit to fix failing tests, re-run the full suite. The parenthetical 'seed-from-existing incremental re-authoring' further pins the operation. It is clearly distinguishable from the from-scratch sibling it names.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit routing guidance: 'Use this to fix a specific finding without re-authoring the whole section; use aethis_generate_and_test for a from-scratch rebuild.' That names an alternative and the condition selecting it. It does not, however, disambiguate from the very close siblings aethis_refine_sections and aethis_refine_fields, which an agent must also choose between.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.