Skip to main content
Glama
Aethis-ai

aethis-mcp

Official
by Aethis-ai

aethis_set_tests

Destructive

Replace an existing project's complete test suite after field discovery, swapping all legacy cases with reviewed ones (up to 500). Destructive: confirm the API supports replacement first.

Instructions

Replace the complete reviewed test suite for an existing project after field discovery. Requires 1 to 100 legacy cases, or up to 500 with contract_version: 1, and replaces prior tests without creating a project or changing its sources, fields, or guidance. This is destructive. The target API must advertise replacement support before any write. If the response is interrupted, inspect the project before approving another replacement.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
project_idYesExisting project ID whose complete test suite will be replaced
test_casesYesThe complete authoritative reviewed suite (1-500 cases); this replaces existing tests
contract_versionNoRequired when acceptance expectations or expected_review_bindings are supplied.
expected_review_bindingsNoOptional review-binding catalogue; omit for no assertion or use {} to assert zero bindings.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.22.0

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=false, but the description adds what that means operationally: prior tests are replaced wholesale, cardinality limits (1-100 legacy, up to 500 with contract_version: 1), a capability gate on the target API, and a recovery instruction after interruption. This is meaningful context beyond the structured hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and scope, followed by limits, the destructive warning, the capability gate, and the recovery note in descending order of importance. Dense but each sentence carries a distinct constraint; slight clause-stacking in the second sentence keeps it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, open-world mutation with nested objects and no output schema, the description covers prerequisites, scope of destruction, limits, and post-interruption recovery. It does not describe permission or auth requirements beyond the capability advertisement, and gives no hint about the response shape, which is a modest remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds semantics the schema does not: the 1-100 vs up-to-500 split tied to contract_version. That nuance clarifies how the two caps interact, which is not derivable from maxItems: 500 alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (replace) and resource (the complete reviewed test suite) scoped to an existing project, and explicitly bounds what it does NOT do: 'without creating a project or changing its sources, fields, or guidance.' That negation cleanly separates it from siblings like aethis_refine_fields or aethis_generate_and_test.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives real preconditions: 'after field discovery' and 'The target API must advertise replacement support before any write.' It also handles the failure path ('If the response is interrupted, inspect the project before approving another replacement'). It stops short of naming an alternative tool when replacement support is absent, so no explicit routing for that case.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.