Skip to main content
Glama
ginsonko

ap-aesthetics

by ginsonko

ap_tune_parameters

Read-onlyIdempotent

Evaluate a finite parameter grid against independent human labels: fit and select on train/validation, report test once, and return the candidate without adopting it.

Instructions

Evaluate an explicit finite parameter grid against independent human labels. Fit/select on train/validation, report test once, return candidate without adoption.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
datasetYes
optionsYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.2.0

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/idempotentHint/non-destructive, so safety is covered. The description adds real behavioral context beyond that: the train/validation fit-select discipline, the single test report to avoid leakage, and the explicit 'return candidate without adoption' contract, which clarifies that nothing is mutated despite being a tuning run.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly packed sentences, front-loaded with the core action and then the workflow guarantees. No filler; every clause (grid, labels, split discipline, non-adoption) earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The workflow prose is reasonably complete for a train/val/test tuning tool, and no output schema exists to explain returns. But with two fully undocumented nested-object parameters and no detail on the option schema or result shape, the definition is thin for a fairly complex grid-search operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and both parameters are undocumented nested objects. The description gestures at them ('explicit finite parameter grid' -> options, 'independent human labels' -> dataset) but never explains their expected structure or required keys, leaving the agent to guess the grid format.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb+resource: it evaluates a finite parameter grid against human labels and reports test performance. It is clearly a tuning/selection operation distinct from cataloging or applying calibrations. It does not, however, name any sibling (e.g., ap_evaluate or ap_fit_calibration) to sharpen the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied. The phrase 'return candidate without adoption' hints that a separate adoption step exists (e.g., ap_apply_calibration), but the description never states explicitly when to prefer this tool over ap_evaluate or ap_compare. An agent must infer the routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.