Skip to main content
Glama
petjal

oklo-aurora-mcp

by petjal

evaluate_physics_surrogate_tool

Predict peak cladding temperature and pressure drop via a least-squares surrogate fit, with validation metrics, extrapolation flags, and solver residuals.

Instructions

Predicts peak cladding temperature and pressure drop with a least-squares surrogate fit to the analytical T/H solver. Reports held-out validation metrics, flags extrapolation outside the training envelope, and returns the residual against the solver.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
temp_cNo
burnup_gwd_tNo
flow_rate_kg_sNo
assembly_power_kwNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.4.0

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does meaningful work: it discloses the surrogate/approximation nature, that held-out validation metrics are reported, that out-of-envelope inputs are flagged as extrapolation, and that a residual against the solver is returned. It does not state what happens when extrapolation is detected (warning vs. refusal) or any accuracy/permission limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence that front-loads the prediction, then appends reporting behaviors (metrics, extrapolation flag, residual). Every clause carries information and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return-value detail is not required, and the description is honest about the surrogate's limitations. However, for a numerical surrogate with four undocumented, zero-covered inputs, the description gives no hint of the training envelope boundaries or valid input ranges, which is the key information an agent needs to avoid triggering extrapolation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for all four parameters, so the description must compensate and it does not — no mention of temp_c, burnup_gwd_t, flow_rate_kg_s, or assembly_power_kw, their units, or their valid ranges. The parameter names themselves encode units and are self-explanatory, but the description adds no semantic value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb (predicts) and concrete outputs (peak cladding temperature, pressure drop) plus the method (least-squares surrogate fit to the analytical T/H solver), which clearly separates it from a full-solver sibling. It never names a sibling tool, so the differentiation is by technique rather than by explicit routing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Describing itself as a 'surrogate fit to the analytical T/H solver' implies the when-to-use case (fast approximate evaluation instead of the analytical solver), but there is no explicit guidance, no stated prerequisites, and no named alternative among the ten sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.