Skip to main content
Glama
putervision

agent-reasoning-mcp

by putervision

ask_score

Read-onlyIdempotent

Evaluate an entity, plan, or action on a bounded continuous scale against weighted criteria. Returns normalized score, criterion breakdown, and evaluation confidence.

Instructions

Evaluate an entity, plan, or action on a bounded continuous scale against weighted criteria. Use ask_score instead of ask_noul when evaluating continuous numeric quality or fitness rather than binary truth.

Returns normalized score within scale bounds, criterion breakdown, and evaluation confidence.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
scaleNo[min, max] range (defaults to [0.0, 1.0])
metricYesMetric name
targetYesSubject to evaluate
projectYesTarget project slug
criteriaNoEvaluation criteria
state_packNoOptional explicit StatePack

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv0.3.1

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, and non-destructive, so the safety profile is covered. The description adds valuable context beyond them by disclosing the return shape (normalized score within scale bounds, criterion breakdown, evaluation confidence), which matters since no output schema exists. It does not mention auth or rate-limit behavior, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two well-formed sentences, with the core purpose front-loaded and the sibling comparison second. No filler, though the return-value sentence could be marginally tighter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only evaluation tool with a rich 6-parameter schema and no output schema, the description covers purpose, routing, and return contents. Minor gaps remain around the optional state_pack and criteria weighting semantics, but nothing needed to call it correctly is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all six parameters are self-documented; the baseline is 3. The description reinforces 'weighted criteria' and 'scale bounds' conceptually but adds no syntax, format, or default details beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (evaluate) and resource (entity, plan, or action) with the exact evaluation mode: a bounded continuous scale against weighted criteria. It explicitly names the sibling it differs from (ask_noul), so an agent can distinguish them without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit routing rule: use ask_score instead of ask_noul when the question is continuous numeric quality/fitness rather than binary truth. This is a when-to-use plus an alternative, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.