Skip to main content
Glama

Catherine Ives-Yim: writing and assessments

Score an assessment

score_assessment

Runs the same deterministic rules engine the website uses for one of the assessments and returns the headline, module scores, flags and the plan. No model is involved: the result is computed from published rules, exactly as a visitor to the site would get. Answers: scale questions take 0 to 3 (worst to best anchor) or "dk" for don't know; single-choice questions take the option value; multi-select questions take an array of option values, e.g. {"CTX-markets": ["eu", "uk"]}. Unknown question ids or invalid values are rejected with an error listing them. Unanswered questions lower confidence rather than the score, and the result carries an "evidence" field: with no scored questions answered the headline is "Not enough answers", and below half confidence the headline is marked provisional. The result is a diagnostic to guide a conversation, not advice.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
packYes
answersYesQuestion id to answer.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changed
    • changedInput schema / properties / pack / enum
      Previous value: -[
      -  "ladder",
      -  "snapshot",
      -  "strategy",
      -  "cra"
      -]New value: +[
      +  "ladder",
      +  "snapshot",
      +  "strategy",
      +  "cra",
      +  "cto"
      +]
  2. First observed

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses that no model is involved, that results are deterministic, that unknown ids or invalid values are rejected with an error listing them, that unanswered questions lower confidence rather than score, and the exact headline behavior for empty and sub-half-confidence results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded: the core purpose and return shape come first, then answer formats, then validation and confidence behavior. Every sentence adds information, though the run-on answer-format sentence could be split for readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description properly names the returned fields (headline, module scores, flags, plan, evidence). For a computation tool with nested free-form answers, this is nearly complete; only the pack enum meaning is left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%, and the description compensates strongly by specifying answer formats per question type (0-3 or "dk" for scale, option value for single-choice, array for multi-select) with a concrete example. The pack enum values are left unexplained, though they are largely self-descriptive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — runs the deterministic rules engine for one assessment pack and returns headline, module scores, flags and the plan. It also distinguishes itself from the sibling list_* tools by making clear it computes a result rather than enumerating resources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage context is implied (score one of the assessments, diagnostic to guide a conversation, not advice), but there is no explicit statement of when to call this versus list_assessments or how it relates to other siblings. The agent must infer the trigger condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources