Skip to main content
Glama

Softician Notary Exam Prep

Grade a practice answer

grade_practice_answer
Read-only

Grade a response to a practice question previously returned by get_sample_question, and return whether it was correct plus the official explanation.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
responseYesThe letter choice being graded.
questionIdYesThe questionId returned by get_sample_question.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
correctYes
explanationYes
correctChoiceYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true and destructiveHint=false, so no contradiction exists. The description adds useful behavioral context by clarifying that the tool returns correctness and the official explanation, and it also notes the dependency on a prior get_sample_question result, which goes beyond the structured metadata.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that immediately conveys the action, target, and outcome. Every phrase contributes essential information: what to grade, the source of the question, and what will be returned.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter grading tool with an output schema, the description is fully adequate. It covers the purpose, source dependency, and expected return value without unnecessary elaboration. The presence of an output schema means detailed return-field documentation is not needed in the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema provides 100% parameter coverage: questionId is described as the ID returned by get_sample_question, and response has an enum of A-D. The description adds no additional meaning to these parameters, so per the baseline for high schema coverage, a score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('grade') and identifies the exact resource ('response to a practice question') and expected outcome ('return whether it was correct plus the official explanation'). It clearly distinguishes this tool from its sibling get_sample_question, which provides sample questions rather than grading them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear contextual guidance by stating the input must be a response to a question 'previously returned by get_sample_question', which establishes the intended sequence of usage. It does not explicitly state when not to use the tool or name alternative grading tools, but the sibling relationship is evident enough for a focused practice-answer workflow.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

A4.5/5.0
Disambiguation5/5

The two tools have clearly distinct purposes: one fetches a question, the other grades an answer. There is no overlap or ambiguity between them.

Naming Consistency5/5

Both tools follow a consistent verb_noun pattern: get_sample_question and grade_practice_answer. The naming is uniform and predictable.

Tool Count4/5

With only 2 tools, the set is slightly below the typical 3-15 range, but it is appropriate for the narrow scope of question retrieval and grading. Each tool serves a necessary role in the workflow.

Completeness5/5

The tools cover the full lifecycle of the intended interaction: fetch a practice question, then grade the answer and receive an explanation. There are no obvious gaps for the server's stated purpose.