Skip to main content
Glama

Quantabble — AP Science Misconception Diagnostics

diagnose_response

Evaluate a student's written explanation in AP Chemistry, AP Physics 1, high school chemistry, high school physics, college chemistry, or college physics (algebra-based). Returns whether a misconception was detected in their reasoning, a ready-to-deliver tutor response you can say to the student verbatim, and a follow-up probe question to ask next. Use this whenever a student writes an explanation and you need to know what to say next. Do not use for single-word or number-only answers.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
schema_idNoOptional. Target a specific schema (e.g. "AP_CHEM_1_1") if you already know the topic. If omitted, the API routes automatically across all 130 schemas.
student_responseYesThe student's own words — their explanation, reasoning, or answer in free-response form. At least one complete sentence with a stated reason works best.

Schema Changelog

Changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. First observed

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses what the tool returns (misconception flag, verbatim-ready tutor response, follow-up probe) and sets expectations for input quality. It does not claim to modify state, and 'Returns' implies a read-only evaluation. This is adequate transparency for a diagnostic tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, front-loaded with the core purpose, followed by return-value preview and usage guidance. Every sentence contributes value, with no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description appropriately describes the return values and gives domain-specific guidance. It is complete enough for an agent to call the tool correctly. A minor gap is not explaining what constitutes a 'misconception' across the listed subjects, but that is not essential for invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The tool description repeats the notion of a 'written explanation' but does not add semantic details about parameters beyond what the schema already provides. It does not compensate for any schema gaps because there are none.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Evaluate') and resource ('a student's written explanation') while enumerating the exact subject domains (AP Chemistry, AP Physics 1, etc.). It also lists concrete outputs (misconception detection, tutor response, follow-up probe), making the tool's function unmistakable and distinct from the sibling get_coverage, which is about coverage rather than diagnosis.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an explicit trigger ('whenever a student writes an explanation and you need to know what to say next') and a clear exclusion ('Do not use for single-word or number-only answers'). It does not name a specific alternative tool, but the sibling get_coverage is contextually different enough that the guidance is sufficient for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

A4.3/5.0
Disambiguation5/5

get_coverage is purely a scope-checking metadata tool, while diagnose_response performs the core diagnostic workflow. Their purposes are clearly separated, so there is no realistic risk of an agent choosing the wrong one.

Naming Consistency5/5

Both tools follow the same verb_noun snake_case pattern: diagnose_response and get_coverage. The names accurately describe their actions and are fully consistent with each other.

Tool Count3/5

Two tools is on the thin side for a server, though both tools serve clear and necessary roles for a narrow diagnostics use case. The count is borderline rather than fully well-scoped.

Completeness4/5

The core workflow is covered: check coverage, diagnose a written explanation, and receive both a tutor response and follow-up probe. Minor gaps exist, such as no support for multiple-choice diagnosis or misconception reference details, but agents can work around these.

Resources