Quantabble — AP Science Misconception Diagnostics
Server Details
Misconception detection for AP Chemistry & Physics 1. 90% catch rate vs 24% baseline.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP
- URL
TDQS
Scored across 2 tools
get_coverage is purely a scope-checking metadata tool, while diagnose_response performs the core diagnostic workflow. Their purposes are clearly separated, so there is no realistic risk of an agent choosing the wrong one.
Both tools follow the same verb_noun snake_case pattern: diagnose_response and get_coverage. The names accurately describe their actions and are fully consistent with each other.
Two tools is on the thin side for a server, though both tools serve clear and necessary roles for a narrow diagnostics use case. The count is borderline rather than fully well-scoped.
The core workflow is covered: check coverage, diagnose a written explanation, and receive both a tutor response and follow-up probe. Minor gaps exist, such as no support for multiple-choice diagnosis or misconception reference details, but agents can work around these.
Available Tools
2 toolsdiagnose_responseAInspect
Evaluate a student's written explanation in AP Chemistry, AP Physics 1, high school chemistry, high school physics, college chemistry, or college physics (algebra-based). Returns whether a misconception was detected in their reasoning, a ready-to-deliver tutor response you can say to the student verbatim, and a follow-up probe question to ask next. Use this whenever a student writes an explanation and you need to know what to say next. Do not use for single-word or number-only answers.
| Name | Required | Description | Default |
|---|---|---|---|
| schema_id | No | Optional. Target a specific schema (e.g. "AP_CHEM_1_1") if you already know the topic. If omitted, the API routes automatically across all 130 schemas. | |
| student_response | Yes | The student's own words — their explanation, reasoning, or answer in free-response form. At least one complete sentence with a stated reason works best. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses what the tool returns (misconception flag, verbatim-ready tutor response, follow-up probe) and sets expectations for input quality. It does not claim to modify state, and 'Returns' implies a read-only evaluation. This is adequate transparency for a diagnostic tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, front-loaded with the core purpose, followed by return-value preview and usage guidance. Every sentence contributes value, with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description appropriately describes the return values and gives domain-specific guidance. It is complete enough for an agent to call the tool correctly. A minor gap is not explaining what constitutes a 'misconception' across the listed subjects, but that is not essential for invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The tool description repeats the notion of a 'written explanation' but does not add semantic details about parameters beyond what the schema already provides. It does not compensate for any schema gaps because there are none.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Evaluate') and resource ('a student's written explanation') while enumerating the exact subject domains (AP Chemistry, AP Physics 1, etc.). It also lists concrete outputs (misconception detection, tutor response, follow-up probe), making the tool's function unmistakable and distinct from the sibling get_coverage, which is about coverage rather than diagnosis.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an explicit trigger ('whenever a student writes an explanation and you need to know what to say next') and a clear exclusion ('Do not use for single-word or number-only answers'). It does not name a specific alternative tool, but the sibling get_coverage is contextually different enough that the guidance is sufficient for selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_coverageAInspect
Returns the list of subjects and unit counts covered by Quantabble. Use this to check whether a student's topic is within scope before calling diagnose_response.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must disclose behavioral traits. It states the output is a list (subjects and unit counts) but does not explicitly confirm it is a read-only, side-effect-free operation. While likely safe, best practice would be to state it's non-destructive. Lacks explicit behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences: first states what it returns, second gives usage guidance. Front-loaded with primary purpose, no extraneous information. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, parameterless tool with no output schema, the description is fully sufficient. It covers what the tool returns, its purpose, and its relationship to the sibling tool. No gaps identified given the complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool has no parameters and schema description coverage is 100%, so baseline is 4. The description does not need to explain parameters, but it adds no extra detail beyond what the schema already provides. Adequate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states it returns a list of subjects and unit counts, specifying the resource (Quantabble's coverage). It distinguishes from sibling tool 'diagnose_response' by explicitly linking its use as a prerequisite check, making the purpose and scope unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Directly states when to use the tool: 'Use this to check whether a student's topic is within scope before calling diagnose_response.' This provides clear context and an explicit alternative, guiding the agent on proper invocation order.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
- First observed
diagnose_response - First observed
get_coverage
Related MCP Connectors
AP study content, practice, FRQ feedback, progress, plans, and student tools for AI apps.
AI visibility checks, software recommendations and tool comparisons from measured AI answer data
AI-powered LMS course builder: 89 tools, 17 skills, SCORM/xAPI export, agentic UI
Pronunciation scoring, speech-to-text, and text-to-speech for language learning
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables AI assistants to access educational domain tools including grading, cognitive diagnosis, knowledge tracing, learning path recommendations, and sentiment analysis through FastAPI-powered small models.2-
- AlicenseBqualityDmaintenanceAn MCP server for tracking student learning via a knowledge graph, with spaced repetition scheduling and intelligent queries to identify gaps, misconceptions, and optimal next steps.71MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI assistants to manage student profiles, track assessments, calculate topic mastery, identify learning gaps, and recommend focus areas. Integrates with Claude Desktop and Claude Code for interactive learning analytics.1MIT
- AlicenseBqualityAmaintenanceA 66-tool reasoning pipeline that intercepts prompts to classify intent, check past mistakes, and generate execution plans, enabling any LLM to think harder and avoid repeating errors.31MIT