Skip to main content
Glama
woodfordintegrations-collab

k12-lesson-toolkit-standards-server

k12-lesson-toolkit-for-claude-code

Standards-grounded K-12 lesson tooling for Claude Code: an MCP standards server built from openly-licensed public data, two reference wikis, and a renderer that ships a Teacher Edition and a Student Edition a teacher can actually hand out.

*"The skills are written to make effective use of the Learning Commons Knowledge Graph Claude connector that is included for all teachers using Claude for Teachers. This connector provides Claude with access to academic standards across all 50 states and progressions beneath them; adapt the skills if your environment does not have access to these tools."*

— anthropics/k12-teacher-skills

This is that adaptation.

Not affiliated with, endorsed by, or supported by Anthropic. Built on their Apache-2.0 licensed k12-teacher-skills, which remains theirs. Anthropic's copyright and NOTICE are preserved wherever their code is vendored here.


The problem

Anthropic's two K-12 teaching skills ground their output in a Learning Commons connector that ships only inside Claude for Teachers. Outside that account the connector is absent.

The skills handle that honestly, and it is worth being precise about this rather than overselling the gap. They check whether the connector's tools are available, they state that they are fully functional without it, they forbid inventing citations, and when the connector is missing they stamp the lesson plan with a visible footer:

"Generated without the Learning Commons Knowledge Graph. Standards and misconceptions reflect general best practice."

That is good engineering. It is also a real cost. A teacher outside Claude for Teachers gets a lesson plan explicitly marked as not grounded in the standard it names — the graceful degradation, every time, forever. The disclaimer is honest precisely because the grounding is genuinely absent.

This project supplies that grounding from data anyone can ship, so the skills take their connected path instead of their fallback path. Same skills, same contract, no gated account.

Related MCP server: Shiranui

What this adds, and what each addition is for

A standards server the skills already know how to use. It registers the same seven tool names the skills look for:

find_standard_statement                     find_curriculum_lessons
find_standards_progression_from_standard    find_materials_for_lesson
find_misconceptions_for_standard            list_standards_for_mathematical_practice
find_learning_components_from_standard

The skills branch on whether those names are present, so nothing here forks, patches or shims them. Use it to: install Anthropic's k12-teacher-skills exactly as published, register this server, and get lesson plans grounded in the actual standard text instead of the not-grounded footer.

The standards themselves, in the repository. 2,303 California mathematics standards as a filtered derivative of the Learning Commons public export, CC BY 4.0, each record keeping its own license and attributionStatement. No account, no API key, no gated connector. Use it to: clone and have grounding work offline, immediately.

Published curriculum aligned to those standards. 3,301 Illustrative Mathematics lessons and 12,599 activities and assessments, CC BY 4.0. Use it to: ask "what has someone already written for this standard?" and get real lesson names back, then pull the activities inside one.

Teacher Edition and Student Edition documents. The renderer emits both as editable .docx, answer keys in one and not the other, with figures carrying screen-reader alt text — upstream's block vocabulary has no image type at all, which is workable for prose subjects and useless for geometry. Use it to: hand a student the Student Edition without hand-editing anything out.

Two reference wikis, 208 pages. docs/sourcing/ answers may I use this source? per host, with the verbatim licence sentence, URL, HTTP status and date behind every verdict — because licences expire, and two this field relies on were withdrawn inside six months during 2026 while OER indexes went on publishing the old answer. It also keeps citing, quoting and adapting apart, which is where most open-education licence trouble starts: ShareAlike and NoDerivatives do not touch citation, and quoting does not trigger ShareAlike. docs/udl/ answers does this design remove a barrier, or does it change what I am measuring? Use them to: settle a sourcing question in a minute, and as standalone reading — they are the part most likely to be useful even if you never run the server.

A slice builder for other states and subjects. The export covers 52 jurisdictions and 4 subjects. Use it to: build Texas Science or New York ELA instead of California mathematics.

Getting it running

You do not need to run any of this yourself. Clone the repository, open it in Claude Code (or any coding agent), and say:

Read SETUP.md and set this up, then verify it and tell me what works.

SETUP.md is written for the agent, not for you. It carries the install steps, the expected numbers at every stage, how to register the MCP server, and — the part that matters — the three ways this system returns a confident empty answer that looks like a real one. It ends by telling the agent exactly what to report back.

If you would rather drive it by hand, SETUP.md is perfectly readable; it is just longer than you need.

What is tested, and what is not

Status

The seven MCP tools against the shipped data

Tested. 115 automated tests, run on every change.

The shipped data reproducing from its own script

Tested. Re-extracting California mathematics reproduces data/ca-math/ record for record.

Teacher and Student .docx output

Tested. The worked example was built end to end and both editions are in examples/.

Figures on a non-macOS machine

Partly. The example was rebuilt with the macOS rasterizer disabled and matched exactly — but on macOS. Never run on Linux or Windows.

Anthropic's skills taking their connected path

Not tested end to end. The tool names and response shapes are pinned by tests/test_contract.py against the contract extracted from the skills. Nobody has watched the real skills run against this server and confirmed the footer disappears.

Any state or subject other than California mathematics

Not tested as content. The extractor is tested and builds other slices; no unit has been written from one.

Misconceptions

Empty by measurement. The upstream export contains none, anywhere. That tool returns [] and the skills fall back to their own knowledge.

Layout

Path

What

src/k12_toolkit/mcp/

the 7-tool MCP standards server

src/k12_toolkit/ingest/

builds the SQLite store from the shipped export

src/k12_toolkit/docgen/

Teacher/Student edition renderer, vendored from upstream plus a figure block

data/ca-math/

the standards export, CC BY 4.0 — see LICENSE-DATA

docs/sourcing/

reference wiki: what you may cite, quote or adapt, per host, with dates

docs/udl/

reference wiki: Universal Design for Learning as a design discipline

docs/reference/

the engine map, the export schema, the sourcing verdict, the document standard

docs/design/

why the MCP server has the shape it has

scripts/

extract_standards.py builds a slice for any state and subject; extract_curriculum.py adds the curriculum layer

tests/

the suite

SETUP.md

install and verification, written for an agent

Both wikis start at their own README.md and read as standalone references.

The worked example

The two finished documents are in examples/hs-geometry-similarity-trig/ — open those first; they are what a teacher actually receives.

hs-geometry-similarity-trig is the full source: a complete two-week grade 9-10 geometry unit built with this toolkit, CC BY 4.0, with ten lesson packages, three quizzes, a final exam, a parallel-form practice exam, 73 accessibility-validated figures, and a construct-register row for every assessable item written before the item existed.

What it cost to make

Measured from commit timestamps, not recalled:

Phase

Wall clock

Standards research, licence sweep, reference-wiki build, unit design

6h 22m

Unit build — 10 lessons, 5 instruments, 73 figures

2h 59m

Document rendering — Teacher and Student editions

1h 05m

First unit, starting from nothing

10h 26m

Output: 143 files, 285,583 words of markdown, 73 validated SVGs, two rendered editions.

A second unit is a projection, not a measurement. Only one has been built, so treat the number accordingly. It would reuse the wikis, the figure validator, the renderer, the document standard and the construct-register schema, and reuse none of the topic analysis, unit design or content. On that basis, roughly 5 hours. If that turns out wrong, the honest number will replace it here.

Building a different state or subject

California mathematics is shipped because a clone has to stay small, not because it is the only slice that works — the whole export would be roughly 280 MB. It covers 52 jurisdictions and 4 subjects, and scripts/extract_standards.py builds any pair.

It needs the Learning Commons knowledge-graph export (nodes.jsonl + relationships.jsonl, ~812 MB), which is public and not redistributed here. Ask your agent for it — "build me the Texas Science slice" — and point it at the head of scripts/extract_standards.py, which documents the one non-obvious rule: every slice also carries the Multi-State standards for its subject. That is not optional. A state's own standards hold no progression edges and reach them only through the crosswalk, so a slice built without Multi-State loads perfectly and answers every progression query with an empty list.

Re-running the script for California mathematics reproduces the shipped data/ca-math/ exactly, and --verify data/ca-math checks that — so the script is not asking to be trusted.

Honest limits, in detail

The table above is the summary. These are the five that will bite you specifically, and the first four share a shape: they return an empty list rather than an error. An empty result from this system is a boundary, not an answer about the standard you asked for.

  • A standards code must be in the exact form the store holds. HSG-SRT.C.6 returns rows. G-SRT.6 and HSG-SRT.6 return zero rows and no error — the failure most likely to read as "this standard has no content" when it means "you typed it differently".

  • Zero misconception records, and this one is not fixable here. Measured rather than assumed: across all 247,786 nodes of the export, no property key and no relationship label matches misconception, error or mistake. misconceptions.jsonl is 0 bytes for every slice, not just this one. Every misconception in the worked example was authored and cited by hand.

  • No progressions outside CCSS mathematics. All 757 buildsTowards edges in the export run between Multi-State mathematics standards. A Science or ELA slice loads and answers every other tool, and returns nothing on progressions; the extractor says so at build time rather than letting the tool look empty later.

  • Curriculum reaches a state standard by inference, not by statement. All 561 standards the curriculum aligns to are Multi-State CCSS nodes; zero are California's. A California lookup therefore travels the export's own evidence-based crosswalk to its CCSS twin, and every such result carries alignedVia naming the standard it came through. 525 of California's 1,467 standards reach curriculum this way; the other 942 reach none — most have no crosswalk edge at all. Treat an alignedVia result as a strong suggestion rather than as the publisher's own alignment.

  • Licence verdicts are dated 2026-08-07 and 2026-08-08. See the note at the foot of LICENSE-DATA about why that matters more than it looks.

Three limits listed here previously have been closed and are recorded rather than erased, because in each case the reason for the limit was wrong in an instructive way: find_curriculum_lessons and find_materials_for_lesson were stubs "because the data is not in the public export" — it was, and the join failed only on an identifier-space mismatch that returns empty without raising (docs/reference/sourcing-verdict.md). list_standards_for_mathematical_practice was a stub justified by its consumer rather than by its data; MP1 to MP8 were in the shipped export all along. And the .docx figure path was macOS-only; it now runs a backend chain and was verified end to end with QuickLook disabled.

Licence

Code is Apache-2.0 (LICENSE). Data under data/ca-math/ is CC BY 4.0 and carries a separate set of obligations, including a CCSS rider that is not Creative Commons at all (LICENSE-DATA). Third-party attribution is in NOTICE.

Available Tools

7 tools
find_curriculum_lessonsB

Curriculum lessons aligned to a standard, most directly-taught first.

Each carries its own licence and attribution string; reproduce those, not a general one.

ParametersJSON Schema
NameRequiredDescriptionDefault
authorNo
lessonNameNo
ordinalNameNo
caseIdentifierUUIDNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full disclosure burden. It adds valuable behavioral context: results are ordered 'most directly-taught first', and each lesson carries its own licence/attribution string that must be reproduced rather than replaced with a generic one.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with no filler. The core function and ordering are front-loaded, and the licensing instruction is a necessary compliance detail that earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description omits how to connect the 'standard' concept to the input parameters and gives no guidance on how this lesson-finding tool relates to the standard-focused siblings. The licensing note and output schema help, but the invocation contract is under-specified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the four undocumented parameters, but it does not explain author, lessonName, ordinalName, or caseIdentifierUUID. The phrase 'aligned to a standard' does not map clearly to any visible parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific operation (find) and resource (curriculum lessons), and adds a key selection criterion ('aligned to a standard') plus sort order. This distinguishes it from siblings like find_materials_for_lesson and find_standard_statement, though it does not name them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus the sibling tools. It implies a lesson-search use case, but provides no exclusions, prerequisites, or alternative routing, so the agent must infer when this tool is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_learning_components_from_standardA

Return up to 5 sub-skill component descriptions for the standard.

Pass caseIdentifierUUID or a bare code.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeNo
caseIdentifierUUIDNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It adds useful behavioral context: a limit of up to 5 results and the requirement to pass one of two identifier forms. It does not disclose behavior when both or neither identifiers are supplied, nor what happens on no results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler. The result limit and primary input rule are front-loaded, and every sentence contributes to using the tool correctly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter lookup with an output schema, the essential call shape is present. However, without annotations or alternative tool guidance, an agent cannot clearly determine when to choose this over the progression or misconception siblings, and the schema's lack of required parameters is not resolved by the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It names both parameters and clarifies the key semantic: they are alternatives ('or a bare code'). It stops short of specifying the exact code format, but the core meaning is clear.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear action ('Return') and resource ('sub-skill component descriptions for the standard'), including a result cap ('up to 5'). It does not explicitly contrast with sibling tools, though 'sub-skill component' is distinct enough to set it apart from misconceptions or progressions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when decomposing a standard into sub-skills, and gives an invocation pattern ('Pass caseIdentifierUUID or a bare code'). However, it does not explicitly say when to prefer this tool over siblings such as find_standards_progression_from_standard or find_misconceptions_for_standard.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_materials_for_lessonC

The activities and assessments a lesson is made of, in sequence.

ParametersJSON Schema
NameRequiredDescriptionDefault
materialSourceNo
lessonIdentifierYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It does reveal that results are ordered ('in sequence'), but it says nothing about read-only status, side effects, authentication requirements, errors, or whether it returns a list or a single structure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and contains no filler, but it is under-specified rather than appropriately concise. A noun phrase with no verb and no parameter or usage context does not serve as an effective tool definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although an output schema exists, so return structure may be covered elsewhere, the description still omits usage guidance and parameter semantics. For a simple two-parameter tool this is a noticeable completeness gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description never mentions lessonIdentifier or materialSource. It adds no meaning beyond the parameter names and types already visible in the input schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The tool name supplies the action 'find' and the description identifies a specific resource: 'the activities and assessments a lesson is made of, in sequence.' This is clear and distinguishes the tool from the standards-oriented sibling tools, though the description itself is a noun phrase rather than an explicit verb phrase.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no indication of when to use this tool versus alternatives such as find_curriculum_lessons or the standards-related tools. There are no usage conditions, exclusions, or comparisons to siblings, so the agent must infer appropriateness entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_misconceptions_for_standardC

Return common misconceptions and the teacher move for each.

Pass caseIdentifierUUID or a bare code.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeNo
subjectNo
caseIdentifierUUIDNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the return type (misconceptions and teacher moves) but says nothing about behavior: whether it requires exact matches, what happens if both code and caseIdentifierUUID are provided, whether subject is needed, or any error/edge-case behavior. The phrase 'or a bare code' hints at flexibility but leaves ambiguity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short and front-loaded with the main purpose. The second sentence gives a direct usage hint. It earns its place, though it could be slightly more structured by adding a sentence about subject or alternatives.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is an output schema, the return format is covered. However, the tool has 3 parameters with 0% schema coverage, no annotations, and no guidance on when to use this vs siblings. The description leaves the 'subject' parameter unexplained and doesn't clarify identifier precedence or requiredness. This is incomplete for reliable invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains that code or caseIdentifierUUID can be used as identifiers, but it does not explain the 'subject' parameter at all, nor does it clarify the relationship between the three parameters. With 3 undocumented parameters and only partial guidance, this is a significant gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb and resource: 'Return common misconceptions and the teacher move for each.' This distinguishes it from sibling tools that find standards, progressions, learning components, lessons, and materials. However, it doesn't explicitly name a sibling or contrast itself, so it's clear but not fully differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a usage hint: 'Pass caseIdentifierUUID or a bare code.' This implies the tool is used when you have one of those identifiers and need misconceptions. But it doesn't explain when to choose this over siblings, nor does it clarify the role of the 'subject' parameter or when to use code vs caseIdentifierUUID.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_standards_progression_from_standardB

Return the single prerequisite (backward) or next (forward) standard.

Pass caseIdentifierUUID (as the real connector does) or a bare code.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeNo
directionNo
caseIdentifierUUIDNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description carries burden. It discloses that the tool returns exactly one standard, that direction selects backward/forward, and that either identifier form is accepted. However, it doesn't disclose default direction behavior, what happens on invalid input, or whether both identifiers can be passed together. The 'as the real connector does' remark is cryptic and adds no actionable behavioral detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler; the core operation is front-loaded. The phrase 'as the real connector does' is slightly obscure but doesn't bloat the description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 3 params, no required fields, no enums, no annotations, and an output schema, the description must at least clarify critical inputs. It handles identifier choice but fails to specify the direction parameter, which is central to behavior. The cryptic connector reference and lack of direction semantics leave an agent unable to reliably select the right direction value. The output is covered by schema, so that's not a gap, but parameter completeness is insufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so description must compensate. It explains the alternative identifier forms (caseIdentifierUUID or code) but omits the third parameter, direction, entirely—despite direction being essential to the tool's purpose. It doesn't state allowed values or default behavior, leaving the agent to guess whether to pass 'backward'/'forward' or something else.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Return') and resource ('the single prerequisite ... or next ... standard'). It distinguishes the tool from siblings: siblings find statements, misconceptions, lessons, materials, and lists; this one finds progression. Even without naming siblings, the 'single standard' and 'backward/forward' phrasing makes its unique niche clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool (when you need a standard's immediate prerequisite or next standard) but doesn't explicitly exclude alternatives or name when to choose this over siblings. It gives invocation guidance ('Pass caseIdentifierUUID ... or a bare code') but that's parameter usage, not tool-selection guidance. No mention of alternatives or when not to use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_standard_statementC

Resolve standards by code (prefix) or keywords; returns statement text + UUID.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeNo
keywordsNo
jurisdictionNo
academicSubjectNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavior itself. It states the operation is a read/resolve and names the return fields, but says nothing about how the optional parameters interact, whether results are restricted or ranked, or what happens when both code and keywords are supplied. It is not misleading, just thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence with no filler; the action, inputs, and return are front-loaded. It is concise to the point of omitting needed parameter context, but as a structure it is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter tool with no annotations and 0% schema coverage, the description is too thin. It lacks usage conditions, parameter interactions, and the meaning of jurisdiction/academicSubject. The output schema covers the return shape, so that part is fine, but an agent cannot reliably choose or call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must define all four parameters. It only explains 'code' (prefix) and 'keywords'; 'jurisdiction' and 'academicSubject' are absent, leaving their role as filters or inputs unstated. Since all parameters are optional, the omission is a real gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Resolve'), a resource ('standards'), and the two lookup routes ('code (prefix)' or 'keywords'), and names the return ('statement text + UUID'). This clearly separates it from siblings that fetch progressions, misconceptions, or lessons, though it does not explicitly contrast with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description says what lookup inputs are accepted but gives no guidance on when to choose this tool over find_standards_progression_from_standard or list_standards_for_mathematical_practice, nor any conditions or exclusions. An agent must infer usage from the name and sibling set.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_standards_for_mathematical_practiceC

The eight Standards for Mathematical Practice, MP1 to MP8.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure, but it says nothing about the operation's side effects, return format, or read-only nature. It only describes the content as the eight standards, not what the tool does with them. While a list operation is implicitly read-only, the lack of any explicit behavioral statement leaves a gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one short phrase, which is highly concise and front-loaded. However, it is more of a label than a functional description; it could benefit from a verb to clearly state the action. Still, it earns points for brevity and clarity of the core subject.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the zero-parameter simplicity and the presence of an output schema, the description is minimally adequate. It identifies the content (the eight standards) but does not explain the exact return behavior or when to use this tool over alternatives. For a trivial list tool, this might be acceptable, but it could be more explicit about the action.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline is 4. The description correctly avoids adding any parameter-specific information since none exist. It does not need to compensate for schema gaps because the schema is trivially complete with no properties.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the subject (the eight Standards for Mathematical Practice) and specifies the range MP1-MP8, which adds some detail beyond the tool name. However, it lacks an explicit verb like 'lists' or 'returns', leaving the action implied by the name. This is clear enough but does not explicitly differentiate the operation from siblings that also reference standards.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus the sibling tools (find_standard_statement, find_standards_progression_from_standard, etc.). The description does not mention that this is for retrieving all standards at once, nor does it exclude scenarios where a specific standard is needed. An agent must infer usage from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 7 tool updatesv0.1.0
    • First observedfind_curriculum_lessons
    • First observedfind_learning_components_from_standard
    • First observedfind_materials_for_lesson
    • First observedfind_misconceptions_for_standard
    • First observedfind_standard_statement
    • First observedfind_standards_progression_from_standard
    • First observedlist_standards_for_mathematical_practice

TDQS

B3.3/5.0

Scored across 7 tools

Disambiguation5/5

Each tool targets a distinct aspect of the standards domain—statement resolution, progression, misconceptions, learning components, lessons, materials, and MP listing—so no two tools overlap in purpose. An agent can easily select the right tool based on the desired information.

Naming Consistency4/5

The naming pattern is highly consistent with a 'find_' prefix for most tools, and the single 'list_' for enumerating MP standards is a minor and conventional deviation. All names are in snake_case with clear verb-noun structure.

Tool Count5/5

Seven tools is well within the ideal 3-15 range and appropriately scoped for a K-12 standards server. Each tool serves a clear purpose without redundancy, making the surface feel intentional and manageable.

Completeness4/5

The tools cover the core user journey: resolving a standard, exploring related pedagogical data, and finding aligned lessons and materials. Minor gaps like bulk listing of all standards or search by grade level exist but can be worked around with existing tools.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    A
    quality
    C
    maintenance
    MCP server that exposes skills from the Brazilian National Common Curricular Base (BNCC) with thematic units, knowledge objects, and prioritization layer from Mapa de Foco, enabling lookup, search, listing, and statistics of educational skills.
    5
    14
    -
  • A
    license
    C
    quality
    C
    maintenance
    MCP server for retrieving clinical standard content from the CDISC Library, supporting controlled terminology, ADaM, SDTM, CDASH, SEND metadata, and search.
    31
    7
    MIT
  • F
    license
    Not graded
    quality
    C
    maintenance
    MCP server for the Georgia Department of Education SuitCASE API, enabling AI assistants to search, fetch, and parse Georgia learning standards, course codes, CTAE clusters, and pathways.
    -