Fiveable for AP Teachers
Server Details
AP rubric grading, assignment drafts, question-bank materials, and class insights for teachers.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- thinkfiveable/fiveable-mcp
- GitHub Stars
- 0
TDQS
Scored across 77 tools
With 77 tools, many adjacent operations exist (e.g., get_assignment_status, get_assignment_question_breakdown, get_class_insights, get_reteach_gaps; get_submission, get_essay_feedback, get_grading_results; score_submission, override_submission_part, override_essay_score). Descriptions are detailed enough to distinguish most, but the volume and overlapping analytics/grading concepts still create real misselection risk.
All tool names use snake_case and follow a verb_noun or verb_noun_phrase pattern consistently (list_*, get_*, create_*, prepare_*/confirm_*, etc.). Plural/singular differences (add_students vs update_student) and a few descriptive names (get_my_teacher_workspace) are minor and do not break the convention.
77 tools is far beyond the 25+ threshold for 'too many' and is an extreme surface for a single MCP server. Even with coherent naming, the count forces agents to sift through many similar get/list/insight tools.
The server covers classes, students, assignments, grading sessions, content, insights, and Google Classroom integration with create/read/update/archive/restore operations. Minor gaps exist, such as no delete/archive operation for grading sessions and no explicit connection management for Google services.
Available Tools
77 toolsadd_studentsAdd studentsAIdempotentInspect
Adds students to a class roster from individual entries with names in First Last order and optional emails. Existing roster students are matched rather than duplicated. Emails support assignment notifications and are never returned. Roster additions can be removed.
| Name | Required | Description | Default |
|---|---|---|---|
| classId | Yes | Class id from list_classes. | |
| students | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds genuinely new behavioral context: emails are used only for assignment notifications and are never returned, and roster additions can be removed (reversibility). These are meaningful details beyond the structured hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences with no filler; the core purpose is front-loaded and each subsequent sentence conveys a distinct behavioral fact (dedup matching, email/privacy behavior, removability).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained, and the description still usefully notes emails are never returned. It covers matching, notifications, and reversibility, but omits practical constraints like the 200-item batch cap and what happens if classId is invalid.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: classId is documented ('Class id from list_classes.') and name is documented ('First Last.'), but the email parameter has no schema description. The description fills that gap by stating emails are optional, drive assignment notifications, and are never returned. It restates the First Last format from the schema but adds real meaning for the undocumented email field.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Adds students to a class roster from individual entries.' It also clarifies the roster-matching behavior, which differentiates it from remove_students/restore_students in the sibling list, though it never names an alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the phrasing ('from individual entries', 'existing roster students are matched rather than duplicated'), which tells the agent this is for bulk roster population and is safe to re-invoke. However, there is no explicit when-to-use vs when-not guidance, and no named alternative such as update_student for changing a single existing student.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
approve_essaysApprove essaysADestructiveIdempotentInspect
Marks scored essays approved (the teacher stands behind the score), or unapproves them with approve: false. Accepts essay handles, or all: true for every scored essay not yet approved. Only approved essays go to Google Classroom. The same handles with approve: false reverse approval.
| Name | Required | Description | Default |
|---|---|---|---|
| all | Yes | ||
| approve | Yes | ||
| sessionId | Yes | Grading session id from list_grading_sessions or create_grading_session. | |
| essayHandles | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, so the description's added value is the downstream consequence ('Only approved essays go to Google Classroom') and the explicit reversibility ('The same handles with approve: false reverse approval'). It doesn't explain privileges required or what happens when a handle is already in the target state, but the key side effects are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the primary action, then the selection modes, then the downstream effect and reversal. No filler or repetition of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values needn't be described. For a mutation tool with partial schema coverage, the description covers purpose, both selection modes, side effects, and reversibility; only edge-case behavior (already-approved handles, permission requirements) is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25% (only sessionId and essayHandles items are documented), so the description must carry the rest. It does explain both undocumented booleans: approve: false as the unapprove switch and all: true as the bulk selector, adding meaning beyond the bare schema defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('marks scored essays approved') plus the inverse operation, clearly distinguishing it from siblings like override_essay_score or rescore_essay. The parenthetical '(the teacher stands behind the score)' even clarifies the semantic meaning of approval.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the two invocation modes: pass essay handles, or use all: true for every scored essay not yet approved. This is clear operational guidance, though it does not name sibling tools or state when approval is a prerequisite versus optional in the workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
archive_assignmentArchive or restore draftADestructiveIdempotentInspect
Archives an unpublished draft (it leaves the coursework list; nothing is deleted) or restores an archived draft with archived: false. Published assignments are archived on the website.
| Name | Required | Description | Default |
|---|---|---|---|
| archived | Yes | false restores an archived draft. | |
| assignmentId | Yes | Assignment id from list_coursework. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true and readOnlyHint=false. The description adds real value beyond them by clarifying the nature of the 'destructive' action: it only removes the draft from the coursework list and 'nothing is deleted'. That reassurance plus the published-assignment exclusion is useful behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence covers both modes and the safety semantics, followed by one short boundary sentence. Every clause earns its place, with no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With annotations covering safety and an output schema covering return values, the description needs only to convey scope, and it does. The one residual gap is not addressing the overlapping restore_assignment_draft sibling, which an agent might reasonably expect to be routed to.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the schema itself documents both parameters (archived's 'false restores' semantics and assignmentId's source in list_coursework). The description restates the archived:false meaning but adds no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (archives/restores) on a specific resource (unpublished draft) and distinguishes the two modes via 'archived: false'. It also draws a clear boundary against published assignments, which are handled elsewhere, so an agent can tell what this tool will and will not touch.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use it (unpublished drafts) and when not (published assignments go through the website). However, it never mentions the sibling restore_assignment_draft, leaving an agent to infer whether restore should be done here via archived:false or through that dedicated tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
archive_classArchive or restore classADestructiveIdempotentInspect
Archives a class the teacher created (it leaves the class list; nothing is deleted), or restores an archived one with archived: false. Co-taught classes can only be archived by their owner.
| Name | Required | Description | Default |
|---|---|---|---|
| classId | Yes | Class id from list_classes. | |
| archived | Yes | false restores an archived class. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description discloses that archiving removes the class from the class list, that nothing is deleted, how to reverse it, and the ownership restriction for co-taught classes. These are real behavioral facts not derivable from readOnlyHint/idempotentHint/destructiveHint alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One tightly packed sentence covers both modes, the side effect, and the permission constraint with no filler. The most decision-relevant information (archive vs restore) is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and full schema coverage, the description only needs to add behavioral context, which it does (soft-hide, reversibility, owner-only for co-taught). Minor gaps remain around non-owner failure handling, but nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: classId's source (list_classes) and archived's meaning (false restores) are already documented in the schema. The description restates the archived:false semantics but adds no format or edge-case detail beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource (archive/restore a class) and clarifies the two modes of the single operation. It is clearly distinct from sibling tools like archive_assignment and update_class, so an agent can pick it without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It tells the agent how to select the restore behavior (pass archived: false) and adds an authorization prerequisite: co-taught classes can only be archived by their owner. It does not name an alternative tool or describe failure behavior when the caller is not the owner, so it stops short of fully explicit when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
confirm_correct_answer_keyCorrect an answer keyADestructiveIdempotentInspect
Applies the owner-bound answer-key correction preview identified by previewId and rescores affected attempts in the background, subject to explicit teacher approval. Rejects scores changed since the preview.
| Name | Required | Description | Default |
|---|---|---|---|
| previewId | Yes | The previewId from the matching prepare_ tool, after the teacher said yes to that preview. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructive, idempotent and open-world, but the description adds behavior they cannot express: rescoring happens in the background (asynchronous), the preview is owner-bound, and calls are rejected if scores changed since the preview (optimistic concurrency check). That is substantive disclosure beyond the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, with the operational effect front-loaded and the rejection condition as a tight second sentence. Every clause carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation. The description covers the async execution, the approval gate, and the failure mode, which is everything an agent needs to invoke this destructive, non-idempotent-looking write correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and there is only one parameter, so the schema already documents previewId's length constraints and origin. The description reinforces the owner-binding and approval conditions but adds no format or syntax detail beyond what the schema and its description supply; baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource pair: applies the answer-key correction preview and rescores affected attempts. It also scopes the operation ('owner-bound', 'identified by previewId'), which separates it cleanly from the prepare_correct_answer_key sibling that produces the preview.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description establishes the precondition ('subject to explicit teacher approval') and ties the call to a specific previewId, which implies the prepare-then-confirm workflow. It does not explicitly name prepare_correct_answer_key as the required predecessor or state when-not-to-use, but the context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
confirm_publish_assignmentPublish assignmentADestructiveIdempotentInspect
Publishes exactly the owner-bound preview identified by previewId, subject to explicit teacher approval of that specific preview. Returns each class's student link and a message to post.
| Name | Required | Description | Default |
|---|---|---|---|
| previewId | Yes | The previewId from the matching prepare_ tool, after the teacher said yes to that preview. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=true, openWorldHint=true and idempotentHint=true, so the mutating and repeat-safe nature is covered. The description adds a genuine constraint beyond that (owner-binding plus explicit teacher approval), but says nothing about downstream effects such as student notification or reversibility, which matters for a destructive-flagged publish.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler; the scoping constraint and the approval gate are front-loaded ahead of the return note. Every clause carries weight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description need not explain return values, and its mention of student links and the message to post is harmless but redundant. For a one-parameter confirm tool it is essentially complete, though it could have named the prepare sibling to close the workflow loop.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single previewId parameter is already documented as coming from 'the matching prepare_ tool, after the teacher said yes.' The description's 'owner-bound preview' adds a slight ownership nuance, but is largely redundant with the schema text, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (publishes) and a tightly scoped resource (exactly the owner-bound preview identified by previewId), which cleanly separates it from a generic publish or from the prepare step. It never names prepare_publish_assignment, so the sibling differentiation is implied rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear gating condition: the preview must be owner-bound and the teacher must have explicitly approved that specific preview. What is missing is any naming of the prepare_publish_assignment sibling that produces the previewId, leaving the agent to infer the two-step workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
confirm_push_grades_to_google_classroomSend grades to Google ClassroomADestructiveIdempotentInspect
Exports exactly the grades in the owner-bound previewId, subject to explicit teacher approval of that preview. Rejects session scores changed since the preview.
| Name | Required | Description | Default |
|---|---|---|---|
| previewId | Yes | The previewId from the matching prepare_ tool, after the teacher said yes to that preview. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructive=true, idempotent=true, openWorld=true, readOnly=false, giving the agent a strong safety profile. The description adds valuable context beyond these: it is owner-bound, requires explicit approval of a specific preview, and rejects stale session scores. However, it does not explain what happens on partial failure, rate limits, or whether rejection is atomic, which for a destructive, open-world write operation would raise this to a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action and its key precondition. Every phrase carries weight ('exactly the grades', 'owner-bound', 'explicit teacher approval', 'rejects session scores changed'). It is slightly terse and could be made more scannable with a bullet, but it is efficiently structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 1 parameter, full schema coverage, an output schema, and rich annotations, the description carries a light burden and meets it well: it clarifies the confirmation semantics, the staleness guard, and the approval requirement. The only mild gaps are the absence of a named prerequisite tool and any detail on the failure mode, but the output schema likely covers return values.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and already documents the previewId with its length bounds and the 'from the matching prepare_ tool' provenance. The description's phrase 'owner-bound previewId' adds a semantic nuance about ownership that the schema does not state, but otherwise the parameter meaning is fully carried by the schema, so the baseline 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Exports') and resource ('the grades in the owner-bound previewId') and adds a critical condition ('subject to explicit teacher approval of that preview'), clearly distinguishing this confirmation tool from the related read-only prepare tool. It stops short of naming the sibling 'prepare_push_grades_to_google_classroom' directly, but the 'previewId' parameter and the confirmation semantics make the distinction sufficiently clear for selection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the workflow: use this after a preview has been prepared and the teacher has approved it. However, it never explicitly states when to use this versus alternatives, nor does it name the required prerequisite tool by name. The condition 'after the teacher said yes' is present in the parameter schema description, not in the tool description itself, so usage guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
confirm_release_resultsHand results backADestructiveIdempotentInspect
Releases results exactly as specified by the owner-bound previewId, subject to explicit teacher approval of that specific preview. Release cannot be reversed.
| Name | Required | Description | Default |
|---|---|---|---|
| previewId | Yes | The previewId from the matching prepare_ tool, after the teacher said yes to that preview. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false. The description adds genuinely new context by stating 'Release cannot be reversed', which qualifies the irreversibility beyond the destructive flag, and reiterates the teacher-approval gate. It does not disclose rate limits or downstream effects, but it usefully supplements the annotation profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with the action and the binding identifier front-loaded, followed by the approval condition and irreversibility warning. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, non-reversible mutation with rich annotations and an output schema, the description covers the action, the gate, and the irreversibility. It is sufficient; only the explicit link to prepare_release_results as the preceding step is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already documents previewId's origin and length constraints, so the baseline is 3. The description adds a meaningful nuance by calling the previewId 'owner-bound', signaling scoping/authorization semantics not present in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('releases results') tied to a specific resource and keyed identifier ('owner-bound previewId'), which is clearly distinct from the sibling prepare_release_results. It does not explicitly name the sibling, but the action is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states a precondition ('subject to explicit teacher approval of that specific preview'), which implies this is a post-approval step, and the parameter description references the 'matching prepare_ tool'. However, it never explicitly says to use prepare_release_results first or names any exclusion, leaving the routing largely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
confirm_update_publicationChange a published assignmentADestructiveIdempotentInspect
Applies the owner-bound publication preview identified by previewId, subject to explicit teacher approval of that specific preview.
| Name | Required | Description | Default |
|---|---|---|---|
| previewId | Yes | The previewId from the matching prepare_ tool, after the teacher said yes to that preview. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true and idempotentHint=true, so the mutation/safety profile is already covered. The description adds genuinely new context beyond the annotations: the operation is 'owner-bound' and requires explicit teacher consent for that specific preview, which is important for an agent deciding whether it is authorized to fire it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the gating constraint (teacher approval) is placed where the agent will read it first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the description covers the consent gating and preview binding an agent needs. It is slightly thin on the ordering relationship with the prepare step and on failure modes (e.g., expired or already-consumed preview).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With one parameter at 100% schema description coverage, the schema already explains previewId's format and its link to the matching prepare_ tool. The description adds the 'owner-bound' scoping nuance but no format or sourcing detail beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource: applying a publication preview to change a published assignment. It is distinguishable from prepare_update_publication via the previewId binding, but it never names the sibling that produces the preview, so the workflow pairing must be inferred from the schema description rather than the prose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear invocation precondition — the preview must be applied only after 'explicit teacher approval of that specific preview' — which tells the agent when this tool is appropriate. It stops short of naming the alternative (prepare_update_publication) or stating when not to call it (e.g., no preview produced yet).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_assignment_draftCreate assignment draftAInspect
Builds an assignment draft from multiple-choice units and topics (skipping already assigned questions), Fiveable FRQ or SAQ ids, readings with optional checks, or a diagnostic. Settings start from the first class's defaults. AP Calculus drafts require courseVariant or target classes sharing one course. Returns the draft id and review link. Drafts never use an included assignment slot.
| Name | Required | Description | Default |
|---|---|---|---|
| dueAt | No | Default due time for every class. | |
| title | Yes | ||
| classes | No | Classes to target, each with an optional due date. | |
| sections | Yes | ||
| settings | No | ||
| requestId | Yes | A unique id you generate for this action. Reuse it only when retrying the same action. | |
| availableAt | No | When it opens, ISO with offset. | |
| subjectSlug | Yes | Subject, e.g. apush or ap-bio. | |
| instructions | No | ||
| courseVariant | No | AP Calculus only: ap-calc-ab or ap-calc-bc. Omit to use the course the target classes share. | |
| resultRelease | No | When students see results. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations establish the write profile (readOnlyHint=false, idempotentHint=false, destructiveHint=false), and the description adds real behavioral detail beyond them: repeat-avoidance of already published questions, settings inheritance from the first class, and the fact that drafts never consume an included assignment slot. It omits auth/permission needs and what the review link implies, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five tightly packed sentences, all front-loaded around what can be built and the key constraints; no filler. Slightly dense in the first sentence (four source types in one clause), but every sentence carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter tool with deeply nested, polymorphic sections and an output schema present (so return values needn't be re-explained), the description covers the main branches and the notable edge cases (AP Calc variant, shared-course requirement, slot behavior). Remaining gaps are minor given the rich schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 64% schema coverage, the description usefully compensates by mapping semantics onto several parameters: sections content sources, the courseVariant requirement for AP Calculus, class-derived settings, and the skip-already-assigned behavior tied to avoidRepeats. Undocumented params (title, requestId, instructions, availableAt, dueAt) are largely self-evident from their schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Builds an assignment draft") and immediately enumerates the four content sources the draft can be composed from (MC units/topics, FRQ/SAQ ids, readings, diagnostic). This clearly separates it from siblings like update_assignment_draft and duplicate_assignment without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides conditional context ("AP Calculus drafts require courseVariant or target classes sharing one course", "Settings start from the first class's defaults") but never states when to prefer this tool over update_assignment_draft or duplicate_assignment, nor any when-not conditions. Usage is implied by the create semantics rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_classesCreate classesAInspect
Creates one or more classes in the teacher's workspace. Students join through an assignment's class link or a teacher-managed roster; there are no class codes. Created classes can be archived.
| Name | Required | Description | Default |
|---|---|---|---|
| classes | Yes | ||
| requestId | Yes | A unique id you generate for this action. Reuse it only when retrying the same action. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false), so the description gets credit for the extra context it supplies: students join via an assignment link or teacher-managed roster, there are no class codes, and created classes can be archived. It still omits failure behavior, permission requirements, and the batch cap (maxItems 8 lives only in the schema).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, purpose front-loaded, zero filler; the domain context (no class codes, archivable) is packed into a single compact sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, and the creation/join/archive story is covered. For a non-idempotent mutation, however, there is no mention of permission requirements, partial-failure behavior across the batch, or what happens when a class name collides.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%; the nested class object is well documented (name, subjectSlug, courseVariant) and requestId carries the retry-id explanation, but the top-level classes array has no description. The phrase 'one or more classes' loosely mirrors the minItems/maxItems bounds but adds no syntax or format detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (creates) plus resource (classes) and scope (teacher's workspace, one or more). An agent can distinguish it from list_classes/get_class/update_class/archive_class, though it never names a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use or when-not-to-use guidance, no prerequisites (e.g. must a workspace exist first, does this require ownership), and no routing against sibling tools like add_students or update_class. The joining mechanics are domain context, not usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_grading_sessionCreate grading sessionAInspect
Creates a workspace grading session for finished student work, including paper, Google Docs or an AP Classroom export. Accepts a Fiveable FRQ or SAQ id, or the teacher's own question in prompt. A session uses an included assignment when grading starts unless the account has unlimited assignments. Returns the new session id.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | ||
| prompt | No | The teacher's own question, when not a Fiveable one. | |
| requestId | Yes | A unique id you generate for this action. Reuse it only when retrying the same action. | |
| subjectSlug | Yes | e.g. apush or ap-bio. | |
| fiveableQuestionId | No | frq:… or saq:… from search_question_bank, or frq: plus the FRQ id from get_frq. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the mutation/safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false). The description adds one genuine behavioral trait beyond them: assignment consumption ('uses an included assignment when grading starts unless the account has unlimited assignments'). It adds little else about permissions or side effects, so a middling score is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with the core purpose front-loaded, followed by input modes and quota/return notes. No filler, though the mention of return id partly duplicates the output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with a nested prompt object and an existing output schema, the description adequately covers purpose, accepted input forms, quota behavior, and return. The main gap is that it does not position this step relative to start_grading or the import tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 80%, so the baseline is 3. The description goes further by clarifying the mutual exclusivity of the two question sources ('Accepts a Fiveable FRQ or SAQ id, or the teacher's own question in prompt'), which the schema alone does not state, adding real meaning to fiveableQuestionId vs prompt.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Creates) and resource (grading session) with scope qualifiers ('for finished student work, including paper, Google Docs or an AP Classroom export'). It implicitly separates itself from start_grading by noting the session is used 'when grading starts,' though it never names the sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Context is implied ('for finished student work') but the description gives no explicit when-to-use versus alternatives such as import_essays_from_files or start_grading. The agent must infer that this is the prerequisite step before grading begins.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
duplicate_assignmentDuplicate assignmentAInspect
Copies a teacher-owned assignment, including archived or prior-year work, into a new draft with the same questions and settings. The resulting draft can be retargeted or archived.
| Name | Required | Description | Default |
|---|---|---|---|
| requestId | Yes | A unique id you generate for this action. Reuse it only when retrying the same action. | |
| assignmentId | Yes | Assignment id from list_coursework. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (not read-only, not idempotent, non-destructive), and the description adds real context beyond them: the source may be archived/prior-year and the result lands as a retargetable draft. It does not restate the retry semantics of requestId, but the added behavioral detail is genuine.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, and the most decision-relevant fact (it copies into a new draft) is front-loaded. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained. The description covers the source scope and the resulting artifact's state, which is enough for a two-parameter, well-documented tool; only the relationship to sibling draft-creation tools is left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and both parameters are documented in-schema, including that assignmentId comes from list_coursework. The description adds no parameter-level information beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (copies) and resource (a teacher-owned assignment) and describes the output (a new draft with the same questions and settings). It does not explicitly differentiate itself from close siblings like create_assignment_draft or restore_assignment_draft, which would be needed for a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'including archived or prior-year work' implies when this is appropriate (reusing existing content rather than authoring fresh), but no alternative tool is named and no condition is spelled out as when-to-use vs when-not. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_grading_resultsExport grading resultsAInspect
Creates a Google Sheet of the session's results in the teacher's own Drive (one row per student, scores by part, feedback links) and returns its link. The rows never pass through chat. Needs Google connected with Sheets access; otherwise returns the link to connect it.
| Name | Required | Description | Default |
|---|---|---|---|
| sessionId | Yes | Grading session id from list_grading_sessions or create_grading_session. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations declaring readOnlyHint=false, idempotentHint=false and openWorldHint=true, the description still adds substantial context: it creates an external artifact in the teacher's Drive, the data does not transit chat, Google must be connected with Sheets access, and the failure path returns a link to connect. That covers side effects, auth prerequisites, and degraded-mode behavior beyond what the annotations state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action and artifact, followed by the auth caveat. Every clause carries information (contents, destination, return value, prerequisite, failure behavior) with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter export tool whose annotations already declare mutability and open-world behavior, the description covers the artifact destination, contents, return value, and the Google connection prerequisite. An output schema exists, so return-shape detail is not required, and nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single sessionId parameter, and the schema already documents its origin (list_grading_sessions or create_grading_session) and its 24-char hex pattern. The description adds no further parameter meaning, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Creates a Google Sheet of the session's results') plus the exact contents of the artifact (one row per student, scores by part, feedback links) and the return value (its link). It also implicitly differentiates from the sibling get_grading_results by stressing 'The rows never pass through chat.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear context for use: it explains the side effect (Sheet in the teacher's own Drive) and the contrast with tools that return rows in chat, which helps an agent pick this over get_grading_results. It doesn't name that sibling explicitly, and it doesn't cover other alternatives, so it stops short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_assignmentGet assignmentBRead-onlyIdempotentInspect
Returns one assignment's sections, question counts, scoring and timing settings, published classes, due dates and student links. Cheap.
| Name | Required | Description | Default |
|---|---|---|---|
| assignmentId | Yes | Assignment id from list_coursework. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly, idempotent, non-destructive behavior, so the safety profile is covered. The description adds the cost hint ('Cheap.') and hints at breadth of return, but omits behavior on missing/invalid ids or what 'student links' resolve to. A modest addition beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the payload description and closing on the one differentiating cost signal. Nothing is wasted, even the fragment 'Cheap.' carries routing value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Because an output schema exists, the enumeration of return fields is largely redundant, and the description compensates with no usage routing or error behavior. It is adequate to invoke the tool but leaves the when/why gaps unfilled.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With a single parameter at 100% schema coverage, the schema fully documents assignmentId (including its origin in list_coursework). The description adds no format, validity, or fallback detail beyond it, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (get one assignment) and enumerates what it returns: sections, question counts, scoring/timing settings, published classes, due dates, student links. This distinguishes it well from thin siblings like get_assignment_status or get_assignment_question_breakdown, though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no named alternative. The trailing 'Cheap.' is the only routing signal, hinting the agent should prefer this over costlier calls, but the agent must infer the rest.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_assignment_question_breakdownGet assignment question breakdownARead-onlyIdempotentInspect
Per-question results for one assignment: percent correct and how many students picked each answer choice (identifying which distractors drew responses), plus FRQ part-by-part full-credit rates. Use it for 'which questions went worst?'. Moderate cost.
| Name | Required | Description | Default |
|---|---|---|---|
| assignmentId | Yes | Assignment id from list_coursework. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, and closed-world, so the safety profile is covered. The description adds behavior beyond that: a 'Moderate cost' hint that helps the agent budget calls, plus a precise characterization of what the payload contains (distractor-level counts, FRQ partial credit).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences that front-load the resource and metrics, then the use case, then the cost note. No filler, though the middle clause enumerating metrics is somewhat dense and slightly overlaps what the output schema likely carries.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return-value exposition is unnecessary, and the description supplies purpose, a use case, and a cost signal for a single-parameter read tool. The only missing element is routing guidance toward sibling analytics tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single parameter's description already points at list_coursework as the id source. The phrase 'for one assignment' reinforces that assignmentId scopes the call, but adds no format or validity detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource and scope ('per-question results for one assignment') plus the exact metrics returned: percent correct, per-choice selection counts, and FRQ part-by-part full-credit rates. This clearly separates it from siblings like get_assignment (metadata), get_grading_results (submission-level), and get_class_insights (class-level).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use it for which questions went worst?' gives a concrete triggering scenario, which is genuinely helpful. However, it names no alternatives and no exclusion conditions, so the agent must infer on its own whether get_grading_results or get_class_insights is the better call for adjacent analytic questions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_assignment_statusGet assignment statusARead-onlyIdempotentInspect
Where each student is on one assignment: in progress, submitted, scored, or waiting for review, with their score when results are available. Students appear by handle and short name. Use it for 'who hasn't finished?' or 'how did period 5 do?'. Cheap.
| Name | Required | Description | Default |
|---|---|---|---|
| assignmentId | Yes | Assignment id from list_coursework. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover safety (readOnly, idempotent, non-destructive, closed-world). The description adds genuine context beyond them: a cost hint ('Cheap') and output composition (students keyed by handle and short name, score only when results exist). No auth or rate-limit detail, but for a cheap read tool that is acceptable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: output shape first, then query examples, then a one-word cost signal. Nothing is redundant and the most important information (what you get back) is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained, and the description still summarizes them usefully. Annotations carry the safety profile and the single param is fully documented, so the only real gap is the absence of explicit sibling routing (e.g. versus get_grading_progress).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter, and schema coverage is 100% — the schema itself documents assignmentId and points to list_coursework as the source. The description adds no further parameter semantics, so the baseline 3 for full schema coverage applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (per-student status on a single assignment) and enumerates the exact states returned (in progress, submitted, scored, waiting for review), plus score availability. It is distinguishable from siblings like get_grading_progress by the explicit student-status framing, though it never names an alternative to route against.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides concrete usage context via example questions ('who hasn't finished?' or 'how did period 5 do?'), which makes the intended scenarios clear. It stops short of naming when-not-to-use or an explicit alternative tool, so it lands just under full marks.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_cheatsheetRead an official Fiveable cheatsheetARead-onlyIdempotentInspect
Reads bounded pages from a prepared official Fiveable PDF, returning page citations, extracted text and actual page images for formulas and diagrams. Image-only pages have no guessed transcription. Unavailable artifacts return the approved PDF link. This read never starts processing.
| Name | Required | Description | Default |
|---|---|---|---|
| page | Yes | ||
| limit | Yes | ||
| cheatsheetId | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover safety (readOnly, idempotent, non-destructive), so the bar is lower. The description adds real behavior beyond them: image-only pages are never guessed at, unavailable artifacts fall back to the approved PDF link, and the read never triggers processing. Return format specifics are left to the output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly packed sentences with no filler, front-loading what the tool does and following with fallback behavior. Slightly dense but every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be re-explained, and the safety profile comes from annotations. What is missing is prerequisite context (where cheatsheetId comes from) and any parameter guidance for a 0%-coverage schema, leaving the definition adequate but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and three parameters are required, so the description must carry the load. 'Bounded pages' gestures at page/limit semantics but never explains the cheatsheetId format (cs_ + 24 hex chars), the 1-32 page range, the 1-3 limit cap, or why limits exist. The schema alone leaves an agent guessing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Reads bounded pages from a prepared official Fiveable PDF') and specifies the distinctive output (page citations, extracted text, page images). It is clearly distinguishable from the sibling list_cheatsheets, which enumerates rather than reads content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the name and the pixel-level output described, but the description never says when to prefer this tool over siblings like get_study_guide or get_guided_notes, nor does it state that a cheatsheetId must first come from list_cheatsheets. Adequate but with a clear routing gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_classGet classARead-onlyIdempotentInspect
One class in the teacher's workspace: subject, roster (students by handle and short name), the class's assignment defaults, and the student links of assignments currently open to it. Students join a class by opening one of those links; there are no class codes. Cheap.
| Name | Required | Description | Default |
|---|---|---|---|
| classId | Yes | Class id from list_classes. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and closed-world, so the safety profile is covered. The description adds genuine domain context beyond that: what the payload contains, that students join via assignment links and 'there are no class codes', and a cost hint ('Cheap'). It stops short of error behavior for a bad classId.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence front-loads the resource and payload contents, followed by a short domain note and the terse 'Cheap.' No filler, though the telegraphic fragments are slightly choppy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity single-id read with full schema coverage, an output schema, and rich annotations, the definition covers purpose, payload shape, and a cost cue. The only real omission is explicit routing against sibling getters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
One parameter with 100% schema description coverage, including the 24-hex pattern and its origin ('Class id from list_classes'). The description adds no parameter-level meaning beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('One class in the teacher's workspace') and enumerates the payload: subject, roster with handles, assignment defaults, and student links for open assignments. This distinguishes it from list_classes and get_class_insights by scope (single class vs. collection vs. analytics).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: it is a single-item getter, and 'Cheap' hints it should be preferred over heavier alternatives, but no when/when-not or sibling routing (e.g., vs. get_class_insights or get_assignment) is stated. The classId lookup chain to list_classes lives in the schema, not the description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_class_insightsGet class insightsARead-onlyIdempotentInspect
How a class (or all the teacher's classes) is doing on Fiveable assignments: accuracy and FRQ points with change over time, each assignment's turn-in rate and average, and each student's average with flags (missing work, dropping, improving), weakest first. Scope by class, subject, assignment or time window. Moderate cost.
| Name | Required | Description | Default |
|---|---|---|---|
| days | Yes | Look back this many days. 0 means this school year. | |
| classId | No | Limit to one class (id from list_classes). | |
| subjectSlug | No | Limit to one subject, e.g. apush or ap-bio. | |
| assignmentId | No | Limit to one assignment (id from list_coursework). |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotentHint/non-destructive, so the safety profile is covered. The description adds genuinely new behavioral context beyond the annotations: a 'Moderate cost' signal and the ordering convention ('weakest first'), which help an agent decide whether and how to call it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences with the returned content front-loaded, followed by the scoping and cost note. No redundant restatement of the name or obvious filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described; the description still covers content, scope, ordering, and cost. It is essentially complete, missing only explicit routing to sibling insight tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema coverage the baseline is 3, and the description goes further by naming the four scoping dimensions (class, subject, assignment, time window) that map onto classId, subjectSlug, assignmentId, and days, reinforcing how each parameter narrows the query.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource (class insights) and enumerates exactly what is returned: accuracy and FRQ points over time, per-assignment turn-in rate and average, and per-student averages with flags. It is clearly class-scoped, implicitly distinguishing it from get_student_insights and get_skill_insights, though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the scoping dimensions ('Scope by class, subject, assignment or time window') and gives a cost signal ('Moderate cost'), which impliedly guides use. But it never states when to prefer this over get_student_insights, get_skill_insights, or get_reteach_gaps, leaving alternative selection to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_content_sectionsGet relevant study guide sectionsARead-onlyIdempotentInspect
Returns a small set of Fiveable study-guide passages relevant to one question. Suitable for narrow explanations, research questions and checking student notes. One call consumes one full-content preview for a caller without existing access.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | The concept or question the student needs help with. | |
| intent | No | Privacy-safe reason for the request. Use explain, quiz, notes, frq, or research; never send the student's raw prompt for analytics. | |
| guideSlug | No | Optional exact guide reference from get_subject_outline or search_content. | |
| subjectSlug | Yes | Fiveable subject slug, e.g. "ap-bio". A course name such as "AP Biology" also resolves. | |
| maximumTokens | Yes | ||
| maximumSections | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior, so the safety profile is covered. The description adds a genuinely useful non-obvious trait not in structured data: the quota cost ('One call consumes one full-content preview for a caller without existing access'). It does not cover caching or pagination behavior, but the budget disclosure is real added value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with no filler, front-loading what is returned before the usage hint and the cost caveat. Slightly terse rather than over-specified, but every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and annotations cover the safety profile. The description closes the main remaining gaps (intended narrow use, quota consumption) though it leaves the two undocumented numeric parameters and sibling routing unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, with query, intent, guideSlug and subjectSlug documented in-schema while maximumTokens and maximumSections carry only defaults. The description's 'one question' and 'small set' loosely gesture at the query and section limits but add no format or limits detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Returns ... Fiveable study-guide passages') and scopes it with 'relevant to one question' and 'a small set'. This is clear, but it never differentiates itself from close siblings like get_study_guide or search_content, leaving the agent to infer the distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives usable scenarios ('narrow explanations, research questions and checking student notes'), which implies a narrow-lookup role versus broad retrieval. However it names no alternative tool and states no explicit when-not condition, so the routing decision against siblings remains implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_essay_feedbackGet essay feedbackARead-onlyIdempotentInspect
One essay's score and feedback part by part, what the teacher changed, and whether it's approved. Set includeText only when the teacher wants to read the response itself.
| Name | Required | Description | Default |
|---|---|---|---|
| sessionId | Yes | Grading session id from list_grading_sessions or create_grading_session. | |
| essayHandle | Yes | Essay handle from get_grading_results, e.g. E-3b91c0. | |
| includeText | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive/closed-world behavior, so the safety profile is covered. The description adds one useful behavioral nuance—that includeText pulls the actual response text—but says nothing about provenance or size implications of that flag.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loading what is returned and then handling the one ambiguous flag. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Read-only retrieval tool with an output schema, so return shape need not be described. The includeText behavior is clarified; the only minor gap is no indication of how this relates to the other retrieval siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%; sessionId and essayHandle are documented in the schema, while includeText is an undescribed boolean. The description compensates for exactly that gap by explaining when to set includeText, adding meaning the schema lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (get) and resource (essay feedback) and enumerates the returned content: score, part-by-part feedback, teacher changes, approval status. It is clear what this returns, though it does not explicitly distinguish itself from near-siblings like get_grading_results or get_submission.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only conditional guidance is parameter-level: 'Set includeText only when the teacher wants to read the response itself.' There is no when-to-use-this-vs-alternatives framing against siblings such as get_submission or get_frq, so usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_exam_optionsFind practice exams and unit testsARead-onlyIdempotentInspect
Find Fiveable exam formats, access requirements and website links without questions or answer keys. Choose a subject for version choices; add unitNumber to verify that unit's current bank and standard format.
| Name | Required | Description | Default |
|---|---|---|---|
| unitNumber | No | ||
| subjectSlug | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent, non-destructive, and closed-world behavior. The description adds useful behavioral context beyond that: it specifies what is returned (formats, access requirements, links) and explicitly excludes questions and answer keys. It does not address auth or rate limits, but those are less critical here.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences with no wasted words. Purpose and exclusion are front-loaded, followed immediately by parameter usage guidance. Structure is clear and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a 2-parameter, optional, read-only tool with an output schema, the description is nearly complete: it states the return content, clarifies the role of both parameters, and notes the exclusion of questions/answer keys. A minor gap is that it does not specify behavior when called with no parameters, but that is not critical.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the burden. It does explain the semantic role of both parameters: subjectSlug selects a subject for version choices, and unitNumber verifies that unit's current bank and standard format. It does not repeat format constraints (pattern, max) already enforced by the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Find') and resources ('Fiveable exam formats, access requirements and website links'), and adds an explicit exclusion ('without questions or answer keys') that helps distinguish it from question-bank siblings. It does not name an alternative sibling, so differentiation is implicit rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives parameter-level guidance: use subjectSlug for version choices and add unitNumber to verify a unit's current bank and format. However, it does not state when to use this tool versus alternatives, nor when not to use it, leaving usage context implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_frqGet a free-response questionARead-onlyIdempotentInspect
Returns one full, original Fiveable FRQ prompt: stimulus, background and every lettered part with its point value. The returned prompt is eligible for Fiveable's saved scoring workflow after the student writes a response. It consumes one full-content preview unless existing access covers this FRQ. Moderate cost — 1-4k tokens.
| Name | Required | Description | Default |
|---|---|---|---|
| frqId | Yes | FRQ id from list_frqs. | |
| intent | No | Privacy-safe reason for the request. Use explain, quiz, notes, frq, or research; never send the student's raw prompt for analytics. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses a side effect: 'It consumes one full-content preview unless existing access covers this FRQ.' This contradicts idempotentHint=true, since repeated calls without access would consume additional previews. Although it also adds useful cost context, the contradiction rule requires a score of 1.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences: what is returned, why it matters for scoring, and the cost/access consequence. Every sentence earns its place and the key operation is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value details need not be repeated. The description covers the prompt contents, eligibility for scoring, preview consumption, and token cost, which is complete enough for an agent to decide whether and how to call it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with frqId described as coming from list_frqs and intent having a privacy-safe enum and instruction. The description adds no parameter-specific meaning beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Returns one full, original Fiveable FRQ prompt' and specifies the content (stimulus, background, lettered parts with point values). This clearly distinguishes it from siblings like list_frqs and get_frq_attempt_history.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear usage context: the returned prompt is intended for Fiveable's saved scoring workflow after a response is written, and the consumption warning tells the agent to consider access before calling. It does not name alternatives or state when not to use, so it lacks explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_grading_progressGet grading progressBRead-onlyIdempotentInspect
Returns grading counts (scored, remaining, failed and approved) and background-job status (running, done or stopped). Cheap.
| Name | Required | Description | Default |
|---|---|---|---|
| sessionId | Yes | Grading session id from list_grading_sessions or create_grading_session. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, non-destructive, and closed-world, so safety is covered. The description adds useful context on what is returned (counts and job status) and the 'Cheap' hint, but omits pagination, rate limits, or detailed status semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One compact sentence that front-loads the return values and ends with a useful 'Cheap.' signal. No waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given annotations fully cover safety profile and an output schema exists (so return structure need not be detailed), the description is nearly complete. It could be stronger with explicit usage guidelines, but is sufficient for a simple read-only status check.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the single sessionId parameter is fully described in the schema (including source tools). The description adds no parameter-specific meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (returns) and resource (grading counts and background-job status). It is reasonably distinct from siblings like get_grading_results, but the description does not explicitly differentiate itself from those more detailed data-fetching tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use guidance or alternatives are provided. The description says 'Cheap' but does not explain that this makes it suitable for frequent polling, nor does it mention when to use get_grading_results instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_grading_resultsGet grading resultsCRead-onlyIdempotentInspect
Class results for a grading session: average, score distribution, how the class did on each rubric part, common gaps, and a table of essays (handle, short name, score, status, what to check). Filter to the ones that need review, drafts, approved or failed. Scores marked AI draft aren't approved yet.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | Yes | ||
| filter | Yes | all | |
| offset | Yes | ||
| sessionId | Yes | Grading session id from list_grading_sessions or create_grading_session. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and a closed world, so the safety profile is fully covered. The description does add real semantics beyond that: "Scores marked AI draft aren't approved yet," which tells the agent how to read status values. It does not mention pagination or result-size behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is front-loaded with the output inventory, which is good, but that inventory substantially duplicates the existing output schema and consumes most of the text. The trailing filter fragment is choppy and the draft/approval note is buried at the end rather than structured with the filter it modifies.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Because an output schema exists, the long return-value inventory is redundant rather than necessary. What is actually missing is routing (when to use this vs export_grading_results or list_submissions_to_review) and paging behavior, so the agent can call it but has weak guidance on choosing it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25% (sessionId alone is documented). The description's filter phrasing merely restates enum values already present in the schema and adds nothing for offset or limit, whose paging semantics are the least self-evident of the three undocumented parameters. With low coverage and no compensating detail, it falls below the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It names a specific resource (class results for a grading session) and enumerates the payload: average, distribution, rubric-part breakdown, gaps, and an essay table. That clearly separates it from a single-essay tool, but no sibling is named, so the agent must infer the boundary against get_grading_progress, export_grading_results, or list_submissions_to_review.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Filter to the ones that need review, drafts, approved or failed" describes parameter values, not when to reach for this tool over its siblings. There is no stated precondition (e.g. a completed session) and no mention of export_grading_results or get_grading_progress as alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_grading_rubricGet grading rubricAInspect
Returns the session's Fiveable scoring guidelines point by point on the AP rubric for its question type. If absent, starts rubric generation and returns pending status; generation takes about a minute. Rubrics use the website's grading pipeline. Cheap once written.
| Name | Required | Description | Default |
|---|---|---|---|
| sessionId | Yes | Grading session id from list_grading_sessions or create_grading_session. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes beyond the annotations by disclosing the real side effect (a missing rubric triggers generation), the returned pending status, the ~1 minute latency, and the cost profile ('cheap once written'). This also explains why readOnlyHint=false and idempotentHint=false are set, which the structured fields alone would leave ambiguous. It does not cover auth requirements or failure modes, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with what is returned and immediately followed by the important non-obvious behavior. Phrasing like 'Cheap once written' and 'Fiveable scoring guidelines' is terse/jargon-y but each sentence carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return-value detail is unnecessary, and the description covers the tricky asynchronous path (pending status plus generation time) that an agent must handle. Missing only failure/permission context for a session that is not yet gradeable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter and schema description coverage is 100%, with the schema itself explaining the sessionId format and its origin (list_grading_sessions or create_grading_session). The description adds the notion that the rubric is scoped to the session's question type but nothing about the parameter itself, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Returns the session's ... scoring guidelines point by point on the AP rubric for its question type'), which is concrete enough to distinguish it from get_grading_results and get_grading_progress. It does not name the closest sibling, update_grading_rubric, or otherwise route the agent explicitly, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: fetch the rubric for a grading session, and the description notes the fallback when no rubric exists. There is no explicit when-to-use versus update_grading_rubric or get_grading_results, and no stated prerequisites such as needing an active grading session.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_guided_notesGet guided notesARead-onlyIdempotentInspect
AMSCO textbook guided-notes prompts with answer keys for one chapter, for subjects that have AMSCO notes (for example AP US History). Without a chapter it lists the chapters. Use for guided notes handouts and reading checks. Cheap.
| Name | Required | Description | Default |
|---|---|---|---|
| chapterSlug | No | Chapter slug from the chapter list, e.g. chapter-1-1. | |
| subjectSlug | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover safety (readOnly, idempotent, non-destructive, closed-world), so the bar is lower. The description usefully adds two behaviors annotations cannot carry: omitting a chapter returns the chapter list instead of notes, and the response includes answer keys. The trailing "Cheap" hints at cost, a genuine extra signal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the resource and scope before the conditional behavior and use case. "Cheap." is a fragment but earns its place as a cost hint; overall there is little waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return-value detail is not required. The description covers the subject constraint, the chapter-optional branch, and the intended use, leaving only the subjectSlug format undocumented — a minor gap for a two-parameter read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50% (chapterSlug documented, subjectSlug not), so the description must compensate. It explains that chapterSlug is optional and that its absence switches to chapter-listing mode, and it constrains subjectSlug to AMSCO-capable subjects. That meaningfully extends the schema, though it does not give the subject-slug format the way the schema does for chapter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb+resource: fetching AMSCO textbook guided-notes prompts with answer keys for a chapter, and scopes it to subjects that have AMSCO notes (e.g. AP US History). It is distinguishable from generic siblings like get_study_guide and get_guide_reading_questions via the AMSCO framing, though it never names an alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use for guided notes handouts and reading checks" gives a use case, and the subject restriction (only subjects with AMSCO notes) narrows applicability. However, it does not say when to prefer this over the many adjacent content tools (get_guide_reading_questions, get_study_guide, get_cheatsheet), so the guidance remains implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_guide_reading_questionsGet guide reading questionsARead-onlyIdempotentInspect
Returns reading-check questions and keys for one Fiveable guide: the curated set when available, otherwise generated questions (slower). Accepts a guide id. Suitable for exit tickets, reading quizzes and guided reading handouts.
| Name | Required | Description | Default |
|---|---|---|---|
| guideId | Yes | Study guide id. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, not open-world), yet the description adds genuinely new behavior: a curated set when available, otherwise slower generated questions. This fallback/slowness detail is real added value beyond the annotations, though it omits latency or determinism implications for the generated path.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler, front-loaded on what is returned and the fallback condition. The intended-use clause is slightly appendage-like but still earns its place by signaling purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists so return values need not be explained, and the description covers the key behavioral nuance (curated vs generated). For a one-parameter read tool this is largely complete, with only minor gaps around how the fallback affects speed or result shape.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema description coverage and a single required guideId, the schema already documents the parameter fully. The description only repeats 'Accepts a guide id' without adding format, valid ranges, or sourcing constraints, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Returns reading-check questions and keys for one Fiveable guide') and clarifies the scope is a single guide. It does not explicitly differentiate itself from siblings like get_questions_for_materials or get_study_guide, but the resource is narrow enough that an agent can distinguish it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names intended uses ('exit tickets, reading quizzes and guided reading handouts'), which implies when to reach for it, but gives no exclusions or explicit alternatives to compare against, such as get_questions_for_materials. Usage is implied rather than guided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_key_termGet a key termARead-onlyIdempotentInspect
Returns the full definition, context and related terms for one AP key term. Use when a student asks "what does X mean" for AP vocabulary. Cheap — usually under 1k tokens.
| Name | Required | Description | Default |
|---|---|---|---|
| intent | No | Privacy-safe reason for the request. Use explain, quiz, notes, frq, or research; never send the student's raw prompt for analytics. | |
| termSlug | Yes | Key term slug or the term itself, e.g. "natural-selection". | |
| subjectSlug | Yes | Fiveable subject slug, e.g. "ap-bio" or "apush". The course name or a close approximation such as "ap-biology" also resolves. Use list_subjects to find it. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds the cost/latency note ('Cheap — usually under 1k tokens') and clarifies the tool returns exactly one term's data, which is useful behavioral context beyond the annotations. No contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three short sentences with zero fluff. It front-loads the core function, then the use case, then a cost hint. Every sentence earns its place and no redundant information appears.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity, full schema coverage, and an output schema, the description covers purpose, trigger, and cost. It doesn't explicitly discuss when not to use it (e.g., for listing multiple terms) but that's a minor gap. Overall it's sufficient for an agent to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with each parameter already documented (intent enum, termSlug, subjectSlug). The description does not add parameter-specific semantics beyond what the schema provides, so the baseline 3 applies. It doesn't detract, but it also doesn't enhance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Returns the full definition, context and related terms for one AP key term.' It clearly distinguishes this from list_key_terms by emphasizing 'one AP key term' and ties to a concrete use case ('what does X mean'), making the tool's purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an explicit trigger condition: 'Use when a student asks "what does X mean" for AP vocabulary.' It does not name alternatives or when-not-to-use, but the context is clear enough to route an agent correctly. The sibling list provides obvious contrast with list_key_terms, so a 4 is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_my_teacher_workspaceGet my teacher workspaceARead-onlyIdempotentInspect
Orientation for the signed-in teacher's Fiveable workspace: classes, coursework awaiting review or due this week, included assignments remaining, and configured Google connections. Cheap.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so safety is covered. The description adds genuinely new context: the payload is scoped to the signed-in teacher (implicit auth/identity requirement, no arguments needed) and the cost hint 'Cheap' signals this is a low-latency call safe to make frequently.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single colon-led sentence front-loads the tool's identity and then lists its contents compactly, with the cost note appended in one word. Nothing is redundant and nothing needed is missing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be described, and the description adequately frames this as a zero-argument aggregate read for the current teacher. It stops just short of saying anything about freshness or whether the aggregated counts are cached.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, which is the baseline case; the description makes clear the scope is derived from the signed-in user rather than from inputs, so an agent understands no arguments are required or expected.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete aggregation ('Orientation for the signed-in teacher's Fiveable workspace') and enumerates the four content buckets it returns: classes, coursework awaiting review or due, remaining included assignments, and Google connections. That distinguishes it from narrow siblings like list_classes or list_coursework, though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The word 'Orientation' implies a session-start overview, and 'Cheap' hints it can be called liberally, but there is no explicit statement of when to use this versus list_classes/list_coursework/list_submissions_to_review, nor any exclusion criteria. Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_questions_for_materialsGet questions for materialsARead-onlyIdempotentInspect
Returns up to 40 bank questions with answer keys, explanations, complete stimuli and Fiveable attribution for classroom materials. Stimuli are required context for their questions; answer keys are teacher-only material. Counts against the content budget.
| Name | Required | Description | Default |
|---|---|---|---|
| purpose | No | What the teacher is making. Used only to improve Fiveable; changes nothing in the result. | |
| questionIds | Yes | ||
| subjectSlug | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive, closed-world behavior, so the bar is lower. The description still adds real value beyond them: the 40-item cap, that stimuli are required context, that answer keys are teacher-only material, and that the call consumes the content budget.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, zero filler, and the most decision-relevant facts (what comes back, the 40 cap, the budget cost) are front-loaded. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return formatting need not be described. Combined with the cap, budget cost, stimulus and teacher-only answer-key notes, this is nearly complete; only ID validation and error/partial-result behavior are unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%, so the description must compensate, and it partially does by surfacing the 40-item limit and the composite return. It says nothing about subjectSlug or the purpose enum, leaving a meaningful gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (returns) and resource (bank questions with answer keys, explanations, stimuli, attribution) plus a hard cap of 40. It is clearly a by-ID content-fetch tool, distinct from search_question_bank, though it never names the sibling that produces the IDs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the schema (not the description) tells the agent IDs come from search_question_bank, and the budget note hints that calls should be deliberate. There is no explicit when-to-use versus search_question_bank or get_guided_notes/get_study_guide guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_reteach_gapsGet reteach gapsARead-onlyIdempotentInspect
Ranks topics, units, CED skills and FRQ rubric rows where a class lost the most points on Fiveable assignments. Optional filters cover class, subject, assignment and time window. Moderate cost.
| Name | Required | Description | Default |
|---|---|---|---|
| days | Yes | Look back this many days. 0 means this school year. | |
| classId | No | Limit to one class (id from list_classes). | |
| subjectSlug | No | Limit to one subject, e.g. apush or ap-bio. | |
| assignmentId | No | Limit to one assignment (id from list_coursework). |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=false and destructiveHint=false, so the safety profile is covered. The description's one addition is 'Moderate cost', which is genuinely useful beyond the annotations, but it stops there — nothing about rate limits, result size or how the ranking is computed, and an output schema exists so return values are already handled elsewhere.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, purpose front-loaded, then filters, then cost. No filler; the cost sentence is arguably the only one that could be trimmed, but it carries real information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only analytics tool with a full input schema and an output schema, the description covers purpose, filters and cost adequately. The remaining gap is positioning relative to the other insight tools, which an agent picking among ~60 siblings would benefit from.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (days, classId, subjectSlug, assignmentId) is already documented in the schema. The description only paraphrases the filter set, adding no format or syntax detail beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (ranks) plus the resources ranked (topics, units, CED skills, FRQ rubric rows) and the criterion (points lost on Fiveable assignments). Clear and concrete, but it never names a sibling, which matters here because get_class_insights, get_skill_insights and get_assignment_question_breakdown sit in adjacent territory.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It conveys scope ('optional filters cover class, subject, assignment and time window') and a cost signal, implying it is used for diagnosing class-wide weakness. But there is no explicit when-to-use or when-not, and no routing to alternatives such as get_class_insights or get_skill_insights.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_skill_insightsGet skill insightsARead-onlyIdempotentInspect
Class progress on AP course skills (for example causation or contextualization): each skill's verdict (solid, improving, stuck, developing, building), average, trend, the students who need support most, and the units to revisit. Says so when evidence is limited instead of guessing. Works best scoped to one subject. Moderate cost.
| Name | Required | Description | Default |
|---|---|---|---|
| days | Yes | Look back this many days. 0 means this school year. | |
| classId | No | Limit to one class (id from list_classes). | |
| subjectSlug | No | Limit to one subject, e.g. apush or ap-bio. | |
| assignmentId | No | Limit to one assignment (id from list_coursework). |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), so the description's extra value is in disclosing output shape, the graceful "says so when evidence is limited" behavior, and a cost signal. These are genuinely useful beyond the annotations, though "moderate cost" is vague with no unit or threshold.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose, then the return contents, then behavioral caveats and scoping advice. Dense but every clause earns its place; no filler or repetition of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the description needn't enumerate return values, yet it does so helpfully, and it adds the degradation behavior and cost note. What remains thin is the relationship to the other insight tools, which matters given three near-siblings in the toolset.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with clear per-parameter docs (days=0 means school year, classId from list_classes, etc.), so the schema carries the burden. The description only reinforces subject scoping, adding little beyond what the schema already provides — baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb+resource (class progress on AP course skills) and enumerates the outputs — per-skill verdict, average, trend, at-risk students, units to revisit — so the agent knows exactly what it returns. It differentiates implicitly by granularity from get_class_insights and get_student_insights, but never names those siblings to make the distinction explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Works best scoped to one subject" and "Moderate cost" give practical usage guidance, but there is no statement of when to prefer this over get_class_insights, get_student_insights, or get_reteach_gaps, and no when-not conditions. The agent is left to infer the routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_student_insightsGet student insightsARead-onlyIdempotentInspect
Returns one student's Fiveable work: average, missing work, assignment scores, weakest units and topics, and skill progress. Accepts an anonymous student handle (S-…). Suitable for conference preparation, parent notes and individual reteaching. Moderate cost.
| Name | Required | Description | Default |
|---|---|---|---|
| days | Yes | Look back this many days. 0 means this school year. | |
| handle | Yes | The student's handle, e.g. S-4f2a1c. | |
| classId | No | Limit to one class (id from list_classes). | |
| subjectSlug | No | Limit to one subject, e.g. apush or ap-bio. | |
| assignmentId | No | Limit to one assignment (id from list_coursework). |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, closed-world behavior, so the safety profile is covered. The description adds a genuinely useful trait beyond the schema: 'Moderate cost', which helps the agent budget calls. It also reinforces the anonymous-handle requirement, though it omits pagination/return-shape notes (partly covered by the output schema).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the returned content so an agent sees the payload scope first, then the input, then the cost. No filler, though the handle restatement is slightly redundant with the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the return values need not be described, and the description still names the key outputs, the input format, and the cost. It is complete enough to call correctly; only explicit sibling routing is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all five parameters, including the S-… handle pattern and the days semantics. The description's 'Accepts an anonymous student handle (S-…)' merely restates schema content, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Returns one student's Fiveable work') and enumerates the concrete outputs (average, missing work, assignment scores, weakest units, skill progress). The scoping word 'one student's' implicitly distinguishes it from the sibling get_class_insights, though it never names the sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides use contexts ('conference preparation, parent notes and individual reteaching'), which implies when the tool is appropriate. However it gives no explicit alternatives (e.g. get_class_insights for aggregate views) and no exclusions, leaving the agent to infer routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_study_guideGet a study guideARead-onlyIdempotentInspect
Returns the full text of one Fiveable AP study guide as Markdown. Accepts the exact canonical guide reference or a close approximation of its title. Suitable when a student asks to learn or review a topic. It consumes one full-content preview for a free caller unless existing access covers this subject and feature. Expensive — a full guide can be 3-8k tokens.
| Name | Required | Description | Default |
|---|---|---|---|
| intent | No | Privacy-safe reason for the request. Use explain, quiz, notes, frq, or research; never send the student's raw prompt for analytics. | |
| guideSlug | Yes | Canonical guide reference from get_subject_outline, content slug, document id, or title, e.g. "/ap-bio/unit-2/cell-structure/study-guide/abc123" or "cell-structure". | |
| subjectSlug | Yes | Fiveable subject slug, e.g. "ap-bio" or "apush". The course name or a close approximation such as "ap-biology" also resolves. Use list_subjects to find it. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint, idempotentHint, and destructiveHint=false. The description adds valuable behavioral context beyond annotations: it consumes one full-content preview for free callers unless access exists, and it warns that a full guide is expensive at 3-8k tokens. This significantly helps an agent anticipate cost and access side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core behavior, followed by suitability and cost context. Every sentence contributes useful information, though it is slightly longer than strictly necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a rich input schema, annotations, and an output schema present, the description covers the important non-obvious context: Markdown output, access/preview consumption, token cost, and when to use the tool. It is complete enough for an agent to invoke this tool appropriately.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds only marginal semantic value over the schema, such as noting that guideSlug accepts a close approximation of a title, but the schema already documents canonical references and title matching.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb and resource: it returns the full text of one Fiveable AP study guide as Markdown. It effectively differentiates this tool from siblings like get_frq or get_cheatsheet by emphasizing 'full text' and 'Markdown', though it does not explicitly name a sibling alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: when a student asks to learn or review a topic. It does not state exclusions or name alternatives like get_cheatsheet for concise review, so it falls just short of full guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_subject_outlineGet subject outlineARead-onlyIdempotentInspect
Returns the full unit and topic outline for one AP subject, including every study guide slug in each unit. Provides the course map for study planning and individual guide retrieval. Moderate cost — roughly 2-5k tokens depending on subject size.
| Name | Required | Description | Default |
|---|---|---|---|
| intent | No | Privacy-safe reason for the request. Use explain, quiz, notes, frq, or research; never send the student's raw prompt for analytics. | |
| subjectSlug | Yes | Fiveable subject slug, e.g. "ap-bio" or "apush". The course name or a close approximation such as "ap-biology" also resolves. Use list_subjects to find it. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds useful behavioral context: the tool returns a complete outline with study guide slugs and gives a token-cost estimate. This goes beyond the structured annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three front-loaded sentences, each earning its place: what it returns, how it is used, and cost considerations. No filler or repetition of schema fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the rich annotations and presence of an output schema, the description is sufficiently complete for selecting and invoking the tool. It covers scope, purpose, and cost. It could have explicitly contrasted itself with get_unit_overview, but the schema already provides subject resolution guidance, so nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both intent and subjectSlug, including examples and the suggestion to use list_subjects. The description adds no parameter-specific semantics, but the baseline of 3 is appropriate because the schema carries the full burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action (returns the full unit and topic outline) and a resource (one AP subject), and distinguishes it from siblings like get_unit_overview and get_study_guide by emphasizing the complete course map and study guide slugs. It is clear what this tool produces and how it differs from adjacent tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies clear usage context: use it for study planning and to locate individual study guide slugs before retrieving guides. It does not explicitly name alternatives or exclusions, but the stated use cases are specific enough to route an agent appropriately among the sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_submissionGet submissionARead-onlyIdempotentInspect
One student's submission on a Fiveable assignment: each section's AI and teacher scores, the FRQ rubric breakdown, overrides and feedback, and whether the student can see results yet. includeText requires that the teacher wants to read the student's responses.
| Name | Required | Description | Default |
|---|---|---|---|
| includeText | Yes | ||
| assignmentId | Yes | Assignment id from list_coursework or list_submissions_to_review. | |
| submissionId | Yes | Submission id from list_submissions_to_review. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint, idempotentHint, destructiveHint=false). The description adds some behavior beyond that: it discloses result-visibility status and the meaning/condition for includeText. Still, no mention of pagination, auth/prerequisite, or side effects, so it only modestly exceeds the annotation baseline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler, and the includeText caveat is placed last where it logically belongs. The enumerated return contents partly duplicate the output schema, slightly reducing the density of value per sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the return details needn't be explained, and the description still pairs the parameter gap it uniquely fills (includeText). An agent has enough to call it correctly; only cross-tool routing guidance is thin.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%; the two id parameters are documented in the schema as coming from list_coursework/list_submissions_to_review. The description compensates for the undocumented includeText flag by explaining its semantics, adding real value beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('One student's submission') and enumerates what it returns (section scores, FRQ rubric breakdown, overrides, feedback, visibility). The singular framing clearly distinguishes it from the plural sibling list_submissions_to_review, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives one conditional usage note for includeText ('requires that the teacher wants to read the student's responses'), which implies when to set it. However it never states when to prefer this tool over alternatives like get_grading_results, so usage is only partially specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_unit_overviewGet unit overviewARead-onlyIdempotentInspect
Returns the overview page for one unit of an AP course: what it covers, how it is weighted on the exam, its topics and study plan. Suitable when a student is reviewing a whole unit. Moderate cost — roughly 1-3k tokens.
| Name | Required | Description | Default |
|---|---|---|---|
| intent | No | Privacy-safe reason for the request. Use explain, quiz, notes, frq, or research; never send the student's raw prompt for analytics. | |
| unitSlug | Yes | Unit slug from get_subject_outline, e.g. "unit-1". | |
| subjectSlug | Yes | Fiveable subject slug, e.g. "ap-bio" or "apush". The course name or a close approximation such as "ap-biology" also resolves. Use list_subjects to find it. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior. The description adds valuable cost information ('Moderate cost — roughly 1-3k tokens'), which goes beyond the structured annotations. No contradictions found.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The core purpose is front-loaded, followed by a concise use-case and cost note. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, idempotent tool with a rich input schema and output schema, the description fully covers purpose, content of the return value, use context, and cost. Nothing essential is missing for an agent to decide whether to call it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema already documents all parameters well, including examples and resolution behavior. The description does not need to add parameter semantics and doesn't, which is acceptable at the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Returns the overview page for one unit of an AP course.' It clearly enumerates what the overview contains (coverage, exam weighting, topics, study plan), making it distinguishable from sibling tools like get_subject_outline or get_study_plan.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Suitable when a student is reviewing a whole unit' provides clear context for when to invoke this tool. It does not explicitly name alternatives or exclusions, stopping short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
import_essays_from_filesImport essays from filesAInspect
Splits teacher-supplied PDFs or Word stacks into one essay per student, including AP Classroom 'FRQ Response PDFs, By Student Page Break' exports. Photos support one essay per image and are read by Fiveable. Accepts up to three ChatGPT attached uploads of 25 MB each, or up to three base64 files totaling about 3 MB. Returns counts and warnings only. Long PDFs can take up to a minute.
| Name | Required | Description | Default |
|---|---|---|---|
| files | Yes | ||
| uploads | Yes | Files the teacher attached in ChatGPT, passed by reference. | |
| sessionId | Yes | Grading session id from list_grading_sessions or create_grading_session. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Excellent behavioral disclosure beyond annotations: it describes the splitting logic (one essay per student), supported formats (PDF, Word, photos), photo handling (one essay per image, read by Fiveable), size limits (25 MB uploads, ~3 MB base64 total, up to 3 files), return value behavior ('Returns counts and warnings only'), and a performance warning ('Long PDFs can take up to a minute'). This is rich, non-obvious context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tightly packed into four sentences, each earning its place with concrete, useful details. It is front-loaded with the core splitting behavior and then flows into format support, size limits, return behavior, and performance note without waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (file parsing, splitting, multiple formats, size/upload limits), the description covers all critical behavioral aspects: input types, splitting logic, size constraints, return value scope, and latency warning. The output schema exists, so return details need not be repeated. This is complete for an agent to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 67% – moderate. The description adds meaning by explaining that 'uploads' are 'ChatGPT attached uploads' and 'files' are teacher-supplied base64 files, but it does not clarify format expectations or constraints for individual parameters beyond what the schema provides. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Import essays from files') and details exactly what it does: 'Splits teacher-supplied PDFs or Word stacks into one essay per student.' It names the specific format ('AP Classroom FRQ Response PDFs, By Student Page Break exports') and distinguishes itself from siblings import_essays_from_google_classroom and import_essays_from_text by being file-based.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it (when you have teacher-supplied PDFs/Word stacks/photos) but does not explicitly state when to use this versus import_essays_from_text or import_essays_from_google_classroom. An agent must infer from the input schema and sibling names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
import_essays_from_google_classroomImport essays from Google ClassroomAInspect
Adds selected Google Classroom turn-in handles to a grading session. Reads authorized Google Docs directly; optional teacher-provided text supports other files or missing Google access. Names come from the Classroom roster. Existing session submissions are skipped. Records the source Classroom assignment for subsequent approved-grade export. Returns counts only.
| Name | Required | Description | Default |
|---|---|---|---|
| turnIns | Yes | ||
| courseId | Yes | Id from list_google_classroom_work. | |
| sessionId | Yes | Grading session id from list_grading_sessions or create_grading_session. | |
| courseWorkId | Yes | Id from list_google_classroom_work. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, and openWorldHint=true. The description goes beyond them by disclosing side effects: duplicate submissions are skipped, the source Classroom assignment is recorded for later grade export, and only counts are returned. This is useful behavioral context an agent cannot get from the annotations alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four compact sentences, front-loaded with the core action, followed by sourcing, skip behavior, side effects, and return shape. No filler; every sentence carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a write tool with four required parameters, an output schema, and moderate nested structure, the description covers sourcing, deduplication, export linkage, and return granularity. It is complete enough to call correctly, with only the lack of explicit sibling routing as a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75% with all five properties described. The description adds meaning beyond the schema by explaining the trade-off between reading an authorized Doc directly versus supplying teacher-provided text, and by clarifying that handles originate from the Classroom roster/list_google_classroom_work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: adding Google Classroom turn-in handles to a grading session. The Classroom-specific scope distinguishes it from import_essays_from_files and import_essays_from_text, though it never names those siblings explicitly, so an agent must infer the routing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It supplies relevant context (names come from the Classroom roster, existing submissions are skipped, and teacher-provided text covers files lacking Google access), which implies when to use it. However, it never states when to prefer this tool over import_essays_from_files or import_essays_from_text, so the selection guidance is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
import_essays_from_textImport essays from textAInspect
Adds individual student responses supplied as pasted, downloaded or transcribed text, with each student's name as written. Photo transcriptions require transcribedFromPhotos and explicit teacherConfirmedTranscription approval of the transcription. Returns counts only.
| Name | Required | Description | Default |
|---|---|---|---|
| essays | Yes | ||
| sessionId | Yes | Grading session id from list_grading_sessions or create_grading_session. | |
| transcribedFromPhotos | Yes | ||
| teacherConfirmedTranscription | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description discloses two real behavioral traits: photo transcriptions require transcribedFromPhotos and explicit teacherConfirmedTranscription approval, and the tool 'Returns counts only'. This adds meaningful context; the mutation/non-idempotent profile itself is already covered by annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action and followed by the conditional approval rule and return behavior. No filler; each sentence carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists and conveniently the description notes that only counts are returned. For a 4-param mutating tool with thin schema coverage, the description supplies the essential approval gate and source-type context, leaving only minor gaps in boolean coupling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 25%, and the two undocumented booleans (transcribedFromPhotos, teacherConfirmedTranscription) are given meaning by the description's statement that photo transcriptions require both. 'With each student's name as written' maps to studentName, but the exact coupling and defaults between the two booleans remain unspecified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (adds) and resource (individual student responses), and scopes the input source to 'pasted, downloaded or transcribed text', which implicitly separates it from import_essays_from_files and import_essays_from_google_classroom. However, it never names those siblings, so the differentiation is left to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied through the input-source phrasing ('pasted, downloaded or transcribed text') and the photo-transcription approval workflow, but there is no explicit when-to-use versus the file/Classroom import siblings or any stated prerequisite like obtaining a sessionId first.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
link_grading_to_classLink grading to a classADestructiveIdempotentInspect
Links a grading session to a workspace class so results feed class insights. Exact roster-name matches link automatically; explicit essay-handle to student-handle matches resolve remaining essays. Returns unmatched essays. A null classId removes the class link.
| Name | Required | Description | Default |
|---|---|---|---|
| classId | Yes | Class id from list_classes, or null to unlink. | |
| matches | Yes | ||
| sessionId | Yes | Grading session id from list_grading_sessions or create_grading_session. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false. The description adds that a null classId removes the class link, clarifying the destructive edge case, and that unmatched essays are returned. It doesn't disclose whether automatic matching overwrites existing links or other side effects, but adds useful context beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three focused sentences: linking purpose, matching behavior, and unlink case. No wasted words, front-loaded with the core action. Could be slightly more structured but efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a mutation tool with a nested matches array and an output schema (presumably for unmatched essays), the description covers purpose, matching logic, and the unlink case. The output schema means return format needn't be explained. Missing details on error handling or rate limits, but complete for the core use case.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 67% schema coverage, the schema documents classId and sessionId origins and constraints. The description clarifies the behavior of classId=null (unlink) and the match resolution logic for the matches array, adding meaning beyond the schema's structural details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (links) and resource relationship (grading session to workspace class) with the purpose (feed class insights). Distinguishes itself from siblings like add_students or create_grading_session by describing the linking action and its two-stage matching process.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the automatic matching via exact roster-name matches and the manual matching via essay/student handles, and notes that a null classId removes the link. Clear context for when to use, though it doesn't name alternative tools or explicitly state prerequisites beyond referencing list_classes and get_grading_results in the schema.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_cheatsheetsList official Fiveable cheatsheetsARead-onlyIdempotentInspect
Find approved Fiveable course and unit PDFs, including course variants, academic year when known, source links and readable-artifact readiness. Does not include user galleries or generated personal visuals.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | Yes | ||
| cursor | No | ||
| unitNumber | No | ||
| subjectSlug | No | ||
| courseVariant | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already signal readOnlyHint, openWorldHint=false, idempotentHint, and non-destructive behavior, so the description's extra context is a bonus. It adds useful behavioral nuance by noting that academic year is included 'when known', mentioning readable-artifact readiness, and explicitly scoping out user-generated content. There is no contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The core purpose is front-loaded, and the second sentence earns its place by drawing a clear boundary around what the tool does not return.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a rich output schema, safety-revealing annotations, and a clear scope statement, the description is largely adequate for a read-only listing tool. The main missing piece is explicit guidance on when to choose this tool over get_cheatsheet or search_content, but the schema and annotations carry much of the remaining context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description bears the burden of explaining the five parameters, but it does not explain limit, cursor, unitNumber, subjectSlug, or courseVariant directly. The phrase 'course and unit PDFs' weakly maps to subjectSlug and unitNumber, and 'course variants' hints at courseVariant, but limit and cursor remain entirely unaddressed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Find') and a specific resource ('approved Fiveable course and unit PDFs'), and enriches it with concrete details about course variants, academic year, source links, and artifact readiness. It also explicitly excludes user galleries and personal visuals, which helps distinguish it from broader content-listing tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear scope and one meaningful exclusion, but it does not name alternative tools or specify when to use this list tool versus fetch tools like get_cheatsheet or search_content. The intended usage is mostly implied by the word 'approved' and the list-oriented verb, rather than explicitly contrasted with siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_classesList my classesARead-onlyIdempotentInspect
Lists the signed-in teacher's classes with subject, student count and class id. Class names such as 'period 3' can be resolved to their ids. Cheap.
| Name | Required | Description | Default |
|---|---|---|---|
| includeArchived | Yes | List archived classes instead of active ones. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, closed-world, so the safety profile is covered. The description adds genuinely new context: the exact fields returned, that human-readable class names can be resolved to ids here, and a cost signal ('Cheap') that affects tool selection.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with zero padding. The resource and its returned fields come first, then the id-resolution utility, then the cost note — well front-loaded for an agent deciding whether to call it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists and the single input parameter is fully documented in the schema, so the description needn't explain returns or the archived flag. For a simple read-only list tool, everything an agent needs to invoke it correctly is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter, and schema description coverage is 100% — the schema fully explains that includeArchived lists archived rather than active classes. The description never mentions archived filtering at all, so it adds nothing beyond the schema; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Lists') and resource ('classes') scoped to 'the signed-in teacher's', and enumerates the returned fields (subject, student count, class id). It does not, however, distinguish itself from siblings like list_subjects or get_class, which an agent scanning 70 tools would benefit from.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the note that class names such as 'period 3' can be resolved to ids suggests this is the lookup step before an operation, and 'Cheap' hints it should be preferred for bulk lookups. No explicit when-to-use, when-not-to-use, or named alternative (e.g., get_class for a single class) is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_courseworkList courseworkARead-onlyIdempotentInspect
Lists the teacher's assignments and grading sessions from their workspace ledger: status, class, due date, roster-based turn-in counts, review counts and class average. The active status filter selects work currently open to students. Optional filters cover class, subject and status. Cheap.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | Yes | ||
| cursor | No | Cursor from a previous page. | |
| status | Yes | Ledger status filter. needs_review shows work waiting for the teacher. | all |
| classId | No | Class id from list_classes. | |
| subjectSlug | No | Limit to one subject, e.g. ap-bio or apush. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint=false, and destructiveHint=false, so the safety profile is fully covered structurally. The description adds the useful note that the tool reads from the 'workspace ledger' and says 'Cheap,' but omits pagination behavior beyond the cursor name and doesn't disclose rate limits or the practical cost of wide status filters.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the resource and returned fields, then filters, then a one-word cost note. No filler, though the dangling 'Cheap.' is terse to the point of being under-explained relative to its placement.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is a read-only filtered list with an output schema, so the description needn't explain return values. It does well to enumerate the columns returned and the filter dimensions. What's missing is any note on pagination limits (limit max 50) or how the status filter interacts with the returned set, which an agent calling this against a large ledger would benefit from.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 80%, so most parameters are already documented (cursor, status enum, classId source, subjectSlug examples). The description only paraphrases 'Optional filters cover class, subject and status' and the meaning of the active status, adding little syntax or format detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (lists) and resource (coursework = assignments and grading sessions) and enumerates the fields returned (status, class, due date, turn-in counts, review counts, class average). It doesn't explicitly differentiate from the nearby list_grading_sessions, list_submissions_to_review, or list_classes siblings, but the scope is clear enough for an agent to identify the tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Mentions that the active status filter selects currently-open work and that filters cover class, subject, and status, which implies when to use each filter. It never states when to prefer this tool over the overlapping siblings (list_grading_sessions, list_submissions_to_review) or when filtering by status is preferable to the plain call.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_frqsList free-response questionsARead-onlyIdempotentInspect
Lists Fiveable's free-response (FRQ) practice prompts for an AP subject, optionally filtered by FRQ type, unit, and excluded IDs. Results use stable ID order; unchanged filters continue pagination. Returns prompt ids and summaries rather than full prompts. Cheap.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | Yes | Maximum prompts to list. | |
| cursor | No | ||
| intent | No | Privacy-safe reason for the request. Use explain, quiz, notes, frq, or research; never send the student's raw prompt for analytics. | |
| unitId | No | Unit ID or public unit slug from get_subject_outline. | |
| frqType | No | Optional FRQ type filter, e.g. "DBQ", "LEQ", "SAQ". Varies by subject. | |
| excludeIds | No | ||
| subjectSlug | Yes | Fiveable subject slug, e.g. "ap-bio". A course name such as "AP Biology" also resolves. | |
| courseVariant | No | AP Calculus only (subject ap-calc): 'ab' hides BC-only units, topics and FRQs; 'bc' includes them. Pass the course of the class or assignment draft you are building for ('ab' for ap-calc-ab, 'bc' for ap-calc-bc). |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the read-only/idempotent/non-destructive profile, and the description adds real behavioral context: stable ID ordering, that pagination is filter-stable, and that the payload is lightweight summaries plus a cost signal ('Cheap').
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences that front-load the resource, then filters, then pagination and return shape; the trailing 'Cheap.' is terse but earns its place as a cost cue.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists so return values need no elaboration, and the description still covers pagination semantics and what is returned. Minor omission: courseVariant and intent are not acknowledged, though the schema documents both.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%, so most parameters are already documented in the schema. The description groups frqType, unit, and excludeIds as filters and explains cursor continuation, but adds no syntax or detail for intent, courseVariant, or limit, so it does not meaningfully exceed the schema baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Lists), a specific resource (Fiveable's free-response practice prompts), and the scope (an AP subject), which clearly separates it from the singular get_frq and the broader search_question_bank.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent what filters are available and that 'unchanged filters continue pagination,' and it hints at the alternative by noting it returns 'prompt ids and summaries rather than full prompts' (i.e. use get_frq for full text). It does not, however, name a sibling or state exclusions directly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_google_classroom_workList Google Classroom workARead-onlyIdempotentInspect
Walks the teacher's Google Classroom: with no ids, their courses; with a course id, its assignments; with a course and assignment id, who turned it in, each with a handle (T-…) and a link to the turned-in file. Use it to grade typed work collected in Classroom. Needs the teacher's Google Classroom connection.
| Name | Required | Description | Default |
|---|---|---|---|
| courseId | No | Id from list_google_classroom_work. | |
| courseWorkId | No | Id from list_google_classroom_work. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, open-world, and non-destructive behavior. The description adds meaningful context beyond that: it requires the teacher's Google Classroom connection and describes the hierarchical traversal and returned handle/link data. It stops short of covering pagination or rate-limit behavior, keeping it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core hierarchical behavior, then gives the use case and auth requirement. Every sentence carries necessary information with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple two-optional-parameter schema, rich annotations, and the presence of an output schema, the description is complete enough for correct invocation. It covers traversal modes, intended use, auth requirement, and returned file handles/links without needing to restate output-schema details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although schema coverage is 100%, the description adds important meaning beyond the schema by explaining how the two optional IDs interact: no IDs yields courses, courseId alone yields assignments, and courseId + courseWorkId yields submission details. This directly clarifies parameter combinations that the schema does not describe.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the exact resource and traversal behavior: with no ids it returns the teacher's courses, with a course id it returns assignments, and with both ids it returns submitters and turned-in files. This hierarchical listing is specific enough to distinguish it from generic siblings like list_coursework or list_classes without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear usage context: 'Use it to grade typed work collected in Classroom.' However, it does not explicitly name alternatives such as list_coursework, get_submission, or import_essays_from_google_classroom, nor does it state when not to use this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_grading_sessionsList grading sessionsARead-onlyIdempotentInspect
The teacher's grading sessions for work done outside Fiveable (stacks of essays graded in the workspace), newest first, with essay, scored and approved counts. Use to resume a stack. Cheap.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the read-only, idempotent, non-destructive profile, so the bar is lower. The description still adds real behavioral context: newest-first ordering, the returned counts (essay, scored, approved), and a cost hint ('Cheap').
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three terse, front-loaded sentences with no filler; scope and ordering come first and the resume guidance follows. 'Cheap' is clipped but earns its place as a cost signal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is unnecessary, and annotations cover the safety profile. The definition is nearly complete for a list tool, with the only real gap being the undocumented limit parameter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions the required 'limit' parameter or its bounds. With a documented parameter and no compensating explanation in either place, the agent gets no added meaning beyond the parameter name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('grading sessions') and adds scope that distinguishes it from siblings: 'for work done outside Fiveable (stacks of essays graded in the workspace).' An agent can tell this apart from create_grading_session or start_grading without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear usage context ('Use to resume a stack') that signals the intended workflow. It does not name explicit alternatives (e.g., list_submissions_to_review) or exclusions, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_key_termsList key termsARead-onlyIdempotentInspect
Lists the key vocabulary terms for an AP subject, paginated alphabetically. Suitable for building flashcards or checking what vocabulary a student should know; results include stable term slugs for full definitions. Cheap.
| Name | Required | Description | Default |
|---|---|---|---|
| page | Yes | 1-indexed page number. | |
| limit | Yes | Terms per page. | |
| intent | No | Privacy-safe reason for the request. Use explain, quiz, notes, frq, or research; never send the student's raw prompt for analytics. | |
| letter | Yes | Filter to terms starting with one letter, or "ALL" for every term. | ALL |
| subjectSlug | Yes | Fiveable subject slug, e.g. "ap-bio" or "apush". The course name or a close approximation such as "ap-biology" also resolves. Use list_subjects to find it. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already establish readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the tool is safe to call repeatedly. The description adds that results are paginated alphabetically, include stable slugs, and are 'cheap' (which suggests low cost or resource usage). It does not mention rate limits or any other behavioral traits beyond the annotations, so it's strong but just shy of perfect.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: it states the main purpose, use cases, and key features (slugs, cheap) in two sentences. Every clause contributes meaning, with no redundant or generic phrasing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a straightforward list tool, and the description plus schema cover all necessary information for a user to call it: the purpose, parameters, and behavior are clearly defined. The only minor gap is that the description doesn't explain what a 'term slug' is or how to use it, but playful hint is given (for full definitions use get_key_term). The output schema exists, so the return format does not need to be described. Overall, it is highly complete for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, meaning all 5 parameters are described in the input schema. The description adds value by mentioning 'stable term slugs' which hints at the subjectSlug parameter, but it doesn't explain any parameter in depth beyond what the schema already provides. Since the schema fully documents each parameter, a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists key vocabulary terms for an AP subject, with alphabetical pagination. It distinguishes itself from get_key_term (singular term vs. list) and list_subjects (which finds subjects), making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states suitable use cases: 'building flashcards or checking what vocabulary a student should know.' It also mentions cost ('Cheap'), which implies it is appropriate for frequent listing. While it doesn't name alternatives, the clear use cases and the existence of get_key_term for individual terms make when to use this tool apparent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_subjectsList AP subjectsARead-onlyIdempotentInspect
Lists every AP subject Fiveable covers, with its canonical subject slug and flags for which features exist for that subject (practice questions, key terms, score calculator). Suitable when a subject slug is not already known. Cheap — roughly 1-2k tokens.
| Name | Required | Description | Default |
|---|---|---|---|
| intent | No | Privacy-safe reason for the request. Use explain, quiz, notes, frq, or research; never send the student's raw prompt for analytics. | |
| search | No | Optional case-insensitive filter on subject name, e.g. "bio" or "history". |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds useful behavioral context: the tool returns canonical slugs and feature flags, and is cheap at roughly 1-2k tokens. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, zero filler. The main function is front-loaded, the useful output details are included, and the cost note is a relevant operational detail for an agent deciding whether to call the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only listing tool with two optional params, full schema coverage, and an output schema, this description is complete. It explains purpose, output contents, when to use it, and cost without missing anything an agent needs to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents both parameters, including the intent enum and the optional search filter. The description doesn't add parameter-level details, which is acceptable given the schema already handles that burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: lists every AP subject Fiveable covers, and specifies the output includes canonical slugs and feature flags. This clearly distinguishes it from sibling list tools like list_cheatsheets or list_key_terms.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says it is suitable when a subject slug is not already known, giving clear context for when to invoke it. It does not name alternative tools or spell out when not to use it, so it stops just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_submissions_to_reviewList work to reviewARead-onlyIdempotentInspect
The teacher's review queue, as on the website inbox: submissions whose free responses need a teacher score or a look at the AI score, imported essays awaiting approval, and assignments ready to hand back. Each row includes the identifiers and review operation for its item. Filter by subject or assignment. Cheap.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | Yes | ||
| cursor | No | ||
| filter | Yes | all | |
| subjectSlug | No | ||
| assignmentId | No | Assignment id from list_coursework or list_submissions_to_review. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent and non-destructive, so the safety profile is covered. The description adds genuinely new behavioral context: the cost signal ('Cheap') and the promise that each row carries identifiers plus the review operation to perform, which tells the agent this is a routing/dispatch listing rather than raw data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three front-loaded sentences covering scope, row shape and filtering, plus a terse cost note. The clipped 'Cheap.' fragment is efficient though slightly informal; no filler otherwise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be spelled out, and the description helpfully flags the row contents. But for a 5-param tool with 20% schema coverage, the unexplained `filter` enum and pagination semantics leave a real gap before an agent can call it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 20% across 5 params, so the description should compensate but largely does not. It gestures at 'filter by subject or assignment' (subjectSlug, assignmentId) but never explains the critical `filter` enum (all/review/handback), the limit default/max, or cursor pagination — the exact values a caller must choose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('teacher's review queue') and enumerates the three item categories it returns (free responses needing scoring, imported essays awaiting approval, assignments ready to hand back). This scope is distinctive enough to separate it from siblings like get_submission or list_coursework.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'as on the website inbox' framing and the item inventory imply when to reach for it, and it notes subject/assignment filtering. However, it names no alternative tools and gives no when-not guidance, so an agent must infer routing from the sibling list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_teaching_resourcesList teaching resourcesBRead-onlyIdempotentInspect
Lists assignment reading resources for one subject: study guides, unit reviews, AMSCO guided notes, key terms and cheatsheets, with resource ids suitable for reading sections. Optional filters cover kind, unit and words. Cheap.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | ||
| unit | No | Exact unit label from the result's unit list. | |
| limit | Yes | ||
| query | No | ||
| offset | Yes | ||
| subjectSlug | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so safety is covered without the description. The description usefully adds a cost signal ('Cheap') and notes resource ids are suitable for reading sections, but says nothing about pagination behavior despite offset/limit params.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the verb/resource/scope, and the terse 'Cheap.' is efficient. No filler, though the filter sentence could be more informative for its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-param list tool with an output schema (so return values needn't be explained), the description covers resource types and filters but omits pagination semantics and, critically, how it relates to the very similar sibling list tools. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 17% (only 'unit' documented), so the description must compensate — but it merely names 'kind, unit and words', which are already the parameter names. It adds no syntax, format, or filtering semantics, and omits limit/offset/subjectSlug entirely.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb+resource+scope: it lists assignment reading resources for a single subject, and enumerates the concrete kinds (study guides, unit reviews, AMSCO guided notes, key terms, cheatsheets) matching the enum. It does not explicitly differentiate itself from overlapping siblings like list_cheatsheets or list_key_terms, which weakens it slightly below a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Only says 'Optional filters cover kind, unit and words' with no guidance on when to choose this over the sibling list_cheatsheets/list_key_terms tools that return subsets, and no prerequisites or exclusions. Usage is implied at best.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
override_essay_scoreAdjust an essay scoreADestructiveIdempotentInspect
Changes the points on one rubric part of a scored essay, with the teacher's reason. The change is logged and the total recalculated, as on the website. Reversal restores the previous part points.
| Name | Required | Description | Default |
|---|---|---|---|
| part | Yes | Rubric part as shown in get_essay_feedback, e.g. "A" or "Thesis". | |
| points | Yes | ||
| reason | Yes | ||
| sessionId | Yes | Grading session id from list_grading_sessions or create_grading_session. | |
| essayHandle | Yes | Essay handle from get_grading_results, e.g. E-3b91c0. | |
| reasonCategory | Yes | teacher-judgment |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, so the safety profile is partly covered. The description adds substantive context beyond them: the change is logged, the total is recalculated, and reversal restores the previous part points — behavior an agent could not infer from the annotations alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action before the logging/recalculation/reversal details. The 'as on the website' clause is slightly informal but does convey parity with the UI behavior; no filler sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and annotations cover the mutation safety profile. The description covers the side effects (logging, recalculation) and reversibility, leaving only auth/permission requirements unstated — a minor gap for this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 50%, and the description names the key concepts (one rubric part, points, teacher's reason), loosely mapping to three of six params. It adds no format or constraint detail beyond the schema, and says nothing about sessionId, essayHandle, or reasonCategory, so it neither compensates for the coverage gap nor exceeds the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (changes/adjusts) and a precise scope (points on one rubric part of a scored essay) plus the reason requirement. It is clearly distinguishable from broad siblings like score_submission or rescore_essay, though it does not explicitly name its closest sibling override_submission_part.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'of a scored essay' implies the tool only applies after an essay has already been scored, which is useful usage context. However, there is no explicit statement of when to prefer this over rescore_essay or override_submission_part, and no prerequisites or exclusions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
override_submission_partAdjust an FRQ part scoreADestructiveIdempotentInspect
Changes the points on one rubric part of an AI-scored FRQ in a submission, with the teacher's reason, as on the website. The section's score is recalculated. Reversal restores the previous part points.
| Name | Required | Description | Default |
|---|---|---|---|
| part | Yes | ||
| itemId | Yes | ||
| points | Yes | ||
| reason | Yes | ||
| assignmentId | Yes | Assignment id from list_coursework or list_submissions_to_review. | |
| submissionId | Yes | Submission id from list_submissions_to_review. | |
| reasonCategory | Yes | teacher-judgment |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true and idempotentHint=true, but the description adds real behavioral context beyond them: the section's score is recalculated as a side effect, and reversal restores the previous part points. That side-effect disclosure is exactly what an agent needs before mutating a grade. It stops short of stating permission requirements or whether the override is reflected downstream in released results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the mutation and its scope, then the recalculation and reversal consequences. Only 'as on the website' is wasted words; otherwise every sentence carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the recalculation/reversal notes cover the main behavioral risks of this mutation. But with 7 required parameters at 29% schema coverage and no description of the 'part' identifier semantics, an agent still lacks enough guidance to construct a correct call confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 29% (three ID fields documented, part/points/reason/reasonCategory not), so the description must compensate. It partially does by implying points, the single 'part' target, and the teacher's reason/reasonCategory pairing, but says nothing about the format or valid range of 'part' or how assignmentId/submissionId/itemId relate to each other.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (changes), resource (points on one rubric part of an AI-scored FRQ in a submission), and scope (a single rubric part). This distinguishes it from the sibling override_essay_score, which operates on the whole essay, though it doesn't name that sibling explicitly. The trailing phrase 'as on the website' is vague filler that doesn't add scope information.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The context ('AI-scored FRQ', 'one rubric part', teacher's reason required) implies when this tool applies, and the reversal sentence hints at the undo path. However, it never states when to choose this over override_essay_score or score_submission, nor any prerequisites such as whether the submission must already be scored or the assignment published.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prepare_correct_answer_keyPreview correcting an answer keyAInspect
Creates an owner-bound confirmation preview for correcting a published multiple-choice key, including how many students gain or lose points. Accepts assignment section and question numbers. Keys and scores remain unchanged until approved preview execution.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | A note for Fiveable with the shared fix. | |
| explanation | No | Why this answer is right, in the teacher's words; defaults to the current explanation. | |
| assignmentId | Yes | Assignment id from list_coursework. | |
| sectionNumber | Yes | 1 for the assignment's first section. | |
| correctAnswers | Yes | The right answer letters, e.g. ["C"]. | |
| questionNumber | Yes | Question number within that section. | |
| shareCorrection | Yes | Also send the fix to Fiveable so the question is corrected for everyone. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, destructiveHint=false, and idempotentHint=false, but the description adds genuinely useful context beyond them: the artifact is owner-bound, it is a preview only, and no keys or scores change until execution is approved. That non-mutation-until-approval behavior is the key fact annotations alone would not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: purpose, accepted inputs, and the non-mutation guarantee, with the most important constraint front-loaded. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and annotations cover the safety profile. The description covers the preview-versus-commit distinction adequately, though it could have mentioned the follow-up confirm tool to complete the workflow picture.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all seven parameters are already documented in the schema, including examples and bounds. The description only restates that it accepts section and question numbers, adding no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Creates an owner-bound confirmation preview for correcting a published multiple-choice key') and adds scope detail (point deltas for students). An agent can distinguish it from confirm_correct_answer_key, which performs the actual correction, without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes clear this is a preview stage that precedes approval ('Keys and scores remain unchanged until approved preview execution'), giving clear usage context. It stops short of naming confirm_correct_answer_key as the follow-up sibling, so the agent must infer the routing from names alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prepare_publish_assignmentPreview publishing an assignmentAInspect
Creates an owner-bound confirmation preview for publishing a draft or republishing with new dates: classes, openings, due dates in the teacher's time zone, result visibility and student-facing changes. Publication remains unchanged. Execution requires approval of this specific preview. With no classes, anyoneWithLink must be true for a public link.
| Name | Required | Description | Default |
|---|---|---|---|
| dueAt | No | Due time for the anyone-with-link publish. | |
| classes | Yes | ||
| timeZone | Yes | The teacher's IANA time zone, e.g. America/Chicago, so the preview shows their local times. | America/New_York |
| assignmentId | Yes | Assignment id from list_coursework. | |
| resultRelease | No | Defaults to the draft's setting. | |
| anyoneWithLink | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=false and destructiveHint=false; the description adds meaningful context beyond that: 'owner-bound confirmation preview', 'Publication remains unchanged', and a hard precondition ('With no classes, anyoneWithLink must be true'). It does not cover preview expiry or error/rejection behavior, so not a full 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences with the core action front-loaded and the required precondition placed last. No filler, though the middle sentences are packed with multiple clauses each.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a prepare/confirm-style tool with 6 params and an output schema (so return values need no prose), the description covers the workflow position, the immutability guarantee, and the empty-classes constraint. Minor gaps remain around what the preview token/preview represents for the follow-up call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, and the description adds semantics not in the schema: the time-zone purpose ('shows their local times'), the result-visibility and student-facing-change scope, and the conditional rule tying anyoneWithLink to an empty classes list. This meaningfully extends the structured fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Creates an owner-bound confirmation preview for publishing a draft or republishing') and makes clear it is the preview step rather than the publish itself, which distinguishes it from the sibling confirm_publish_assignment. An agent can identify the tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when to use it (publishing a draft or republishing with new dates) and states that execution requires approval of this specific preview, implying the confirm step. It stops short of explicitly naming the confirm_publish_assignment alternative, so routing is left partly to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prepare_push_grades_to_google_classroomPreview sending grades to Google ClassroomAInspect
Creates an owner-bound confirmation preview for exporting approved grading-session scores to the source Google Classroom assignment. Includes included/excluded counts and the point scale. No grades are sent until approved preview execution.
| Name | Required | Description | Default |
|---|---|---|---|
| sessionId | Yes | Grading session id from list_grading_sessions or create_grading_session. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, destructiveHint=false, which is a slightly unusual profile for something called a 'preview'; the description resolves that tension by explaining that it creates a bound confirmation artifact and that nothing is transmitted yet. It also discloses the payload contents (included/excluded counts, point scale). It does not state how long the preview remains valid or what approval requires.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and scope, then the no-side-effect guarantee, then the returned contents. No filler and nothing important buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return structure need not be explained, and the description still sketches what the preview contains. The remaining gap is the unstated handoff to the confirmation sibling and any preview lifetime/authorship constraints ('owner-bound' is asserted but not elaborated).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single sessionId parameter already carries its own description pointing to list_grading_sessions/create_grading_session; the description adds no syntax, format, or constraint detail beyond the schema. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (creates) and resource (an owner-bound confirmation preview for exporting grading-session scores to the source Google Classroom assignment), and implicitly contrasts with the confirm_push_grades_to_google_classroom sibling by labeling itself a preview. An agent can tell this apart from the actual send operation without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"No grades are sent until approved preview execution" clearly frames this as the pre-approval step, giving the agent the usage context. It stops short of naming the confirm_push_grades_to_google_classroom sibling as the explicit follow-up, so routing is inferred rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prepare_release_resultsPreview handing results backAInspect
Creates an owner-bound confirmation preview for releasing assignment scores and feedback to every class, including counts of unscored responses. Release status remains unchanged until approved preview execution. A release cannot be taken back.
| Name | Required | Description | Default |
|---|---|---|---|
| assignmentId | Yes | Assignment id from list_coursework. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover safety flags (readOnly=false, destructive=false, idempotent=false, openWorld=false), so the bar is lower, yet the description adds real context: the preview is owner-bound, it reports counts of unscored responses, it leaves release status untouched, and it warns that the eventual release is irreversible. It does not clarify the lifetime/expiry of the preview token or who may approve it, keeping it at a solid 4 rather than 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with the core action front-loaded and no filler. The middle sentence slightly restates the preview semantics already implied by sentence one, but the closing irreversibility warning earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description covers the key non-obvious facts an agent needs (no status change, unscored counts included, owner binding, irreversibility). Only the relationship to the confirm sibling and the approval subject are left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter, assignmentId, and the schema documents it at 100% coverage including its own pointer to list_coursework. The description adds no syntax, format, or sourcing detail beyond that, so it meets the baseline rather than exceeding it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Creates an owner-bound confirmation preview for releasing assignment scores and feedback to every class.' The scope ('every class') and the preview nature clearly separate it from a plain release action. It stops short of naming the sibling confirm_release_results that executes the preview, so differentiation is implied rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Release status remains unchanged until approved preview execution' implies this is the pre-step in a two-phase flow, and an agent can infer that confirm_release_results follows. However, it never explicitly says when to use this tool versus that sibling, nor what qualifies an approver, so usage is inferred rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prepare_update_publicationPreview changing a published assignmentAInspect
Creates an owner-bound confirmation preview for closing or reopening a published assignment, or changing class open/due dates. Links and submissions remain intact. Publication remains unchanged until approved preview execution.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | ||
| changes | Yes | For change_dates only. | |
| timeZone | Yes | The teacher's IANA time zone, e.g. America/Chicago, so the preview shows their local times. | America/New_York |
| assignmentId | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds meaningful context beyond the annotations: 'Links and submissions remain intact' reinforces destructiveHint=false, and 'Publication remains unchanged until approved preview execution' clarifies the mutation is deferred to the confirm step. 'Owner-bound' hints at an ownership/auth requirement. No contradiction with readOnlyHint=false since creating a preview object is a write.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, zero filler, with the core purpose front-loaded and the safety guarantees (data intact, deferred mutation) following in logical order. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the description adequately covers the mutation-safety story. The remaining gap is the 50% parameter coverage, which forces the agent to lean on the schema for the changes/timeZone semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%. The description maps the three actions (close/reopen/change_dates) to their intent, which reflects the enum, but adds little about changes, timeZone, or assignmentId that the schema doesn't already convey. Marginal value over structured fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Creates an owner-bound confirmation preview') and enumerates the exact mutations covered (closing/reopening, changing open/due dates). The two-phase framing ('until approved preview execution') implies the confirm step, cleanly distinguishing it from confirm_update_publication, though it never names that sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The use case is implied (preview the listed changes before approving execution), but there is no explicit when-to-use/when-not-to-use guidance or named alternative tool. An agent can infer routing from context but isn't told the conditions that select this over confirm_update_publication.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preview_assignment_as_studentPreview assignment as a studentARead-onlyIdempotentInspect
What a student will see when they open the assignment: instructions and each section in order. Answers are hidden, as they are for students. Diagnostics can't be previewed. Cheap.
| Name | Required | Description | Default |
|---|---|---|---|
| assignmentId | Yes | Assignment id from list_coursework. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly/idempotent/non-destructive, and the description adds meaningful behavior beyond them: answers are hidden exactly as for students, diagnostics are not previewable, and it signals low cost ('Cheap'). These are genuine operational facts an agent should know before calling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core behavior and kept to one tight sentence plus a terse cost hint. 'Cheap.' is a fragment, but it carries a real signal and everything earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is unnecessary, and the description supplies the key behavioral caveats (hidden answers, no diagnostics preview). Only the missing sibling routing keeps it from being fully complete for a one-param read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and assignmentId's meaning (id from list_coursework) is fully documented in the schema. The description adds no parameter detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('what a student will see when they open the assignment') and enumerates the returned content (instructions and each section in order). The 'as a student' framing implicitly separates it from the teacher-facing get_assignment sibling, so an agent can pick it without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the preview framing, and it notes one boundary ('Diagnostics can't be previewed'), but it never names an alternative (e.g., get_assignment or preview_question) or states when to choose this tool over them. Adequate but with a clear routing gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preview_questionPreview questionBRead-onlyIdempotentInspect
Returns one bank question with its stimulus, choices, marked correct answer and explanation. Stimulus sets include all linked questions and keys. Answer keys are teacher-only material. Cheap.
| Name | Required | Description | Default |
|---|---|---|---|
| questionId | Yes | Question id from search_question_bank, e.g. mcq:64f… or stimulus-set:64f…. | |
| subjectSlug | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior, so the safety profile is covered. The description adds genuine context beyond that: answer keys are teacher-only material (an access/permission note) and the call is cheap (a cost signal), plus it discloses that stimulus sets expand to all linked questions and keys.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the return contents and with zero filler; each sentence (return shape, stimulus-set expansion, teacher-only keys, cost) carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value detail is not strictly required, and the description still adds the teacher-only access caveat and stimulus-set expansion behavior. It is broadly complete for a read-only tool, though it says nothing about permissions required or about pagination/limits implied by 'cheap'.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With only 50% schema description coverage, the description carries part of the burden, yet it says nothing about either parameter. questionId is documented in the schema (pattern + example), but subjectSlug has no description anywhere, and the description does not clarify how these two strings combine.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource (returns one bank question) and enumerates what comes back: stimulus, choices, marked correct answer, and explanation. It is clearly distinguishable from list/search siblings, though it doesn't name them explicitly; the schema description points to search_question_bank as the source of the id.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance, and no alternative is named (e.g., get_frq, get_questions_for_materials, search_question_bank). The single word 'Cheap' hints at a cost tradeoff but leaves the agent to infer when this preview is preferable to a heavier fetch.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_studentsRemove studentsADestructiveInspect
Removes students by handle from a class roster while preserving their past work in the teacher's results. Returns an undoId valid for exact restoration within seven days.
| Name | Required | Description | Default |
|---|---|---|---|
| classId | Yes | Class id from list_classes. | |
| handles | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark this as destructive and non-idempotent. The description adds important context: past work is preserved in the teacher's results, and an undoId enables exact restoration within seven days. This exceeds annotation coverage but omits permission or failure-condition details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the effect and followed by the restoration behavior. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema provided, return-value explanation is unnecessary. The description covers the core mutation effect and reversibility well, though it could mention permission requirements or what happens to invalid handles.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%, but the schema itself describes both parameters (classId source and handle format). The description adds 'by handle' but no syntax, limits, or classId explanation beyond the schema, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (removes), resource (students by handle from a class roster), and distinguishes this from add/restore siblings through the roster-removal scope. An agent can identify the action without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides no explicit when-to-use guidance, prerequisites, or alternatives. It does not mention restore_students, update_student, or class-related siblings, leaving the agent to infer when removal is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
report_content_issueReport a content issue to FiveableAIdempotentInspect
Submits a caller-requested content issue to Fiveable support, which may lead to support follow-up. Requires explicit confirmReport approval. The description contains the reported issue only, excluding conversation history, unrelated notes and personal information. Fiveable resolves content and reporter identity. Matching requestId retries are idempotent.
| Name | Required | Description | Default |
|---|---|---|---|
| contentId | Yes | Exact guide/question/FRQ ID, or the exact canonical key-term path from Fiveable. | |
| requestId | Yes | ||
| contentType | Yes | ||
| description | Yes | Only the issue the student wants reported; no conversation transcript or unrelated personal notes. | |
| confirmReport | Yes | True only after the student explicitly requests sending this issue to Fiveable. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare write semantics and openWorldHint, but the description adds real context beyond them: support may follow up, only the reported issue is transmitted (no conversation history or PII), Fiveable resolves content and reporter identity, and matching requestId retries are idempotent. It repeats the idempotentHint already in annotations, so it falls short of fully additive disclosure, but the data-handling and follow-up behavior are genuinely new information.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with purpose and approval requirement, and each sentence carries distinct information (gating, scope of payload, resolution ownership, idempotency). Slightly dense/packed, but no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive-capable write tool with an output schema present, the description covers the confirmation gate, payload scope, privacy handling, and retry semantics, which is most of what an agent needs. It does not say what happens on failure or how support follow-up is surfaced, but the output schema can carry that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 60%, so the description must carry some weight, and it does: it constrains 'description' to the issue only (excluding transcripts and unrelated personal notes) and clarifies 'confirmReport' as gated on an explicit student request. It also notes idempotency keyed on requestId. contentType and contentId semantics are left to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Submits a caller-requested content issue to Fiveable support') with a clear destination. No sibling tool in the list does anything comparable, so an agent can route to it unambiguously. The purpose is stated in the first clause without requiring the title.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a hard precondition ('Requires explicit confirmReport approval') and scopes intake to caller-requested issues, which implies when the tool may be used. However, it does not contrast with any alternative action (e.g., answering the student directly or searching content first), so usage guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rescore_essayRetry a failed essayAIdempotentInspect
Retries background AI grading for one essay whose scoring failed. Existing successful scores cannot be changed by this operation.
| Name | Required | Description | Default |
|---|---|---|---|
| requestId | Yes | A unique id you generate for this action. Reuse it only when retrying the same action. | |
| sessionId | Yes | Grading session id from list_grading_sessions or create_grading_session. | |
| essayHandle | Yes | Essay handle from get_grading_results, e.g. E-3b91c0. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish it is a non-read-only, non-destructive, idempotent write. The description adds a real behavioral constraint beyond them: the operation only affects failed scoring and leaves successful scores untouched. It does not say what happens (error vs. no-op) if invoked on an already-scored essay, which is the remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, with the retry action front-loaded and the scope limitation immediately after. Nothing repeats the title or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need no explanation, and all three parameters are documented in the schema. The description supplies the failure-retry framing and the non-mutating-of-successful-scores constraint, which is everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with all three required parameters fully documented, including patterns and provenance for sessionId and essayHandle. The description contributes no additional parameter meaning, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Retries background AI grading for one essay'), and the qualifier 'whose scoring failed' plus the scoping sentence implicitly carve it out from the score-changing sibling override_essay_score. It stops short of naming that alternative explicitly, so an agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'for one essay whose scoring failed' gives a clear triggering condition, and 'Existing successful scores cannot be changed by this operation' supplies a when-not condition. No sibling is named as the alternative for the case the exclusion rules out, so the guidance is clear but not fully routed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
restore_assignment_draftRestore assignment draftADestructiveInspect
Restores an assignment draft to its previous state using an undoId from a draft update.
| Name | Required | Description | Default |
|---|---|---|---|
| undoId | Yes | undoId from update_assignment_draft. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false, and 'restores ... to its previous state' is consistent with that mutation profile. The description adds a bit of context about the undo semantics but does not disclose the key behavioral traits an agent would want, such as whether the undoId is single-use/consumed, what exactly is overwritten, or whether the restore is reversible.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence that front-loads the action and its required input, with zero padding or redundancy. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be described, and the annotations carry the safety profile. The origin of undoId is covered, but the description stops short of noting the undoId's lifecycle (single-use), which would make it fully self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the parameter already carries a pattern and the description 'undoId from update_assignment_draft.' The tool description restates that same origin without adding format, validity-window, or edge-case detail, so it does not exceed the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('restores') and resource ('assignment draft') and specifies the mechanism ('previous state using an undoId from a draft update'). It ties itself to the relevant sibling update_assignment_draft, giving the agent a usable anchor, though it doesn't explicitly name what it is not versus other restore/update tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'using an undoId from a draft update' implies it must follow update_assignment_draft, giving implied usage context. However, there is no explicit when-to-use/when-not guidance, no indication of how long the undoId stays valid, and no statement about alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
restore_studentsRestore removed studentsAInspect
Restores removed roster students using a valid undoId, preserving their original roster entries.
| Name | Required | Description | Default |
|---|---|---|---|
| undoId | Yes | undoId from remove_students. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false, and destructiveHint=false. The description adds useful context that the restore is non-destructive ('preserving their original roster entries'), but it does not explain what happens on an invalid/expired undoId or that reusing an undoId after success will fail, which matters given idempotentHint=false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence with the operation front-loaded and no filler. It is appropriately sized, though it is terse enough to leave the idempotency edge case unaddressed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is not required, and there is only one parameter with full schema coverage. The description covers the core action and the non-destructive nature; only edge-case behavior (invalid or reused undoId) is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% (the undoId pattern and its 'from remove_students' origin are documented in the schema), so the baseline is 3. The description mentions the parameter's role but adds no format, validity, or lifetime details beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Restores') and resource ('removed roster students') with the qualifying mechanism ('using a valid undoId'). An agent can distinguish it from remove_students as the inverse operation, though the description never names that sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the requirement of a 'valid undoId' (and the schema note 'undoId from remove_students') signals this must follow a remove_students call, but there is no explicit when-to-use or when-not-to-use guidance, nor any mention of undo-window expiry.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_submission_feedbackSave submission feedbackADestructiveIdempotentInspect
Saves teacher-reviewed written feedback on one submission, overall or per section, without changing scores. Students see it once results are released, including immediately if already released. Returns previous text for restoration.
| Name | Required | Description | Default |
|---|---|---|---|
| assignmentId | Yes | Assignment id from list_coursework or list_submissions_to_review. | |
| itemFeedback | No | ||
| submissionId | Yes | Submission id from list_submissions_to_review. | |
| overallFeedback | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true and idempotentHint=true, and the description adds critical context beyond these: it explains the side effect (students see feedback upon release), the return value ('returns previous text for restoration'), and the non-effect on scores. This is rich, non-contradictory behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each front-loaded and earning its place: purpose, visibility semantics, and return value. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with annotations, the description covers visibility, non-scoring behavior, and return value. It lacks details on parameter constraints (e.g., max lengths, item structure) and doesn't explain error conditions, but output schema handles returns and annotations cover safety hints.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50%, with assignmentId and submissionId documented but overallFeedback and itemFeedback (including nested itemId/text) not described. The description implies overall or per-section feedback, which loosely maps to overallFeedback and itemFeedback, but adds no syntax, format, or constraints beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (saves) and resource (teacher-reviewed written feedback on one submission), and clarifies scope: overall or per section. It also distinguishes itself from scoring siblings like score_submission and override_essay_score by explicitly stating it does not change scores.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clarifies when the feedback is visible to students (once results are released, or immediately if already released) and that it doesn't change scores, which routes the agent away from scoring tools. However, it doesn't explicitly name alternative tools or state when not to use this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
score_submissionScore submissionBDestructiveIdempotentInspect
Sets the teacher's score on sections of one submission (submission item ids), optionally with written feedback. Works in any scoring mode. Says whether the student sees it yet; to undo, set the sections back to the previous scores this returns.
| Name | Required | Description | Default |
|---|---|---|---|
| itemScores | Yes | ||
| assignmentId | Yes | Assignment id from list_coursework or list_submissions_to_review. | |
| itemFeedback | No | ||
| submissionId | Yes | Submission id from list_submissions_to_review. | |
| overallFeedback | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true and idempotentHint=true, so the agent knows it's a destructive but repeatable write. The description adds that it returns whether the student sees the score yet and how to undo, which is helpful beyond the annotations. However, it doesn't explain the destructiveness beyond 'to undo, set the sections back,' leaving the precise mutation scope unclear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three concise sentences, front-loading the core action and then adding operational details. No fluff, though the last clause about undo is slightly crammed in but still earns its place by hinting at reversibility.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (5 params, destructive write) and the presence of an output schema, the description covers the basic action and undo behavior but omits potential side effects, permission requirements, or how scores interact with grading sessions. It's minimally adequate but leaves gaps for an agent to infer safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 40% schema description coverage, the description doesn't compensate with additional meaning for all parameters. It mentions 'sections' and 'submission item ids' mapping to itemScores.itemId, but doesn't clarify points, itemFeedback, overallFeedback, or their constraints. The schema handles required fields and patterns; the description adds little beyond the core concept.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource: 'Sets the teacher's score on sections of one submission (submission item ids), optionally with written feedback.' This clearly distinguishes it from siblings like override_essay_score and override_submission_part by emphasizing item-level scoring. It doesn't explicitly name those siblings, but the detail about submission item ids is strong enough to avoid confusion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives. The phrase 'Works in any scoring mode' implies broad applicability but doesn't tell the agent when to choose score_submission over override_submission_part or override_essay_score. The undo instruction is useful but assumes prior context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_contentSearch Fiveable contentARead-onlyIdempotentInspect
Keyword search across every Fiveable study guide, unit, subject hub and key term. Suitable when the relevant subject or guide is not known. Returns titles and canonical references rather than full content, so it is cheap.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | Yes | Maximum results to return. | |
| query | Yes | What to search for, e.g. "photosynthesis light reactions". | |
| intent | No | Privacy-safe reason for the request. Use explain, quiz, notes, frq, or research; never send the student's raw prompt for analytics. | |
| subjectSlug | No | Optional: restrict the search to one subject. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint safety. The description adds useful behavioral context beyond those: it returns only 'titles and canonical references rather than full content' and notes that the search is cheap. This helps an agent set expectations about cost and output shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, no filler. The main action and scope come first, followed by the use case and a key return-value caveat. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the presence of a full output schema, rich annotations, and a 100%-covered input schema, the description provides the missing context: the cross-cutting search scope, when to prefer it, and the lightweight return format. Nothing essential is left for the agent to infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each parameter is already documented. The description adds general keyword-search framing, but it does not add meaning beyond the schema for query, limit, intent, or subjectSlug. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Keyword search across every Fiveable study guide, unit, subject hub and key term.' This clearly identifies the tool's scope and distinguishes it from sibling get_ and list_ tools that operate on specific known resources.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit guidance about when to use the tool: 'Suitable when the relevant subject or guide is not known.' It implies the alternative is to use more targeted tools when a known resource exists, though it does not name those alternatives explicitly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_question_bankSearch question bankARead-onlyIdempotentInspect
Searches Fiveable's question bank for one subject (the same bank as the assignment builder): multiple choice, stimulus sets, SAQs and FRQs, filtered by unit, topic, skill, format and stimulus, optionally hiding questions this teacher already published. For AP Calculus, courseVariant scopes AB to exclude BC-only units, topics and FRQs. Returns ids, short stems and alignment, never answers. Filters accept exact unit/topic/skill labels from the filter lists in the result. Cheap.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | mcq = single question, question-set = stimulus set, saq, or frq (LEQ, DBQ and other FRQs). | |
| unit | No | Exact unit label from filters.units. | |
| limit | Yes | ||
| query | No | Words to match in the question. | |
| skill | No | Exact skill label from filters.skills. | |
| topic | No | Exact topic label from filters.topics. | |
| format | No | ||
| offset | Yes | ||
| stimulus | No | ||
| subjectSlug | Yes | Subject, e.g. ap-bio or apush. | |
| courseVariant | No | AP Calculus only: ap-calc-ab or ap-calc-bc, the course of the class or draft you are building for. | |
| excludeAssigned | Yes | Hide questions this teacher already published. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), so the description's job is to add beyond that. It does: it discloses that the tool returns 'ids, short stems and alignment, never answers' (a meaningful boundary), that result includes filter lists, and that it is 'Cheap' (a cost cue). These are concrete behavioral traits not derivable from annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded, leading with the core action and resource before layering scoping and return-format details. Every sentence carries information, though the filter enumeration runs slightly long and could compress the parameter list already present in the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter tool with an output schema, this is close to complete: it names the subject scope, the type/filter dimensions, the AP Calc variant rule, the excludeAssigned behavior, and the answer-free return contract. An agent has enough to call it correctly, with only pagination and sibling disambiguation left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 67% schema coverage, the baseline is 3, but the description adds genuine semantic value: it explains that courseVariant scopes AP Calc AB to exclude BC-only units, topics and FRQs, and that unit/topic/skill must be exact labels drawn from the returned filter lists (a cross-call dependency). It does not explain offset/limit pagination, so it stops short of a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Searches Fiveable's question bank') and scopes it to one subject, enumerating the four question types and the filter dimensions. It also grounds the tool by noting it is 'the same bank as the assignment builder,' which helps route the agent. It does not, however, distinguish itself from the sibling search_content, which an agent could easily confuse with it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the reference to the assignment builder and 'optionally hiding questions this teacher already published' convey the intended context, and the note that filters accept exact labels 'from the filter lists in the result' implies chaining. There is no explicit when-to-use vs. when-not, and no named alternative such as search_content or get_questions_for_materials.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_class_assignment_defaultsSet class assignment defaultsADestructiveIdempotentInspect
Sets a class's assignment defaults (scoring mode, timer, attempts, result release, notifications) so new assignments for that class start with them. Replaces the previous defaults and returns them for undo.
| Name | Required | Description | Default |
|---|---|---|---|
| classId | Yes | Class id from list_classes. | |
| defaults | Yes | The class's new defaults; unset fields fall back to Fiveable's. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, so the replacement semantics are partly covered. The description adds genuine value beyond them by spelling out that it 'Replaces the previous defaults and returns them for undo' and by disclosing the undo affordance, which the structured fields do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero waste, and the core action plus its effect are front-loaded before the replacement/undo clause. No padding or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With annotations covering the mutation profile and an output schema present, the description needn't explain return values, and it still flags replacement and undo. Complete enough for an agent to call it correctly, with only the missing usage routing as a gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents classId and every nested defaults field. The description's parenthetical mention of the field names adds only marginal meaning beyond the schema, matching the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Sets a class's assignment defaults') and enumerates the affected fields (scoring mode, timer, attempts, result release, notifications). It clearly conveys scope ('so new assignments for that class start with them'), though it does not explicitly contrast with a sibling, so 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: the description explains the defaults apply to 'new assignments for that class,' which hints at when this tool matters. However, it names no alternative or precondition (e.g., versus update_assignment_draft or update_class), leaving routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_gradingStart gradingAIdempotentInspect
Starts background AI grading for unscored essays on the session's Fiveable rubric, generating that rubric first if needed. Returns immediately with job status. retryFailed includes previously failed essays. Scores remain drafts until teacher approval.
| Name | Required | Description | Default |
|---|---|---|---|
| requestId | Yes | A unique id you generate for this action. Reuse it only when retrying the same action. | |
| sessionId | Yes | Grading session id from list_grading_sessions or create_grading_session. | |
| retryFailed | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true and destructiveHint=false, so the safety profile is covered. The description goes further by disclosing real behavioral traits: it is asynchronous ('returns immediately with job status'), it has a side effect (generating the rubric first if needed), and its output is staged ('scores remain drafts until teacher approval'). Minor gaps remain (no guidance on duplicate concurrent runs), but this is meaningful context beyond structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the core action and then the async, parameter, and approval semantics. Nothing is padded, though the final draft-status sentence could arguably merge with the async sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and annotations cover safety hints. The description supplies the async/background nature, rubric auto-generation, and draft-approval lifecycle, which is sufficient for an agent to invoke this correctly; only concurrency/duplicate-run behavior is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%; sessionId and requestId are already documented in the schema with patterns and sourcing hints. The description adds the one thing the schema omits — that retryFailed controls inclusion of previously failed essays — which is genuine semantic value over the structured data.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Starts background AI grading for unscored essays on the session's Fiveable rubric') with clear scope. It is readily distinguishable from siblings like rescore_essay, score_submission, or get_grading_progress by the word 'background' and the 'unscored essays' scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'unscored essays' scope and the 'retryFailed includes previously failed essays' note imply when to use it, but no alternative tool is named and no explicit when-not guidance is given. An agent must infer that get_grading_progress is the polling counterpart and that rescore_essay is for individual re-grades.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_assignment_draftUpdate assignment draftADestructiveInspect
Changes a draft's title, instructions, AP Calculus course, classes, dates, settings or sections. Existing section ids identify sections to remove. Published assignments cannot be changed here. Returns an undoId for restoring the previous draft.
| Name | Required | Description | Default |
|---|---|---|---|
| dueAt | No | Default due time for every class. | |
| title | No | ||
| classes | No | Replaces the class targets. | |
| settings | No | Merged into the current settings. | |
| addSections | No | ||
| availableAt | No | When it opens, ISO with offset. | |
| assignmentId | Yes | ||
| instructions | No | ||
| courseVariant | No | AP Calculus only: switch the draft to ap-calc-ab or ap-calc-bc. Its multiple choice, free-response and diagnostic sections must be removed in the same call (add new ones there too); readings stay. | |
| resultRelease | No | When students see results. | |
| removeSectionIds | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry the safety profile (destructiveHint=true, idempotentHint=false), and the description adds genuinely non-derivable behavior: existing section ids mean removal, settings are merged rather than replaced, and an undoId is returned for restoring the prior draft. It stops short of describing auth needs or what happens to omitted fields beyond settings.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the mutable surface, then the two operational caveats (published lock, section-id semantics) and the return value. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive 11-parameter mutation with a rich schema and an output schema, the description covers the highest-risk gotchas (published immutability, section removal semantics, undo path). It omits any indication that settings are merged vs. replaced only for the top-level object and says nothing about whether title/instructions are replaced wholesale, a minor residual gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 55% schema coverage across 11 parameters, the description compensates by mapping the mutable fields to parameters and clarifying the tricky ones: 'Existing section ids identify sections to remove' for removeSectionIds, merge semantics for settings, and the courseVariant constraint that matching sections must be removed in the same call. Nested class/section details are left to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('changes') and resource ('a draft') and enumerates the mutable surface: title, instructions, course, classes, dates, settings, sections. An agent can distinguish it from create_assignment_draft and restore_assignment_draft without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear exclusion ('Published assignments cannot be changed here') and implies the undo path via restore_assignment_draft, but never names that sibling or states which tool should be used for published assignments instead. Clear context, no explicit routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_classUpdate classADestructiveIdempotentInspect
Renames a class, changes its subject, or sets an AP Calculus class's course (AB or BC). Returns the previous values for reversal through a subsequent class update.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| classId | Yes | Class id from list_classes. | |
| subjectSlug | No | ||
| courseVariant | No | AP Calculus only: ap-calc-ab or ap-calc-bc. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true and readOnlyHint=false, so the safety profile is covered. The description adds the reversal mechanism (previous values returned for a subsequent update), which is genuinely useful, but says nothing about whether omitted fields are preserved, permission requirements, or side effects on enrolled students.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the operations before the return-value note. Every clause carries information, though the reversal sentence could arguably be trimmed since an output schema exists.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity, four-parameter mutation tool with annotations and an output schema, this covers the operations, the conditional course field, and the reversal path. The main remaining gap is partial-update semantics (what happens to unspecified fields), which an agent would otherwise have to assume.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With only 50% schema description coverage (name and subjectSlug undocumented in the schema), the description compensates by mapping each parameter to its effect: name → rename, subjectSlug → subject change, courseVariant → AB/BC course. It also covers the AP-Calculus-only restriction on courseVariant. It adds no format or constraint detail for the string fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('renames', 'changes', 'sets') tied to the class resource and enumerates the exact mutable fields (name, subject, AP Calculus course variant). It is clearly distinct from siblings like create_classes or archive_class, though it never explicitly contrasts itself with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the concrete conditions under which each capability applies, including the important constraint that courseVariant only makes sense for AP Calculus, and hints that the returned previous values enable reversal via a follow-up call. It stops short of naming alternatives or stating when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_grading_rubricChange grading promptADestructiveIdempotentInspect
Changes a session's Fiveable question or teacher-provided prompt and AP question type before any essay is scored. Regenerates the rubric using that type's AP rubric and point total. Returns the previous prompt for restoration.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | No | ||
| sessionId | Yes | Grading session id from list_grading_sessions or create_grading_session. | |
| fiveableQuestionId | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true and readOnly=false, so the mutation profile is covered externally. The description adds genuinely useful context by disclosing that the rubric is regenerated from the type's AP rubric and point total, and that the previous prompt is returned for restoration. It never states what existing rubric/scoring data is discarded, so the destructive consequence remains under-explained.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with the action front-loaded, followed by the effect and then the return value. No filler, and each sentence contributes a distinct fact.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a nesting-heavy mutation tool, the description covers purpose, timing precondition, side effect and return. With an output schema present it rightly does not enumerate return fields, and the only real gap is silence on what happens to already-existing rubric data.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is low (33%), so the description carries more burden. It clarifies the two prompt modes (a Fiveable question versus a teacher-provided prompt) and the role of the AP question type, which maps onto prompt/text, frqType and fiveableQuestionId. It adds nothing for sessionId or the nested images array, so it only partially compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Changes) and a precise resource: a session's question/teacher prompt and AP question type, plus the follow-on effect of regenerating the rubric. It is clearly distinguishable from read-only siblings like get_grading_rubric, though it never names the alternatives it competes with (e.g. update_assignment_draft).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"before any essay is scored" gives a concrete timing precondition that tells the agent when this tool is applicable, which is more than mere implied usage. It stops short of naming exclusions or the alternative tools to use once scoring has begun, so it lands at clear-context-but-no-exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_studentUpdate studentADestructiveIdempotentInspect
Fixes a student's name on a class roster (the old name is kept as an alias so past work still matches), or sets the email used for assignment notifications. Students with Fiveable accounts can be renamed but their email comes from their account.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Full name as the teacher wants it shown. | |
| No | |||
| handle | Yes | Student handle from get_class, e.g. S-4f2a1c. | |
| classId | Yes | Class id from list_classes. |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | Readable tool result text for clients that consume structured output. |
| status | Yes | Operation status: completed, pending, partial, failed, unavailable, or a domain-specific outcome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructive=true, idempotent=true, readOnly=false, openWorld=false. The description adds meaningful context beyond them: the old name is kept as an alias so past work still matches (qualifying the destructive hint), the email drives assignment notifications, and Fiveable accounts block email changes. It does not describe partial-failure or permission behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences, front-loaded with the rename behavior and the alias side-effect, then the email behavior and its account constraint. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. For a mutation tool whose annotations cover the safety profile, the description supplies the key behavioral facts (alias retention, notification-email purpose, Fiveable constraint) and is nearly complete; only edge cases like what happens when both name and email are omitted are unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75% and the schema itself documents name, handle, and classId. The description touches name and email semantics but adds little beyond the schema's own 'Full name as the teacher wants it shown' and handle/classId descriptions. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States specific verbs and resources ('Fixes a student's name on a class roster', 'sets the email used for assignment notifications'), making the update-vs-add/remove distinction clear against siblings like add_students and remove_students. It stops short of naming those siblings explicitly, so it lands at a strong 4 rather than a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the two update contexts (rename vs. set notification email) and gives a clear conditional constraint: students with Fiveable accounts can be renamed but their email comes from the account. It never names an alternative tool for cases it does not cover, so it is clear context without explicit routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
77 tool updates
- First observed
add_students - First observed
approve_essays - First observed
archive_assignment - First observed
archive_class - First observed
confirm_correct_answer_key - First observed
confirm_publish_assignment - First observed
confirm_push_grades_to_google_classroom - First observed
confirm_release_results - First observed
confirm_update_publication - First observed
create_assignment_draft - First observed
create_classes - First observed
create_grading_session - First observed
duplicate_assignment - First observed
export_grading_results - First observed
get_assignment - First observed
get_assignment_question_breakdown - First observed
get_assignment_status - First observed
get_cheatsheet - First observed
get_class - First observed
get_class_insights - First observed
get_content_sections - First observed
get_essay_feedback - First observed
get_exam_options - First observed
get_frq - First observed
get_grading_progress - First observed
get_grading_results - First observed
get_grading_rubric - First observed
get_guide_reading_questions - First observed
get_guided_notes - First observed
get_key_term - First observed
get_my_teacher_workspace - First observed
get_questions_for_materials - First observed
get_reteach_gaps - First observed
get_skill_insights - First observed
get_student_insights - First observed
get_study_guide - First observed
get_subject_outline - First observed
get_submission - First observed
get_unit_overview - First observed
import_essays_from_files - First observed
import_essays_from_google_classroom - First observed
import_essays_from_text - First observed
link_grading_to_class - First observed
list_cheatsheets - First observed
list_classes - First observed
list_coursework - First observed
list_frqs - First observed
list_google_classroom_work - First observed
list_grading_sessions - First observed
list_key_terms - First observed
list_subjects - First observed
list_submissions_to_review - First observed
list_teaching_resources - First observed
override_essay_score - First observed
override_submission_part - First observed
prepare_correct_answer_key - First observed
prepare_publish_assignment - First observed
prepare_push_grades_to_google_classroom - First observed
prepare_release_results - First observed
prepare_update_publication - First observed
preview_assignment_as_student - First observed
preview_question - First observed
remove_students - First observed
report_content_issue - First observed
rescore_essay - First observed
restore_assignment_draft - First observed
restore_students - First observed
save_submission_feedback - First observed
score_submission - First observed
search_content - First observed
search_question_bank - First observed
set_class_assignment_defaults - First observed
start_grading - First observed
update_assignment_draft - First observed
update_class - First observed
update_grading_rubric - First observed
update_student
Related MCP Connectors
AP study content, practice, FRQ feedback, progress, plans, and student tools for AI apps.
Generate tailored quality criteria and scoring guides from your task descriptions. Refine objectiv…
Your English homework, exercises and flashcards — with your real teacher.
Misconception detection for AP Chemistry & Physics 1. 90% catch rate vs 24% baseline.
Related MCP Servers
- AlicenseAqualityBmaintenanceProvides AI-powered educational tools for teachers including lesson planning, quiz creation, student progress analysis, learning path recommendations, and rubric generation. Enables educators to automate administrative tasks and personalize learning experiences through natural language interactions.58 npmMIT
- FlicenseNot gradedqualityCmaintenanceEnables teachers to use AI agents for lesson preparation, grading, and administrative tasks through tools for class averages, school resources, and assessment prompts.-
- FlicenseNot gradedqualityCmaintenanceEnables AI assistants to interact with Google Classroom to check submission statuses, list students and assignments, and create coursework, announcements, and topics through natural language.-
- AlicenseBqualityCmaintenanceEnables teachers to manage Google Classroom courses, topics, assignments, materials, and announcements locally, including creating drafts, scheduling posts, and attaching files via MCP.211MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.